EconBase
← Back to paper

Asymptotic Theory of Principal Component Analysis for High-Dimensional Time Series Data under a Factor Structure

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

174,943 characters · 9 sections · 141 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Principal Component Analysis .3cm for High-Dimensional Approximate Factor Models in Time Series: Assumptions, Asymptotic Theory, and Identification

center[center omitted — 91 chars of source]
abstractWe consider estimation of large approximate factor models in high-dimensional panels of stationary time series using Principal Component Analysis (PCA). We review the key results establishing the necessary and sufficient conditions for consistency and asymptotic normality of the estimators. We compare two equivalent approaches to PCA and present the asymptotic properties associated with each formulation. Special emphasis is placed on identification, where we discuss the restrictions required to uniquely determine factors and loadings and examine their consequences for statistical inference. Keywords: Large Approximate Factor Model; Principal Component Analysis; Identification.

\thispagestyle{empty}

\footnotetext{Department of Economics - Universit\`a di Bologna. Email: [email removed] \\ I thank Haeran Cho and Philipp Gersing for helpful comments and discussions.}

Introduction

Consider a $T$-dimensional realization of $n$ zero-mean time series: $\{x_{it},\,i=1,\ldots, n,\, t=1,\ldots, T\}$. We say that $x_{it}$ follows an $r$-factor model if

align[align omitted — 127 chars of source]

where $\bm\lambda_i:=(\lambda_{i1}\cdots\lambda_{ir})^\prime$ and $\mathbf F_t:=(F_{1t}\cdots F_{rt})^\prime$ are the $r$-dimensional vectors of loadings for series $i$ and factors, respectively, and $r\ll \min(n,T)$. We call $ e_{it}$ the idiosyncratic component and ${C}_{it}:=\bm \lambda_i^\prime \mathbf F_t$ the common component.

While for small fixed $n$ it can be reasonably assumed that there is no idiosyncratic cross-sectional correlation and, thus, we say that the model is exact, in a high-dimensional setting the idiosyncratic components are likely to be weakly cross-sectionally correlated, since, even if the common factors capture the main covariances, some weaker local covariances are likely to remain unexplained. In this case we say that the factor model is approximate.

In an approximate factor model the common and idiosyncratic component can be identified, provided that the $r$ eigenvalues of the common component covariance matrix diverge as $n\to\infty$, while the $n$ idiosyncratic eigenvalues are bounded for all $n\in\mathbb N$ chamberlainrothschild83. Hence, identification of the model, and then also of the number of factors $r$, is possible only when letting $n\to\infty$, a feature sometimes known as the blessing of dimensionality. However, it is well known that, unless further $r^2$ restrictions are imposed, the loadings and factors are identified only up to post- and pre-multiplication by an invertible $r\times r$ matrix and its inverse.

An approximate factor model is characterized by an eigen-gap in the population covariance matrix that widens as $n \to \infty$. Conversely, if the population covariance exhibits a widening eigen-gap as $n \to \infty$, the data can be said to follow an approximate factor model---see, e.g., chamberlainrothschild83, or gersing2023weak, or BH25, for various proofs. This is precisely the setting in which Principal Component Analysis (PCA) is most effective for dimensionality reduction, as it projects the data onto a subspace spanned by directions of maximum variation, by construction. It is then of no surprise that PCA is the most natural and common way to estimate factor models. Furthermore, unlike classical Maximum Likelihood methods, PCA is a non-parametric approach that relies solely on moment conditions rather than distributional assumptions. Moreover, it can be applied to weakly stationary time series without any modification from its standard formulation for independent data.

Specifically, PCA is formulated as the minimization of a quadratic loss. This optimization yields estimated loadings proportional to the $r$ normalized eigenvectors corresponding to the $r$ largest eigenvalues of the $n\times n$ sample covariance matrix, whose $(i,j)$th entry is given by $T^{-1}\sum_{t=1}^T x_{it} x_{jt}$, $i,j=1,\ldots, n$. The estimated factors are obtained by linear projection of the data onto these estimated loadings.

However, to guarantee the existence of a unique solution, the minimizers of the PCA objective must satisfy an additional set of $r^2$ constraints. There are two standard ways of imposing these constraints, each leading to estimated loadings that are equivalent up to a rescaling of the corresponding eigenvectors.

compactenum• If we impose the estimated loadings to be orthogonal and the estimated factors to be orthonormal, then the eigenvectors are rescaled by the square-root of the corresponding eigenvalues mardia1979multivariate,Bai03,FGLR09,FLM13. • If we impose the estimated loadings to be orthonormal and the estimated factors to be orthogonal, then the eigenvectors are rescaled by $\sqrt n$ stockwatson02JASA,stockwatson02JBES.

Definition A1 is the classical and most popular one and it is our main focus. Its theoretical properties, consistency and asymptotic normality, have been fully established by Bai03, when formulating the estimators using an equivalent, although less intuitive, definition, where the estimated factors are written as $\sqrt T$ times the $r$ normalized eigenvectors corresponding to the $r$ largest eigenvalues of the $T\times T$ matrix with entries given by $n^{-1}\sum_{i=1}^n x_{it} x_{is}$, $t,s=1,\ldots, T$. The estimated loadings are obtained by linear projection of the data onto the estimated factors. We shall denote this definition as B1 (while a fourth definition, B2, is obtained analogously from A2).

Some important issues arise from the above overview. First, while we need $n\to\infty$ to identify the model, the sample size $T$ plays no role for identification of the model. However, we still need $T\to\infty$ for estimation. It follows that, on the one hand, we need to consider double asymptotics $n,T\to\infty$ to ensure consistency, but, on the other hand, it is clear that $n$ and $T$ do not have a symmetric role and should not be treated on equal grounds. This implies that carrying out asymptotic analysis using the classical definition A1 (or A2) rather than the alternative B1 (or B2) could have benefits in terms of interpretation. In fact, there are cases in which considering PCA based on the definition of approach A1 (or A2) is clearly the most natural way forward for developing the theory. First, in presence of non-stationarities in the data, due, e.g., to change-points BCF18, exogenous massacci2017 or endogenous regime shifts urga2024estimation,BM22, unit roots bai04,baing04,OW19,BLL2,barigozzi2022testing, time-varying loadings motta2011locally,bates_plagborgmoller_stock_watson_2013,pelger_xiong_2022,mikkelsen_hillebrand_urga_2018. Second, when dealing with PCA on the spectral density matrix FHLR00,FHLR05,FHLZ17. Third, when considering factor models for matrix and tensor data, which are estimated via PCA on every mode YuHe2022,chang2023modelling,chen2023statistical,chen2024rank,barigozzi2025tail,barigozzi2026tensor. Fourth, when dealing with factors which are pervasive along the time dimension so that the idiosyncratic components are white noise (LamYao2012,ZRY,chen2020constrained; chen2022factor). Fifth, when estimating exact factor models with $n$ is fixed so that only the loadings, but not the factors, can be estimated consistently via PCA anderson1963asymptotic.

A second issue relates to identification. Since, in general, the true loadings and factors are not identified, it is not clear what the target of the estimator is when it comes to defining consistency. This means that no inference can be carried out, unless one can provide a clear theoretical study of the impact and the meaning of any imposed identifying assumptions. This topic is not much explored in the literature of high-dimensional factor models for time series, besides the works by baing13 and BW15.

In this paper we provide a full treatment of PCA for factor models, by reviewing, discussing, completing, and extending the asymptotic existing results, and, thus, also clarifying the above issues. Specifically, we cover the following topics.

Derivation of PCA. We show that the two standard definitions of PC estimators---those based on the classical $n \times n$ sample covariance matrix and their counterparts based on the $T \times T$ covariance matrix---are formally equivalent. Although elementary, this equivalence is often overlooked. Establishing it explicitly clarifies how the PC solution arises directly from the eigenvalue--eigenvector decomposition of these matrices, without resorting to iterative procedures such as alternating least squares. See Section (ref) and Appendix (ref), where we derive the PC estimators in definition A2.

Assumptions. We present the full set of necessary and sufficient assumptions needed to derive the asymptotic properties of PC estimators in a purely time-series framework. These assumptions are fewer in number and more transparent than those typically imposed in Bai03. For each assumption---such as the various moment and dependence conditions---we provide intuitions and motivations to clarify their role in the analysis. See Section (ref) and Appendix (ref), where we review primitive conditions ensuring consistent estimation of the covariance matrix.

Asymptotic theory. We derive the asymptotic properties of the classical PC estimator as computed under definition A1.

compactenum[(i)] • We establish consistency under a minimal set of assumptions, using a direct argument based on Weyl's inequality MK04 and the Davis Kahan theorem yu15. Orthogonality---i.e., uncorrelatedness---between common and idiosyncratic components suffices. See Section (ref). • We show that the sharper convergence rates of Bai03 can be obtained by imposing additional assumptions involving higher-order dependence. See Section (ref). • We provide several alternative representations for the linear transformation matrix linking true and estimated loadings/factors. See Section (ref). • Under the extended assumptions, we prove asymptotic normality and show that, under full identification of loadings and factors, the leading term of the asymptotic expansion coincides with the Ordinary Least Squares (OLS) expansion. See Section (ref). • We further show that the same relationship between PC and OLS holds for the formulation B1 in Bai03. See Section (ref) and Appendix (ref), where asymptotic normality is also established for the PC estimator in definitions A2 and B2.

Identification. We examine several identifying assumptions commonly imposed in exploratory factor analysis. These assumptions define the true loadings and factors as the population PCs of the common component and, in principle, should enable valid inference. However, we show that such restrictions either slow convergence---thereby affecting asymptotic normality---or are overly stringent, as they effectively would hold only if the factors were deterministic. See Section (ref).

Proofs. In Appendix (ref) we report the proofs of four main results---i.e., Propositions (ref), (ref), (ref), and (ref)---which are either simpler than the original ones or entirely new. All other proofs are collected in Appendices (ref) and (ref).

Estimation of factor models via Principal Component Analysis

In this section we present PCA estimation of a factor model. We assume to observe $n$ zero-mean time series over $T$ periods following the factor model, as in (ref), which in vector notation reads as \[ \mathbf x_t = \bm\Lambda\mathbf F_t +\bm e_t,\quad t=1,\ldots, T, \] where $\mathbf x_t:=(x_{1t}\cdots x_{nt})^\prime$ and $\bm e_{t}:=( e_{1t}\cdots e_{nt})^\prime$ are $n$-dimensional vectors of observables and idiosyncratic components, respectively, and $\bm\Lambda:=(\bm\lambda_1\cdots\bm\lambda_n)^\prime$ is the $n\times r$ matrix of factor loadings. We call $\bm{C}_t:=\bm\Lambda\mathbf F_t$ the vector of common components. Furthermore, let $\bm X:=(\mathbf x_1\cdots\mathbf x_T)^\prime$ and $\bm E:=(\bm e_1\cdots\bm e_T)^\prime$ be $T\times n$ matrices of observables and idiosyncratic components, respectively, and $\bm F:=(\mathbf F_1\cdots\mathbf F_T)^\prime$ be the $T\times r$ matrix of factors.

The idea of PCs as the $r$ directions of best fit in an $n$-dimensional space is originally due to pearson1901 and the use of PCA for factor analysis dates back to hotelling1933analysis. The PC estimators of $\bm\Lambda$ and $\bm F$ must be such that they solve:

align[align omitted — 567 chars of source]

where $\bm L:=(\bm\ell_1\cdots\bm\ell_n)^\prime\in\mathbb R^{n\times r}$ and $\bm G:=(\mathbf G_1\cdots\mathbf G_T)^\prime\in\mathbb R^{T\times r}$ are generic loadings and factor matrices.

Clearly, for a given estimator of the loadings the estimator of the factors is obtained by least squares and viceversa. This motivates the classical way to derive the solution of (ref) by using an alternating least squares approach bai2020simpler. However, there is a simpler way to derive the solution. Indeed, we can restate (ref) in two equivalent ways which depend either only on the loadings or only on the factors. First, by substituting ${\bm G}=\bm X{\bm L}({\bm L}^\prime{\bm L})^{-1}$ in (ref), we obtain the equivalent problem:

align[align omitted — 343 chars of source]

which gives an estimator of the loadings, while the factors can be estimated in a second step by linear projection onto such estimator. Second, by substituting ${\bm L}=\bm X^\prime {\bm G}({\bm G}^\prime{\bm G})^{-1}$ in (ref), we obtain yet another equivalent formulation of (ref) or (ref):

align[align omitted — 345 chars of source]

which gives an estimator of the factors, while the loadings can be estimated in a second step by linear projection onto such estimator.

From (ref) and (ref) we see that the PC estimators are related to eigenvectors, but the above solutions are not unique unless we fix the scale of those eigenvectors. Indeed, the original problem in (ref) does not have a unique solution too. For, given any solution $(\widehat{\bm\Lambda}, \widehat{\bm F})$, we can find infinite other equivalent solutions given by $\widehat{\bm\Lambda}{\bm K}$ and $\widehat{\bm F}({\bm K}^\prime)^{-1}$, where ${\bm K}$ is any $r\times r$ invertible matrix. Therefore, to ensure uniqueness of the PCA solution we must solve (ref) when imposing $r^2$ identifying constraints. The two most common choices are: (a) $n^{-1}\bm L'\bm L$ being diagonal with distinct entries sorted in descending order and $T^{-1}\bm G'\bm G= \mathbf I_r$, or (b) $n^{-1}\bm L'\bm L=\mathbf I_r$ and $T^{-1}\bm G'\bm G$ being diagonal with distinct entries sorted in descending order.

Denote the classical the $n\times n$ sample covariance matrix obtained by averaging over time as (note that now $\mathbb{E}[\mathbf x_t]=\mathbf 0_n$ by assumption)

equation[equation omitted — 136 chars of source]

having its $r$ largest eigenvalues of $\widehat{\bm\Gamma}^x$ in the $r\times r$ diagonal matrix $\widehat{\mathbf M}^x$ (sorted in descending order) and the corresponding normalized eigenvectors as the columns of the $n\times r$ matrix $\widehat{\mathbf V}^x$. Then, a solution of (ref) is found as follows.

\paragraph{Approach A1.} If we impose $n^{-1} {{\bm L}^\prime {\bm L}}$ to be diagonal, then by construction each column of ${\bm L} ({\bm L}^\prime{\bm L})^{-1/2}$ is normalized and, by definition of eigenvectors and eigenvalues, the value of the objective function in (ref) must be the sum of the $r$ largest eigenvalues of $\widehat{\bm\Gamma}^x=T^{-1}{\bm X^\prime\bm X}$ divided by $n$, i.e., it must give $n^{-1}\text{tr}({\widehat{\mathbf M}^x})$. So, our estimator $\widehat{\bm\Lambda}$ must be such that $\widehat{\bm\Lambda} (\widehat{\bm\Lambda}^\prime\widehat{\bm\Lambda})^{-1/2}$ is the matrix of normalized eigenvectors corresponding the $r$ largest eigenvalues of $({nT})^{-1}{\bm X^\prime\bm X}$, that is:

align[align omitted — 337 chars of source]

Therefore, from (ref) we must have $\widehat{\bm\Lambda} = \widehat{\mathbf V}^x(\widehat{\bm\Lambda}^\prime\widehat{\bm\Lambda})^{1/2}$. So, by linear projection, the estimator of the factors is obtained as $\widehat{\bm F}= \bm X\widehat{\bm\Lambda}(\widehat{\bm\Lambda}^\prime\widehat{\bm\Lambda})^{-1}=\bm X \widehat{\mathbf V}^x(\widehat{\bm\Lambda}^\prime\widehat{\bm\Lambda})^{-1/2}$. Moreover, by imposing also the identifying constraint $T^{-1}\widehat{\bm F}^\prime\widehat{\bm F}=\mathbf I_r$, it follows that we must have $\widehat{\bm\Lambda}^\prime\widehat{\bm\Lambda} = \widehat{\mathbf M}^x$, which is diagonal as expected.

From the above reasoning it follows that the PC estimators are given by:

align[align omitted — 252 chars of source]

The factors defined in this way are the classical normalized PCs of $\bm X$ as in the definition of PCs dating back to hotelling1933analysis (see also lawleymaxwell71, mardia1979multivariate, jolliffe2002principal). In high-dimensional factor analysis this approach is considered, for example, by FGLR09, who prove consistency although not with the sharpest possible rate. \\

Consider the $T\times T$ sample covariance matrix obtained by averaging over the cross-section (note that now $\mathbb{E}[\bm x_i]=\mathbf 0_T$ by assumption)

equation[equation omitted — 130 chars of source]

having its $r$ largest eigenvalues collected in the $r\times r$ diagonal matrix $\widetilde{\mathbf M}^x$ (sorted in descending order) and the corresponding normalized eigenvectors as columns of the $T\times r$ matrix $\widetilde{\mathbf V}^x$. Then, another solution of (ref), equivalent to A1, is found as follows.

\paragraph{Approach B1.} If we impose $T^{-1}{{\bm G}^\prime{\bm G}}=\mathbf I_r$, then each column of $T^{-1/2}{ {\bm G}}$ is normalized and, by definition of eigenvectors and eigenvalues, the value of the objective function in the above maximization must be the sum of the $r$ largest eigenvalues of $\widetilde{\bm\Gamma}^x=n^{-1}{\bm X\bm X^\prime}$ divided by $T$, i.e., it must give $T^{-1}\text{tr}(\widetilde{\mathbf M}^x)$. So, our estimator $\widetilde{\bm F}$ must be such that $T^{-1/2}{\widetilde{\bm F}}$ is the matrix of normalized eigenvectors corresponding the $r$ largest eigenvalues of $(nT)^{-1}{\bm X\bm X^\prime}$, that is:

align[align omitted — 247 chars of source]

Therefore, the PC estimators of the factors and the loadings (obtained by linear projection) are

align[align omitted — 225 chars of source]

And, as expected, we get $n^{-1}{\widetilde{\bm \Lambda}^\prime\widetilde{\bm \Lambda}}=T^{-1}{\widetilde{\mathbf M}^x}$, which is diagonal. These are the estimators defined by Bai03, who proves consistency and asymptotic normality.

\paragraph{Equivalence.} As long as the number of factors $r$ is such that $r<\min(n,T)$, the $r$ largest eigenvalues of $\bm X^\prime \bm X$ and $\bm X \bm X^\prime$ coincide, i.e.,

equation[equation omitted — 95 chars of source]

and the objective function in (ref) and (ref) has the same value in its minimum. Since approaches A1 and B1 are based on the same identification conditions, the solutions (ref)-(ref) and (ref)-(ref) must then coincide, i.e., it must be that

equation[equation omitted — 131 chars of source]

This is also seen in the following way. Consider the Singular Value Decomposition

equation[equation omitted — 89 chars of source]

where $ \bm U$ and $ \bm V$ are orthogonal matrices of dimensions $T\times r$ and $n\times r$, respectively, and $\bm D$ contains the $r$ largest singular values of $\bm X$. Now, \[ \frac{\widehat{\bm\Gamma}^x}n=\frac{\bm X^\prime \bm X}{nT} = \bm V \bm D^2 \bm V^\prime + \bm E^\prime\bm E\;\text{ and }\; \frac{\widetilde{\bm\Gamma}^x}T=\frac{\bm X \bm X^\prime}{nT} = \bm U \bm D^2 \bm U^\prime + \bm E\bm E ^\prime \] and since the best rank-$r$ approximation of $(nT)^{-1/2}\bm X$ is $ \bm U \bm D\bm V^\prime$ EY36, it is straightforward to see that (using also (ref))

equation[equation omitted — 179 chars of source]

Moreover, under our assumptions on the eigenvalues of the common and idiosyncratic components (see (ref) and (ref) below), it holds that $(nT)^{-1/2}\bm X=(nT)^{-1/2} \bm F\bm\Lambda^\prime+o_{\mathrm P}(1)$. Then, the PC estimators must be such that $ \widehat{\bm F}\widehat{\bm\Lambda}^\prime=\widetilde{\bm F}\widetilde{\bm\Lambda}^\prime= \sqrt{nT} \,\bm U \bm D\bm V^\prime$. By solving (ref) using an alternating least squares approach jointly with (ref) and by imposing orthogonal loadings and orthonormal factors, bai2020simpler show that the estimator of the loadings is $\check{\bm\Lambda}= \sqrt n\, \bm V \bm D$, while the estimator of the factors is $\check{\bm F}\,=\sqrt T\, \bm U$. And from (ref) we see that $\check{\bm\Lambda}=\widehat{ \mathbf V}^{x}(\widehat{\mathbf M}^{x})^{1/2}=\widehat{\bm\Lambda}$ as in approach A1, and $\check{\bm F}=\sqrt T\, \widetilde{\mathbf V}^x=\widetilde{\bm F}$ as in approach B1. Moreover, by linear projection and again from (ref),

align[align omitted — 563 chars of source]

Hence, as stated in (ref), approaches A1 and B1 coincide. Alternative estimators are obtained when imposing orthonormal loadings and orthogonal factors. These are presented in approaches A2 and B2 in Appendix (ref).

\paragraph{Centering and standardizing.} We conclude by noticing that if the observed time series do not have zero mean, then, the above derivation would be the same provided we replaced each element of $\bm X$ with its centered version $x_{it}- \bar{x}_i$, $i=1,\ldots, n$, $t=1,\ldots T$, with $ \bar {x}_i:=T^{-1}\sum_{t=1}^T x_{it}$. Since it is commonly assumed that the factors and the idiosyncratic components have zero mean, i.e., $\mathbb{E}[\mathbf F_t]=\mathbf 0_r$ and $\mathbb{E}[\bm e_{t}]=\mathbf 0_n$, then, when the data has not zero-mean, model (ref) is still valid provided we add a constant, $\alpha_i$, $i=1,\ldots, n$. Clearly, $\bar{x}_i$ is a consistent estimator of $\alpha_i$.

Moreover, to avoid scale effects due to data having different units of measure, each element of $\bm X$ should always be standardized so that, in practice, the above derivation would be applied to $\frac{x_{it}-\bar x_i}{s_i}$, $i=1,\ldots, n$, $t=1,\ldots T$, with $s_i:=T^{-1}\sum_{t=1}^T (x_{it}-\bar x_i)^2$.

Finally, note that both centering and standardizing must always be carried out along the time dimension even when working with a $T\times T$ covariance. This is because the cross-sectional units are unlikely to be identically distributed.

Assumptions

We adapt to the high-dimensional setting the approach taken in most works in time series analysis. First, we define the model for an infinite dimensional process $\{x_{it},\, i\in\mathbb N,\, t\in\mathbb Z\}$. Second, the properties and the assumptions of the model are defined for the $n$-dimensional sub-process $\{x_{it},\, i=1\ldots, n,\, t\in\mathbb Z\}$ either holding for all $n\in\mathbb N$ or in the limit $n\to\infty$. Third, the properties of the estimators are derived for a given $nT$-dimensional realization $\{x_{it},\, i=1\ldots, n,\, t=1,\ldots, T\}$ and hold in the limit $n,T\to\infty$.

The infinite dimensional process follows an $r$-factor model if

align[align omitted — 122 chars of source]

This is the equivalent of (ref), but when we have infinite time series, so now $i\in\mathbb N$, each having zero-mean for simplicity. The common component is defined as $\bm{C}_t:=\bm\Lambda\mathbf F_t$.

We characterize the common component by means of the following assumption.

ass[common component] $\,$ \begin{compactenum}[(a)] • $\lim_{n\to\infty}\Vert n^{-1}\bm\Lambda^\prime\bm\Lambda-\bm\Sigma_{\Lambda}\Vert=0$, where $\bm\Sigma_{\Lambda}$ is $r\times r$ positive definite, and, for all $i\in\mathbb N$, $\Vert\bm\lambda_i\Vert\le M_\Lambda$ for some finite positive real $M_\Lambda$ independent of $i$. • For all $t\in\mathbb Z$, $\mathbb{E}[\mathbf F_{t}]=\mathbf 0_r$ and $\bm\Gamma^F:=\mathbb{E}[\mathbf F_t\mathbf F_t^\prime]$ is $r\times r$ positive definite and $\Vert\bm\Gamma^F\Vert\le M_F$ for some finite positive real $M_F$ independent of $t$. • \begin{inparaenum} • For all $t\in\mathbb Z$, $\mathbb{E}[\Vert \mathbf F_t\Vert^4]\le K_F$ for some finite positive real $K_F$ independent of $t$; \\ • $\mathrm P\text{-}\lim_{T\to\infty}\left\Vert T^{-1}\bm F^\prime\bm F -\bm\Gamma^F\right\Vert=0$. \end{inparaenum} • There exists an integer $N$ such that for all $n> N$, $r$ is a finite positive integer, independent of $n$. \end{compactenum}

Parts (a) and (b), are also assumed in Bai03. They imply that the loadings matrix has asymptotically maximum column rank $r$ (part (a)) and the factors have a finite full-rank covariance matrix (part (b)). Moreover, because of part (a), for any given $n\in\mathbb N$, all the factors have a finite contribution to each series (upper bound on $\Vert\bm\lambda_i\Vert$) and are cross-sectionally pervasive (full-rank of $\bm\Sigma_\Lambda$).

In Lemma (ref)(iv) we prove that parts (a) and (b) imply that the covariance matrix of the common component, $\bm\Gamma^{C}:=\mathbb{E}[\bm{C}_t\bm{C}_t^\prime]=\bm\Lambda\bm\Gamma^F \bm\Lambda^\prime$, has rank $r$ and its $j$th largest eigenvalue $\mu_j^{C}$, $j=1,\ldots, r$, is such that

equation[equation omitted — 110 chars of source]

for some finite positive reals $\underline C_j$ and $\overline C_j$ independent of $n$. This means we consider only pervasive strong factors. Any weaker rate of divergence would imply a specific ordering of the cross-sectional units BH25. Although this might be relevant in some datasets, it is not the subject of this work. In absence of any specific ordering of the cross-sectional items, the reasonable convergence rate to assume in part (a) is then $\sqrt n$, indicating some form of “cross-sectional” stationarity.

Part (c) is assumed also in Bai03. In part (c-i) we assume finite 4th order moments of the factors, and, in particular, we are bounding $\mathbb{E}[\Vert \mathbf F_t\Vert^4]=\sum_{j=1}^r\sum_{k=1}^r \mathbb{E}[F_{jt}^2F_{kt}^2]$. Part (c-ii) is very general, it simply says that the sample covariance matrix of the factors is a consistent estimator of its population counterpart $\bm\Gamma^F$. In principle, we do not have to require weak stationarity for part (c-ii) to hold, and we can have $\bm\Gamma^F$ that depends on $t$ as long as $M_F$ in part (b) is independent of $t$ (hence the reason for stating it explicitly). This is, for example the case when we are in presence of structural breaks/change-points (see, e.g., BCF18 and duan2022quasi) or regime shifts (massacci2017,BM22).\footnote{Unit roots are however excluded since in that case $\{\mathbf F_t\}$ does not have a finite covariance, and everything that is considered in this paper would hold for properly differenced data.} This assumption is a high-level one and it requires introducing the sample size $T$. In Appendix (ref) we show that if the process $\{\mathbf F_t\}$ is either linear, or has summable 4th order cumulants, or has bounded physical dependence wu05, or is strongly mixing, then part (c-ii) holds, possibly even in mean-square. There it is also shown that the convergence rate is $\sqrt T$, as it should be expected.

Notice that, on the one hand, we only consider non-random factor loadings for simplicity and by analogy with classical factor models.\footnote{As in Bai03, the generalization to the case of random loadings is possible provided that: we assume convergence in probability in part (a), for all $i\in\mathbb N$, $\mathbb{E}[\Vert \bm\lambda_i\Vert^4]\le K_\Lambda$, for some finite positive real $K_\Lambda$ independent of $i$, and $\{\bm\lambda_i,\, i\in\mathbb N\}$ is an independent sequence.} But, on the other hand, we explicitly treat the factors as random variables. This has important implications when it comes to understanding the implications of various identifying assumptions (see Section (ref)).

Part (d) implies the existence of a finite number of factors. In particular, the number of common factors, $r$, is identified only for $n\to\infty$. Here $N$ is the minimum number of series we need to be able to identify $r$ so that $r\le N$. Without loss of generality hereafter when we say “for all $n\in\mathbb N$” we always mean that $n>N$ so that $r$ can be identified. In practice, we must always work with $n$ such that $r<n$. Moreover, because PCA is based on eigenvalues of a matrix $n\times n$ estimated using $T$ observations then we must also have samples of size $T$ such that $r<T$. Therefore, sometimes it is directly assumed that $r<\min(n,T)$.

To characterize the idiosyncratic component, we make the following assumptions.

ass[idiosyncratic component] $\,$ \begin{compactenum}[(a)] • For all $i,j\in\mathbb N$, all $t\in\mathbb Z$, and all $k\in\mathbb Z$, $\mathbb{E}[ e_{it}]= 0$, $\gamma^ e_{ij,k}:=\mathbb{E}[ e_{it} e_{j,t-k}]$ with $\gamma_{ii,0}^ e:=\mathbb{E}[ e_{it}^2]$. • For all $i,j\in\mathbb N$, all $t\in\mathbb Z$, and all $k\in\mathbb Z$, $\gamma^ e_{ij,k}:=\mathbb{E}[ e_{it} e_{j,t-k}]$ such that $\vert \gamma^ e_{ij,k}\vert\le \rho^{\vert k\vert} M_{ij}$, where $\rho$ and $M_{ij}$ are finite positive reals independent of $t$ such that $0\le \rho <1$, $M_{ii}\le C_ e$, $\sum_{j=1,i\ne j}^n M_{ij}\le M_ e$, and $\sum_{i=1, j\ne i}^n M_{ij}\le M_ e$ for some finite positive reals $C_ e$ and $M_{ e}$ independent of $i$, $j$, and $n$. • \begin{inparaenum} • For all $i=1,\ldots, n$, all $t=1,\ldots, T$, and all $n,T\in\mathbb N$, $\mathbb{E}[ e_{it}^4]\le Q_ e$ for some finite positive real $\phantom {(i)}\, Q_ e$ independent of $i$ and $t$;\\ • for all $j=1,\ldots, n$, all $s=1,\ldots, T$, and all $n,T\in\mathbb N$, \[ \mathbb{E}\left[\left\vert\frac 1{\sqrt{nT}} \sum_{i=1}^n\sum_{t=1}^T\left\{ e_{is} e_{jt}-\mathbb{E}[ e_{is} e_{jt}]\right\} \right\vert^2\right] \le K_ e \] for some finite positive real $K_ e$ independent of $j$, $s$, $n$, and $T$. \end{inparaenum} \end{compactenum}

By part (a), we assume that the idiosyncratic components are cross-sectionally heteroskedastic but serially homoskedastic. It follows that for any $n\in\mathbb N$ the idiosyncratic covariance matrix is defined as $\bm\Gamma^ e:=\mathbb{E}[\bm e_t\bm e_t^\prime]$ with entries $\gamma_{ij,0}^ e$, $i,j=1,\ldots,n$.

Part (b) has a twofold purpose. First, it limits the degree of serial correlation of the idiosyncratic components, implying weak stationarity of idiosyncratic components. Indeed, they have summable autocovariances (see Lemma (ref)(iii)): \[ \sup_{n,T\in\mathbb N}\max_{i=1,\ldots, n}\frac 1{T}\sum_{t,s=1}^T \vert\mathbb{E}_{}[ e_{it} e_{is}]\vert \le \frac{C_ e(1+\rho)}{1-\rho}. \] The same condition follows directly from Bai03 although it is not stated explicitly. Second, part (b) also limits the degree of cross-sectional correlation between idiosyncratic components, which is usually assumed in approximate factor models. Moreover, it implies the usual conditions for PC estimation (see Lemmas (ref)(i), (ref)(ii), (ref)(i), (ref)(ii)):

align[align omitted — 840 chars of source]

which coincide with the conditions required by Bai03. Furthermore, part (b) directly implies also that

align[align omitted — 269 chars of source]

where (ref) is the first condition in Bai03, and (ref) is the first condition in Bai03.

Notice that the formulation of part (b) is inspired by FHLZ17 and it implicitly bounds the L1 norm of $\bm\Gamma^ e$ as in FLM13. Furthermore, in Lemma (ref)(v) we prove that the largest eigenvalue of $\bm\Gamma^ e$ is such that (this is an instance of the Gershgorin circle theorem)

equation[equation omitted — 83 chars of source]

where $C_ e$ and $M_ e$ are defined in Assumption (ref)(b). Condition (ref) was originally assumed by chamberlainrothschild83. This is the essence of an approximate factor model as opposed to an exact factor model where it is imposed the restrictive assumption of $\bm\Gamma^ e$ being diagonal.

Part (c-i) assumes finite and summable 4th order moments of the idiosyncratic components, it is weaker than what assumed by Bai03 where finite 8th order moments are required. Part (c-ii) gives summability conditions over the cross-section and time dimensions for the 4th order cumulants of $\{ e_{it}\}$. Indeed, it can be equivalently written as:

align[align omitted — 501 chars of source]

If we set $j=i$ in (ref) we get a similar condition to what is assumed by Bai03, where, however, $T=1$, and the sums of 8th cross-cumulants are bounded.

Part (c-ii) has two more important consequences. First, it implies that (see Lemma (ref)(iii)):

align[align omitted — 263 chars of source]

This result is crucial to prove asymptotic normality of the estimated loadings. In particular, without (ref) consistency of the loadings would hold but with a slower rate, $\min(\sqrt n,\sqrt T)$ which is slower than $\min(n,\sqrt T)$ derived in Proposition (ref). Such a slower rate is what prevents the proof of asymptotic normality of the estimated loadings to go through.

Second, part (c-ii) implies also that (setting $n=1$ therein):

align[align omitted — 482 chars of source]

which means that the sample (auto)covariances between $\{ e_{it}\}$ and $\{ e_{jt}\}$ are $\sqrt T$-consistent estimators of their population counterparts. In particular, by setting $s=t$ in (ref), we see that we can consistently estimate the $(i,j)$th entry of the idiosyncratic covariance matrix $\bm\Gamma^ e$. Like for the factors, this assumption is a high-level one and requires introducing the sample size $T$. As for the factors, in Appendix (ref) we provide a series of possible primitive conditions on the process $\{ e_{it}\}$ which guarantee part (c-ii) to hold.

Last, we notice that if we allowed for serial heteroskedasticity so that we had $\gamma_{ij,t,t-k}^ e:=\mathbb{E}[ e_{it} e_{j,t-k}]$, then all our proofs would still hold provided in part (b) we required $\sup_{t\in\mathbb Z} \vert \gamma_{ij,t,t-k}^ e\vert\le \rho^{|k|} M_{ij}$, and then we replaced throughout $\gamma_{ij,k}^ e$ with $\gamma_{ij,t,t-k}^ e$.

We then make the following identifying assumption.

ass[distinct eigenvalues] For all $n\in\mathbb N$ and all $j=1,\ldots, r-1$, $\mu_j^{C} > \mu_{j+1}^{C}$.

This can be equivalently stated by requiring $\overline C_j<\underline C_{j-1}$, $j=2,\ldots, r$, in (ref). Since the $r$ eigenvalues of $\bm\Sigma_\Lambda\bm\Gamma^F$ are given by $\lim_{n\to\infty} n^{-1}\mu_j^{C}$, $j=1,\ldots,r$, then, Assumption (ref) implies that they are also distinct (see Lemma (ref)), as also required in Bai03. Note that the $r$ eigenvalues of $(\bm\Sigma_\Lambda)^{1/2}\bm\Gamma^F(\bm\Sigma_\Lambda)^{1/2}$ or equivalently of $(\bm\Gamma^F)^{1/2}\bm\Sigma_\Lambda(\bm\Gamma^F)^{1/2}$ are also distinct because they coincide with those of $\bm\Sigma_\Lambda\bm\Gamma^F$.

Distinct eigenvalues are not needed to estimate consistently the space spanned by the eigenvectors since we can apply a recent version of Davis Kahan theorem yu15 which relies only on the existence of a positive eigen-gap between the $r$th and $(r+1)$th eigenvalues of $\bm\Gamma^x$ and $\bm\Gamma^{C}$. The former matrix has an eigen-gap widening with $n$ as shown in (ref) below, while the latter one is of rank $r$, so in both cases eigen-gap is always strictly positive. However, in order to identify the space spanned by the loadings, we need also the inverse of the $r\times r$ matrix of normalized eigenvectors of $(\bm\Gamma^F)^{1/2}\bm\Sigma_\Lambda(\bm\Gamma^F)^{1/2}$ and this requires distinct eigenvalues since, in turn, they imply that the eigenvectors are identified and thus are linearly independent, so that the eigenvector matrix is invertible (see the proof of Proposition (ref)). Distinct eigenvalues are also crucial for identification (see Section (ref)).

We then turn to the issue of dependence between common and idiosyncratic components. Given that PCA deals only with covariances it would seem enough to just assume contemporaneous orthogonality of the common and idiosyncratic components, i.e., that for all $i,j\in\mathbb N$ and all $t\in\mathbb Z$

equation[equation omitted — 62 chars of source]

This is indeed true if we aim to prove only consistency of the PC estimators and we make primitive assumptions on the processes $\{\mathbf F_t\}$ and $\{ e_{it}\}$ as those in Appendix (ref) (see the proofs of Proposition (ref) and Lemma (ref)(i-b)).

Furthermore, orthogonality in (ref) is also enough to asymptotically identify the common component. Indeed, because of Weyl's inequality MK04, (ref), (ref), and (ref) imply that $\bm\Gamma^x:=\mathbb{E}[\mathbf x_t\mathbf x_t^\prime]=\bm\Gamma^{C}+\bm\Gamma^ e$ and the existence of an eigen-gap in the eigenvalues $\mu_j^x$, $j=1,\ldots,n$, of $\bm\Gamma^x$ (see Lemma (ref)(vi)):

align[align omitted — 208 chars of source]

where $\underline C_r$ and $\overline C_r$ are defined in (ref) and $M_ e$ is defined in Assumption (ref)(b). Note that the viceversa is also true, i.e., if (ref) holds, then (ref)-(ref) hold (chamberlainrothschild83, and gersing2023weak, or BH25 for more recent and complete proofs). This is the result upon which most methods for determining the number of factors $r$ rely on baing02,ABC10,ahnhorenstein13. Last, notice that for this result we do not need Assumption (ref) to hold so we can have non-distinct leading eigenvalues of $\bm\Gamma^x$.

If we want to prove asymptotic normality of the PC estimators, then additional conditions besides orthogonality and relating to 4th-order dependencies are needed. To this end, we can assume the following.

ass[independence] The processes $\{ e_{it},\, i\in\mathbb N,\, t\in\mathbb Z\}$ and $\{F_{jt},\, j=1,\ldots, r,\, t\in\mathbb Z\}$ are mutually independent.

This assumption obviously implies (ref) but is much stronger and apparently even more strict than what is usually assumed. Because of this, it deserves some careful discussion. Specifically, it has three main implications.

First, from Assumptions (ref)(c-ii) and (ref) it follows that (see (ref) in the proof of Proposition (ref)):

align[align omitted — 268 chars of source]

where $\bm\varepsilon_i:=( e_{i1}\cdots e_{iT})^\prime$. Condition (ref) coincides with what is assumed by Bai03. This inequality is the analogous for the factors of condition (ref), which we derived for the loadings from Assumption (ref)(c-ii). It is a crucial requirement as it is necessary to prove asymptotic normality of the estimated factors. In particular, without (ref) consistency of the factors would hold, but with a rate, $\min(\sqrt n,\sqrt T)$ slower than the rate $\min(\sqrt n, T)$ derived in Proposition (ref). Such a slower rate is what prevents the proof of asymptotic normality of the estimated factors to go through. While proving (ref) is easy since the loadings are deterministic, to prove (ref) it must be that $\mathbb{E}[ e_{i_1t} e_{i_2t} e_{i_1s_1} e_{i_2s_2}F_{ks_1}F_{ks_2}]=\mathbb{E}[ e_{i_1t} e_{i_2t} e_{i_1s_1} e_{i_2s_2}]\,\mathbb{E}[F_{ks_1}F_{ks_2}]$ for all $i_1,i_2=1,\ldots, n$, all $t,s_1,s_2=1,\ldots, T$, and all $k=1,\ldots, r$, this is why we need independence in Assumption (ref) if we want to prove (ref), rather than just assuming it.

Second, Assumption (ref) implies that there exists a finite positive real $M_{F e}$ such that (see the proof of Lemma (ref)(i))

align[align omitted — 179 chars of source]

which, in turn, implies the same condition as in Bai03, indeed,

equation[equation omitted — 326 chars of source]

Notice that both (ref) and (ref) allow for weak dependence between the factors and the idiosycratic components, but just at the level of 4th-order moments, while they still imply contemporaneous orthogonality. Indeed, when letting $T\to\infty$, by Chebychev's inequality, these conditions imply that $T^{-1}\sum_{t=1}^T\mathbf F_t e_{it}=O_{\mathrm P}(T^{-1/2})$. Moreover, under the assumptions on serial dependence of the factors and the idiosyncratic components allowing for Assumptions (ref)(c-ii) and (ref)(c-ii) to hold, we have the Law of Large Numbers $\Vert T^{-1}\sum_{t=1}^T\mathbf F_t e_{it}-\mathbb{E}[\mathbf F_t e_{it}]\Vert=O_{\mathrm P}(T^{-1/2})$ (this is implied also by Assumption (ref)(a) below). Hence, by uniqueness of the limit, it follows that we must have $\mathbb{E}[\mathbf F_t e_{it}] = \mathbf 0_r$ for all $i\in\mathbb N$ and $t\in\mathbb Z$, and, thus, contemporaneous orthogonality in (ref) still holds.

Crucially, the moment condition in (ref) is needed to prove not only asymptotic normality, but also consistency of the PC estimators (see Lemma (ref)(i-a)) so either it is assumed or it must be derived from Assumption (ref). Or it could be ensured by making further primitive assumptions on the processes $\{\mathbf F_t\}$ and $\{ e_{it}\}$, as those discussed in Appendix (ref).

Third, from Assumptions (ref)(a), (ref)(b), and (ref) it follows that (see (ref) in the proof of Lemma (ref)):

align[align omitted — 210 chars of source]

which is the same condition as in Bai03. Once again condition (ref) is crucial for proving asymptotic normality, thus, it must be either assumed or derived from Assumption (ref).

The trade-off is then clear. Either we make Assumption (ref), which, although strong, is intuitive and it implies orthogonality in (ref) as well as the required moment inequalities in (ref), (ref), and (ref). Or we directly assume the less intuitive (ref), (ref), and (ref) as in Bai03, which imply also orthogonality in (ref). Here we choose the former approach in order to keep working with more intuitive assumptions. Notice that our choice coincides with baing04, where, in presence of unit roots both in the factors and the idiosyncratic components, they propose to first estimate the loadings and the factors via PCA on the differenced data and, to this end, assume independence for simplicity.\footnote{ In the case of random loadings baing04 also assume that the sequence $\{\lambda_{ij},\,i\in\mathbb N,\, j=1,\ldots, r\}$ is independent of the factor and idiosyncratic processes, so that (ref) still holds.}

In order to prove asymptotic normality of the PC estimators it is common to introduce the following assumptions.

ass[Central limit theorems] $\,$ \begin{compactenum}[(a)] • For all $i\in\mathbb N$, as $T\to\infty$, $ \frac 1{\sqrt T}\sum_{t=1}^T \mathbf F_t e_{it} \to_d \mathcal N\left(\mathbf 0_r, \bm\Phi_i \right), $ where $\bm\Phi_i:=\lim_{T\to\infty}\frac 1 T\sum_{t,s=1}^T \mathbb{E}_{}[\mathbf F_t\mathbf F_s^\prime e_{it} e_{is}]$. • For all $t\in\mathbb Z$, as $n\to\infty$, $ \frac 1{\sqrt n}\sum_{i=1}^n \bm\lambda_i e_{it} \to_d \mathcal N\left(\mathbf 0_r, \bm\Gamma_t \right), $ where $\bm\Gamma_t:=\lim_{n\to\infty}\frac 1 n\sum_{i,j=1}^n \bm\lambda_i\bm\lambda_j^\prime\mathbb{E}_{}[ e_{it} e_{jt}]$. \end{compactenum}

Part (a) is also assumed by Bai03. It can be derived from more primitive assumptions. For example, we could assume strong mixing factors and idiosyncratic components and such that, for all $t\in\mathbb Z$ and all $i\in\mathbb N$, we strengthen Assumptions (ref)(c-i) and (ref)(c-i) to: $\mathbb{E}[\Vert\mathbf F_{t}\Vert^{4+\epsilon}]\le K_F$ and $\mathbb{E}[| e_{it}|^{4+\epsilon}]\le Q_ e$, for some $\epsilon>0$. Then, $\{\mathbf F_t e_{it}\}$ is also strong mixing because of bradley05 and with finite 4th order moments because of Assumption (ref) and part (a) would follow from ibra62. For alternative primitive conditions see also Appendix (ref). Notice also that, under Assumption (ref), in part (a) we actually have $\bm\Phi_i:=\lim_{T\to\infty}T^{-1}\sum_{t,s=1}^T \mathbb{E}_{}[\mathbf F_t\mathbf F_s^\prime]\,\mathbb{E}[ e_{it} e_{is}]$.

Part (b) is also assumed by Bai03. Clearly, it holds if we assumed cross-sectionally uncorrelated idiosyncratic components, i.e., if $\bm\Gamma^ e$ were diagonal as in an exact factor model. In general, to derive part (b) from primitive assumptions we could either introduce an ordering of the $n$ cross-sectional items and a related notion of spatial dependence and derive it from the properties of stationary mixing random fields bolthausen1982, or of cross-sectional martingale difference sequences KP13. Or we could apply results on exchangeable sequences, which are instead independent of the ordering, and are in turn obtained by virtue of the Hewitt-Savage-de Finetti theorem (austern2022limit, and bolthausen1984). In any case we would also need to strengthen Assumption (ref)(c-i) by asking that for all $t\in\mathbb Z$, all $i\in\mathbb N$, $\mathbb{E}[| e_{it}|^{4+\epsilon}]\le Q_ e$, for some $\epsilon>0$. In the case of random loadings we should also require that for all $i\in\mathbb N$, $ \mathbb{E}[\Vert \bm\lambda_i\Vert^{4+\epsilon}]\le K_\Lambda$, for some $\epsilon>0$, and we would have: $\bm\Gamma_t:=\lim_{n\to\infty}n^{-1}\sum_{i,j=1}^n \mathbb{E}[\bm\lambda_i\bm\lambda_j^\prime]\,\mathbb{E}_{}[ e_{it} e_{jt}]$.

ass[Rates] As $n,T\to\infty$, $\sqrt T/n\to0$ and $\sqrt n/T\to0$.

This assumption is needed only for proving asymptotic normality. It is assumed also in Bai03. It is in fact very mild, and it allows for the typical case of $n\asymp T$.

To summarize, in Table (ref) we show the relation between our Assumptions (ref)-(ref) and those made by Bai03. As discussed in detail above, the latter are implied by ours, which are fewer and more intuitive. Notice that three assumptions we make are not explicitly stated in Bai03, but hold also therein implicitly. Indeed: (i) the number of factors $r$ must be independent of $n$ and $T$ (Assumption (ref)(d)), (ii) the idiosyncratic components must have zero mean (Assumption (ref)(a)), and (iii) the relative rates of divergence between $n$ and $T$ must always be imposed to prove asymptotic normality (Assumption (ref)).

table[table omitted — 2,066 chars of source]

Finally, similar assumptions are found in other works on PCA for factor models with some important differences, though. First, stockwatson02JASA assume, like here, only summable 4th-order moments, but do not impose any other cross-moment conditions since they just prove consistency but not asymptotic normality. Second, FLM13 impose stronger assumptions by requiring exponentially decaying tails of the distributions of factors and idiosyncratic components. This is done in order to derive uniform consistency for loadings and factors, which, in turn, is needed to prove consistency of the resulting high-dimensional covariance matrix having a low-rank (factor driven) plus sparse (idiosyncratic) structure.

Consistent estimation of the loadings and factors spaces - part 1

Our first result is derived using only the properties of the sample covariance and its eigenvalues and eigenvectors (see Appendix (ref) for a proof).

propUnder Assumptions (ref) through (ref), contemporaneous orthogonality as in (ref) and either of the following: \begin{inparaenum} • Assumptions (ref) or (ref) or (ref) in Appendix (ref) on serial dependence of the processes; • the moment conditions in (ref) and Assumption (ref)(c-ii) holding with the rate $\sqrt T$; \end{inparaenum} it follows that, as $n,T\to\infty$,\footnote{We say parts (b) and (d) hold uniformly in $i$ or $t$ when the constants involved in the $O_{\mathrm P}$ bound do not depend on $i$ and $t$, it does not mean uniform consistency, which would imply statements like $\max_{i=1,\ldots, n}\Vert\widehat{\bm \lambda}_i^\prime-\bm\lambda_i^\prime\bm{\mathcal H}\Vert=o_{\mathrm P}(1)$ nor $\max_{t=1,\ldots,T}\Vert\widehat{\mathbf F}_t-\bm{\mathcal H}^{-1}\mathbf F_t\Vert=o_{\mathrm P}(1)$.} \begin{compactenum}[(a)] • $\min(n,\sqrt T)\left\Vert\frac{\widehat{\bm\Lambda}-\bm\Lambda\bm{\mathcal H}}{\sqrt n}\right\Vert= O_{\mathrm P}(1)$; • $ {\min(\sqrt n,\sqrt T)} \left\Vert\widehat{\bm \lambda}_i^\prime-\bm\lambda_i^\prime\bm{\mathcal H}\right\Vert= O_{\mathrm P}(1)$, uniformly in $i$; • $\min(\sqrt n,\sqrt T)\left\Vert\frac {\widehat{\bm F}-\bm F(\bm{\mathcal H}^{-1})^\prime}{\sqrt T}\right\Vert= O_{\mathrm P}(1)$; • ${\min(\sqrt n,\sqrt T)} \left\Vert\widehat{\mathbf F}_t-\bm{\mathcal H}^{-1}\mathbf F_t\right\Vert= O_{\mathrm P}(1)$, uniformly in $t$; \end{compactenum} where $\bm{\mathcal H} :=(\bm\Gamma^F)^{1/2} \mathbf K{\mathbf J}$ is an $r\times r$ finite and positive definite matrix, with $\mathbf K$ having as columns the normalized eigenvectors of $(\bm\Gamma^F)^{1/2}(n^{-1}\bm \Lambda^\prime\bm \Lambda) (\bm\Gamma^F)^{1/2}$, and $\mathbf J$ being an $r\times r$ diagonal matrix with entries $\pm 1$ which depend on $n$ and $T$.

The bound derived in part (a) is tighter than the one derived by FGLR09 and bai2020simpler where the rate is $\min(\sqrt n,\sqrt T)$. Parts (b), (c), and (d) follow directly from part (a). In particular, part (c) coincides with the result by bai2020simpler.

The proof makes use of the Cauchy-Schwarz inequality, therefore, the implied rates are not the sharpest possible. Hence this result does not allow to derive asymptotic normality. However, besides providing an intuitive and quick proof of consistency of PC, which is often enough for many purposes, Proposition (ref) is also needed for proving successive results, where the consistency rates are refined, provided additional moment conditions are imposed (see Section (ref)).

Part (a) is based essentially on a result on the $n\times n$ sample covariance matrix $\widehat{\bm\Gamma}^x$ which in turn is determined by the serial dependence of the considered stochastic process, as explained in Appendix (ref). Its proof is based on four main steps proved in Lemma (ref). First, we prove that the rescaled sample covariance matrix $n^{-1}\widehat{\bm\Gamma}^x$ is a $\sqrt T$-consistent estimator of the rescaled population covariance matrix $n^{-1}\bm\Gamma^x$. This result holds under a minimal and mild set of assumptions, namely Assumptions (ref) and (ref) jointly with contemporaneous orthogonality as in (ref) plus either the conditions (A) (see Lemma (ref)(i-b)) or the conditions (B) (see Lemma (ref)(i-a)). We stress that under either (A) or (B) Assumption (ref) of independence between factors and idiosyncratic components is not needed.

Second, we prove that the rescaled sample covariance $n^{-1}\widehat{\bm\Gamma}^x$ actually converges to the rescaled common component covariance $n^{-1}\bm\Gamma^{C}$ with rate $\min(n,\sqrt T)$ (see Lemma (ref)(ii)). This result follows from Assumption (ref) which implies that the idiosyncratic covariance matrix $\bm\Gamma^ e$ has bounded eigenvalues for all $n\in\mathbb N$ (see (ref)). It is clear that, in fact, Assumption (ref)(b) is not strictly needed since it is enough to make the weaker assumption of bounded idiosyncratic eigenvalues.

Third, we prove that the matrices of sorted eigenvalues, $\widehat{\mathbf M}^x$, and corresponding normalized eigenvectors, $\widehat{\mathbf V}^x$, of $\widehat{\bm\Gamma}^x$ converge to the matrices of sorted eigenvalues, ${\mathbf M^{C}}$, and corresponding normalized eigenvectors, ${\mathbf V^{C}}$, of $\bm\Gamma^{C}$. Specifically, from Weyl's inequality we get $n^{-1}\Vert \widehat{\mathbf M}^x-\mathbf M^{C}\Vert=O_{\mathrm P}(\max(n^{-1},T^{-1/2}))$ (see Lemma (ref)(iii) and MK04), while from Davis Kahan theorem we get $\Vert \widehat{\mathbf V}^x-\mathbf V^{C}\widehat{\mathbf J}\Vert =O_{\mathrm P}(\max(n^{-1},T^{-1/2}))$, where $\mathbf J$ is an $r\times r$ diagonal matrix with entries $\pm 1$ depending on $n$ and $T$, and accounting for the fact that a sample and a population eigenvector identify asymptotically the same one-dimensional subspaces, but might point in opposite directions (see Lemma (ref)(iv) and yu15).

Fourth, since $\bm\Gamma^{C} =\bm\Lambda\bm\Gamma^F\bm\Lambda^\prime = \mathbf V^{C}\mathbf M^{C}\mathbf V^{{C}\prime}$, the columns of the true loadings matrix $\bm\Lambda$ must span the same space as the normalized eigenvectors of the common component covariance matrix. In other words, it must be that $\bm\Lambda(\bm\Gamma^F)^{1/2}\mathbf K=\mathbf V^{C}(\mathbf M^{C})^{1/2}$ for some invertible $r\times r$ matrix $\mathbf K$. In particular, we prove that the columns of $\mathbf K$ are the normalized eigenvectors of $(\bm\Gamma^F)^{1/2} (n^{-1}\bm\Lambda^\prime\bm\Lambda) (\bm\Gamma^F)^{1/2}$ (see Lemma (ref)(iii))), i.e., $\mathbf K$ is a rotation.

Summing up, given that the PC estimator of the loadings is $\widehat{\bm\Lambda}=\widehat{\mathbf V}^x(\widehat {\mathbf M}^x)^{1/2}$ (see (ref)), part (a) follows, with $\bm{\mathcal H}=(\bm\Gamma^F)^{1/2} \mathbf K{\mathbf J}$ being a finite and invertible linear transformation. So the true loadings are recovered up to: (i) a scale given by $(\bm\Gamma^F)^{1/2}$; (ii) a rotation $\mathbf K$, and (iii) a diagonal matrix of signs ${\mathbf J}$. Note that both $(\bm\Gamma^F)^{1/2}$ and $\mathbf K$ do not depend on the sample size $T$, but depend only on population quantities as well as $n$. The only dependence on $T$ is in the sign matrix ${\mathbf J}$, but such dependence can be easily fixed (see Section (ref)).

The following result gives a series of more intuitive asymptotic expressions for $\bm{\mathcal H}$ (see Appendix (ref) for a proof).

propUnder Assumptions (ref) through (ref), contemporaneous orthogonality as in (ref) and either of the following: \begin{inparaenum} • Assumptions (ref) or (ref) or (ref) in Appendix (ref) on serial dependence of the processes; • the moment conditions in (ref) and Assumption (ref)(c-ii) holding with the rate $\sqrt T$; \end{inparaenum} it follows that, as $n,T\to\infty$, \begin{compactenum} • $\min(n,\sqrt T)\left\Vert\bm{\mathcal H} - (\bm\Lambda^\prime\bm\Lambda)^{-1}\bm\Lambda^\prime\widehat{\bm\Lambda} \right\Vert = O_{\mathrm P}(1)$; • $\min(\sqrt n,\sqrt T)\left\Vert\bm{\mathcal H}^{-1} - \widehat{\bm F}^\prime \bm F (\bm F^\prime\bm F)^{-1} \right\Vert = O_{\mathrm P}(1)$; • $\min(n,\sqrt T)\left\Vert\bm{\mathcal H}^{-1} - (\widehat{\bm\Lambda}^\prime\widehat{\bm\Lambda})^{-1}\widehat{\bm\Lambda}^\prime{\bm\Lambda} \right\Vert = O_{\mathrm P}(1)$, where we recall that $\widehat{\bm\Lambda}^\prime\widehat{\bm\Lambda}=\widehat{\mathbf M}^x$; • $\min(\sqrt n,\sqrt T)\left\Vert\bm{\mathcal H} - {\bm F}^\prime \widehat{\bm F} (\widehat{\bm F}^\prime\widehat{\bm F})^{-1} \right\Vert = O_{\mathrm P}(1)$, where we recall that $\widehat{\bm F}^\prime \widehat{\bm F} = T\, \mathbf I_r$. \end{compactenum}

The proof is an immediate consequence of Proposition (ref)(a) and (ref)(c) and since the assumptions made are the same as those made for Proposition (ref), these results are directly applicable to the consistency results proved therein.

Proposition (ref) shows that asymptotically the matrix $\bm{\mathcal H}$ is obtained by linear projection of the estimated loadings onto the true loadings (part (a)) or of the true factors onto the estimated ones (part (d)). Similarly, asymptotically $\bm{\mathcal H}^{-1}$ is obtained by linear projection of the estimated factors onto the true factors (part (b)) or by linear projection of the true loadings onto the estimated ones (part (c)). The derived rates are not the sharpest possible for two reasons. First, the rates in part (a) and (b) are obtained by using the bound $n^{-1}\Vert\bm\Lambda^\prime(\widehat{\bm\Lambda}-\bm\Lambda\bm{\mathcal H})\Vert\le n^{-1/2}\Vert \bm\Lambda\Vert\, n^{-1/2}\Vert\widehat{\bm\Lambda}-\bm\Lambda\bm{\mathcal H} \Vert$, which, therefore, inherits the bound obtained for the loadings in Proposition (ref)(a). Second, the rates in parts (b) and (d) are slower than those in parts (a) and (c) since the rates in Proposition (ref)(c) are not the sharpest possible. As shown, in Proposition (ref) below, it is possible to define a different transformation matrix, which is asymptotically equivalent to $\bm{\mathcal H}$, but which converges to the projection matrices in Proposition (ref) at a faster rate.

Finally, consider the following spectral decomposition:

equation[equation omitted — 133 chars of source]

where $\bm\Upsilon_0$ is the $r\times r$ matrix having as columns the normalized eigenvectors of $(\bm\Gamma^F)^{1/2}\bm\Sigma_\Lambda(\bm\Gamma^F)^{1/2}$, and $\bm V_0$ is the $r\times r$ matrix of corresponding eigenvalues sorted in decreasing order. Then, we prove a result, which is the analogous of the result in Bai03, and is needed for the next section. It gives the limit of the space spanned by the estimated loadings (see Appendix (ref) for a proof).

propUnder Assumptions (ref) through (ref), contemporaneous orthogonality as in (ref) and either of the following: \begin{inparaenum} • Assumptions (ref) or (ref) or (ref) in Appendix (ref) on serial dependence of the processes; • the moment conditions in (ref); \end{inparaenum} it follows that, as $n,T\to\infty$, $ \left\Vert \frac{\widehat{\bm\Lambda}^\prime\bm\Lambda}{n}-\bm V_0\bm{\mathcal J}_0\bm\Upsilon_0^\prime (\bm\Gamma^F)^{-1/2}\right\Vert = o_{\mathrm P}(1), $ where the columns of $\bm\Upsilon_0$ are the normalized eigenvectors of $(\bm\Gamma^F)^{1/2}\bm\Sigma_\Lambda(\bm\Gamma^F)^{1/2}$, $\bm V_0$ is the $r\times r$ diagonal matrix containing the corresponding eigenvalues, and $\bm{\mathcal J}_0$ is an $r\times r$ diagonal matrix with entries $\pm 1$ independent of $n$ and $T$. If Assumption (ref)(a) holds with rate $\sqrt n$ and Assumption (ref)(c-ii) holds with rate $\sqrt T$, then the above holds with rate $\min(\sqrt n,\sqrt T)$. Furthermore, $\bm V_0\bm{\mathcal J}_0\bm\Upsilon_0^\prime (\bm\Gamma^F)^{-1/2}$ is finite and positive definite.

Consistent estimation of the loadings and factors spaces - part 2

Let us start from the estimated loadings. By definition of $\widehat{\bm\Lambda}$ in (ref) $\widehat{\mathbf V}^x=\widehat{\bm\Lambda}(\widehat{\mathbf M}^x)^{-1/2}$. Thus, from (ref), we obtain

equation[equation omitted — 133 chars of source]

Then, substituting $\bm X^\prime\bm X=(\bm\Lambda\bm F^\prime+\bm E^\prime)^\prime(\bm F\bm\Lambda^\prime+\bm E)$ into (ref), and letting

equation[equation omitted — 179 chars of source]

we obtain

align[align omitted — 326 chars of source]

Notice that, as $n,T\to\infty$, $\widehat{\mathbf H}$ is well defined because of Lemma (ref)(i). Taking the $i$th row of (ref), we get

align[align omitted — 553 chars of source]

The following bounds hold for the terms in (ref) (see Appendix (ref) for a proof).

lemUnder Assumptions (ref) through (ref), as $n,T\to\infty$, \begin{inparaenum}[(a)] • $\sqrt {nT}\left\Vert \text{\upshape (1.a)}\right\Vert = O_{\mathrm P}(1)$; • $\sqrt T \left\Vert \text{\upshape (1.b)}\right\Vert = O_{\mathrm P}(1)$; • $\min(n,\sqrt{nT})\left\Vert \text{\upshape (1.c)}\right\Vert = O_{\mathrm P}(1)$; \end{inparaenum} uniformly in $i$.

Consistency follows immediately (see Appendix (ref) for a proof).

propUnder Assumptions (ref) through (ref), as $n,T\to\infty$, \begin{compactenum}[(a)] • $\min(n,\sqrt {T})\left\Vert \widehat{\bm\lambda}_i-\widehat{\mathbf H}^\prime{\bm\lambda}_i\right\Vert = O_{\mathrm P} (1)$, uniformly in $i$; • $\min(n,\sqrt {T}) \left\Vert \frac {\widehat{\bm\Lambda}-{\bm\Lambda}\widehat{\mathbf H}}{\sqrt n}\right\Vert=O_{\mathrm P}(1)$. \end{compactenum}

Part (a) refines the result in Proposition (ref)(b). Part (b) has the same rate as in Proposition (ref)(a). As shown in Proposition (ref)(a) below, the matrix $\widehat{\mathbf H}$ is asymptotically equivalent to the projection matrix of the estimated loadings onto the true ones.

Turning to the estimated factors, from (ref) and (ref), we have

align[align omitted — 492 chars of source]

Then, substituting on the right-hand-side of (ref) $\bm{\Lambda}$ with $(\bm\Lambda-\widehat{\bm\Lambda}\widehat{\mathbf H}^{-1})+\widehat{\bm\Lambda}\widehat{\mathbf H}^{-1},$ in the first term and $\widehat{\bm\Lambda}$ with $ \widehat{\bm\Lambda}-\bm\Lambda\widehat{\mathbf H}+\bm\Lambda\widehat{\mathbf H} $ in the second term, we obtain

align[align omitted — 363 chars of source]

Notice that, as $n,T\to\infty$, $\widehat{\mathbf H}^{-1}$ is well defined because of Lemma (ref)(ii). Taking the $t$th row of (ref), we get

align[align omitted — 519 chars of source]

The following bounds hold for the terms in (ref) (see Appendix (ref) for a proof).

lemUnder Assumptions (ref) through (ref), as $n,T\to\infty$, \begin{inparaenum}[(a)] • $\min(n,\sqrt{nT},T)\Vert \text{2.a}\Vert=O_{\mathrm {P}}(1)$; • $\min(n,\sqrt{nT},T)\Vert \text{2.b}\Vert=O_{\mathrm {P}}(1)$; • $\sqrt n\Vert \text{2.c}\Vert=O_{\mathrm {P}}(1)$; \end{inparaenum} uniformly in $t$.

Consistency follows immediately (see Appendix (ref) for a proof).

propUnder Assumptions (ref) through (ref), as $n,T\to\infty$ \begin{compactenum}[(a)] • $\min(\sqrt n,{T})\left\Vert \widehat{\mathbf F}_t-\widehat{\mathbf H}^{-1}{\mathbf F}_t \right\Vert = O_{\mathrm P} (1)$, uniformly in $t$; • $\min(\sqrt n, {T}) \left\Vert \frac{\widehat{\bm F}-{\bm F}(\widehat{\mathbf H}^{-1})^\prime}{\sqrt T}\right\Vert=O_{\mathrm P}(1)$. \end{compactenum}

Parts (a) and (b) refine the results in Proposition (ref)(c)-(ref)(d). As shown in Proposition (ref)(b) below, the matrix $\widehat{\mathbf H}^{-1}$ is asymptotically equivalent to the projection matrix of the estimated loadings onto the true ones.

Finally, we derive a result analogous to the result in baing02 and Bai03 (see Appendix (ref) for a proof).

propUnder Assumptions (ref) through (ref), as $n,T\to\infty$, $$ \min(n,T)\left\{\frac 1T\sum_{t=1}^T \left\Vert \widehat{\mathbf F}_t-\widehat{\mathbf H}^{-1}{\mathbf F}_t \right\Vert^2\right\}=O_{\mathrm P}(1). $$

This result is often needed when using the estimated quantities for further analysis, e.g., in factor augmented regressions baing06. It is also needed to prove consistency for the selection of the number of factors based on Information Criteria baing02,ABC10.

The role of $\widehat{\mathbf H}$ and its relation with $\pmb{\mathcal H}$

To understand the meaning of $\widehat{\mathbf H}$ in Propositions (ref) and (ref), we can derive a series of more intuitive asymptotic expressions (see Appendix (ref) for a proof).

propUnder Assumptions (ref) through (ref), as $n,T\to\infty$, \begin{compactenum}[(a)] • $\min(n,\sqrt{nT}, {T}) \left\Vert \widehat{\mathbf H} - (\bm\Lambda^\prime\bm\Lambda)^{-1} {\bm\Lambda}^\prime\widehat{\bm\Lambda} \right\Vert=O_{\mathrm P}(1)$; • $\min( n,\sqrt{nT}, {T}) \left\Vert\widehat{\mathbf H}^{-1}- \widehat{\bm F}^\prime{\bm F}(\bm F^\prime\bm F)^{-1} \right\Vert=O_{\mathrm P}(1)$; • $ \min(n,\sqrt{nT}, {T})\left\Vert \widehat{\mathbf H}^{-1} - (\widehat{\bm\Lambda}^\prime\widehat{\bm\Lambda})^{-1} \widehat{\bm\Lambda}^\prime{\bm\Lambda} \right\Vert=O_{\mathrm P}(1)$, where we recall that $\widehat{\bm\Lambda}^\prime\widehat{\bm\Lambda}=\widehat{\mathbf M}^x$; • $\min(n,\sqrt{nT}, {T}) \left\Vert \widehat{\mathbf H} - {\bm F}^\prime\widehat{\bm F}(\widehat{\bm F}^\prime\widehat{\bm F})^{-1}\right\Vert= O_{\mathrm P}(1)$, where we recall that $\widehat{\bm F}^\prime \widehat{\bm F} = T\, \mathbf I_r$; • $\min(n,\sqrt{nT}, {T})\left\Vert \widehat{\mathbf H}^{-1} -({\bm\Lambda}^\prime\widehat{\bm\Lambda})^{-1} \bm\Lambda^\prime\bm\Lambda \right\Vert=O_{\mathrm P}(1)$; • $\min( n,\sqrt{nT}, {T}) \left\Vert\widehat{\mathbf H}- \bm F^\prime\bm F (\widehat{\bm F}^\prime{\bm F})^{-1} \right\Vert=O_{\mathrm P}(1)$; • $\min( n,\sqrt{nT}, {T}) \left\Vert \frac{(\widehat{\mathbf H}^{-1})^\prime\widehat{\bm\Lambda}^\prime\widehat{\bm\Lambda}\widehat{\mathbf H}^{-1}}n-\frac{\bm\Lambda^\prime\bm\Lambda}n\right\Vert=O_{\mathrm P}(1)$, where we recall that $\widehat{\bm\Lambda}^\prime\widehat{\bm\Lambda}=\widehat{\mathbf M}^x$; • $\min( n,\sqrt{nT}, {T}) \left\Vert \frac{\widehat{\mathbf H} \widehat{\bm F}^\prime\widehat{\bm F}\widehat{\mathbf H}^\prime}T- \frac{{\bm F}^\prime{\bm F}}T\right\Vert=O_{\mathrm P}(1)$, where we recall that $\widehat{\bm F}^\prime \widehat{\bm F} = T\, \mathbf I_r$. \end{compactenum}

This result, which is similar to what proved by bai2020simpler, shows that asymptotically the matrix $\widehat{\mathbf H}$ is obtained by linear projection of the estimated loadings onto the true loadings (part (a)) or of the true factors onto the estimated ones (part (d)). Similarly, asymptotically $\widehat{\mathbf H}^{-1}$ is obtained by linear projection of the estimated factors onto the true factors (part (b)) or by linear projection of the true loadings onto the estimated ones (part (c)). Part (e)-(h) are less intuitive, but still useful for proving further results.

Moreover, $\widehat{\mathbf H}$ has also the following asymptotic expansion (see Appendix (ref) for a proof).

propUnder Assumptions (ref) through (ref), as $n,T\to\infty$, \begin{compactenum} • $\min(n,\sqrt{T})\left\Vert \widehat{\mathbf H} - \left(\frac{\bm F^\prime\bm F}{T}\right)^{1/2}\widehat{\mathbf Q}\right\Vert = O_{\mathrm P}(1)$; • $\min(n,\sqrt{T})\left\Vert\widehat{\mathbf H}^{-1} - \widehat{\mathbf Q}^\prime \left(\frac{\bm F^\prime\bm F}T\right)^{-1/2}\right\Vert =O_{\mathrm P}(1)$; \end{compactenum} where the columns of $\widehat{\mathbf Q}$ are the normalized eigenvectors of $(T^{-1}\bm F^\prime\bm F)^{1/2}(n^{-1}\bm \Lambda^\prime\bm \Lambda) (T^{-1}\bm F^\prime\bm F)^{1/2}$.

According to this result and Propositions (ref) and (ref), the loadings and the factors can be consistently estimated up to: (i) a scale $(T^{-1}{\bm F^\prime\bm F})^{1/2}$ and (ii) a rotation $\widehat{\mathbf Q}$. The proof of this result rests on four intermediate relevant results, which we summarize here. First and second, from Proposition (ref), it follows that (see also (ref) and (ref) in the proof of Proposition (ref))

align[align omitted — 489 chars of source]

Third, (ref) and (ref) imply (see also (ref) in the proof of Proposition (ref))

align[align omitted — 323 chars of source]

Last, the $r$ non-zero eigenvalues of $n^{-1} \bm\Lambda(T^{-1}\bm F^\prime\bm F) \bm\Lambda^\prime$, collected into the diagonal $r\times r$ matrix $n^{-1}\widehat{\mathbf M}^C$ coincide with the eigenvalues of $(T^{-1}\bm F^\prime\bm F)^{1/2}(n^{-1}\bm \Lambda^\prime\bm \Lambda) (T^{-1}\bm F^\prime\bm F)^{1/2}$, and are such that (see also (ref) in the proof of Proposition (ref))

equation[equation omitted — 192 chars of source]

Then, part (b) follows directly from (ref) and (ref).

We then prove a result which is the counterpart of Propositions (ref) and it gives the limit of the space spanned by the estimated factors. It follows essentially from the definition of $\widehat{\mathbf H}$ in (ref) and Propositions (ref) and (ref)(f) (see Appendix (ref) for a proof).

propUnder Assumptions (ref) through (ref), as $n,T\to\infty$, $ \left\Vert \frac{\widehat{\bm F}^\prime\bm F}{T}-\bm{\mathcal J}_0\bm\Upsilon_0^\prime (\bm\Gamma^F)^{1/2}\right\Vert = o_{\mathrm P}(1), $ where the columns of $\bm\Upsilon_0$ are the normalized eigenvectors of $(\bm\Gamma^F)^{1/2}\bm\Sigma_\Lambda(\bm\Gamma^F)^{1/2}$, and $\bm{\mathcal J}_0$ is an $r\times r$ diagonal matrix with entries $\pm 1$ independent of $n$ and $T$. If Assumption (ref)(a) holds with rate $\sqrt n$ and Assumption (ref)(c-ii) holds with rate $\sqrt T$, then the above holds with rate $\min(\sqrt n,\sqrt T)$. Furthermore, $\bm{\mathcal J}_0\bm\Upsilon_0^\prime (\bm\Gamma^F)^{1/2}$ is finite and positive definite.

Finally, we can derive useful limits for $\widehat{\mathbf H}$, which depends only on population quantities (see Appendix (ref) for a proof).

propUnder Assumptions (ref) through (ref), as $n,T\to\infty$, \begin{inparaenum} • $\left\Vert\widehat{\mathbf H}- (\bm\Gamma^F)^{1/2}\bm\Upsilon_0 \bm{\mathcal J}_0 \right\Vert =o_{\mathrm P}(1)$; • $\left\Vert\widehat{\mathbf H}^{-1} - \bm{\mathcal J}_0\bm\Upsilon_0^\prime (\bm\Gamma^F)^{-1/2} \right\Vert =o_{\mathrm P}(1)$; \end{inparaenum} where the columns of $\bm\Upsilon_0$ are the normalized eigenvectors of $(\bm\Gamma^F)^{1/2}\bm\Sigma_\Lambda(\bm\Gamma^F)^{1/2}$, and $\bm{\mathcal J}_0$ is an $r\times r$ diagonal matrix with entries $\pm 1$ independent of $n$ and $T$. If Assumption (ref)(a) holds with rate $\sqrt n$ and Assumption (ref)(c-ii) holds with rate $\sqrt T$, then the above holds with rate $\min(\sqrt n,\sqrt T)$.

Part (a) is analogous to the result in bai2020simpler. According to this result and Propositions (ref) and (ref), the loadings and the factors are consistently estimated up to: (i) $(\bm\Gamma^F)^{1/2}$, (ii) a rotation, $\bm{\Upsilon}_0$, and (iii) a diagonal matrix of signs $\bm{\mathcal J}_0$.

Proposition (ref) can be proved either by directly combining Propositions (ref)(d) and (ref), or from Proposition (ref), by noticing that Assumptions (ref)(a) and (ref)(c-ii) imply $\Vert(T^{-1}\bm F^\prime\bm F)^{1/2}(n^{-1}\bm \Lambda^\prime\bm \Lambda) (T^{-1}\bm F^\prime\bm F)^{1/2}-(\bm\Gamma^F)^{1/2} \bm\Sigma_\Lambda(\bm\Gamma^F)^{1/2}\Vert =o_{\mathrm P}(1)$, and, therefore, by Davis Kahan theorem yu15, the corresponding normalized eigenvectors satisfy: $\Vert\widehat{\mathbf Q}-\bm\Upsilon_0\bm{\mathcal J}_0\Vert=o_{\mathrm P}(1)$.

Proposition (ref) and (ref) convey the same message. The difference is in the way we approximate $\widehat{\mathbf H}$. In Proposition (ref) the rate is $\min(n,\sqrt{T})$, however, the limiting quantity is still random and depends on $n$ and $T$. In Proposition (ref) we obtain a deterministic limiting quantity independent of $n$ and $T$, but the rate is slower and based on the rates at which Assumptions (ref)(a) and (ref)(c-ii) are satisfied, which are, obviously, $\sqrt n$ and $\sqrt T$, respectively.

We can now compare $\bm{\mathcal H}$ and $\widehat{\mathbf H}$ by means of the results in Propositions (ref) and (ref). The main difference is that while $\bm{\mathcal H}$ depends only on population quantities (but for an irrelevant sign indeterminacy), $\widehat{\mathbf H}$ depends both on population and estimated quantities. Such difference has a series of important consequences: (i) $\widehat{\mathbf H}$ is positive definite and finite only asymptotically as $n,T\to\infty$ (see Lemma (ref)), while $\bm{\mathcal H}$ is always positive definite and finite (see Lemma (ref)); (ii) $\widehat{\mathbf H}$ can be interpreted only asymptotically (see Propositions (ref), (ref), and (ref)), while $\bm{\mathcal H}$ has a well defined expression for all $n$ and $T$ as given in Proposition (ref); (iii) any identification restriction imposed on the loadings and/or the factors will constrain $\widehat{\mathbf H}$ only asymptotically, while we can derive exact expressions for $\bm{\mathcal H}$ (see Section (ref)).

Still, $\bm{\mathcal H}$ and $\widehat{\mathbf H}$ are asymptotically equivalent. This can be seen in two ways. First, by comparing Proposition (ref)(a) and Proposition (ref)(a), it follows that

align[align omitted — 576 chars of source]

It is important to stress that $\bm{\mathcal H}$ approximates the projection matrix of the estimated loadings onto the true loadings at a slower rate than $\widehat{\mathbf H}$. Indeed, while to prove Proposition (ref)(a) we are able to use a tighter bound than the one obtained for the loadings in Proposition (ref)(b), by bounding $ n^{-1}\Vert\bm\Lambda^\prime(\widehat{\bm\Lambda}-\bm\Lambda\widehat{\mathbf H})\Vert$, we already noticed that to prove Proposition (ref)(a), we can only use a looser bound, which therefore inherits the weaker bound obtained for the loadings in Proposition (ref)(a). This result has important implications when it comes to discussing the behaviour of $\bm{\mathcal H}$ and $\widehat{\mathbf H}$ under various identification conditions (see Section (ref)).

Second, since the columns of $\mathbf K$ in the definition of $\bm{\mathcal H}$ (see Proposition (ref)) are the normalized eigenvectors of $(\bm\Gamma^F)^{1/2} (n^{-1}\bm\Lambda^\prime\bm\Lambda) (\bm\Gamma^F)^{1/2}$, then, by Assumption (ref)(a) and Davis Kahan theorem yu15 it must hold that $\lim_{n\to\infty}\Vert \mathbf K-\bm\Upsilon_0\bm{\mathcal J}^*\Vert =0$ for some $r\times r$ diagonal matrix $\bm{\mathcal J}^*$ with entries $\pm 1$ and independent of $n$, and letting $\bm{\mathcal J}_0$ be such that $\Vert\bm{\mathcal J}^*\mathbf J -\bm{\mathcal J}_0\Vert=o_{\mathrm P}(1)$ as $n,T\to\infty$, we get $\Vert \mathbf K\mathbf J-\bm\Upsilon_0\bm{\mathcal J}_0\Vert=o_{\mathrm P}(1)$ (see (ref) and (ref) in the proof of Proposition (ref)). As a consequence of this result, of the definition of $\bm{\mathcal H}$ in Proposition (ref), and of Proposition (ref)(a), it immediately follows that

align[align omitted — 296 chars of source]

thus, providing a proof of the asymptotic equivalence of $\bm{\mathcal H}$ and $\widehat{\mathbf H}$ alternative to the proof in (ref). This proof, however, relies on a looser bound than the one used in (ref), indeed, it would deliver the slower rate $\min(\sqrt n,\sqrt T)$ inherited from Proposition (ref).

Asymptotic normality

In this section we prove asymptotic normality of the PC estimators. In particular, Theorems (ref) and (ref) proved below directly show the relationship with the unfeasible OLS estimators. The analogous results given in Bai03 are not so intuitive, but more intuitive asymptotic expansions, equivalent to those derived in this section, can be easily derived (see Section (ref) for details).

It is crucial to stress that, in order to prove asymptotic normality without need of imposing too restrictive bounds on the rates of divergence of $n$ and $T$, we have to make use of the rates derived in Propositions (ref) and (ref), which are sharper than those derived in Proposition (ref). This, however, requires the use of the matrix $\widehat{\mathbf H}$ rather than $\bm{\mathcal H}$.

Asymptotic normality of the PC estimator of the loadings follows from Lemma (ref).

theoremUnder Assumptions (ref) through (ref), as $n,T\to\infty$, for any given $i=1,\ldots,n$, $$ \sqrt T(\widehat{\bm\lambda}_i-\widehat{\mathbf H}^\prime{\bm\lambda}_i) \to_d\mathcal N\left(\mathbf 0_r, \bm\Upsilon_0^\prime(\bm\Gamma^F)^{1/2}\bm\Theta_i^{\text{\tiny \upshape OLS}} (\bm\Gamma^F)^{1/2} \bm\Upsilon_0 \right), $$ with $\bm\Upsilon_0$ defined in (ref) and $\bm\Theta_i^{\text{\tiny \upshape OLS}}:=(\bm\Gamma^F)^{-1} \bm\Phi_i (\bm\Gamma^F)^{-1}$ with $\bm\Phi_i$ defined in Assumption (ref)(a).

Proof of Theorem (ref). From the transposed of (ref) for any $i=1,\ldots, n$, we get

align[align omitted — 1,345 chars of source]

because of Proposition (ref)(a) and Lemma (ref), the definition of $\widehat{\mathbf H}$ in (ref), and since $\Vert n ({\widehat{\mathbf M}^x})^{-1}\Vert=O_{\mathrm P}(1)$ and $\Vert\bm{\mathcal H}\Vert=O(1)$ by Lemmas (ref)(iv) and (ref)(i), respectively.

Define \[ \widehat{\bm\lambda}_i^{\text{\tiny OLS}} := \left(\frac{\bm F^\prime\bm F}{T}\right)^{-1} \left(\frac 1{\sqrt T}\sum_{t=1}^T \mathbf F_t x_{it}\right), \] which is the unfeasible OLS estimator of $\bm\lambda_i$ when regressing $x_{it}$ onto $\mathbf F_t$. Then, since by Assumption (ref) $\sqrt T/n\to 0$ as $n,T\to\infty$, from (ref), using Proposition (ref)(a),

align[align omitted — 670 chars of source]

Moreover, as $T\to\infty$,

equation[equation omitted — 291 chars of source]

by Assumptions (ref)(b), (ref)(c-ii), and (ref)(a), and Slutsky's theorem. The proof follows by substituting (ref) into (ref), using again Slutsky's theorem, and since $\bm{\mathcal J}_0$ being a diagonal sign matrix plays no role in the asymptotic covariance. $\Box$\\

Moving to the PC estimator of the factors, asymptotic normality follows from Lemma (ref).

theoremUnder Assumptions (ref) through (ref), as $n,T\to\infty$, for any given $t=1,\ldots,T$, $$ \sqrt n(\widehat{\mathbf F}_t-\widehat{\mathbf H}^{-1}{\mathbf F}_t) \to_d\mathcal N\left(\mathbf 0_r, \bm\Upsilon_0^\prime(\bm\Gamma^F)^{-1/2}\bm\Pi_t^{\text{\tiny \upshape OLS}}(\bm\Gamma^F)^{-1/2}\bm\Upsilon_0\right), $$ with $\bm\Upsilon_0$ defined in (ref) and $\bm\Pi_t^{\text{\tiny \upshape OLS}}:=(\bm\Sigma_\Lambda)^{-1}\bm\Gamma_t(\bm\Sigma_\Lambda)^{-1}$ with $\bm\Gamma_t$ defined in Assumption (ref)(b).

Proof of Theorem (ref). From the transposed of (ref), for any $t=1,\ldots, T$, we get

align[align omitted — 1,483 chars of source]

because of Proposition (ref)(b) and Lemma (ref), the definition of $\widehat{\bm\Lambda}$ in (ref), and since $\Vert n({\widehat{\mathbf M}^x})^{-1}\Vert=O_{\mathrm P}(1)$ by Lemma (ref)(iv). Define, \[ \widehat{\mathbf F}_t^{\text{\tiny OLS}}:=\left(\frac{\bm\Lambda^\prime\bm\Lambda}{n}\right)^{-1}\left(\frac 1{\sqrt n}\sum_{i=1}^n \bm\lambda_i x_{it}\right), \] which is the unfeasible OLS estimator of $\mathbf F_t$ when regressing $x_{it}$ onto $\bm\lambda_i$. Then, since by Assumption (ref) $\sqrt n/T\to 0$ as $n,T\to\infty$, from (ref), using Proposition (ref)(b),

align[align omitted — 682 chars of source]

Moreover, as $n\to\infty$,

equation[equation omitted — 297 chars of source]

by Assumptions (ref)(a) and (ref)(b), and Slutsky's theorem. The proof follows by substituting (ref) into (ref), using again Slutsky's theorem, and since $\bm{\mathcal J}_0$ being a diagonal sign matrix plays no role in the asymptotic covariance. $\Box$\\

It is important to stress that Theorems (ref) and (ref), as well as the analogous theorems in Bai03, cannot be used to make inference on the loadings or the factors. This is due to the presence of the unknown matrix $\widehat{\mathbf H}$, which depends on the unknown true factors and loadings. Notice that even when $r=1$, so that $\widehat{\mathbf H}$ is a scalar, still factors and loadings are in general consistently estimated only up to an unknown scale.

In particular, the presence of $\widehat{\mathbf H}$ has two effects. First, the location of the asymptotic distribution of $\widehat{\bm\lambda}_i$ and $\widehat{\mathbf F}_t$ are random. Second, the asymptotic covariance matrix cannot be estimated consistently, since any such estimator would require consistent estimators of $\bm\Gamma^F$ and/or $\bm\Sigma_\Lambda$, but these cannot be estimated consistently. Indeed, the candidate estimators $T^{-1}\widehat{\bm F}^\prime \widehat{\bm F}$ and $n^{-1}\widehat{\bm \Lambda}^\prime \widehat{\bm \Lambda}$ have probability limits which still depend on $\widehat{\mathbf H}$ (see Proposition (ref)(g) and (ref)(h)).

sidewaystable[htbp] \caption{Asymptotic normality} \scriptsize{ \begin{tabular}{l | l | c | c | c } \hline \hline &&&&\\ (ref) $\sqrt T(\widehat{\bm\lambda}_i-\widehat{\mathbf H}^\prime{\bm\lambda}_i)\to_d\mathcal N\left(\mathbf 0_r,\mathbf H_{0,A}^\prime \bm\Theta_i^{\text{OLS}}\mathbf H_{0,A}\right) $&&&&\\ (ref) $\sqrt T(\widehat{\bm\lambda}_i-\widehat{\mathbf H}^\prime{\bm\lambda}_i)\to_d\mathcal N\left(\mathbf 0_r,\mathbf H_{1,A}^\prime \bm\Theta_i^{\text{OLS}}\mathbf H_{1,A}\right) $&$\widehat{\mathbf H}$ & $\left(\frac{\bm F^\prime\bm F}{T}\right)^{1/2}\widehat{\mathbf Q}$ & $\mathbf H_{0,A}:=(\bm \Gamma^F)^{1/2} \bm\Upsilon_0 \bm{\mathcal J}_0$ & $\mathbf H_{1,A}:=(\bm\Sigma_\Lambda)^{-1/2}\bm\Upsilon_1 \bm{\mathcal J}_1\bm V_0^{1/2}$\\[3pt] (ref) $\sqrt n(\widehat{\mathbf F}_t-\widehat{\mathbf H}^{-1}{\mathbf F}_t)\to_d\mathcal N\left(\mathbf 0_r,\mathbf H_{0,A}^{-1} \bm\Pi_t^{\text{OLS}}(\mathbf H_{0,A}^{-1})^\prime\right)$& rate & $\min(n,\sqrt{T})$ & $\min(\sqrt n,\sqrt T)$ & $\min(\sqrt n,\sqrt T)$\\[3pt] (ref) $\sqrt n(\widehat{\mathbf F}_t-\widehat{\mathbf H}^{-1}{\mathbf F}_t)\to_d\mathcal N\left(\mathbf 0_r,\mathbf H_{1,A}^{-1} \bm\Pi_t^{\text{OLS}}(\mathbf H_{1,A}^{-1})^\prime\right)$& &Proposition (ref) & Proposition (ref) & (ref)\\[3pt] \hline &&&&\\[-3pt] (ref) $\sqrt T(\widetilde{\bm\lambda}_i-\widetilde{\mathbf H}^{-1}{\bm\lambda}_i)\to_d\mathcal N\left(\mathbf 0_r,\mathbf H_{0,B}^{-1}\bm\Theta_i^{\text{OLS}}(\mathbf H_{0,B}^{-1})^\prime\right)$&&&&\\ (ref) $\sqrt T(\widetilde{\bm\lambda}_i-\widetilde{\mathbf H}^{-1}{\bm\lambda}_i)\to_d\mathcal N\left(\mathbf 0_r,\mathbf H_{1,B}^{-1}\bm\Theta_i^{\text{OLS}}(\mathbf H_{1,B}^{-1})^\prime\right)$&$\widetilde{\mathbf H}$ &$ \left(\frac{\bm F^\prime\bm F}{T}\right)^{-1/2} \widehat{\mathbf Q}$& $\mathbf H_{0,B}:=(\bm\Gamma^F)^{-1/2} \bm\Upsilon_0\bm{\mathcal J}_0$ & $\mathbf H_{1,B}:=(\bm\Sigma_\Lambda)^{1/2}\bm\Upsilon_1 \bm{\mathcal J}_1 \bm V_0^{-1/2}$\\[3pt] (ref) $\sqrt n(\widetilde{\mathbf F}_t-\widetilde{\mathbf H}^{\prime}{\mathbf F}_t)\to_d\mathcal N\left(\mathbf 0_r,\mathbf H_{0,B}^\prime\bm\Pi_t^{\text{OLS}} \mathbf H_{0,B}\right)$& rate & $\min(n,\sqrt{T})$ & $\min(\sqrt n,\sqrt T)$ & $\min(\sqrt n,\sqrt T)$\\[3pt] (ref) $\sqrt n(\widetilde{\mathbf F}_t-\widetilde{\mathbf H}^{\prime}{\mathbf F}_t)\to_d\mathcal N\left(\mathbf 0_r,\mathbf H_{1,B}^\prime\bm\Pi_t^{\text{OLS}} \mathbf H_{1,B}\right)$&&(ref)&(ref)& (ref)\\[3pt] \hline &&&&\\[-3pt] (ref) $\sqrt T(\wideparen{\bm\lambda}_i-\wideparen{\mathbf H}^{\prime}{\bm\lambda}_i)\to_d\mathcal N\left(\mathbf 0_r,\mathbf H_{0,C}^{\prime}\bm\Theta_i^{\text{OLS}}\mathbf H_{0,C}\right)$&&&&\\ (ref) $\sqrt T(\wideparen{\bm\lambda}_i-\wideparen{\mathbf H}^{\prime}{\bm\lambda}_i)\to_d\mathcal N\left(\mathbf 0_r,\mathbf H_{1,C}^{\prime}\bm\Theta_i^{\text{OLS}}\mathbf H_{1,C}\right)$&$\wideparen{\mathbf H}$ &$ \left(\frac{\bm\Lambda^\prime\bm\Lambda}{n}\right)^{-1/2} \wideparen {\mathbf Q}$& $\mathbf H_{0,C}:=(\bm\Gamma^F)^{1/2} \bm\Upsilon_0\bm{\mathcal J}_0\bm V_0^{-1/2}$ & $\mathbf H_{1,C}:=(\bm\Sigma_\Lambda)^{-1/2} \bm\Upsilon_1\bm{\mathcal J}_1$\\[3pt] (ref) $\sqrt n(\wideparen{\mathbf F}_t-\wideparen{\mathbf H}^{-1}{\mathbf F}_t ) \to_d\mathcal N\left(\mathbf 0_r, \mathbf H_{0,C}^{-1}\bm\Pi_t^{\text{\tiny \upshape OLS}}(\mathbf H_{0,C}^{-1})^\prime\right)$& rate & $\min(n,\sqrt{T})$ & $\min(\sqrt n,\sqrt T)$ & $\min(\sqrt n,\sqrt T)$\\[3pt] (ref) $\sqrt n(\wideparen{\mathbf F}_t-\wideparen{\mathbf H}^{-1}{\mathbf F}_t ) \to_d\mathcal N\left(\mathbf 0_r, \mathbf H_{1,C}^{-1}\bm\Pi_t^{\text{\tiny \upshape OLS}}(\mathbf H_{1,C}^{-1})^\prime\right)$&&(ref)&(ref)& (ref)\\[3pt] \hline &&&&\\[-3pt] (ref) $\sqrt T(\bar{\bm\lambda}_i-\bar{\mathbf H}^{-1}{\bm\lambda}_i)\to_d\mathcal N\left(\mathbf 0_r,\mathbf H_{0,D}^{-1}\bm\Theta_i^{\text{OLS}}(\mathbf H_{0,D}^{-1})^\prime\right)$ &&&&\\ (ref) $\sqrt T(\bar{\bm\lambda}_i-\bar{\mathbf H}^{\prime}{\bm\lambda}_i)\to_d\mathcal N\left(\mathbf 0_r,\mathbf H_{1,D}^{-1}\bm\Theta_i^{\text{OLS}}(\mathbf H_{1,D}^{-1})^\prime\right)$&$\bar{\mathbf H}$ &$ \left(\frac{\bm\Lambda^\prime\bm\Lambda}{n}\right)^{1/2} \wideparen {\mathbf Q}$& $\mathbf H_{0,D}:=(\bm\Gamma^F)^{-1/2} \bm\Upsilon_0\bm{\mathcal J}_0\bm V_0^{1/2}$ & $\mathbf H_{1,D}:=(\bm\Sigma_\Lambda)^{1/2} \bm\Upsilon_1\bm{\mathcal J}_1$\\[3pt] (ref) $\sqrt n(\bar{\mathbf F}_t-\bar{\mathbf H}^{\prime}{\mathbf F}_t ) \to_d\mathcal N\left(\mathbf 0_r, \mathbf H_{0,D}^{\prime}\bm\Pi_t^{\text{\tiny \upshape OLS}}\mathbf H_{0,D}\right)$& rate & $\min(n,\sqrt{T})$ & $\min(\sqrt n,\sqrt T)$ & $\min(\sqrt n,\sqrt T)$\\[3pt] (ref) $\sqrt n(\bar{\mathbf F}_t-\bar{\mathbf H}^{\prime}{\mathbf F}_t ) \to_d\mathcal N\left(\mathbf 0_r, \mathbf H_{1,D}^{\prime}\bm\Pi_t^{\text{\tiny \upshape OLS}}\mathbf H_{1,D}\right)$ &&(ref)&(ref)& (ref)\\ &&&&\\ \hline \hline \end{tabular} \begin{tabular}{p{.8\textwidth}} $\widehat{\mathbf Q}$ are normalized eigenvectors of $(T^{-1}\bm F^\prime\bm F)^{1/2}(n^{-1}\bm \Lambda^\prime\bm \Lambda) (T^{-1}\bm F^\prime\bm F)^{1/2}$; \\ $\wideparen {\mathbf Q}$ are the normalized eigenvectors of $(n^{-1}\bm \Lambda^\prime\bm \Lambda)^{1/2} (T^{-1}\bm F^\prime\bm F) (n^{-1}\bm \Lambda^\prime\bm \Lambda)^{1/2}$; \\ $\bm\Upsilon_0$ are normalized eigenvectors of $(\bm\Gamma^F)^{1/2}\bm\Sigma_\Lambda(\bm\Gamma^F)^{1/2}$; $\bm\Upsilon_1$ are normalized eigenvectors of $(\bm\Sigma_\Lambda)^{1/2}\bm\Gamma^F(\bm\Sigma_\Lambda)^{1/2}$;\\ $\bm V_0=\lim_{n\to\infty} n^{-1}\mathbf M^{C}$ are eigenvalues of $(\bm\Gamma^F)^{1/2}\bm\Sigma_\Lambda(\bm\Gamma^F)^{1/2}$ and $(\bm\Sigma_\Lambda)^{1/2}\bm\Gamma^F(\bm\Sigma_\Lambda)^{1/2}$;\\ $\bm{\mathcal J}_0$ and $\bm{\mathcal J}_1$ are diagonal with entries $\pm 1$. \end{tabular} }

Equivalence with approach B1

The PC estimators studied so far coincide with those studied by Bai03, indeed, by (ref), $\widetilde{\bm\Lambda}=\widehat{\bm\Lambda}$ and $\widetilde{\bm F}=\widehat{\bm F}$. The purpose of this section is then to show that, by following the proofs of Bai03 we can derive the same asymptotic expansions as in Theorems (ref) and (ref), showing the relation between the PC and the unfeasible OLS estimators. These are more interpretable than the original results. Notice that, by virtue of Table (ref), the results in Bai03 can be proved under the same assumptions made in this paper so there is no need to prove them again here.\footnote{Although the estimators are identical, in this section we keep using the notation $\widetilde{\bm \Lambda}$ and $\widetilde{\bm F}$ to highlight that their properties are derived using a different approach with respect to the one used so far.}

\paragraph{Consistency.} First of all notice that Proposition (ref) still holds, and its proof using $\widetilde{\bm\Lambda}$ and $\widetilde{\bm F}$ is a special case of the proof by FLM13, when no uniform bounds are computed.\footnote{Note that the proof by FLM13 follows the same steps as in Bai03, but it derives slower rates since fewer cross-moment assumptions are imposed.} Now, by definition of eigenvectors, it holds that

equation[equation omitted — 131 chars of source]

By using $\bm X=\bm F\bm\Lambda^\prime+\bm E$ and taking the $t$th row of (ref), for any $t=1,\ldots,T$,

align[align omitted — 578 chars of source]

which is the expansion given in Bai03 and where

equation[equation omitted — 186 chars of source]

Notice that $\Vert\widetilde{\mathbf H}\Vert= O_{\mathrm P}(1)$ and $\Vert\widetilde{\mathbf H}^{-1}\Vert= O_{\mathrm P}(1)$ because, as $n,T\to\infty$, all its terms tend to finite and positive definite quantities (see Assumption (ref)(a), Proposition (ref), and Lemma (ref)(iii) jointly with Lemmas (ref)(i) and (ref)(iii)).

Then, from Bai03, $\Vert \text{(1.a)}\Vert= O_{\mathrm P}(\max(n^{-1},(nT)^{-1/2}) )$, $\Vert \text{(1.b)}\Vert= O_{\mathrm P}( n^{-1/2})$, and $\Vert \text{(1.c)}\Vert= O_{\mathrm P}(\max(n^{-1},(nT)^{-1/2}, T^{-1}) )$. Therefore, the estimated factors are consistent: $$ \left\Vert \widehat{\mathbf F}_t-\widetilde{\mathbf H}^{\prime}{\mathbf F}_t \right\Vert = O_{\mathrm P} \left(\max\left(\frac 1{\sqrt n},\frac 1T\right)\right), $$ which coincides with Proposition (ref), but when stated using $\widetilde{\mathbf H}$ in place of $\widehat{\mathbf H}$.

From Bai03, for any $i=1,\ldots, n$,

align[align omitted — 471 chars of source]

which is the obtained from the linear projection of $\bm X^\prime$ onto the estimated factors $\widetilde{\bm F}$ in a similar way as we obtained the linear projection of $\bm X$ onto the estimated loadings $\widehat{\bm \Lambda}$ giving (ref). Then, from Bai03, $\Vert \text{(2.a)}\Vert= O_{\mathrm P}(\max(n^{-1},(nT)^{-1/2},T^{-1}) )$, $\Vert \text{(2.b)}\Vert= O_{\mathrm P}(\max(n^{-1},(nT)^{-1/2},T^{-1}) )$, and $\Vert \text{(2.c)}\Vert= O_{\mathrm P}(T^{-1/2} )$. Therefore, the estimated loadings are consistent: $$ \left\Vert \widetilde{\bm\lambda}_i-\widetilde{\mathbf H}^{-1}{\bm\lambda}_i\right\Vert = O_{\mathrm P} \left(\max\left(\frac 1{ n},\frac 1{\sqrt T}\right)\right), $$ which coincides with Proposition (ref), but when stated using $\widetilde{\mathbf H}$ in place of $\widehat{\mathbf H}$.

\paragraph{The role of $ \widetilde{\mathbf H}$.} First, from Bai03,

align[align omitted — 225 chars of source]

which is the analogous of Proposition (ref)(b). This immediately implies also the analogous of Propositions (ref)(d), (ref)(f), and (ref)(h).

Second, following the same reasoning used to prove Proposition (ref), but this time using (ref), it follows that

align[align omitted — 280 chars of source]

from which it follows that we must have

align[align omitted — 437 chars of source]

where $\widehat {\mathbf Q}$ are the normalized eigenvectors of $(T^{-1}\bm F^\prime\bm F)^{1/2}(n^{-1}\bm \Lambda^\prime\bm \Lambda) (T^{-1}\bm F^\prime\bm F)^{1/2}$, which has $n^{-1}\widehat{\mathbf M}^C$ as eigenvalues, and these, in turn, are such that $n^{-1}\Vert \widehat{\mathbf M}^C-\widehat{\mathbf M}^x\Vert = O_{\mathrm P}(\max(n^{-1},T^{-1/2}))$ with $T^{-1}\widetilde{\mathbf M}^x=n^{-1}\widehat{\mathbf M}^x$. Results (ref) and (ref) are the analogous of Proposition (ref). Once again, we see that the loadings and the factors can be consistently estimated up to: (i) a scale $(T^{-1}{\bm F^\prime\bm F})^{-1/2}$ and (ii) a rotation $\widehat{\mathbf Q}$.

We can then derive a limit for $\widetilde{\mathbf H}$ which depends only on population quantities. This can be done in two ways. First, by Assumptions (ref)(a) and (ref)(c-ii), we also have $\Vert(T^{-1}\bm F^\prime\bm F)^{1/2}(n^{-1}\bm \Lambda^\prime\bm \Lambda) (T^{-1}\bm F^\prime\bm F)^{1/2}-(\bm\Gamma^F)^{1/2} \bm\Sigma_\Lambda(\bm\Gamma^F)^{1/2}\Vert =o_{\mathrm P}(1)$. Therefore, by Davis Kahan theorem yu15, the corresponding normalized eigenvectors satisfy: $\Vert\widehat{\mathbf Q}-\bm\Upsilon_0\bm{\mathcal J}_0\Vert=o_{\mathrm P}(1)$, which, jointly with (ref) and Assumption (ref)(c-ii), implies

equation[equation omitted — 302 chars of source]

which is the analogous of Proposition (ref).

Second, consider the following spectral decomposition:

equation[equation omitted — 139 chars of source]

where $\bm\Upsilon_1$ is the $r\times r$ matrix having as columns the normalized eigenvectors of $(\bm\Sigma_\Lambda)^{1/2}\bm\Gamma^F(\bm\Sigma_\Lambda)^{1/2}$, and $\bm V_0$ is the $r\times r$ matrix of corresponding eigenvalues sorted in decreasing order, which coincide with those of $(\bm\Gamma^F)^{1/2}\bm\Sigma_\Lambda ( \bm\Gamma^F)^{1/2}$, and are given by $\bm V_0=\lim_{n\to\infty} n^{-1}\mathbf M^{C}$. Then, from Bai03

equation[equation omitted — 193 chars of source]

where $\bm{\mathcal J}_1$ is an $r \times r$ diagonal matrix with entries $\pm 1$.\footnote{Notice that the sign indeterminacy due to $\bm{\mathcal J}_1$ is not present in the original proof since the sign of the eigenvectors is implicitly fixed.} Notice that the limiting quantity in (ref) is finite and positive definite because of Assumption (ref)(a) and Lemmas (ref)(i) and (ref)(iii). From (ref), (ref), Assumption (ref)(a), and Lemma (ref)(iv),

equation[equation omitted — 338 chars of source]

If Assumption (ref)(a) holds with rate $\sqrt n$ and Assumption (ref)(c-ii) holds with rate $\sqrt T$, then (ref), (ref), and (ref) hold with rate $\min(\sqrt n,\sqrt T)$.

\paragraph{A new limit for $\widehat{\mathbf H}$.} From (ref) and (ref) and uniqueness of the limits,

align[align omitted — 326 chars of source]

And, from Proposition (ref)(b) and (ref), it follows that we also have

equation[equation omitted — 204 chars of source]

Therefore, from (ref), (ref), Assumption (ref)(c-ii), and Lemma (ref)(iv),

equation[equation omitted — 191 chars of source]

and, since from (ref), $\bm\Gamma^F(\bm\Sigma_\Lambda)^{1/2}=(\bm\Sigma_\Lambda)^{-1/2}\bm\Upsilon_1\bm V_0\bm\Upsilon_1^\prime$, from (ref) we obtain

equation[equation omitted — 339 chars of source]

Once again, if Assumption (ref)(a) holds with rate $\sqrt n$ and Assumption (ref)(c-ii) holds with rate $\sqrt T$, the rate in (ref)-(ref) is $\min(\sqrt n,\sqrt T)$.

\paragraph{The relation between $\widetilde{\mathbf H}$ and $\widehat{\mathbf H}$.} Since $\widetilde{\bm F}=\widehat{\bm F}$ from Proposition (ref)(b) and (ref), we see that we must have

align[align omitted — 362 chars of source]

It follows that, by substituting (ref) into Proposition (ref)(a), we obtain its analogous:

align[align omitted — 246 chars of source]

which immediately implies also the analogous of Propositions (ref)(c), (ref)(e), and (ref)(g).

\paragraph{Asymptotic normality.} From (ref) and Bai03, if $\sqrt n/T\to 0$, as $n,T\to\infty$,

align[align omitted — 1,001 chars of source]

where the first line is the only asymptotic expansion reported by Bai03, and for the subsequent equalities we used also (ref), (ref), and (ref), or (ref). The last expression in (ref) coincides with (ref) in the proof of Theorem (ref), hence, by Slutsky's theorem,

equation[equation omitted — 269 chars of source]

Moreover, using the second last line of (ref), by Slutsky's theorem, we also have

equation[equation omitted — 304 chars of source]

where $\bm\Pi_t^{\text{\tiny \upshape OLS}}=(\bm\Sigma_\Lambda)^{-1}\bm\Gamma_t(\bm\Sigma_\Lambda)^{-1}$ with $\bm\Gamma_t$ is defined in Assumption (ref)(b). Furthermore, letting $\bm{\mathcal Q}:=\bm V_0^{1/2}\bm\Upsilon_1^\prime(\bm\Sigma_\Lambda)^{-1/2}$ the asymptotic covariance matrix of $\widetilde{\mathbf F}_t$ can also be written as $(\bm V_0)^{-1}\bm {\mathcal Q}\bm\Gamma_t\bm {\mathcal Q}^\prime (\bm V_0)^{-1}$, which is the expression given in Bai03. Notice that from (ref) it follows that we can state Thoreom (ref) also as

equation[equation omitted — 296 chars of source]

Similarly, from (ref) and Bai03, if $\sqrt T/n\to 0$, as $n,T\to\infty$,

align[align omitted — 1,310 chars of source]

where the first line is the only asymptotic expansion reported by Bai03, and for the subsequent equalities we used also the fact that $T^{-1}\widetilde{\mathbf F}^\prime\widetilde{\mathbf F}=\mathbf I_r$ by definition, (ref) which implies $\Vert T^{-1}\widetilde{\bm F}^\prime\widetilde{\bm F}-T^{-1} \widetilde{\mathbf H}^{\prime}{\bm F}^\prime{\bm F}\widetilde{\mathbf H} \Vert=o_{\mathrm P}(1)$, (ref), and (ref), or (ref). The last expression in (ref) coincides with (ref) in the proof of Theorem (ref), hence, by Slutsky's theorem,

equation[equation omitted — 264 chars of source]

Moreover, using the second last line of (ref), by Slutsky's theorem, we also have

equation[equation omitted — 304 chars of source]

where $\bm\Theta_i^{\text{\tiny \upshape OLS}}=(\bm\Gamma^F)^{-1}\bm\Phi_i(\bm\Gamma^F)^{-1}$ with $\bm\Phi_i$ is defined in Assumption (ref)(a). Furthermore, from the second line of (ref), using (ref), we also have that the asymptotic covariance matrix of $\widetilde{\bm\lambda}_i$ can also be written as $\bm V_0^{-1/2}\bm\Upsilon_1^\prime (\bm\Sigma_\Lambda)^{1/2} \bm\Phi_i (\bm\Sigma_\Lambda)^{1/2}\bm\Upsilon_1\bm V_0^{-1/2}$ and by letting $\bm {\mathcal Q}:=(\bm V_0)^{1/2}\bm\Upsilon_1^\prime(\bm\Sigma_\Lambda)^{-1/2}$ this is equivalent to $(\bm {\mathcal Q}^{\prime})^{-1} \bm\Phi_i\bm {\mathcal Q}^{-1}$, which is the expression given in Bai03. Notice that from (ref) it follows that we can state Thoreom (ref) also as

equation[equation omitted — 304 chars of source]

Identification and inference

\paragraph{Constraints in exploratory factor analysis.} In this section, we consider a series of identifying restriction where either the factors or the loadings are assumed to be orthonormal. These are statistical restrictions in the sense that the identified factors have no direct interpretation for the considered data. In other words, we are considering only exploratory factor analysis, but not confirmatory factor analysis. Let us first summarize the identifying conditions that are typically found in the literature.\footnote{ Other identifying conditions related to confirmatory factor analysis exist in the literature. In particular, an appealing one is to assume $\bm\Lambda^\prime=[\mathbf I_r | \bm\Lambda^{\prime}_1]$ with $\bm\Lambda_1$ being $N-r\times r$ and full and leave $\bm\Gamma^F$ and, especially, $T^{-1}\bm F^\prime \bm F$ unrestricted. This is tantamount as assuming the the first $r$ elements of the observables are noisy measures of the latent factors, i.e., $x_{jt}= F_{jt}+ e_{jt}$, $j=1,\ldots, r$ PF86,AFP87. }

compactenum$n^{-1}\bm\Lambda^\prime\bm\Lambda$ diagonal for all $n\in\mathbb N$ (AR56; baing13). • $n^{-1}\bm\Lambda^\prime(\text{diag}(\bm\Gamma^ e))^{-1}\bm\Lambda$ diagonal for all $n\in\mathbb N$ (AR56; lawleymaxwell71; mardia1979multivariate; baili12,baili16). • $\bm\Gamma^F=\mathbf I_r$ (hotelling1933analysis; AR56; lawleymaxwell71; mardia1979multivariate; RT82; jolliffe2002principal; anderson2003introduction). • $T^{-1}\bm F^\prime\bm F=\mathbf I_r$ for all $T\in\mathbb N$ (anderson2003introduction; baili12,baili16; onatski2012asymptotics; baing13; Freyaldenhoven2022).

Two comments follow from inspection of the above works. First, when considering maximum likelihood estimation of an exact factor model, for the loadings it is usually imposed (B) rather than (A). Indeed, the maximum likelihood estimates of loadings and idiosyncratic variances depend on each other so they are usually jointly constrained. Typically, (A) is instead imposed when considering PC estimation, since in that case we are not interested in estimating the idiosyncratic covariance. Hence we will focus only on (A). Classical works consider only the fixed $n$ case, so it is natural to assume that (A) holds for all $n\in\mathbb N$. However, when we allow $n\to\infty$, (A) might seem too restrictive, and it is interesting to consider also its limiting case, i.e., $\bm\Sigma_\Lambda=\mathbf I_r$.

The choice between (C) and (D) depends on the nature of the factors. The distinction between random and deterministic factors can be found only in the context of maximum likelihood estimation (anderson2003introduction; AFP87; baili12), where it is shown that, regardless of such distinction, asymptotic normality of the maximum likelihood estimator of the loadings can be derived provided we impose (A) or, more often, (B) and either (C) if factors are random or (D) if factors are deterministic. So in maximum likelihood estimation the distinction is irrelevant. However, this distinction is extremely relevant in PC estimation since, if the factors are random, (D) is not realistic and (C) should be imposed instead. Nevertheless, we will also consider (D) in the following to clearly show its implications.

\paragraph{Identification of the true loadings and factors.} The next results show the implications of various identifying constraints for the true loadings and factors (see Appendix (ref) for a proof).

propUnder Assumptions (ref) and (ref), \begin{compactenum} • if $\bm\Gamma^F=\mathbf I_r$ and $\bm\Lambda$ is unrestricted, then: \begin{compactenum} • $\bm\lambda_i^\prime=\mathbf v_i^{C\prime}(\mathbf M^{C})^{1/2}\mathbf K^\prime$ for all $i=1,\ldots, n$ and all $n\in\mathbb N$; • $\mathbf F_t=\mathbf K(\mathbf M^{C})^{-1/2} \mathbf V^{{C}\prime}\bm{C}_t$ for all $t\in\mathbb Z$ and all $n\in\mathbb N$; • $\lim_{n\to\infty}\Vert \bm\lambda_i^\prime-\bm p_{i0}^\prime (\bm V_0)^{1/2} \bm\Upsilon_0^\prime \Vert=0$ for all $i\in\mathbb N$; • $\text{\upshape m.s.-}\lim_{n\to\infty}\Vert \mathbf F_t-\bm\Upsilon_0(\bm V_0)^{-1/2} \bm W_{\infty}^\prime\bm C_{t,\infty}\Vert=0$ for all $t\in\mathbb Z$; \end{compactenum} • if $\bm\Gamma^F=\mathbf I_r$ and $n^{-1}\bm\Lambda^\prime\bm \Lambda$ is diagonal for all $n\in\mathbb N$, then: \begin{compactenum} • $\bm\lambda_i^\prime=\mathbf v_i^{C\prime}(\mathbf M^{C})^{1/2}\bm S$ for all $i=1,\ldots, n$ and all $n\in\mathbb N$; • $\mathbf F_t=\bm S(\mathbf M^{C})^{-1/2}\mathbf V^{{C}\prime}\bm{C}_t$ for all $t\in\mathbb Z$ and all $n\in\mathbb N$; • $\lim_{n\to\infty}\Vert \bm\lambda_i^\prime-\bm p_{i0}^\prime (\bm V_0)^{1/2} \bm {\mathcal S}_0 \Vert=0$ for all $i\in\mathbb N$; • $\text{\upshape m.s.-}\lim_{n\to\infty}\Vert\mathbf F_t-\bm{\mathcal S}_0(\bm V_0)^{-1/2} \bm W_{\infty}^{\prime}\bm{C}_{t,\infty}\Vert=0$ for all $t\in\mathbb Z$; \end{compactenum} • if $\bm\Gamma^F=\mathbf I_r$ and $\bm\Sigma_\Lambda$ is diagonal, then: \begin{compactenum} • $\lim_{n\to\infty}\Vert \bm\lambda_i^\prime-\bm p_{i0}^{\prime}(\bm V_0)^{1/2}\bm{\mathcal S}_0\Vert=0$ for all $i\in\mathbb N$; • $\text{\upshape m.s.-}\lim_{n\to\infty}\Vert\mathbf F_t-\bm{\mathcal S}_0(\bm V_0)^{-1/2} \bm W_{\infty}^{\prime}\bm{C}_{t,\infty}\Vert=0$ for all $t\in\mathbb Z$; \end{compactenum} • if $T^{-1} \bm F^\prime\bm F=\mathbf I_r$ for all $T\in\mathbb N$ and $\bm\Lambda$ is unrestricted, then: \begin{compactenum} • $\bm\lambda_i^\prime=\widehat{\mathbf v}_i^{C\prime}(\widehat{\mathbf M}^{C})^{1/2}\widehat{\mathbf Q}^\prime$ for all $i=1,\ldots, n$ and all $n,T\in\mathbb N$; • $\mathbf F_t=\widehat{\mathbf Q}(\widehat{\mathbf M}^{C})^{-1/2} \widehat{\mathbf V}^{{C}\prime}\bm{C}_t=0$ for all $t=1,\ldots, T$ and all $n,T\in\mathbb N$; • $\lim_{T\to\infty}\Vert \bm\lambda_i^\prime-{\mathbf v}_i^{C\prime}({\mathbf M}^{C})^{1/2}{\mathbf K}^\prime\Vert =0$ for all $i=1,\ldots, n$ and all $n\in\mathbb N$; • $\text{\upshape m.s.-}\lim_{T\to\infty}\Vert\mathbf F_t-{\mathbf K}(\widehat{\mathbf M}^{C})^{-1/2} {\mathbf V}^{{C}\prime}\bm{C}_t\Vert=0$ for all $t\in\mathbb N$ and all $n\in\mathbb N$; • $\lim_{n,T\to\infty}\Vert \bm\lambda_i^\prime-\bm p_{i0}^\prime (\bm V_0)^{1/2} \bm\Upsilon_0^\prime \Vert=0$ for all $i\in\mathbb N$; • $\text{\upshape m.s.-}\lim_{n,T\to\infty}\Vert \mathbf F_t- \bm\Upsilon_0(\bm V_0)^{-1/2} \bm W_{\infty}^\prime\bm C_{t,\infty}\Vert=0$ for all $t\in\mathbb N$; \end{compactenum} • if $T^{-1} \bm F^\prime\bm F=\mathbf I_r$ and $n^{-1}\bm\Lambda^\prime\bm \Lambda$ is diagonal for all $n,T\in\mathbb N$, then: \begin{compactenum} • $\bm\lambda_i^\prime=\widehat{\mathbf v}_i^{C\prime}(\widehat{\mathbf M}^{C})^{1/2}\widehat{\bm S}$ for all $i=1,\ldots, n$ and all $n,T\in\mathbb N$; • $\mathbf F_t=\widehat{\bm S}(\widehat{\mathbf M}^{C})^{-1/2} \widehat{\mathbf V}^{{C}\prime}\bm{C}_t=0$ for all $t=1,\ldots, T$ and all $n,T\in\mathbb N$; • $\lim_{T\to\infty}\Vert \bm\lambda_i^\prime-{\mathbf v}_i^{C\prime}({\mathbf M}^{C})^{1/2}{\bm S}^\prime\Vert =0$ for all $i=1,\ldots, n$ and all $n\in\mathbb N$; • $\text{\upshape m.s.-}\lim_{T\to\infty}\Vert\mathbf F_t-{\bm S}(\widehat{\mathbf M}^{C})^{-1/2} {\mathbf V}^{{C}\prime}\bm{C}_t\Vert=0$ for all $t\in\mathbb N$ and all $n\in\mathbb N$; • $\lim_{n,T\to\infty}\Vert \bm\lambda_i^\prime-\bm p_{i0}^\prime (\bm V_0)^{1/2} \bm{\mathcal S}_0 \Vert=0$ for all $i\in\mathbb N$; • $\text{\upshape m.s.-}\lim_{n,T\to\infty}\Vert \mathbf F_t- \bm{\mathcal S}_0(\bm V_0)^{-1/2} \bm W_{\infty}^\prime\bm C_{t,\infty}\Vert=0$ for all $t\in\mathbb N$; \end{compactenum} • if $T^{-1} \bm F^\prime\bm F=\mathbf I_r$ for all $T\in\mathbb N$ and $\bm\Sigma_\Lambda$ is diagonal, then: \begin{compactenum} • $\lim_{n,T\to\infty}\Vert \bm\lambda_i^\prime-\bm p_{i0}^{\prime}(\bm V_0)^{1/2}\bm{\mathcal S}_0\Vert=0$ for all $i\in\mathbb N$; • $\text{\upshape m.s.-}\lim_{n,T\to\infty}\Vert\mathbf F_t-\bm{\mathcal S}_0(\bm V_0)^{-1/2} \bm W_{\infty}^\prime\bm C_{t,\infty}\Vert=0$ for all $t\in\mathbb N$; \end{compactenum} • if $\bm\Sigma_\Lambda=\mathbf I_r$ and $\bm F$ is unrestricted, then: \begin{compactenum} • $\lim_{n\to\infty} \left\Vert\bm\lambda_i^\prime - \bm p_{i0}^{\prime} \bm\Upsilon_1^\prime \right \Vert=0$ for all $i\in\mathbb N$; • $\text{\upshape m.s.-}\lim_{n\to\infty}\Vert\mathbf F_t-\bm \Upsilon_1 \bm W_{\infty}^\prime\bm C_{t,\infty}\Vert =0$ for all $t\in\mathbb Z$; \end{compactenum} • if $\bm\Sigma_\Lambda=\mathbf I_r$ and $T^{-1} \bm F^\prime\bm F$ is diagonal for all $T\in\mathbb N$, then: \begin{compactenum} • $\text{\upshape P-}\lim_{n,T\to\infty}\Vert\bm\lambda_i^\prime- \bm p_{i0}^\prime \bm{\mathcal S}_0\Vert=0$ for all $i\in\mathbb N$; • $\text{\upshape P-}\lim_{n,T\to\infty}\Vert\mathbf F_t-\bm {\mathcal S}_0 \bm W_{\infty}^\prime\bm C_{t,\infty}\Vert =0$ for all $t\in\mathbb N$; \end{compactenum} • if $\bm\Sigma_\Lambda=\mathbf I_r$ and $\bm\Gamma^F$ is diagonal, then: \begin{compactenum} • $\lim_{n\to\infty}\Vert\bm\lambda_i^\prime- \bm p_{i0}^\prime \bm{\mathcal S}_0\Vert=0$ for all $i\in\mathbb N$; • $\text{\upshape m.s.-}\lim_{n\to\infty}\Vert\mathbf F_t-\bm {\mathcal S}_0 \bm W_{\infty}^\prime\bm C_{t,\infty}\Vert =0$ for all $t\in\mathbb Z$; \end{compactenum} • if $n^{-1}\bm\Lambda^\prime\bm \Lambda=\mathbf I_r$ for all $n\in\mathbb N$ and $\bm F$ is unrestricted, then: \begin{compactenum} • $\bm\lambda_i^\prime= n \mathbf v_i^{C\prime} (\mathbf M^{C})^{-1/2}\mathbf K^\prime(\bm\Gamma^F)^{1/2}$ for all $i=1,\ldots, n$ and all $n\in\mathbb N$; • $\mathbf F_t= (\bm\Gamma^F)^{1/2}\mathbf K(\mathbf M^{C})^{-1/2}\mathbf V^{C\prime}\bm C_t$ for all $t\in\mathbb Z$ and all $n\in\mathbb N$; • $\lim_{n\to\infty} \left\Vert\bm\lambda_i^\prime - \bm p_{i0}^{\prime} \bm\Upsilon_1^\prime \right \Vert=0$ for all $i\in\mathbb N$; • $\text{\upshape m.s.-}\lim_{n\to\infty}\Vert\mathbf F_t-\bm \Upsilon_1 \bm W_{\infty}^\prime\bm C_{t,\infty}\Vert =0$ for all $t\in\mathbb Z$; \end{compactenum} • if $n^{-1}\bm\Lambda^\prime\bm \Lambda=\mathbf I_r$ and $T^{-1} \bm F^\prime\bm F$ is diagonal for all $n,T\in\mathbb N$, then: \begin{compactenum} • $\bm\lambda_i^\prime =\sqrt n \widehat{\mathbf v}_i^{C\prime} \widehat{\bm S}$ for all $i=1,\ldots, n$ and all $n,T\in\mathbb N$; • $\mathbf F_t=n^{-1/2}\widehat{\bm S} \widehat{\mathbf V}^{C\prime}\bm C_t$ for all $t=1,\ldots,T$ and all $n,T\in\mathbb N$; • $\text{\upshape P-}\lim_{T\to\infty}\Vert \bm\lambda_i^\prime -\sqrt n {\mathbf v}_i^{C\prime} {\bm S}\Vert=0$ for all $i=1,\ldots, n$ and all $n\in\mathbb N$; • $\text{\upshape P-}\lim_{T\to\infty}\Vert\mathbf F_t-n^{-1/2}{\bm S} {\mathbf V}^{C\prime}\bm C_t\Vert=0$ for all $t\in\mathbb N$ and all $n\in\mathbb N$; • $\text{\upshape P-}\lim_{n,T\to\infty}\Vert \bm\lambda_i^\prime-\bm p_{i0}^\prime \bm{\mathcal S}_0\Vert=0$ for all $i\in\mathbb N$; • $\text{\upshape P-}\lim_{n,T\to\infty}\Vert \mathbf F_t-\bm {\mathcal S}_0 \bm W_{\infty}^\prime\bm C_{t,\infty}\Vert=0$ for all $t\in\mathbb N$; \end{compactenum} • if $n^{-1}\bm\Lambda^\prime\bm \Lambda=\mathbf I_r$ for all $n\in\mathbb N$ and $\bm\Gamma^F$ is diagonal, then: \begin{compactenum} • $\bm\lambda_i^\prime = \sqrt n \mathbf v_i^{C\prime}\bm S$ for all $i=1,\ldots, n$ and all $n\in\mathbb N$; • $\mathbf F_t=n^{-1/2}\bm S \mathbf V^{C\prime}\bm C_t$ for all $t\in\mathbb Z$ and all $n\in\mathbb N$; • $\lim_{n\to\infty}\Vert \bm\lambda_i^\prime-\bm p_{i0}^\prime \bm{\mathcal S}_0\Vert=0$ for all $i\in\mathbb N$; • $\text{\upshape m.s.}\lim_{n\to\infty}\Vert \mathbf F_t-\bm {\mathcal S}_0 \bm W_{\infty}^\prime\bm C_{t,\infty}\Vert=0$ for all $t\in\mathbb Z$; \end{compactenum} \end{compactenum} and if we replace Assumption (ref)(c-ii) with Assumption (ref), or (ref), or (ref) in Appendix (ref), the limits in (II.b) and (V.b) hold also in mean-square. We used the following notation: the columns of $\mathbf V^C$ and $\widehat{\mathbf V}^C$ are the normalized eigenvectors of $\bm\Lambda\bm\Gamma^F\bm\Lambda$ (with eigenvalues $\mathbf M^C$) and of $\bm\Lambda(T^{-1}\bm F^\prime\bm F)\bm\Lambda$ (with eigenvalues $\widehat{\mathbf M}^C$), respectively, the columns of ${\mathbf K}$ and $\widehat{\mathbf Q}$ are the normalized eigenvectors of $(\bm\Gamma^F)^{1/2}(n^{-1}\bm \Lambda^\prime\bm \Lambda) (\bm\Gamma^F)^{1/2}$ (with eigenvalues $n^{-1}\mathbf M^C$) and of $(T^{-1}\bm F^\prime\bm F)^{1/2}(n^{-1}\bm \Lambda^\prime\bm \Lambda) (T^{-1}\bm F^\prime\bm F)^{1/2}$ (with eigenvalues $n^{-1}\widehat{\mathbf M}^C$), respectively, the columns of $\bm\Upsilon_0$ and $\bm\Upsilon_1$ are the normalized eigenvectors of $(\bm\Gamma^F)^{1/2}\bm \Sigma_\Lambda (\bm\Gamma^F)^{1/2}$ (with eigenvalues $\bm V_0$) and of $(\bm \Sigma_\Lambda)^{1/2}\bm\Gamma^F(\bm \Sigma_\Lambda)^{1/2}$ (with eigenvalues $\bm V_0$), respectively, $\bm S$ is an $r\times r$ diagonal matrix with entries $\pm 1$ depending only on $n$, $\widehat{\bm S}$ is an $r\times r$ diagonal matrix with entries $\pm 1$ depending on $n$ and $T$, $\bm{\mathcal S}_0$ is an $r\times r$ diagonal matrix with entries $\pm 1$ independent of $n$ and $T$; $\bm p_{i0}:=\lim_{n\to\infty}\sqrt n \mathbf v_i^C$, $i\in\mathbb N$, $\bm W_{\infty}:=\lim_{n\to\infty}n^{-1/2}\bm P_{n,\infty}$, with $\bm P_{n,\infty}:=(\bm p_{10}\cdots \bm p_{n0})^\prime$, and $\bm{C}_{t,\infty}:=\text{\upshape m.s.-}\lim_{n\to\infty} n^{-1/2}\bm C_t$, $t\in\mathbb Z$, such that, as $n\to\infty$, the following hold: $\Vert \bm p_{i0}\Vert = O(1)$, $n^{-1/2}\Vert\bm P_{n,\infty}\Vert = O(1)$, $\Vert\bm W_{\infty}\Vert=O(1)$, and $\Vert \bm C_{t,\infty}\Vert = O_{\mathrm{m.s.}}(1)$.

First note that, unless we constrain $T^{-1}\bm F^\prime\bm F$, $T$ plays no role, since in population the factors are a stochastic process so $t\in\mathbb Z$ and then we consider only cases with either $n$ fixed or $n\to\infty$ (parts (I.a), (II.a), (III.a), (I.b), (III.b), (IV.b), (VI.b)). If instead we constrain also $T^{-1}\bm F^\prime\bm F$, then we either consider the case of $n$ and $T$ fixed or $n,T\to\infty$ (parts (IV.a), (V.a), (VI.a), (II.b), (V.b)). Moreover, in three cases we can derive limits for fixed $n$ and $T\to\infty$ (parts (IV.a), (V.a), (V.b)).

Overall it is clear that $n$ and $T$ have no symmetric roles. This is apparent by looking in detail at the implications of the various conditions considered, which are also summarized in Table (ref).

compactenum• From parts (I.a), (IV.a), (I.b), and (IV.b) it is obvious that it is not enough to impose orthonormal factors or loadings to achieve identification. This is because the imposed conditions are only $r(r+1)/2$ so identification can be achieved only up to a rotation accounting for the remaining $r(r-1)/2$ degrees-of-freedom. Note that part (I.b) is the weakest since identification requires also $n\to\infty$. • Parts (II.a) and (VI.b) allow to identify the loadings and the factors up to a sign for any given fixed $n$ and regardless of $T$. In particular, the factors are identified as the population PCs of the common component. These are the classical identification conditions for time series. • Parts (V.a) and (V.b) allow to identify the loadings and the factors up to a sign for any given fixed $n$ and $T$. Provided we let $n,T\to\infty$, these have the same implications as parts (II.a) and (VI.b), respectively. But, given that they impose stronger conditions on the factors, they also imply that, for any fixed $n$ and $T$, the true loadings are random while the true factors are the sample PCs of the common component. However, these restrictions are not credible in time series. Indeed, if $\bm F$ is stochastic, then $\mathrm P(T^{-1}\bm F^\prime\bm F=\mathbf I_r)=0$, so for the restrictions to hold the factors have to be deterministic, implying that the common components, and so the loadings, are deterministic too. For these reasons parts (V.a) and (V.b), should never be considered in a time series context. • Parts (VI.a) and (II.b) not only require deterministic factors to hold, but do not even allow for identification unless we also require $n,T\to\infty$, so also these constraints should never be considered in a time series context. • Finally, parts (III.a) and (III.b) require $n\to\infty$ for achieving identification and do not depend on $T$. As shown below, in this case the factors are identified as the PCs of the infinite dimensional process of common components. They are the most realistic conditions which should be employed consistently with the fact that the approximate factor model is identified only when $n\to\infty$.

Following gersing_id, who firstly derived it, part (III.a) can be reformulated as follows. First, let $\bm\Lambda_{n,\infty}$ be the $n\times r$ matrix with rows $\bm\lambda_{i,\infty}^\prime:=\lim_{n\to\infty} \mathbf v_i^{C\prime}(\mathbf M^C)^{1/2}$. Then, by definition of $\bm p_{i0}$ and Lemma (ref)(i),

equation[equation omitted — 95 chars of source]

and part (III.a.i) is equivalent to $\lim_{n\to\infty}\Vert \bm\lambda_i^\prime -\bm\lambda_{i,\infty}^\prime \bm{\mathcal S}_0 \Vert=0.$ Note that the sequence $\{\bm\Lambda_{n,\infty},\, n\in\mathbb N\}$ is nested as $n$ grows since the rows of $\bm\Lambda_{n,\infty}$ do not depend on $n$.

Second, by definition of $\bm W_\infty$ and $\bm p_{i0}$,

align[align omitted — 343 chars of source]

Now let $\mathbf F_{t,\infty}:=\text{m.s.-}\lim_{n\to\infty}(\bm\Lambda_{n,\infty}^\prime \bm\Lambda_{n,\infty})^{-1}\bm\Lambda_{n,\infty}^\prime \bm C_t$. Then, since $\bm\Lambda_{n,\infty} = \bm P_{n,\infty} (\bm V_0)^{1/2}$ for all $n\in\mathbb N$, from (ref) it follows that $n^{-1}\bm\Lambda_{n,\infty}^\prime \bm\Lambda_{n,\infty}=\bm V_0$ for all $n\in\mathbb N$, and also

equation[equation omitted — 195 chars of source]

So part (III.a.ii) is equivalent to $\text{m.s.-}\lim_{n\to\infty}\Vert \mathbf F_t-\bm{\mathcal S}_0\mathbf F_{t,\infty}\Vert=0$. Note that $\mathbf F_{t,\infty}$ does not depend on $n$ and it is obtained as a weighted average of infinitely many common components. Specifically, the elements of $\mathbf F_{t,\infty}$ are the normalized PCs of the infinite dimensional vector $\bm C_{t,\infty}$ which is the mean-squared limit of the $n$-dimensional vector $n^{-1/2}\mathbf C_t$. Indeed, by definition of $\bm W_{\infty}$ and $\bm C_{t,\infty}$, and by Lemma (ref)(i),

align[align omitted — 663 chars of source]

and, because of (ref) and (ref), the columns of $\bm W_{\infty}$ are the normalized eigenvectors of the infinite dimensional matrix $\mathbb{E}[\bm C_{t,\infty}\bm C_{t,\infty}^\prime]=\lim_{n\to\infty} n^{-1} \bm\Gamma^C$ which is the covariance matrix of the infinite dimensional process $\bm C_{t,\infty}$. An analogous reasoning could be done for part (III.b), but in that case the factors would be identified with the non-normalized PCs of $\bm C_{t,\infty}$, having covariance matrix $\bm V_0$.

Finally, since the columns of $\mathbf V^C$ and $\bm\Lambda$ span the same space, then $\bm\Lambda=\mathbf V^{C}\mathbf V^{C\prime}\bm\Lambda$, so that $\bm\lambda_i^\prime=\mathbf v_i^{C\prime}\mathbf V^{C\prime}\bm\Lambda$ and $C_{it} = \bm\lambda_i^\prime \mathbf F_t= \mathbf v_i^{C\prime}\mathbf V^{C\prime} \bm\Lambda \mathbf F_t= \mathbf v_i^{C\prime}\mathbf V^{C\prime}\bm C_t$ for all $i=1,\ldots, n$ and all $n\in\mathbb N$. Therefore, since $C_{it}$ does not depend on $n$ and, by definition, $\bm P_{n,\infty}=\sqrt n\mathbf V^{C}$ for all $n\in\mathbb N$, from (ref) and (ref), we obtain

align[align omitted — 573 chars of source]

Hence, in general, for all $i\in\mathbb N$, there exists an invertible $r\times r$ matrix $\bm{\mathcal H}_{\infty}$, independent of $n$ and $T$, such that

equation[equation omitted — 207 chars of source]

From part (I.a), we see that, if we just impose orthonormal factors, then $\bm{\mathcal H}_{\infty}= \lim_{n\to\infty}\mathbf K=\bm\Upsilon_0$, where $\bm\Upsilon_0$ is a rotation, while from parts (II.a) or (III.a), we see that, if we also impose orthogonal loadings, then $\bm{\mathcal H}_{\infty}=\bm{\mathcal S}_0$. For more details we refer to gersing_id,gersing_lag.

sidewaystable[htbp] \caption{Identifying constraints and their implications} \scriptsize{ \begin{tabular}{l | ll | l | l | l | l} \hline \hline &&&&&&\\[-8pt] & \multicolumn{2}{c|}{Constraints}&$n,T$&$\bm\lambda_i^\prime$ &$\mathbf F_t$ & $\widehat{\mathbf H}$ \\[1pt] \hline &&&&&&\\[-8pt] I.a& $\bm\Gamma^F=\mathbf I_r$ & $\bm\Lambda$ unrestricted & $n$ fixed & $\mathbf v_i^{C\prime}(\mathbf M^{C})^{1/2}\mathbf K^\prime$ & $\mathbf K(\mathbf M^{C})^{-1/2} \mathbf V^{{C}\prime}\bm{C}_t$ & $ \widehat{\mathbf H}^\prime\widehat{\mathbf H}=\mathbf I_r+ O_{\mathrm P}\left(\frac 1n,\frac 1{\sqrt T}\right)$\\[3pt] &&& $n\to\infty$ & $ \bm p_{i0}^\prime (\bm V_0)^{1/2}\bm\Upsilon_0^\prime$ & ${\text{\tiny m.s.}} \bm\Upsilon_0(\bm V_0)^{-1/2}\bm W_{\infty}^\prime \bm C_{t,\infty}$ &\\[3pt] \hline &&&&&&\\[-8pt] II.a& $\bm\Gamma^F=\mathbf I_r$ & $\frac{\bm\Lambda^\prime\bm \Lambda}n$ diagonal & $n$ fixed & $\mathbf v_i^{C\prime}(\mathbf M^{C})^{1/2}\bm S$ & $\bm S(\mathbf M^{C})^{-1/2} \mathbf V^{{C}\prime}\bm{C}_t$ & $\widehat{\mathbf H}=\mathbf J+ O_{\mathrm P}\left(\frac 1n,\frac 1{\sqrt T}\right)$\\[3pt] &&& $n\to\infty$ & $ \bm p_{i0}^\prime (\bm V_0)^{1/2}\bm{\mathcal S}_0$ & ${\text{\tiny m.s.}} \bm{\mathcal S}_0(\bm V_0)^{-1/2}\bm W_{\infty}^\prime \bm C_{t,\infty}$ &\\[3pt] \hline &&&&&&\\[-8pt] III.a&$\bm\Gamma^F=\mathbf I_r$ & $\bm\Sigma_\Lambda$ diagonal & $n\to\infty$ & $ \bm p_{i0}^\prime (\bm V_0)^{1/2}\bm{\mathcal S}_0$ & ${\text{\tiny m.s.}} \bm{\mathcal S}_0(\bm V_0)^{-1/2}\bm W_{\infty}^\prime \bm C_{t,\infty}$ & $\widehat{\mathbf H}=\mathbf J+ O_{\mathrm P}\left(\frac 1{\sqrt n},\frac 1{\sqrt T}\right)$\\[3pt] \hline &&&&&&\\[-8pt] IV.a& $\frac{ \bm F^\prime\bm F}T=\mathbf I_r$ & $\bm\Lambda$ unrestricted & $n,T$ fixed & $\widehat{\mathbf v}_i^{C\prime}(\widehat{\mathbf M}^C)^{1/2} \widehat{\mathbf Q}^\prime$ & $\widehat{\mathbf Q}(\widehat{\mathbf M}^{C})^{-1/2} \widehat{\mathbf V}^{{C}\prime}\bm{C}_t$ & $\widehat{\mathbf H}^\prime\widehat{\mathbf H}=\mathbf I_r+ O_{\mathrm P}\left(\frac 1n,\frac 1{ T}\right)$\\[3pt] &&& $T\to\infty$ & $\mathbf v_i^{C\prime}(\mathbf M^{C})^{1/2}\mathbf K^\prime$ & $\text{\tiny m.s.} \mathbf K(\mathbf M^{C})^{-1/2} \mathbf V^{{C}\prime}\bm{C}_t$ &\\[3pt] &&& $n,T\to\infty$ & $\bm p_{i0}^\prime (\bm V_0)^{1/2}\bm\Upsilon_0^\prime$ & ${\text{\tiny m.s.}} \bm\Upsilon_0(\bm V_0)^{-1/2}\bm W_{\infty}^\prime \bm C_{t,\infty}$ &\\[3pt] \hline &&&&&&\\[-8pt] V.a& $\frac{ \bm F^\prime\bm F}T=\mathbf I_r$ & $\frac{\bm\Lambda^\prime\bm \Lambda}n$ diagonal & $n,T$ fixed & $\widehat{\mathbf v}_i^{C\prime}(\widehat{\mathbf M}^C)^{1/2} \widehat{\bm S}$ & $\widehat{\bm S}(\widehat{\mathbf M}^{C})^{-1/2} \widehat{\mathbf V}^{{C}\prime}\bm{C}_t$ & $\widehat{\mathbf H}=\mathbf J+ O_{\mathrm P}\left(\frac 1n,\frac 1{ T}\right)$\\[3pt] &&&$T\to\infty$ & $\mathbf v_i^{C\prime}(\mathbf M^{C})^{1/2}\bm S$ & $\text{\tiny m.s.} \bm S(\mathbf M^{C})^{-1/2} \mathbf V^{{C}\prime}\bm{C}_t$ &\\[3pt] &&& $n,T\to\infty$ & $ \bm p_{i0}^\prime (\bm V_0)^{1/2}\bm{\mathcal S}_0$ & ${\text{\tiny m.s.}} \bm{\mathcal S}_0(\bm V_0)^{-1/2}\bm W_{\infty}^\prime \bm C_{t,\infty}$ &\\[3pt] \hline &&&&&&\\[-8pt] VI.a& $\frac{ \bm F^\prime\bm F}T=\mathbf I_r$ & $\bm\Sigma_\Lambda$ diagonal & $n,T\to\infty$ & $\bm p_{i0}^\prime (\bm V_0)^{1/2}\bm{\mathcal S}_0$ & ${\text{\tiny m.s.}}\bm{\mathcal S}_0(\bm V_0)^{-1/2}\bm W_{\infty}^\prime \bm C_{t,\infty}$ & $\widehat{\mathbf H}=\mathbf J+ O_{\mathrm P}\left(\frac 1{\sqrt n},\frac 1{ T}\right)$\\[3pt] \hline \hline &&&&&&\\[-8pt] I.b& $\bm\Sigma_\Lambda=\mathbf I_r$ & $\bm F$ unrestricted & $n\to\infty$ & $\bm p_{i0}^{\prime} \bm\Upsilon_1^\prime$ & ${\text{\tiny m.s.}} \bm \Upsilon_1 \bm W_{\infty}^\prime\bm C_{t,\infty}$ & $\widehat{\mathbf H}^\prime\widehat{\mathbf H}=\frac{\widehat{\mathbf M}^x}{n}+ O_{\mathrm P}\left(\frac 1{\sqrt n},\frac 1{T}\right)$\\[3pt] \hline &&&&&\\[-8pt] II.b& $\bm\Sigma_\Lambda=\mathbf I_r$ & $\frac{ \bm F^\prime\bm F}T$ diagonal & $n,T\to\infty$ & $\text{\tiny m.s.} \bm p_{i0}^\prime \bm{\mathcal S}_0$ & $\text{\tiny m.s.} \bm {\mathcal S}_0 \bm W_{\infty}^\prime\bm C_{t,\infty}$ &$\widehat{\mathbf H}=\mathbf J\left(\frac{\widehat{\mathbf M}^x}n\right)^{1/2}+ O_{\mathrm P}\left(\frac 1{\sqrt n},\frac 1{T}\right)$\\[3pt] \hline &&&&&\\[-8pt] III.b& $\bm\Sigma_\Lambda=\mathbf I_r$ & $\bm\Gamma^F$ diagonal & $n\to\infty$ & $\bm p_{i0}^\prime \bm{\mathcal S}_0$ & $\text{\tiny m.s.} \bm {\mathcal S}_0 \bm W_{\infty}^\prime\bm C_{t,\infty}$ &$\widehat{\mathbf H}=\mathbf J\left(\frac{\widehat{\mathbf M}^x}n\right)^{1/2}+ O_{\mathrm P}\left(\frac 1{\sqrt n},\frac 1{\sqrt T}\right)$\\[3pt] \hline &&&&&\\[-8pt] IV.b& $\frac{\bm\Lambda^\prime\bm \Lambda}n=\mathbf I_r$ & $\bm F$ unrestricted & $n$ fixed & $n \mathbf v_i^{C\prime} (\mathbf M^{C})^{-1/2}\mathbf K^\prime(\bm\Gamma^F)^{1/2}$ & $ (\bm\Gamma^F)^{1/2}\mathbf K(\mathbf M^{C})^{-1/2}\mathbf V^{C\prime}\bm C_t$ & $\widehat{\mathbf H}^\prime\widehat{\mathbf H}=\frac{\widehat{\mathbf M}^x}{n}+ O_{\mathrm P}\left(\frac 1n,\frac 1{T}\right)$ \\[3pt] &&& $n\to\infty$ & $\bm p_{i0}^{\prime} \bm\Upsilon_1^\prime$ & $\text{\tiny m.s.}\bm \Upsilon_1 \bm W_{\infty}^\prime\bm C_{t,\infty}$ & \\[3pt] \hline &&&&&\\[-8pt] V.b& $\frac{\bm\Lambda^\prime\bm \Lambda}n=\mathbf I_r$ & $\frac{ \bm F^\prime\bm F}T$ diagonal & $n,T$ fixed & $\sqrt n \widehat{\mathbf v}_i^{C\prime} \widehat{\bm S}$ & $\frac 1{\sqrt n}\widehat{\bm S} \widehat{\mathbf V}^{C\prime}\bm C_t$ &$\widehat{\mathbf H}=\mathbf J\left(\frac{\widehat{\mathbf M}^x}n\right)^{1/2}+ O_{\mathrm P}\left(\frac 1{ n},\frac 1{ T}\right)$\\[3pt] &&& $T\to\infty$ & $\text{\tiny m.s.} \sqrt n {\mathbf v}_i^{C\prime} {\bm S}$ & $\text{\tiny m.s.} \frac 1{\sqrt n}{\bm S} {\mathbf V}^{C\prime}\bm C_t$ & \\[3pt] &&& $n,T\to\infty$ & $\text{\tiny m.s.} \bm p_{i0}^\prime \bm{\mathcal S}_0$ & $\text{\tiny m.s.} \bm {\mathcal S}_0 \bm W_{\infty}^\prime\bm C_{t,\infty}$ & \\[3pt] \hline &&&&&\\[-8pt] VI.b& $\frac{\bm\Lambda^\prime\bm \Lambda}n=\mathbf I_r$ & $\bm\Gamma^F$ diagonal & $n$ fixed & $\sqrt n \mathbf v_i^{C\prime}\bm S$ & $\frac 1{\sqrt n}\bm S \mathbf V^{C\prime}\bm C_t$ &$\widehat{\mathbf H}=\mathbf J\left(\frac{\widehat{\mathbf M}^x}n\right)^{1/2}+ O_{\mathrm P}\left(\frac 1{ n},\frac 1{\sqrt T}\right)$\\[3pt] &&& $n\to\infty$ & $ \bm p_{i0}^\prime \bm{\mathcal S}_0$ & $\text{\tiny m.s.} \bm {\mathcal S}_0 \bm W_{\infty}^\prime\bm C_{t,\infty}$ & \\[3pt] \hline \hline \end{tabular} } \begin{tabular}{p{.77\textwidth}} $\mathbf V^C$ normalized eigenvectors of $\bm\Lambda \bm\Gamma^F \bm\Lambda ^\prime$ with eigenvalues $\mathbf M^C$; $\mathbf K$ normalized eigenvectors of $(\bm\Gamma^F)^{1/2}(n^{-1}\bm \Lambda^\prime\bm \Lambda) (\bm\Gamma^F)^{1/2}$ with eigenvalues $n^{-1}\mathbf M^C$;\\ $\bm \Upsilon_0$ normalized eigenvectors of $(\bm\Gamma^F)^{1/2}\bm \Sigma_\Lambda (\bm\Gamma^F)^{1/2}$ with eigenvalues $\bm V_0$; $\bm\Upsilon_1$ normalized eigenvectors of $(\bm \Sigma_\Lambda)^{1/2}\bm\Gamma^F(\bm \Sigma_\Lambda)^{1/2}$ with eigenvalues $\bm V_0$;\\ $\widehat{\mathbf V}^C$ normalized eigenvectors of $\bm\Lambda(T^{-1}\bm F^\prime\bm F) \bm\Lambda ^\prime$ with eigenvalues $\widehat{\mathbf M}^C$; $\widehat{\mathbf Q}$ normalized eigenvectors of $(T^{-1}\bm F^\prime\bm F)^{1/2}(n^{-1}\bm \Lambda^\prime\bm \Lambda) (T^{-1}\bm F^\prime\bm F)^{1/2}$ with eigenvalues $\widehat{\mathbf M}^C$;\\ $\bm S$ is an $r\times r$ diagonal matrix with entries $\pm 1$ depending only on $n$; $\bm{\mathcal S}_0$ is an $r\times r$ diagonal matrix with entries $\pm 1$ independent of $n$ and $T$; $\widehat{\bm S}$ is an $r\times r$ diagonal matrix with entries $\pm 1$ depending on $n$ and $T$;\\ $\bm p_{i0}:=\lim_{n\to\infty}\sqrt n \mathbf v_i^C$, $i\in\mathbb N$; $\bm W_{\infty}:=\lim_{n\to\infty}n^{-1/2}\bm P_{n,\infty}$, with $\bm P_{n,\infty}:=(\bm p_{10}\cdots \bm p_{n0})^\prime$; $\bm{C}_{t,\infty}:=\text{\upshape m.s.-}\lim_{n\to\infty} n^{-1/2}\bm C_t$. \end{tabular}

\paragraph{Identification of $\widehat{\mathbf H}$ and $\widetilde{\mathbf H}$.} The following result states the behavior of $\widehat{\mathbf H}$ defined in (ref) and $\widetilde{\mathbf H}$ defined in (ref), when imposing further identifying constraints (see Appendix (ref) for a proof).

propUnder Assumptions (ref) through (ref), with Assumption (ref)(a) holding with rate $\sqrt n$ and Assumption (ref)(c-ii) holding with rate $\sqrt T$, as $n,T\to\infty$, \begin{compactenum} • if $\bm\Gamma^F=\mathbf I_r$ and $\bm\Lambda$ is unrestricted, $\min(n,\sqrt T) \Vert \widehat{\mathbf H}^\prime\widehat{\mathbf H}-\mathbf I_r\Vert=O_{\mathrm P}(1)$; • if $\bm\Gamma^F=\mathbf I_r$ and $n^{-1}\bm\Lambda^\prime\bm \Lambda$ is diagonal for all $n\in\mathbb N$, $\min(n,\sqrt T)\Vert\widehat{\mathbf H}-\mathbf J\Vert=O_{\mathrm P}(1)$; • if $\bm\Gamma^F=\mathbf I_r$ and $\bm\Sigma_\Lambda$ is diagonal, $\min(\sqrt n,\sqrt T)\Vert\widehat{\mathbf H}-\mathbf J\Vert=O_{\mathrm P}(1)$; • if $T^{-1}\bm F^\prime\bm F=\mathbf I_r$ for all $T\in\mathbb N$ and $\bm\Lambda$ is unrestricted, $\min(n, T)\Vert \widehat{\mathbf H}^\prime\widehat{\mathbf H} -\mathbf I_r \Vert = O_{\mathrm P}(1)$; • if $T^{-1} \bm F^\prime \bm F=\mathbf I_r$ and $n^{-1}\bm\Lambda^\prime\bm \Lambda$ is diagonal for all $n,T\in\mathbb N$, $\min(n, T)\Vert\widehat{\mathbf H}-\mathbf J\Vert=O_{\mathrm P}(1)$; • if $T^{-1} \bm F^\prime \bm F=\mathbf I_r$ for all $T\in\mathbb N$ and $\bm\Sigma_\Lambda$ is diagonal, $\min(\sqrt n, T)\Vert\widehat{\mathbf H}-\mathbf J\Vert=O_{\mathrm P}(1)$; • if $\bm\Sigma_\Lambda=\mathbf I_r$ and $\bm F$ is unrestricted, $\min(\sqrt n, T)\Vert \widehat{\mathbf H}^{^\prime}\widehat{\mathbf H} - n^{-1}\widehat{\mathbf M}^x \Vert = O_{\mathrm P}(1)$; • if $\bm\Sigma_\Lambda=\mathbf I_r$ and $T^{-1}\bm F^\prime \bm F$ is diagonal for all $T\in\mathbb N$, $\min(\sqrt n, T)\Vert\widehat{\mathbf H}-\mathbf J(n^{-1}\widehat{\mathbf M}^x)^{1/2}\Vert=O_{\mathrm P}(1)$; • if $\bm\Sigma_\Lambda=\mathbf I_r$ and $\bm\Gamma^F$ is diagonal, $\min(\sqrt n, \sqrt T)\Vert\widehat{\mathbf H}-\mathbf J(n^{-1}\widehat{\mathbf M}^x)^{1/2}\Vert=O_{\mathrm P}(1)$; • if $n^{-1}\bm\Lambda^\prime\bm \Lambda=\mathbf I_r$ for all $n\in\mathbb N$ and $\bm F$ is unrestricted, $\min(n,T)\Vert \widehat{\mathbf H}^{^\prime}\widehat{\mathbf H} - n^{-1}\widehat{\mathbf M}^x \Vert = O_{\mathrm P}(1)$; • if $n^{-1}\bm\Lambda^\prime\bm \Lambda=\mathbf I_r$ and $T^{-1} \bm F^\prime \bm F$ is diagonal for all $n,T\!\in\mathbb N$, $\min(n,T)\Vert\widehat{\mathbf H}-\mathbf J(n^{-1}\widehat{\mathbf M}^x)^{1/2}\Vert=O_{\mathrm P}(1)$; • if $n^{-1}\bm\Lambda^\prime\bm \Lambda=\mathbf I_r$ for all $n\in\mathbb N$ and $\bm\Gamma^F$ is diagonal, $\min(n,\sqrt T)\Vert\widehat{\mathbf H}-\mathbf J(n^{-1}\widehat{\mathbf M}^x)^{1/2}\Vert=O_{\mathrm P}(1)$; \end{compactenum} where $\mathbf J$ is an $r\times r$ diagonal matrix with entries $\pm 1$, depending on $n$ and $T$. Moreover, all statements hold also for $\widetilde{\mathbf H}$.

As an immediate consequence of Proposition (ref) notice that, by means of (ref), (ref), and (ref), we can derive also the properties of $\wideparen {\mathbf H}$ and $\bar{\mathbf H}$, which are the matrices determining identification under approaches A.2 and B.2. In particular, the results for $\widehat {\mathbf H}$ and $\widetilde{\mathbf H}$ under the constraints (a) in Proposition (ref) hold for $\wideparen {\mathbf H}$ and $\bar{\mathbf H}$ but under the constraints (b).

Although in some cases we impose non-asymptotic identifying conditions, we can derive only asymptotic results. Indeed, $\widehat{\mathbf H}$ depends on estimated quantities so its properties hold only asymptotically. The different convergence rates depend on the columns of the loadings being exactly or just asymptotically orthonormal.

Under parts (II.a), (III.a), (V.a) and (VI.a), we see that $\widehat{\mathbf H}$ can be reduced to an asymptotically diagonal matrix of signs $\mathbf J$. The sign indeterminacy, given by $\mathbf J$, can always be fixed, for example, as follows. Let $\iota_j$ be the first index of the $j$th column of $\bm\Lambda$ such that $\lambda_{\iota_j j}\ne 0$, $j=1,\ldots, r$. Then, we assume $\lambda_{\iota_j j}> 0$, for $j=1,\ldots,r$. This condition, which is also imposed, although implicitly, by Bai03, implies that $\mathbf J=\mathbf I_r$ in Proposition (ref).

Given that the assumptions made are the same as those under which Propositions (ref) and (ref) hold, the results in Proposition (ref) are directly applicable to the consistency results proved therein. Specifically, under all (a) cases the PC estimators are consistent for the true loadings and factors as defined in Proposition (ref). Indeed, by Propositions (ref) and (ref),

align[align omitted — 678 chars of source]

However, the term $\eta_{nT}$ which is given in Proposition (ref) depends on the imposed constraints. This has important implications for inference.

Under part (II.a) $\eta_{nT}=\max(n^{-1},T^{-1/2})$ this means that $ \sqrt T(\widehat{\mathbf H}^\prime-\mathbf J) = O_{\mathrm P}(1)$ and the CLT for the loadings in Theorem (ref) cannot hold, but $\sqrt n(\widehat{\mathbf H}^{-1}-\mathbf J)= O_{\mathrm P}(\sqrt {n/T})$, so the CLT for the factors in Theorem (ref) can still hold provided $n/T\to 0$, as $n,T\to\infty$, which is a stronger constraint. Under part (VI.a) $\eta_{nT}=\max(n^{-1/2},T^{-1})$ which implies that Theorem (ref) for the factors cannot hold but Theorem (ref) for the loadings can still hold provided $T/n\to 0$, as $n,T\to\infty$. Under part (III.a) $\eta_{nT}=\max(n^{-1/2},T^{-1/2})$ and neither Theorem (ref) nor Theorem (ref) can hold and we conjecture that if any asymptotic normality can be proved the asymptotic covariance will be larger than in those theorems.

Still, under (II.a), (III.a) or (VI.a) we can test for hypothesis of the form $\text H_0: \mathbf R \bm\lambda_i=\mathbf 0_q$ with $\mathbf R$ being $q\times r$ and full-rank for some $q\le r$. Indeed, in these cases, under $\text H_0$, by following (ref) we get:

align[align omitted — 180 chars of source]

because ${\text{P-lim}}_{n,T\to\infty} \widehat{\mathbf H}^\prime=\mathbf I_r$, and we can immediately build a Wald statistics for $\text H_0$. Similarly, we can build pointwise $(1-\alpha)$-confidence intervals for the factors given by $\{\widehat{F}_{jt}\pm z_{1-\alpha/2}{T^{-1/2} [\widehat{\bm \Pi}_t^{\text{\tiny OLS}}]_{jj}^{1/2}} \}$, $j=1,\ldots, r$.

However, to make inference for general hypothesis, we have to resort to the constraints in part (V.a), which is also considered by baing13. In this case, following again (ref) and (ref) and combining them with (ref) and (ref),

align[align omitted — 528 chars of source]

where used also the fact that $\Vert\bm\lambda_i\Vert = O(1)$, because of Assumption (ref)(a), and $\Vert \mathbf F_t\Vert = O_{\mathrm P}(1)$, because of Lemma (ref)(ii).

It is then clear that under condition (V.a), neither the location nor the covariance of the asymptotic distribution depend on $\widehat{\mathbf H}$, or its limit, anymore. Still, the price to be paid is not irrelevant. Indeed, condition (V.a) is stated for a sample quantity, and, thus, in order for it to hold, we should either consider the factors as a deterministic sequence or we should interpret the CLTs in (ref) and (ref) as conditional on a realization of the factors $\mathbf F_1,\ldots, \mathbf F_T$ (see also the comments made above and in baili12, in the case of maximum likelihood estimation). Under this interpretation, we can then impose condition (V.a) and use (ref) and (ref) to make inference not only on the loadings, but also on the factors since now $\mathbf F_t$ is no more random.

\paragraph{Identification of $\bm{\mathcal H}$.} To conclude, recall from Proposition (ref) that $\bm{\mathcal H}$ depends on $T$ only through a sign matrix $\mathbf J$. If we let $\widehat{\mathbf V}^{x}_j$ and $\mathbf V_j^{C}$, $j=1,\ldots, r$, be the columns of $\widehat{\mathbf V}^{x}$ and $\mathbf V^{C}$, respectively, we can assume, without loss of generality, that we always choose sample eigenvectors $\widehat{\mathbf V}^x_j$ such that $\widehat{\mathbf V}^{x\prime}_j \mathbf V_j^{C}\ge 0$ for all $j=1,\ldots, r$. Hereafter, we work under this condition and, therefore, $\mathbf J=\mathbf I_r$ and $\bm{\mathcal H}$ depends only on population quantities. We then derive the implications for $\bm{\mathcal H}$ as defined in Proposition (ref) (see Appendix (ref) for a proof).\footnote{If we did not set $\mathbf J=\mathbf I_r$, then $\bm S$ in Proposition (ref)(II.a) would depend also on $T$ while convergence in Proposition (ref)(III.a) would be in probability only, thus requiring also $T\to\infty$. Moreover, $\bm S$ and $\bm{\mathcal S}_0$ in Proposition (ref) would still be diagonal matrices with entries $\pm 1$, but, in general, they would be different from $\bm S$ and $\bm{\mathcal S}_0$ in Proposition (ref).}

propUnder Assumptions (ref) and (ref), and if $\mathbf J=\mathbf I_r$, \begin{compactenum} • if $\bm\Gamma^F=\mathbf I_r$ and $\bm\Lambda$ is unrestricted, then, for all $n\in\mathbb N$, $\bm{\mathcal H}^\prime \bm{\mathcal H}=\mathbf I_r$; • if $\bm\Gamma^F=\mathbf I_r$ and $n^{-1}\bm\Lambda^\prime\bm \Lambda$ is diagonal for all $n\in\mathbb N$, then, for all $n\in\mathbb N$, $\bm{\mathcal H}=\bm S$; • if $\bm\Gamma^F=\mathbf I_r$ and $\bm\Sigma_\Lambda$ is diagonal, then, $\lim_{n\to\infty}\Vert\bm{\mathcal H}-\bm{\mathcal S}_0\Vert=0$; \end{compactenum} where $\bm S$ is an $r\times r$ diagonal matrix with entries $\pm 1$ depending only on $n$, and $\bm{\mathcal S}_0$ is an $r\times r$ diagonal matrix with entries $\pm 1$ independent of $n$ and $T$.

Here, we consider only the case of orthonormal factors and orthogonal loadings, which is consistent with the PC estimators considered under approach A.1 or, equivalently, B.1. The case of orthonormal loadings and orthogonal factors, consistent with approaches A.2 and B.2, can be easily derived from the proof or Proposition (ref), hence it is omitted. Furthermore, for simplicity we also do not cover the cases in which we constraint $T^{-1}\bm F^\prime\bm F$.

Part (I.a) shows that if we restrict only the factors to be orthonormal, then $\bm{\mathcal H}$ is an orthogonal matrix. This is straightforward from its definition in Proposition (ref). Under parts (II.a) and (III.a) we see that $\bm{\mathcal H}$ can be reduced exactly or, at least, asymptotically to a diagonal matrix of signs. By comparing Propositions (ref) and (ref), from (ref), we also see that, $\bm{\mathcal H}_{\infty}=\lim_{n\to\infty}\bm{\mathcal H}$.

\singlespacing {{ {.2cm} }}

\setcounter{section}{0} \setcounter{subsection}{0} \setcounter{equation}{0} \gdef\thesection{ \Alph{section}} \gdef\thesubsection{\Alph{section}.\arabic{subsection}} \gdef\thefigure{\Alph{section}\arabic{figure}} \gdef\theequation{\Alph{section}\arabic{equation}} \gdef\varphible{\Alph{section}\arabic{table}} \gdef\arabic{footnote}{\Alph{section}\arabic{footnote}}