The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
174,933 characters
Principal Component Analysis .3cm for High-Dimensional Approximate Factor Models in Time Series: Assumptions, Asymptotic Theory, and Identification
\maketitle
\begin{center}\vspace{-1.5cm}
Matteo Barigozzi$^\dag$ \\[.1cm]
\small This version: \today
\end{center}
\begin{abstract}
We consider estimation of large approximate factor models in high-dimensional panels of stationary time series using Principal Component Analysis (PCA). We review the key results establishing the necessary and sufficient conditions for consistency and asymptotic normality of the estimators. We compare two equivalent approaches to PCA and present the asymptotic properties associated with each formulation. Special emphasis is placed on identification, where we discuss the restrictions required to uniquely determine factors and loadings and examine their consequences for statistical inference.
\vspace{0.5cm}
\noindent \textit{Keywords:}
Large Approximate Factor Model; Principal Component Analysis; Identification.
\end{abstract}
\thispagestyle{empty}
\footnotetext{Department of Economics - Universit\`a di Bologna. Email: [email removed] \\
I thank Haeran Cho and Philipp Gersing for helpful comments and discussions.}
\section{Introduction}
Consider a $T$-dimensional realization of $n$ zero-mean time series: $\{x_{it},\,i=1,\ldots, n,\, t=1,\ldots, T\}$. We say that $x_{it}$ follows an $r$-factor model if
\begin{align}
x_{it}&=\bm\lambda_i^\prime \mathbf F_t + e_{it}, \quad i=1,\ldots, n, \quad t=1,\ldots, T,\label{eq:SDFM1R_base}
\end{align}
where $\bm\lambda_i:=(\lambda_{i1}\cdots\lambda_{ir})^\prime$ and $\mathbf F_t:=(F_{1t}\cdots F_{rt})^\prime$ are the $r$-dimensional vectors of loadings for series $i$ and factors, respectively, and $r\ll \min(n,T)$. We call $ e_{it}$ the idiosyncratic component and
${C}_{it}:=\bm \lambda_i^\prime \mathbf F_t$ the common component.
While for small fixed $n$ it can be reasonably assumed that there is no idiosyncratic cross-sectional correlation and, thus, we say that the model is \textit{exact}, in a high-dimensional setting the idiosyncratic components are likely to be weakly cross-sectionally correlated, since, even if the common factors capture the main covariances, some weaker local covariances are likely to remain unexplained. In this case we say that the factor model is \textit{approximate}.
In an approximate factor model the common and idiosyncratic component can be identified, provided that the $r$ eigenvalues of the common component covariance matrix diverge as $n\to\infty$, while the $n$ idiosyncratic eigenvalues are bounded for all $n\in\mathbb N$ \citep{chamberlainrothschild83}. Hence, identification of the model, and then also of the number of factors $r$, is possible only when letting $n\to\infty$, a feature sometimes known as the \textit{blessing of dimensionality}. However, it is well known that, unless further $r^2$ restrictions are imposed, the loadings and factors are identified only up to post- and pre-multiplication by an invertible $r\times r$ matrix and its inverse.
An approximate factor model is characterized by an eigen-gap in the population covariance matrix that widens as $n \to \infty$. Conversely, if the population covariance exhibits a widening eigen-gap as $n \to \infty$, the data can be said to follow an approximate factor model---see, e.g., \citet[Theorem 4]{chamberlainrothschild83}, or \citet[Theorem 2]{gersing2023weak}, or \citet[Proposition 1]{BH25}, for various proofs. This is precisely the setting in which Principal Component Analysis (PCA) is most effective for dimensionality reduction, as it projects the data onto a subspace spanned by directions of maximum variation, by construction.
It is then of no surprise that PCA is the most natural and common way to estimate factor models. Furthermore, unlike classical Maximum Likelihood methods, PCA is a non-parametric approach that relies solely on moment conditions rather than distributional assumptions. Moreover, it can be applied to weakly stationary time series without any modification from its standard formulation for independent data.
Specifically, PCA is formulated as the minimization of a quadratic loss. This optimization yields estimated loadings proportional to the $r$ normalized eigenvectors corresponding to the $r$ largest eigenvalues of the $n\times n$ sample covariance matrix, whose $(i,j)$th entry is given by $T^{-1}\sum_{t=1}^T x_{it} x_{jt}$, $i,j=1,\ldots, n$. The estimated factors are obtained by linear projection of the data onto these estimated loadings.
However, to guarantee the existence of a unique solution, the minimizers of the PCA objective must satisfy an additional set of $r^2$ constraints. There are two standard ways of imposing these constraints,
each leading to estimated loadings that are equivalent up to a rescaling of the corresponding eigenvectors.
\begin{compactenum}
\item [A1.] If we impose the estimated loadings to be orthogonal and the estimated factors to be orthonormal, then the eigenvectors are rescaled by the square-root of the corresponding eigenvalues \citep[see, e.g.,][]{mardia1979multivariate,Bai03,FGLR09,FLM13}.
\item [A2.] If we impose the estimated loadings to be orthonormal and the estimated factors to be orthogonal, then the eigenvectors are rescaled by $\sqrt n$ \citep[see, e.g.,][]{stockwatson02JASA,stockwatson02JBES}.
\end{compactenum}
Definition A1 is the classical and most popular one and it is our main focus. Its theoretical properties, consistency and asymptotic normality, have been fully established by \citet{Bai03}, when formulating the estimators using an equivalent, although less intuitive, definition, where the estimated factors are written as $\sqrt T$ times the $r$ normalized eigenvectors corresponding to the $r$ largest eigenvalues of the $T\times T$ matrix with entries given by $n^{-1}\sum_{i=1}^n x_{it} x_{is}$, $t,s=1,\ldots, T$. The estimated loadings are obtained by linear projection of the data onto the estimated factors. We shall denote this definition as B1 (while a fourth definition, B2, is obtained analogously from A2).
Some important issues arise from the above overview. First, while we need $n\to\infty$ to identify the model, the sample size $T$ plays no role for identification of the model. However, we still need $T\to\infty$ for estimation. It follows that, on the one hand, we need to consider double asymptotics $n,T\to\infty$ to ensure consistency, but, on the other hand, it is clear that $n$ and $T$ do not have a symmetric role and should not be treated on equal grounds. This implies that carrying out asymptotic analysis using the classical definition A1 (or A2) rather than the alternative B1 (or B2) could have benefits in terms of interpretation. In fact, there are cases in which considering PCA based on the definition of approach A1 (or A2) is clearly the most natural way forward for developing the theory. First, in presence of non-stationarities in the data, due, e.g., to change-points \citep{BCF18}, exogenous \citep{massacci2017} or endogenous regime shifts \citep{urga2024estimation,BM22}, unit roots \citep{bai04,baing04,OW19,BLL2,barigozzi2022testing}, time-varying loadings \citep{motta2011locally,bates_plagborgmoller_stock_watson_2013,pelger_xiong_2022,mikkelsen_hillebrand_urga_2018}. Second, when dealing with PCA on the spectral density matrix \citet{FHLR00,FHLR05,FHLZ17}.
Third, when considering factor models for matrix and tensor data, which are estimated via PCA on every mode \citep{YuHe2022,chang2023modelling,chen2023statistical,chen2024rank,barigozzi2025tail,barigozzi2026tensor}.
Fourth, when dealing with factors which are pervasive along the time dimension so that the idiosyncratic components are white noise (\citealp{LamYao2012,ZRY,chen2020constrained}; \citealp{chen2022factor}). Fifth, when estimating exact factor models with $n$ is fixed so that only the loadings, but not the factors, can be estimated consistently via PCA \citep{anderson1963asymptotic}.
A second issue relates to identification. Since, in general, the true loadings and factors are not identified, it is not clear what the target of the estimator is when it comes to defining consistency. This means that no inference can be carried out, unless one can provide a clear theoretical study of the impact and the meaning of any imposed identifying assumptions. This topic is not much explored in the literature of high-dimensional factor models for time series, besides the works by \citet{baing13} and \citet{BW15}.
In this paper we provide a full treatment of PCA for factor models, by reviewing, discussing, completing, and extending the asymptotic existing results, and, thus, also clarifying the above issues. Specifically, we cover the following topics.\smallskip
\noindent
\textit{Derivation of PCA.} We show that the two standard definitions of PC estimators---those based on the classical $n \times n$ sample covariance matrix and their counterparts based on the $T \times T$ covariance matrix---are formally equivalent. Although elementary, this equivalence is often overlooked. Establishing it explicitly clarifies how the PC solution arises directly from the eigenvalue--eigenvector decomposition of these matrices, without resorting to iterative procedures such as alternating least squares. See Section~\ref{sec:ABCD} and Appendix \ref{sec:otherPCA}, where we derive the PC estimators in definition A2. \smallskip
\noindent
\textit{Assumptions.} We present the full set of necessary and sufficient assumptions needed to derive the asymptotic properties of PC estimators in a purely time-series framework. These assumptions are fewer in number and more transparent than those typically imposed in \citet{Bai03}. For each assumption---such as the various moment and dependence conditions---we provide intuitions and motivations to clarify their role in the analysis. See Section \ref{sec:ass} and Appendix \ref{sec:hannan}, where we review primitive conditions ensuring consistent estimation of the covariance matrix. \smallskip
\noindent
\textit{Asymptotic theory.} We derive the asymptotic properties of the classical PC estimator as computed under definition A1.
\begin{compactenum}[(i)]
\item We establish consistency under a minimal set of assumptions, using a direct argument based on Weyl's inequality \citep[see, e.g.,][Theorem 1]{MK04} and the Davis Kahan theorem \citep[see, e.g.,][Corollary 1]{yu15}. Orthogonality---i.e., uncorrelatedness---between common and idiosyncratic components suffices. See Section \ref{sec:cons1}.
\item We show that the sharper convergence rates of \citet[Theorems 1 and 2]{Bai03} can be obtained by imposing additional assumptions involving higher-order dependence. See Section \ref{sec:cons2}.
\item We provide several alternative representations for the linear transformation matrix linking true and estimated loadings/factors. See Section \ref{sec:Hhat}.
\item Under the extended assumptions, we prove asymptotic normality and show that, under full identification of loadings and factors, the leading term of the asymptotic expansion coincides with the Ordinary Least Squares (OLS) expansion.
See Section \ref{sec:AN}.
\item We further show that the same relationship between PC and OLS holds for the formulation B1 in \citet{Bai03}. See Section \ref{sec:cmppca} and Appendix \ref{sec:cmppcaCD}, where asymptotic normality is also established for the PC estimator in definitions A2 and B2. \smallskip
\end{compactenum}
\noindent
\textit{Identification.} We examine several identifying assumptions commonly imposed in exploratory factor analysis. These assumptions define the true loadings and factors as the population PCs of the common component and, in principle, should enable valid inference. However, we show that such restrictions either slow convergence---thereby affecting asymptotic normality---or are overly stringent, as they effectively would hold only if the factors were deterministic. See Section \ref{sec:II}.\smallskip
\noindent
\textit{Proofs.} In Appendix \ref{app:mainproof} we report the proofs of four main results---i.e., Propositions \ref{prop:L}, \ref{prop:HHAT}, \ref{corol:K00}, and \ref{corol:H00}---which are either simpler than the original ones or entirely new.
All other proofs are collected in Appendices \ref{sec:otherproofs} and \ref{sec:lemma}.\smallskip
\section{Estimation of factor models via Principal Component Analysis}\label{sec:ABCD}
In this section we present PCA estimation of a factor model. We assume to observe $n$ zero-mean time series over $T$ periods following the factor model, as in \eqref{eq:SDFM1R_base}, which in vector notation reads as
\[
\mathbf x_t = \bm\Lambda\mathbf F_t +\bm e_t,\quad t=1,\ldots, T,
\]
where $\mathbf x_t:=(x_{1t}\cdots x_{nt})^\prime$ and $\bm e_{t}:=( e_{1t}\cdots e_{nt})^\prime$ are $n$-dimensional vectors of observables and idiosyncratic components, respectively, and $\bm\Lambda:=(\bm\lambda_1\cdots\bm\lambda_n)^\prime$ is the $n\times r$ matrix of factor loadings. We call $\bm{C}_t:=\bm\Lambda\mathbf F_t$ the vector of common components.
Furthermore, let $\bm X:=(\mathbf x_1\cdots\mathbf x_T)^\prime$ and $\bm E:=(\bm e_1\cdots\bm e_T)^\prime$ be $T\times n$ matrices of observables and idiosyncratic components, respectively, and $\bm F:=(\mathbf F_1\cdots\mathbf F_T)^\prime$ be the $T\times r$ matrix of factors.
The idea of PCs as the $r$ directions of best fit in an $n$-dimensional space is originally due to \citet{pearson1901} and the use of PCA for factor analysis dates back to \citet{hotelling1933analysis}. The PC estimators of $\bm\Lambda$ and $\bm F$ must be such that they solve:
\begin{align}
&\min_{{\bm L},{\bm {G}}} \frac 1{nT}\sum_{i=1}^n\sum_{t=1}^T (x_{it}-
{\bm \ell}_i^\prime {\mathbf G}_t)^2=\min_{{\bm L},{\bm {G}}} \frac 1{nT}\left\Vert \bm X-
{\bm G}\bm L^\prime\right\Vert ^2_F\nonumber\\
&=
\min_{{\bm L},{\bm {G}}} \frac 1{nT} \text{tr}\left\{
\left(\bm X-
{\bm G}\,{\bm L}^\prime \right)
\left(\bm X-
{\bm G}\,{\bm L}^\prime \right)^\prime
\right\}=
\min_{{\bm L},{\bm {G}}} \frac 1{nT} \text{tr}\left\{
\left(\bm X-
{\bm G}\,{\bm L}^\prime \right)^\prime
\left(\bm X-
{\bm G}\,{\bm L}^\prime \right)
\right\},\label{eq:minimizza}
\end{align}
where $\bm L:=(\bm\ell_1\cdots\bm\ell_n)^\prime\in\mathbb R^{n\times r}$ and $\bm G:=(\mathbf G_1\cdots\mathbf G_T)^\prime\in\mathbb R^{T\times r}$ are generic loadings and factor matrices.
Clearly, for a given estimator of the loadings the estimator of the factors is obtained by least squares and viceversa. This motivates the classical way to derive the solution of \eqref{eq:minimizza} by using an alternating least squares approach \citep[and references therein]{bai2020simpler}. However, there is a simpler way to derive the solution. Indeed, we can restate \eqref{eq:minimizza} in two equivalent ways which depend either only on the loadings or only on the factors. First, by substituting ${\bm G}=\bm X{\bm L}({\bm L}^\prime{\bm L})^{-1}$ in \eqref{eq:minimizza}, we obtain the equivalent problem:
\begin{align}
\min_{{\bm L}}\frac 1{nT} \text{tr}\left\{
\bm X\left(\mathbf I_n-{\bm L}({\bm L}^\prime{\bm L})^{-1}{\bm L}^\prime \right)\bm X^\prime
\right\}
=\max_{{\bm L}}\frac 1{n} \text{tr}
\left\{
({\bm L}^\prime{\bm L})^{-1/2}
{\bm L}^\prime
\frac{\bm X^\prime\bm X}{T}
{\bm L}
({\bm L}^\prime{\bm L})^{-1/2}
\right\},\label{eq:maxforni}
\end{align}
which gives an estimator of the loadings, while the factors can be estimated in a second step by linear projection onto such estimator.
Second, by substituting ${\bm L}=\bm X^\prime {\bm G}({\bm G}^\prime{\bm G})^{-1}$ in \eqref{eq:minimizza}, we obtain yet another equivalent formulation of \eqref{eq:minimizza} or \eqref{eq:maxforni}:
\begin{align}
\min_{{\bm G}}\frac 1{nT} \text{tr}\left\{
\bm X^\prime\left(\mathbf I_T-{{\bm G}({\bm G}^\prime{\bm G})^{-1}{\bm G}^\prime} \right)\bm X
\right\}
=\max_{{\bm G}}\frac 1{T} \text{tr}
\left\{
({\bm G}^\prime{\bm G})^{-1/2}{{\bm G}^\prime}
\frac{\bm X\bm X^\prime}{n}
{{\bm G}}({\bm G}^\prime{\bm G})^{-1/2}
\right\},\label{eq:maxbai}
\end{align}
which gives an estimator of the factors, while the loadings can be estimated in a second step by linear projection onto such estimator.
From \eqref{eq:maxforni} and \eqref{eq:maxbai} we see that the PC estimators are related to eigenvectors, but the above solutions are not unique unless we fix the scale of those eigenvectors.
Indeed, the original problem in \eqref{eq:minimizza} does not have a unique solution too. For, given any solution $(\widehat{\bm\Lambda}, \widehat{\bm F})$, we can find infinite other equivalent solutions given by $\widehat{\bm\Lambda}{\bm K}$ and $\widehat{\bm F}({\bm K}^\prime)^{-1}$, where ${\bm K}$ is any $r\times r$ invertible matrix. Therefore, to ensure uniqueness of the PCA solution we must solve \eqref{eq:minimizza} when imposing $r^2$ identifying constraints. The two most common choices are: (a) $n^{-1}\bm L'\bm L$ being diagonal with distinct entries sorted in descending order and $T^{-1}\bm G'\bm G= \mathbf I_r$, or (b) $n^{-1}\bm L'\bm L=\mathbf I_r$ and $T^{-1}\bm G'\bm G$ being diagonal with distinct entries sorted in descending order.
Denote the classical the $n\times n$ sample covariance matrix obtained by averaging over time as (note that now $\mathbb{E}[\mathbf x_t]=\mathbf 0_n$ by assumption)
\begin{equation}\label{eq:nncov}
\widehat{\bm\Gamma}^x:=\frac 1T\sum_{t=1}^T \mathbf x_t\mathbf x_t^\prime = \frac{\bm X^\prime\bm X}{T},
\end{equation}
having its $r$ largest eigenvalues of $\widehat{\bm\Gamma}^x$ in the $r\times r$ diagonal matrix $\widehat{\mathbf M}^x$ (sorted in descending order) and the corresponding normalized eigenvectors as the columns of the $n\times r$ matrix $\widehat{\mathbf V}^x$.
Then, a solution of \eqref{eq:maxforni} is found as follows.
\paragraph{Approach A1.}
If we impose $n^{-1} {{\bm L}^\prime {\bm L}}$ to be diagonal, then by construction each column of ${\bm L}
({\bm L}^\prime{\bm L})^{-1/2}$ is normalized and, by definition of eigenvectors and eigenvalues, the value of the objective function in \eqref{eq:maxforni} must be the sum of the $r$ largest eigenvalues of $\widehat{\bm\Gamma}^x=T^{-1}{\bm X^\prime\bm X}$ divided by $n$, i.e., it must give $n^{-1}\text{tr}({\widehat{\mathbf M}^x})$. So, our estimator $\widehat{\bm\Lambda}$ must be such that $\widehat{\bm\Lambda}
(\widehat{\bm\Lambda}^\prime\widehat{\bm\Lambda})^{-1/2}$ is the matrix of normalized eigenvectors corresponding the $r$ largest eigenvalues of $({nT})^{-1}{\bm X^\prime\bm X}$, that is:
\begin{align}
&(\widehat{\bm\Lambda}^\prime\widehat{\bm\Lambda})^{-1/2}
\widehat{\bm\Lambda}^\prime
\frac{\bm X^\prime\bm X}{nT}
\widehat{\bm\Lambda}
(\widehat{\bm\Lambda}^\prime\widehat{\bm\Lambda})^{-1/2}=\frac{\widehat{\mathbf M}^x}{n}=\widehat{\mathbf V}^{x\prime}
\frac{\bm X^\prime\bm X}{nT}
\widehat{\mathbf V}^x.\label{eq:evecXX2}
\end{align}
Therefore, from \eqref{eq:evecXX2} we must have $\widehat{\bm\Lambda} = \widehat{\mathbf V}^x(\widehat{\bm\Lambda}^\prime\widehat{\bm\Lambda})^{1/2}$. So, by linear projection, the estimator of the factors is obtained as $\widehat{\bm F}= \bm X\widehat{\bm\Lambda}(\widehat{\bm\Lambda}^\prime\widehat{\bm\Lambda})^{-1}=\bm X \widehat{\mathbf V}^x(\widehat{\bm\Lambda}^\prime\widehat{\bm\Lambda})^{-1/2}$. Moreover, by imposing also the identifying constraint $T^{-1}\widehat{\bm F}^\prime\widehat{\bm F}=\mathbf I_r$, it follows that we must have
$\widehat{\bm\Lambda}^\prime\widehat{\bm\Lambda} = \widehat{\mathbf M}^x$, which is diagonal as expected.
From the above reasoning it follows that the PC estimators are given by:
\begin{align}
&\widehat{\bm\Lambda}:=\widehat{\mathbf V}^x(\widehat{\mathbf M}^x)^{1/2},\label{eq:estL}\\
&\widehat{\bm F}:=\bm X\widehat{\mathbf V}^x(\widehat{\mathbf M}^x)^{-1/2} = \bm X\widehat{\bm \Lambda}(\widehat{\mathbf M}^x)^{-1}.\label{eq:estF}
\end{align}
The factors defined in this way are the classical normalized PCs of $\bm X$ as in the definition of PCs dating back to \citet{hotelling1933analysis} (see also \citealp[Chapter 4]{lawleymaxwell71}, \citealp[Chapter 9.3]{mardia1979multivariate}, \citealp[Chapter 7.2]{jolliffe2002principal}). In high-dimensional factor analysis this approach is considered, for example, by \citet{FGLR09}, who prove consistency although not with the sharpest possible rate.
\\
Consider the $T\times T$ sample covariance matrix obtained by averaging over the cross-section (note that now $\mathbb{E}[\bm x_i]=\mathbf 0_T$ by assumption)
\begin{equation}\label{eq:TTcov}
\widetilde{\bm\Gamma}^x:=\frac 1n\sum_{i=1}^n \bm x_i\bm x_i^\prime = \frac{\bm X\bm X^\prime}{n},
\end{equation}
having its $r$ largest eigenvalues collected in the $r\times r$ diagonal matrix $\widetilde{\mathbf M}^x$ (sorted in descending order) and the corresponding normalized eigenvectors as columns of the $T\times r$ matrix $\widetilde{\mathbf V}^x$.
Then, another solution of \eqref{eq:maxbai}, equivalent to A1, is found as follows.
\paragraph{Approach B1.} If we impose $T^{-1}{{\bm G}^\prime{\bm G}}=\mathbf I_r$, then each column of $T^{-1/2}{ {\bm G}}$ is normalized and, by definition of eigenvectors and eigenvalues, the value of the objective function in the above maximization must be the sum of the $r$ largest eigenvalues of $\widetilde{\bm\Gamma}^x=n^{-1}{\bm X\bm X^\prime}$ divided by $T$, i.e., it must give $T^{-1}\text{tr}(\widetilde{\mathbf M}^x)$.
So, our estimator $\widetilde{\bm F}$ must be such that $T^{-1/2}{\widetilde{\bm F}}$ is the matrix of normalized eigenvectors corresponding the $r$ largest eigenvalues of $(nT)^{-1}{\bm X\bm X^\prime}$, that is:
\begin{align}
&\frac{\widetilde{\bm F}^\prime}{\sqrt T}
\frac{\bm X\bm X^\prime}{nT}
\frac{\widetilde{\bm F}}{\sqrt T}
=\frac{\widetilde{\mathbf M}^x}{T}=\widetilde{\mathbf V}^{x\prime}
\frac{\bm X\bm X^\prime}{nT}
\widetilde{\mathbf V}^x.\nonumber
\end{align}
Therefore, the PC estimators of the factors and the loadings (obtained by linear projection) are
\begin{align}
&\widetilde{\bm F}:=\widetilde{\mathbf V}^x\sqrt T,\label{eq:estFbai}\\
&\widetilde{\bm \Lambda}:=\frac{\bm X^\prime \widetilde{\mathbf V}^x}{\sqrt T} = \frac{\bm X^\prime \widetilde{\bm F}}{T},\label{eq:estLbai}
\end{align}
And, as expected, we get $n^{-1}{\widetilde{\bm \Lambda}^\prime\widetilde{\bm \Lambda}}=T^{-1}{\widetilde{\mathbf M}^x}$, which is diagonal. These are the estimators defined by \citet{Bai03}, who proves consistency and asymptotic normality.
\paragraph{Equivalence.} As long as the number of factors $r$ is such that $r<\min(n,T)$, the $r$ largest eigenvalues of
$\bm X^\prime \bm X$ and $\bm X \bm X^\prime$ coincide, i.e.,
\begin{equation}\label{eq:evaleq}
\frac{\widehat{\mathbf M}^x}n=\frac{\widetilde{\mathbf M}^x}T.
\end{equation}
and the objective function in \eqref{eq:maxforni} and \eqref{eq:maxbai} has the same value in its minimum. Since approaches A1 and B1 are based on the same identification conditions, the solutions \eqref{eq:estL}-\eqref{eq:estF} and \eqref{eq:estFbai}-\eqref{eq:estLbai} must then coincide, i.e., it must be that
\begin{equation}\label{eq:stessi}
\widehat{\bm F}=\widetilde{\bm F} \; \text{ and }\; \widehat{\bm \Lambda}=\widetilde{\bm \Lambda}.
\end{equation}
This is also seen in the following way.
Consider the Singular Value Decomposition
\begin{equation}
\frac{\bm X}{\sqrt{nT}} = \bm U \bm D\bm V^\prime + \bm E, \label{eq:SVD}
\end{equation}
where $ \bm U$ and $ \bm V$ are orthogonal matrices of dimensions $T\times r$ and $n\times r$, respectively, and $\bm D$ contains the $r$ largest singular values of $\bm X$.
Now,
\[
\frac{\widehat{\bm\Gamma}^x}n=\frac{\bm X^\prime \bm X}{nT} = \bm V \bm D^2 \bm V^\prime + \bm E^\prime\bm E\;\text{ and }\;
\frac{\widetilde{\bm\Gamma}^x}T=\frac{\bm X \bm X^\prime}{nT} = \bm U \bm D^2 \bm U^\prime + \bm E\bm E ^\prime
\]
and since the best rank-$r$ approximation of $(nT)^{-1/2}\bm X$ is $ \bm U \bm D\bm V^\prime$ \citep{EY36}, it is straightforward to see that (using also \eqref{eq:evaleq})
\begin{equation}\label{eq:giovine}
\bm D^2 =\frac{\widehat{\mathbf M}^x}n= \frac{\widetilde{\mathbf M}^x}T,\quad \bm U=\widetilde{\mathbf V}^x, \quad \bm V=\widehat{\mathbf V}^x.
\end{equation}
Moreover, under our assumptions on the eigenvalues of the common and idiosyncratic components (see \eqref{eq:lindiv} and \eqref{eq:evalidio} below), it holds that $(nT)^{-1/2}\bm X=(nT)^{-1/2} \bm F\bm\Lambda^\prime+o_{\mathrm P}(1)$. Then, the PC estimators must be such that $ \widehat{\bm F}\widehat{\bm\Lambda}^\prime=\widetilde{\bm F}\widetilde{\bm\Lambda}^\prime= \sqrt{nT} \,\bm U \bm D\bm V^\prime$. By solving \eqref{eq:minimizza} using an alternating least squares approach jointly with \eqref{eq:SVD} and by imposing orthogonal loadings and orthonormal factors, \citet{bai2020simpler} show that the estimator of the loadings is $\check{\bm\Lambda}= \sqrt n\, \bm V \bm D$, while the estimator of the factors is $\check{\bm F}\,=\sqrt T\, \bm U$. And from \eqref{eq:giovine} we see that $\check{\bm\Lambda}=\widehat{ \mathbf V}^{x}(\widehat{\mathbf M}^{x})^{1/2}=\widehat{\bm\Lambda}$ as in approach A1, and
$\check{\bm F}=\sqrt T\, \widetilde{\mathbf V}^x=\widetilde{\bm F}$ as in approach B1. Moreover, by linear projection and again from \eqref{eq:giovine},
\begin{align}
\widehat {\bm F} &= \bm X\widehat{\bm\Lambda}(\widehat{\bm\Lambda}^\prime\widehat{\bm\Lambda})^{-1}=
\bm X \widehat{ \mathbf V}^{x}(\widehat{\mathbf M}^{x})^{-1/2}=
\sqrt {nT}\, \bm U\bm D\bm V^\prime \bm V ( \sqrt n \bm D)^{-1} = \sqrt T \bm U = \widetilde{\bm F},\nonumber\\
\widetilde{\bm\Lambda} &=\bm X^\prime \widetilde{\bm F}(\widetilde{\bm F}^\prime\widetilde{\bm F})^{-1} =\frac{\bm X^\prime \widetilde{\mathbf V}^{x} }{\sqrt T}=\frac{\sqrt{nT}\,\bm V\bm D\bm U^\prime \bm U}{\sqrt T}
=\sqrt n \bm V\bm D= \widehat{\bm\Lambda}.\nonumber
\end{align}
Hence, as stated in \eqref{eq:stessi}, approaches A1 and B1 coincide.
Alternative estimators are obtained when imposing orthonormal loadings and orthogonal factors. These are presented in approaches A2 and B2 in Appendix \ref{sec:otherPCA}.
\paragraph{Centering and standardizing.}
We conclude by noticing that if the observed time series do not have zero mean, then, the above derivation would be the same provided we replaced each element of $\bm X$ with its centered version $x_{it}- \bar{x}_i$, $i=1,\ldots, n$, $t=1,\ldots T$, with $
\bar {x}_i:=T^{-1}\sum_{t=1}^T x_{it}$. Since it is commonly assumed that the factors and the idiosyncratic components have zero mean, i.e., $\mathbb{E}[\mathbf F_t]=\mathbf 0_r$ and $\mathbb{E}[\bm e_{t}]=\mathbf 0_n$, then, when the data has not zero-mean, model \eqref{eq:SDFM1R_base} is still valid provided we add a constant, $\alpha_i$, $i=1,\ldots, n$. Clearly, $\bar{x}_i$ is a consistent estimator of $\alpha_i$.
Moreover, to avoid scale effects due to data having different units of measure, each element of $\bm X$ should always be standardized so that, in practice, the above derivation would be applied to $\frac{x_{it}-\bar x_i}{s_i}$, $i=1,\ldots, n$, $t=1,\ldots T$, with $s_i:=T^{-1}\sum_{t=1}^T (x_{it}-\bar x_i)^2$.
Finally, note that both centering and standardizing must always be carried out along the time dimension even when working with a $T\times T$ covariance. This is because the cross-sectional units are unlikely to be identically distributed.
\section{Assumptions}\label{sec:ass}
We adapt to the high-dimensional setting the approach taken in most works in time series analysis. First, we define the model for an infinite dimensional process $\{x_{it},\, i\in\mathbb N,\, t\in\mathbb Z\}$. Second, the properties and the assumptions of the model are defined for the $n$-dimensional sub-process $\{x_{it},\, i=1\ldots, n,\, t\in\mathbb Z\}$ either holding for all $n\in\mathbb N$ or in the limit $n\to\infty$. Third, the properties of the estimators are derived for a given $nT$-dimensional realization $\{x_{it},\, i=1\ldots, n,\, t=1,\ldots, T\}$ and hold in the limit $n,T\to\infty$.
The infinite dimensional process follows an $r$-factor model if
\begin{align}
x_{it}&=\bm\lambda_i^\prime \mathbf F_t + e_{it}, \quad i\in\mathbb N, \quad t\in\mathbb Z.\label{eq:SDFM1R}
\end{align}
This is the equivalent of \eqref{eq:SDFM1R_base}, but when we have infinite time series, so now $i\in\mathbb N$, each having zero-mean for simplicity. The common component is defined as $\bm{C}_t:=\bm\Lambda\mathbf F_t$.
We characterize the common component by means of the following assumption.
\begin{ass}[\textsc{common component}]\label{ass:common} $\,$
\begin{compactenum}[(a)]
\item $\lim_{n\to\infty}\Vert n^{-1}\bm\Lambda^\prime\bm\Lambda-\bm\Sigma_{\Lambda}\Vert=0$, where $\bm\Sigma_{\Lambda}$ is $r\times r$ positive definite, and, for all $i\in\mathbb N$,
$\Vert\bm\lambda_i\Vert\le M_\Lambda$ for some finite positive real $M_\Lambda$ independent of $i$.
\item For all $t\in\mathbb Z$, $\mathbb{E}[\mathbf F_{t}]=\mathbf 0_r$ and $\bm\Gamma^F:=\mathbb{E}[\mathbf F_t\mathbf F_t^\prime]$ is $r\times r$ positive definite and $\Vert\bm\Gamma^F\Vert\le M_F$ for some finite positive real $M_F$ independent of $t$.
\item
\begin{inparaenum}
\item [(i)] For all $t\in\mathbb Z$, $\mathbb{E}[\Vert \mathbf F_t\Vert^4]\le K_F$ for some finite positive real $K_F$ independent of $t$; \\
\item [(ii)] $\mathrm P\text{-}\lim_{T\to\infty}\left\Vert T^{-1}\bm F^\prime\bm F -\bm\Gamma^F\right\Vert=0$.
\end{inparaenum}
\item There exists an integer $N$ such that for all $n> N$, $r$ is a finite positive integer, independent of $n$.
\end{compactenum}
\end{ass}
Parts (a) and (b), are also assumed in \citet[Assumptions A and B]{Bai03}. They imply that the loadings matrix has asymptotically maximum column rank $r$ (part (a)) and the factors have a finite full-rank covariance matrix (part (b)). Moreover, because of part (a), for any given $n\in\mathbb N$, all the factors have a finite contribution to each series (upper bound on $\Vert\bm\lambda_i\Vert$) and are cross-sectionally pervasive (full-rank of $\bm\Sigma_\Lambda$).
In Lemma \ref{lem:Gxi}(iv) we prove that parts (a) and (b) imply that the covariance matrix of the common component, $\bm\Gamma^{C}:=\mathbb{E}[\bm{C}_t\bm{C}_t^\prime]=\bm\Lambda\bm\Gamma^F \bm\Lambda^\prime$, has rank $r$ and its $j$th largest eigenvalue $\mu_j^{C}$, $j=1,\ldots, r$, is such that
\begin{equation}\label{eq:lindiv}
\underline C_j\!\le \lim_{n\to\infty} \frac{\mu_{j}^{C}}n\le\! \overline C_j,
\end{equation}
for some finite positive reals $\underline C_j$ and $\overline C_j$ independent of $n$. This means we consider only pervasive strong factors. Any weaker rate of divergence would imply a specific ordering of the cross-sectional units \citep{BH25}. Although this might be relevant in some datasets, it is not the subject of this work. In absence of any specific ordering of the cross-sectional items, the reasonable convergence rate to assume in part (a) is then $\sqrt n$, indicating some form of ``cross-sectional'' stationarity.
Part (c) is assumed also in \citet[Assumption A]{Bai03}. In part (c-i) we assume finite 4th order moments of the factors, and, in particular, we are bounding
$\mathbb{E}[\Vert \mathbf F_t\Vert^4]=\sum_{j=1}^r\sum_{k=1}^r \mathbb{E}[F_{jt}^2F_{kt}^2]$.
Part (c-ii) is very general, it simply says that the sample covariance matrix of the factors is a consistent estimator of its population counterpart $\bm\Gamma^F$. In principle, we do not have to require weak stationarity for part (c-ii) to hold, and we can have $\bm\Gamma^F$ that depends on $t$ as long as $M_F$ in part (b) is independent of $t$ (hence the reason for stating it explicitly). This is, for example the case when we are in presence of structural breaks/change-points (see, e.g., \citealp{BCF18} and \citealp{duan2022quasi}) or regime shifts (\citealp{massacci2017,BM22}).\footnote{Unit roots are however excluded since in that case $\{\mathbf F_t\}$ does not have a finite covariance, and everything that is considered in this paper would hold for properly differenced data.} This assumption is a high-level one and it requires introducing the sample size $T$. In Appendix \ref{sec:hannan} we show that if the process $\{\mathbf F_t\}$ is either linear, or has summable 4th order cumulants, or has bounded physical dependence \citep{wu05}, or is strongly mixing, then part (c-ii) holds, possibly even in mean-square. There it is also shown that the convergence rate is $\sqrt T$, as it should be expected.
Notice that, on the one hand, we only consider non-random factor loadings for simplicity and by analogy with classical factor models.\footnote{As in \citet{Bai03}, the generalization to the case of random loadings is possible provided that: we assume convergence in probability in part (a), for all $i\in\mathbb N$, $\mathbb{E}[\Vert \bm\lambda_i\Vert^4]\le K_\Lambda$, for some finite positive real $K_\Lambda$ independent of $i$, and $\{\bm\lambda_i,\, i\in\mathbb N\}$ is an independent sequence.} But, on the other hand, we explicitly treat the factors as random variables. This has important implications when it comes to understanding the implications of various identifying assumptions (see Section \ref{sec:II}).
Part (d) implies the existence of a finite number of factors. In particular, the number of common factors, $r$, is identified only for $n\to\infty$. Here $N$ is the minimum number of series we need to be able to identify $r$ so that $r\le N$. Without loss of generality hereafter when we say ``for all $n\in\mathbb N$'' we always mean that $n>N$ so that $r$ can be identified. In practice, we must always work with $n$ such that $r<n$. Moreover, because PCA is based on eigenvalues of a matrix $n\times n$ estimated using $T$ observations then we must also have samples of size $T$ such that $r<T$.
Therefore, sometimes it is directly assumed that $r<\min(n,T)$.
To characterize the idiosyncratic component, we make the following assumptions.
\begin{ass}[\textsc{idiosyncratic component}]\label{ass:idio}
$\,$
\begin{compactenum}[(a)]
\item For all $i,j\in\mathbb N$, all $t\in\mathbb Z$, and all $k\in\mathbb Z$, $\mathbb{E}[ e_{it}]= 0$, $\gamma^ e_{ij,k}:=\mathbb{E}[ e_{it} e_{j,t-k}]$ with $\gamma_{ii,0}^ e:=\mathbb{E}[ e_{it}^2]$.
\item For all $i,j\in\mathbb N$, all $t\in\mathbb Z$, and all $k\in\mathbb Z$, $\gamma^ e_{ij,k}:=\mathbb{E}[ e_{it} e_{j,t-k}]$ such that
$\vert \gamma^ e_{ij,k}\vert\le \rho^{\vert k\vert} M_{ij}$, where $\rho$ and $M_{ij}$ are finite positive reals independent of $t$ such that $0\le \rho <1$, $M_{ii}\le C_ e$, $\sum_{j=1,i\ne j}^n M_{ij}\le M_ e$, and $\sum_{i=1, j\ne i}^n M_{ij}\le M_ e$ for some finite positive reals $C_ e$ and $M_{ e}$ independent of $i$, $j$, and $n$.
\item
\begin{inparaenum}
\item [(i)] For all $i=1,\ldots, n$, all $t=1,\ldots, T$, and all $n,T\in\mathbb N$, $\mathbb{E}[ e_{it}^4]\le Q_ e$ for some finite positive real $\phantom {(i)}\, Q_ e$ independent of $i$ and $t$;\\
\item [(ii)] for all $j=1,\ldots, n$, all $s=1,\ldots, T$, and all $n,T\in\mathbb N$,
\[
\mathbb{E}\left[\left\vert\frac 1{\sqrt{nT}} \sum_{i=1}^n\sum_{t=1}^T\left\{ e_{is} e_{jt}-\mathbb{E}[ e_{is} e_{jt}]\right\} \right\vert^2\right]
\le K_ e
\]
for some finite positive real $K_ e$ independent of $j$, $s$, $n$, and $T$.
\end{inparaenum}
\end{compactenum}
\end{ass}
By part (a), we assume that the idiosyncratic components are cross-sectionally heteroskedastic but serially homoskedastic. It follows that for any $n\in\mathbb N$ the idiosyncratic covariance matrix is defined as $\bm\Gamma^ e:=\mathbb{E}[\bm e_t\bm e_t^\prime]$ with entries $\gamma_{ij,0}^ e$, $i,j=1,\ldots,n$.
Part (b) has a twofold purpose. First, it limits the degree of serial correlation of the idiosyncratic components, implying weak stationarity of idiosyncratic components. Indeed, they have summable autocovariances (see Lemma \ref{lem:Gxi}(iii)):
\[
\sup_{n,T\in\mathbb N}\max_{i=1,\ldots, n}\frac 1{T}\sum_{t,s=1}^T \vert\mathbb{E}_{}[ e_{it} e_{is}]\vert \le \frac{C_ e(1+\rho)}{1-\rho}.
\]
The same condition follows directly from \citet[Assumption C.4]{Bai03} although it is not stated explicitly.
Second, part (b) also limits the degree of cross-sectional correlation between idiosyncratic components, which is usually assumed in approximate factor models. Moreover, it implies the usual conditions for PC estimation (see Lemmas \ref{lem:Gxi}(i), \ref{lem:Gxi}(ii), \ref{lem:GxiBAI}(i), \ref{lem:GxiBAI}(ii)):
\begin{align}
&\sup_{n,T\in\mathbb N}\frac 1{nT}\sum_{i,j=1}^n\sum_{t,s=1}^T \vert\mathbb{E}_{}[ e_{it} e_{js}]\vert \le \frac{(C_ e+M_ e)(1+\rho)}{1-\rho},\label{eq:bai}\\
&\sup_{n,T\in\mathbb N}\max_{t=1,\ldots, T}\frac 1{n}\sum_{i,j=1}^n \vert \mathbb{E}_{}[ e_{it} e_{jt}] \vert \le C_ e+M_{ e},
\label{eq:baiC3}\\
&\sup_{n,T\in\mathbb N} \frac 1{T}\sum_{t,s=1}^T \left \vert\frac 1n \sum_{i=1}^n \mathbb{E}_{}[ e_{it} e_{is}]\right \vert \le\frac{C_ e (1+\rho)}{1-\rho}, \label{eq:baiC2}\\
&\sup_{n,T\in\mathbb N} \max_{t=1,\ldots, T}\sum_{s=1}^T \left \vert\frac 1n \sum_{i=1}^n \mathbb{E}_{}[ e_{it} e_{is}]\right\vert \le\frac{C_ e (1+\rho-\rho^t)}{1-\rho}, \label{eq:baiE1}\\
&\sup_{n,T\in\mathbb N} \max_{\substack{i=1,\ldots, n\\t=1,\ldots, T}} \sum_{j=1}^n \vert \mathbb{E}[ e_{it} e_{jt}] \vert\le C_ e+M_ e, \label{eq:baiE2}
\end{align}
which coincide with the conditions required by \citet[Assumptions C.4, C.3, C.2, E.1, and E.2, respectively]{Bai03}.
Furthermore, part (b) directly implies also that
\begin{align}
&\sup_{n,T\in\mathbb N} \max_{s=1,\ldots,T} \left\vert \frac 1n \sum_{i=1}^n \mathbb{E}[ e_{is}^2]\right\vert \le C_ e,\label{eq:baiC2bis}\\
&\sup_{T\in\mathbb N}
\max_{t=1,\ldots, T} \vert \mathbb{E}_{}[ e_{it} e_{jt}] \vert \le M_{ij},\label{eq:baiC3bis}
\end{align}
where \eqref{eq:baiC2bis} is the first condition in \citet[Assumption C.2]{Bai03},
and \eqref{eq:baiC3bis} is the first condition in \citet[Assumption C.3]{Bai03}.
Notice that the formulation of part (b) is inspired by \citet{FHLZ17} and it implicitly bounds the L1 norm of $\bm\Gamma^ e$ as in \citet{FLM13}.
Furthermore, in Lemma \ref{lem:Gxi}(v) we prove that the largest eigenvalue of $\bm\Gamma^ e$ is such that (this is an instance of the Gershgorin circle theorem)
\begin{equation}\label{eq:evalidio}
\sup_{n\in\mathbb N}\mu_{1}^ e \le C_ e+M_ e,
\end{equation}
where $C_ e$ and $M_ e$ are defined in Assumption \ref{ass:idio}(b). Condition \eqref{eq:evalidio} was originally assumed by \citet{chamberlainrothschild83}. This is the essence of an approximate factor model as opposed to an exact factor model where it is imposed the restrictive assumption of $\bm\Gamma^ e$ being diagonal.
Part (c-i) assumes finite and summable 4th order moments of the idiosyncratic components, it is weaker than what assumed by \citet[Assumption C.1]{Bai03} where finite 8th order moments are required. Part (c-ii) gives summability conditions over the cross-section and time dimensions for the 4th order cumulants of $\{ e_{it}\}$. Indeed, it can be equivalently written as:
\begin{align}
\sup_{n,T\in\mathbb N}\max_{\substack{j=1,\ldots, n\\s=1,\ldots, T}}
\mathbb{E}&\left[\left\vert\frac 1{\sqrt{nT}} \sum_{i=1}^n\sum_{t=1}^T\left\{ e_{is} e_{jt}-\mathbb{E}[ e_{is} e_{jt}]\right\} \right\vert^2\right]\nonumber\\
&=\sup_{n,T\in\mathbb N}\max_{\substack{j=1,\ldots, n\\s=1,\ldots, T}}\frac 1{nT}\sum_{i,k=1}^n\sum_{t,u=1}^T \left\{\mathbb{E}\left[ e_{is} e_{jt} e_{ks} e_{ju}\right]-\mathbb{E}[ e_{is} e_{jt}]\mathbb{E}[ e_{ks} e_{ju}]\right\}\le K_ e.\label{eq:bordellobai}
\end{align}
If we set $j=i$ in \eqref{eq:bordellobai} we get
a similar condition to what is assumed by \citet[Assumption C.5]{Bai03}, where, however, $T=1$, and the sums of 8th cross-cumulants are bounded.
Part (c-ii) has two more important consequences. First, it implies that (see Lemma \ref{lem:LLN}(iii)):
\begin{align}
&\sup_{n,T\in\mathbb N}\mathbb{E}\left[\left\Vert\frac 1{n\sqrt {T}}\sum_{t=1}^T\left\{ \bm\Lambda^\prime\bm e_t\bm e_t^\prime-\bm\Lambda^\prime\mathbb{E}[\bm e_t\bm e_t^\prime]\right\}\right\Vert_F^2\right] \le r M_\Lambda K_ e.\label{eq:frascati24}
\end{align}
This result is crucial to prove asymptotic normality of the estimated loadings. In particular, without \eqref{eq:frascati24} consistency of the loadings would hold but with a slower rate, $\min(\sqrt n,\sqrt T)$ which is slower than $\min(n,\sqrt T)$ derived in Proposition \ref{prop:L2}. Such a slower rate is what prevents the proof of asymptotic normality of the estimated loadings to go through.
Second, part (c-ii) implies also that (setting $n=1$ therein):
\begin{align}
\sup_{n,T\in\mathbb N}
\max_{\substack{i,j=1,\ldots,n\\s=1,\ldots, T}}
&\mathbb{E}\left[\left\vert\frac 1{\sqrt {T}} \sum_{t=1}^T\left\{ e_{is} e_{jt}-\mathbb{E}[ e_{is} e_{jt}]\right\} \right\vert^2\right]\nonumber\\
&=\sup_{n,T\in\mathbb N}
\max_{\substack{i,j=1,\ldots,n\\s=1,\ldots, T}}\frac 1{T}\sum_{t=1}^T\sum_{u=1}^T\left\{\mathbb{E}[ e_{is} e_{jt} e_{is} e_{ju}]-\mathbb{E}[ e_{is} e_{jt}]\mathbb{E}[ e_{is} e_{ju}]
\right\}\le {K_ e}, \label{eq:bordellobai2}
\end{align}
which means that the sample (auto)covariances between $\{ e_{it}\}$ and $\{ e_{jt}\}$ are $\sqrt T$-consistent estimators of their population counterparts. In particular, by setting $s=t$ in \eqref{eq:bordellobai2}, we see that we can consistently estimate the $(i,j)$th entry of the idiosyncratic covariance matrix $\bm\Gamma^ e$.
Like for the factors, this assumption is a high-level one and requires introducing the sample size $T$. As for the factors, in Appendix \ref{sec:hannan} we provide a series of possible primitive conditions on the process $\{ e_{it}\}$ which guarantee part (c-ii) to hold.
Last, we notice that if we allowed for serial heteroskedasticity so that we had $\gamma_{ij,t,t-k}^ e:=\mathbb{E}[ e_{it} e_{j,t-k}]$, then all our proofs would still hold provided in part (b) we required $\sup_{t\in\mathbb Z} \vert \gamma_{ij,t,t-k}^ e\vert\le \rho^{|k|} M_{ij}$, and then we replaced throughout $\gamma_{ij,k}^ e$ with $\gamma_{ij,t,t-k}^ e$.
We then make the following identifying assumption.
\begin{ass}[\textsc{distinct eigenvalues}]
\label{ass:eval}
For all $n\in\mathbb N$ and all $j=1,\ldots, r-1$, $\mu_j^{C} > \mu_{j+1}^{C}$.
\end{ass}
This can be equivalently stated by requiring $\overline C_j<\underline C_{j-1}$, $j=2,\ldots, r$, in \eqref{eq:lindiv}. Since the $r$ eigenvalues of $\bm\Sigma_\Lambda\bm\Gamma^F$ are given by $\lim_{n\to\infty} n^{-1}\mu_j^{C}$, $j=1,\ldots,r$, then, Assumption \ref{ass:eval} implies that they are also distinct (see Lemma \ref{lem:Vzero}), as also required in \citet[Assumption G]{Bai03}. Note that the $r$ eigenvalues of
$(\bm\Sigma_\Lambda)^{1/2}\bm\Gamma^F(\bm\Sigma_\Lambda)^{1/2}$ or equivalently of
$(\bm\Gamma^F)^{1/2}\bm\Sigma_\Lambda(\bm\Gamma^F)^{1/2}$ are also distinct because they coincide with those of $\bm\Sigma_\Lambda\bm\Gamma^F$.
Distinct eigenvalues are not needed to estimate consistently the space spanned by the eigenvectors since we can apply a recent version of Davis Kahan theorem \citep[Corollary 1]{yu15} which relies only on the existence of a positive eigen-gap between the $r$th and $(r+1)$th eigenvalues of $\bm\Gamma^x$ and $\bm\Gamma^{C}$. The former matrix has an eigen-gap widening with $n$ as shown in \eqref{eq:evalidioX} below, while the latter one is of rank $r$, so in both cases eigen-gap is always strictly positive. However, in order to identify the space spanned by the loadings, we need also the inverse of the $r\times r$ matrix of normalized eigenvectors of $(\bm\Gamma^F)^{1/2}\bm\Sigma_\Lambda(\bm\Gamma^F)^{1/2}$ and this requires distinct eigenvalues since, in turn, they imply that the eigenvectors are identified and thus are linearly independent, so that the eigenvector matrix is invertible (see the proof of Proposition \ref{prop:KKK}). Distinct eigenvalues are also crucial for identification (see Section \ref{sec:II}).
We then turn to the issue of dependence between common and idiosyncratic components. Given that PCA deals only with covariances it would seem enough to just assume contemporaneous orthogonality of the common and idiosyncratic components, i.e., that for all $i,j\in\mathbb N$ and all $t\in\mathbb Z$
\begin{equation}\label{eq:ortho}
\mathbb{E}[{C}_{it} e_{jt}]=0.
\end{equation}
This is indeed true if we aim to prove only consistency of the PC estimators and we make primitive assumptions on the processes $\{\mathbf F_t\}$ and $\{ e_{it}\}$ as those in Appendix \ref{sec:hannan} (see the proofs of Proposition \ref{prop:L} and Lemma \ref{lem:covarianze}(i-b)).
Furthermore, orthogonality in \eqref{eq:ortho} is also enough to asymptotically identify the common component. Indeed, because of Weyl's inequality \citep[Theorem 1]{MK04}, \eqref{eq:lindiv}, \eqref{eq:evalidio}, and \eqref{eq:ortho} imply that $\bm\Gamma^x:=\mathbb{E}[\mathbf x_t\mathbf x_t^\prime]=\bm\Gamma^{C}+\bm\Gamma^ e$ and the existence of an eigen-gap in the eigenvalues $\mu_j^x$, $j=1,\ldots,n$, of $\bm\Gamma^x$ (see Lemma \ref{lem:Gxi}(vi)):
\begin{align}
&\underline C_r\le \lim\inf_{n\to\infty} \frac{ \mu_{r}^x}n\le\lim\sup_{n\to\infty} \frac{\mu_{r}^x}n\le \overline C_r\;\text{ and }\;\sup_{n\in\mathbb N}\mu_{r+1}^x \le M_ e,\label{eq:evalidioX}
\end{align}
where $\underline C_r$ and $\overline C_r$ are defined in \eqref{eq:lindiv} and
$M_ e$ is defined in Assumption \ref{ass:idio}(b). Note that the viceversa is also true, i.e., if \eqref{eq:evalidioX} holds, then \eqref{eq:lindiv}-\eqref{eq:evalidio} hold (\citealp[see][Theorem 4, for the original proof]{chamberlainrothschild83}, and \citealp[Theorem 2]{gersing2023weak}, or \citealp[Proposition 1]{BH25} for more recent and complete proofs). This is the result upon which most methods for determining the number of factors $r$ rely on \citep[see, e.g.,][]{baing02,ABC10,ahnhorenstein13}. Last, notice that for this result we do not need Assumption \ref{ass:eval} to hold so we can have non-distinct leading eigenvalues of $\bm\Gamma^x$.
If we want to prove asymptotic normality of the PC estimators, then additional conditions besides orthogonality and relating to 4th-order dependencies are needed. To this end, we can assume the following.
\begin{ass}[\textsc{independence}]
\label{ass:ind} The processes
$\{ e_{it},\, i\in\mathbb N,\, t\in\mathbb Z\}$ and $\{F_{jt},\, j=1,\ldots, r,\, t\in\mathbb Z\}$ are mutually independent.
\end{ass}
This assumption obviously implies \eqref{eq:ortho} but is much stronger and apparently even more strict than what is usually assumed. Because of this, it deserves some careful discussion. Specifically, it has three main implications.
First, from Assumptions \ref{ass:idio}(c-ii) and \ref{ass:ind} it follows that (see \eqref{eq:IVb1} in the proof of Proposition \ref{prop:F2}):
\begin{align}
\sup_{n,T\in\mathbb N}\mathbb{E}\left[\left\Vert\frac 1{\sqrt n T}\sum_{i=1}^n\left\{ \bm\varepsilon_i\bm\varepsilon_i^\prime\bm F-\mathbb{E}[\bm\varepsilon_i\bm\varepsilon_i^\prime]\bm F\right\}\right\Vert_F^2\right] \le r M_F K_ e,\label{eq:frascati24b}
\end{align}
where $\bm\varepsilon_i:=( e_{i1}\cdots e_{iT})^\prime$. Condition \eqref{eq:frascati24b} coincides with what is assumed by \citet[Assumption F.1]{Bai03}. This inequality is the analogous for the factors of condition \eqref{eq:frascati24}, which we derived for the loadings from Assumption \ref{ass:idio}(c-ii).
It is a crucial requirement as it is necessary to prove asymptotic normality of the estimated factors. In particular, without \eqref{eq:frascati24b} consistency of the factors would hold, but with a rate, $\min(\sqrt n,\sqrt T)$ slower than the rate $\min(\sqrt n, T)$ derived in Proposition \ref{prop:F2}. Such a slower rate is what prevents the proof of asymptotic normality of the estimated factors to go through.
While proving \eqref{eq:frascati24} is easy since the loadings are deterministic, to prove \eqref{eq:frascati24b} it must be that $\mathbb{E}[ e_{i_1t} e_{i_2t} e_{i_1s_1} e_{i_2s_2}F_{ks_1}F_{ks_2}]=\mathbb{E}[ e_{i_1t} e_{i_2t} e_{i_1s_1} e_{i_2s_2}]\,\mathbb{E}[F_{ks_1}F_{ks_2}]$ for all $i_1,i_2=1,\ldots, n$, all $t,s_1,s_2=1,\ldots, T$, and all $k=1,\ldots, r$, this is why we need independence in Assumption \ref{ass:ind} if we want to prove \eqref{eq:frascati24b}, rather than just assuming it.
Second, Assumption \ref{ass:ind} implies that there exists a finite positive real $M_{F e}$ such that
(see the proof of Lemma \ref{lem:LLN}(i))
\begin{align}\label{eq:oslo2}
&\sup_{n,T\in\mathbb N}\max_{i=1,\ldots, n}\mathbb{E}\left[\left\Vert\frac 1{\sqrt {T}}\sum_{t=1}^T\mathbf F_t e_{it}\right\Vert^2\right] \le M_{F e},
\end{align}
which, in turn, implies
the same condition as in \citet[Assumption D]{Bai03}, indeed,
\begin{equation}\label{eq:BaiD}
\sup_{n,T\in\mathbb N}\mathbb{E}\left[\frac 1n\sum_{i=1}^n \left\Vert\frac 1{\sqrt {T}}\sum_{t=1}^T\mathbf F_t e_{it} \right\Vert^2\right] \le
\sup_{n,T\in\mathbb N} \max_{i=1,\ldots, n} \mathbb{E}\left[\left\Vert\frac 1{\sqrt {T}}\sum_{t=1}^T\mathbf F_t e_{it} \right\Vert^2\right] \le M_{F e}.
\end{equation}
Notice that both \eqref{eq:oslo2} and \eqref{eq:BaiD} allow for weak dependence between the factors and the idiosycratic components, but just at the level of 4th-order moments, while they still imply contemporaneous orthogonality. Indeed, when letting $T\to\infty$, by Chebychev's inequality, these conditions imply that $T^{-1}\sum_{t=1}^T\mathbf F_t e_{it}=O_{\mathrm P}(T^{-1/2})$. Moreover, under the assumptions on serial dependence of the factors and the idiosyncratic components allowing for Assumptions \ref{ass:common}(c-ii) and \ref{ass:idio}(c-ii) to hold, we have the Law of Large Numbers
$\Vert T^{-1}\sum_{t=1}^T\mathbf F_t e_{it}-\mathbb{E}[\mathbf F_t e_{it}]\Vert=O_{\mathrm P}(T^{-1/2})$ (this is implied also by Assumption \ref{ass:CLT}(a) below). Hence, by uniqueness of the limit, it follows that we must have $\mathbb{E}[\mathbf F_t e_{it}] = \mathbf 0_r$ for all $i\in\mathbb N$ and $t\in\mathbb Z$, and, thus, contemporaneous orthogonality in \eqref{eq:ortho} still holds.
Crucially, the moment condition in \eqref{eq:oslo2} is needed to prove not only asymptotic normality, but also consistency of the PC estimators (see Lemma \ref{lem:covarianze}(i-a)) so either it is assumed or it must be derived from Assumption \ref{ass:ind}. Or it could be ensured by making further primitive assumptions on the processes $\{\mathbf F_t\}$ and $\{ e_{it}\}$, as those discussed in Appendix \ref{sec:hannan}.
Third, from Assumptions \ref{ass:common}(a), \ref{ass:idio}(b), and \ref{ass:ind} it follows that (see \eqref{eq:1a3} in the proof of Lemma \ref{prop:load}):
\begin{align}
\sup_{n,T\in\mathbb N}\mathbb{E}\left[\left\Vert\frac 1{\sqrt{nT}}\sum_{t=1}^T\sum_{j=1}^n\mathbf F_t{\bm\lambda}_j^\prime e_{jt} \right\Vert_F^2\right]
\le r^2M_\Lambda^2M_{F e},
\label{eq:BaiF2}
\end{align}
which is the same condition as in \citet[Assumption F.2]{Bai03}. Once again condition \eqref{eq:BaiF2} is crucial for proving asymptotic normality, thus, it must be either assumed or derived from Assumption \ref{ass:ind}.
The trade-off is then clear. Either we make Assumption \ref{ass:ind}, which, although strong, is intuitive and it implies orthogonality in \eqref{eq:ortho} as well as the required moment inequalities in \eqref{eq:frascati24b}, \eqref{eq:oslo2}, and \eqref{eq:BaiF2}. Or we directly assume the less intuitive \eqref{eq:frascati24b}, \eqref{eq:BaiD}, and \eqref{eq:BaiF2} as in \citet[Assumptions F.1, D, and F.2, respectively]{Bai03}, which imply also orthogonality in \eqref{eq:ortho}. Here we choose the former approach in order to keep working with more intuitive assumptions.
Notice that our choice coincides with \citet[Assumption D]{baing04}, where, in presence of unit roots both in the factors and the idiosyncratic components, they propose to first estimate the loadings and the factors via PCA on the differenced data and, to this end, assume independence for simplicity.\footnote{
In the case of random loadings \citet{baing04} also assume that the sequence $\{\lambda_{ij},\,i\in\mathbb N,\, j=1,\ldots, r\}$ is independent of the factor and idiosyncratic processes, so that \eqref{eq:frascati24} still holds.}
In order to prove asymptotic normality of the PC estimators it is common to introduce the following assumptions.
\begin{ass}[\textsc{Central limit theorems}]
\label{ass:CLT}
$\,$
\begin{compactenum}[(a)]
\item For all $i\in\mathbb N$, as $T\to\infty$,
$
\frac 1{\sqrt T}\sum_{t=1}^T
\mathbf F_t e_{it}
\to_d
\mathcal N\left(\mathbf 0_r, \bm\Phi_i
\right),
$
where $\bm\Phi_i:=\lim_{T\to\infty}\frac 1 T\sum_{t,s=1}^T
\mathbb{E}_{}[\mathbf F_t\mathbf F_s^\prime e_{it} e_{is}]$.
\item For all $t\in\mathbb Z$, as $n\to\infty$,
$
\frac 1{\sqrt n}\sum_{i=1}^n
\bm\lambda_i e_{it}
\to_d
\mathcal N\left(\mathbf 0_r, \bm\Gamma_t
\right),
$
where $\bm\Gamma_t:=\lim_{n\to\infty}\frac 1 n\sum_{i,j=1}^n
\bm\lambda_i\bm\lambda_j^\prime\mathbb{E}_{}[ e_{it} e_{jt}]$.
\end{compactenum}
\end{ass}
Part (a) is also assumed by \citet[Assumption F.4]{Bai03}. It can be derived from more primitive assumptions. For example, we could assume strong mixing factors and idiosyncratic components and such that, for all $t\in\mathbb Z$ and all $i\in\mathbb N$, we strengthen Assumptions \ref{ass:common}(c-i) and \ref{ass:idio}(c-i) to: $\mathbb{E}[\Vert\mathbf F_{t}\Vert^{4+\epsilon}]\le K_F$ and $\mathbb{E}[| e_{it}|^{4+\epsilon}]\le Q_ e$, for some $\epsilon>0$.
Then, $\{\mathbf F_t e_{it}\}$ is also strong mixing because of \citet[Theorem 5.1.a]{bradley05} and with finite 4th order moments because of Assumption \ref{ass:ind} and part (a) would follow from \citet[Theorem 1.4]{ibra62}. For alternative primitive conditions see also Appendix \ref{sec:hannan}. Notice also that, under Assumption \ref{ass:ind}, in part (a) we actually have $\bm\Phi_i:=\lim_{T\to\infty}T^{-1}\sum_{t,s=1}^T
\mathbb{E}_{}[\mathbf F_t\mathbf F_s^\prime]\,\mathbb{E}[ e_{it} e_{is}]$.
Part (b) is also assumed by \citet[Assumption F.3]{Bai03}. Clearly, it holds if we assumed cross-sectionally uncorrelated idiosyncratic components, i.e., if $\bm\Gamma^ e$ were diagonal as in an exact factor model. In general, to derive part (b) from primitive assumptions we could either introduce an ordering of the $n$ cross-sectional items and a related notion of spatial dependence and derive it from the properties of stationary mixing random fields \citep{bolthausen1982}, or of cross-sectional martingale difference sequences \citep{KP13}. Or we could apply results on exchangeable sequences, which are instead independent of the ordering, and are in turn obtained by virtue of the Hewitt-Savage-de Finetti theorem (\citealp[Theorem 4]{austern2022limit}, and \citealp{bolthausen1984}). In any case we would also need to strengthen Assumption \ref{ass:idio}(c-i) by asking that for all $t\in\mathbb Z$, all $i\in\mathbb N$, $\mathbb{E}[| e_{it}|^{4+\epsilon}]\le Q_ e$, for some $\epsilon>0$. In the case of random loadings we should also require that for all $i\in\mathbb N$, $
\mathbb{E}[\Vert \bm\lambda_i\Vert^{4+\epsilon}]\le K_\Lambda$, for some $\epsilon>0$, and we would have:
$\bm\Gamma_t:=\lim_{n\to\infty}n^{-1}\sum_{i,j=1}^n
\mathbb{E}[\bm\lambda_i\bm\lambda_j^\prime]\,\mathbb{E}_{}[ e_{it} e_{jt}]$.
\begin{ass}[\textsc{Rates}]
\label{ass:rates}
As $n,T\to\infty$,
$\sqrt T/n\to0$ and $\sqrt n/T\to0$.
\end{ass}
This assumption is needed only for proving asymptotic normality. It is assumed also in \citet[Theorems 1 and 2]{Bai03}. It is in fact very mild, and it allows for the typical case of $n\asymp T$.
To summarize, in Table \ref{tab:ass} we show the relation between our Assumptions \ref{ass:common}-\ref{ass:rates} and those made by \citet{Bai03}. As discussed in detail above, the latter are implied by ours, which are fewer and more intuitive. Notice that three assumptions we make are not explicitly stated in \citet{Bai03}, but hold also therein implicitly. Indeed: (i) the number of factors $r$ must be independent of $n$ and $T$ (Assumption \ref{ass:common}(d)), (ii) the idiosyncratic components must have zero mean (Assumption \ref{ass:idio}(a)), and (iii) the relative rates of divergence between $n$ and $T$ must always be imposed to prove asymptotic normality (Assumption \ref{ass:rates}).
\begin{table}[t!]
\caption{Assumptions for PCA}\label{tab:ass}
\centering
\scriptsize{
\begin{tabular}{l | l}
\hline
\hline
\citet{Bai03} & This paper\\
\hline\\[-5pt]
Assumption A & Assumptions \ref{ass:common}(b), \ref{ass:common}(c-i), \ref{ass:common}(c-ii).\\[3pt]
Assumption B&Assumption \ref{ass:common}(a).\\[3pt]
Assumption C.1 & Assumption \ref{ass:idio}(c-i), but only for 4th order moments need to be finite.\\[3pt]
Assumption C.2 & Conditions \eqref{eq:baiC2}, implied by Assumption \ref{ass:idio}(b), and \eqref{eq:baiC2bis}, implied by Assumption \ref{ass:idio}(b) (see Lemma \ref{lem:GxiBAI}(i)).\\[3pt]
Assumption C.3 & Conditions \eqref{eq:baiC3}, implied by Assumption \ref{ass:idio}(b), and \eqref{eq:baiC3bis}, implied by Assumption \ref{ass:idio}(b) (see Lemma \ref{lem:Gxi}(ii)).\\[3pt]
Assumption C.4 & Condition \eqref{eq:bai}, implied by Assumption \ref{ass:idio}(b) (see Lemma \ref{lem:Gxi}(i)).\\[3pt]
Assumption C.5 & Assumption \ref{ass:idio}(c-ii), but only 4th order cumulants need to be summable.\\[3pt]
Assumption D & Condition \eqref{eq:oslo2}, implied by Assumption \ref{ass:ind} (see Lemma \ref{lem:LLN}(i)).\\[3pt]
Assumption E.1 & Condition \eqref{eq:baiE1}, implied by Assumption \ref{ass:idio}(b) (see Lemma \ref{lem:GxiBAI}(ii)).\\[3pt]
Assumption E.2& Condition \eqref{eq:baiE2}, implied by Assumption \ref{ass:idio}(b) (see Lemma \ref{lem:GxiBAI}(iii)).\\[3pt]
Assumption F.1& Condition \eqref{eq:frascati24b}, implied by Assumptions \ref{ass:common}(b), \ref{ass:idio}(c-ii), and \ref{ass:ind} (see \eqref{eq:IVb1} in the proof of Proposition \ref{prop:F2}).\\[3pt]
Assumption F.2& Condition \eqref{eq:BaiF2}, implied by Assumptions \ref{ass:common}(a), \ref{ass:idio}(b), and \ref{ass:ind} (see \eqref{eq:1a3} in the proof of Lemma \ref{prop:load}).\\[3pt]
Assumption F.3& Assumption \ref{ass:CLT}(b).\\[3pt]
Assumption F.4& Assumption \ref{ass:CLT}(a).\\[3pt]
Assumption G& Assumption \ref{ass:eval}.\\[3pt]
N.A. & Assumptions \ref{ass:common}(d), \ref{ass:idio}(a), and \ref{ass:rates}.\\
\hline
\hline
\end{tabular}
}
\end{table}
Finally, similar assumptions are found in other works on PCA for factor models with some important differences, though. First, \citet{stockwatson02JASA} assume, like here, only summable 4th-order moments, but do not impose any other cross-moment conditions since they just prove consistency but not asymptotic normality. Second, \citet{FLM13} impose stronger assumptions by requiring exponentially decaying tails of the distributions of factors and idiosyncratic components. This is done in order to derive uniform consistency for loadings and factors, which, in turn, is needed to prove consistency of the resulting high-dimensional covariance matrix having a low-rank (factor driven) plus sparse (idiosyncratic) structure.
\section{Consistent estimation of the loadings and factors spaces - part 1}\label{sec:cons1}
Our first result is derived using only the properties of the sample covariance and its eigenvalues and eigenvectors (see Appendix \ref{prop:Lproof} for a proof).
\begin{prop}\label{prop:L}
Under Assumptions \ref{ass:common} through \ref{ass:eval}, contemporaneous orthogonality as in
\eqref{eq:ortho} and either of the following:
\begin{inparaenum}
\item [(A)] Assumptions \ref{ass:Wold} or \ref{ass:Hannan} or \ref{ass:Wu} in Appendix \ref{sec:hannan} on serial dependence of the processes;
\item [(B)] the moment conditions in \eqref{eq:oslo2} and Assumption \ref{ass:common}(c-ii) holding with the rate $\sqrt T$;
\end{inparaenum}
it follows that, as $n,T\to\infty$,\footnote{We say parts (b) and (d) hold uniformly in $i$ or $t$ when the constants involved in the $O_{\mathrm P}$ bound do not depend on $i$ and $t$, it does not mean uniform consistency, which would imply statements like
$\max_{i=1,\ldots, n}\Vert\widehat{\bm \lambda}_i^\prime-\bm\lambda_i^\prime\bm{\mathcal H}\Vert=o_{\mathrm P}(1)$ nor
$\max_{t=1,\ldots,T}\Vert\widehat{\mathbf F}_t-\bm{\mathcal H}^{-1}\mathbf F_t\Vert=o_{\mathrm P}(1)$.}
\begin{compactenum}[(a)]
\item $\min(n,\sqrt T)\left\Vert\frac{\widehat{\bm\Lambda}-\bm\Lambda\bm{\mathcal H}}{\sqrt n}\right\Vert= O_{\mathrm P}(1)$;
\item $ {\min(\sqrt n,\sqrt T)} \left\Vert\widehat{\bm \lambda}_i^\prime-\bm\lambda_i^\prime\bm{\mathcal H}\right\Vert= O_{\mathrm P}(1)$, uniformly in $i$;
\item $\min(\sqrt n,\sqrt T)\left\Vert\frac {\widehat{\bm F}-\bm F(\bm{\mathcal H}^{-1})^\prime}{\sqrt T}\right\Vert= O_{\mathrm P}(1)$;
\item ${\min(\sqrt n,\sqrt T)} \left\Vert\widehat{\mathbf F}_t-\bm{\mathcal H}^{-1}\mathbf F_t\right\Vert= O_{\mathrm P}(1)$, uniformly in $t$;
\end{compactenum}
where $\bm{\mathcal H} :=(\bm\Gamma^F)^{1/2} \mathbf K{\mathbf J}$ is an $r\times r$ finite and positive definite matrix, with $\mathbf K$ having as columns the normalized eigenvectors of $(\bm\Gamma^F)^{1/2}(n^{-1}\bm \Lambda^\prime\bm \Lambda) (\bm\Gamma^F)^{1/2}$, and $\mathbf J$ being an $r\times r$ diagonal matrix with entries $\pm 1$ which depend on $n$ and $T$.
\end{prop}
The bound derived in part (a) is tighter than the one derived by \citet[Proposition P]{FGLR09} and \citet[Proposition 1.ii]{bai2020simpler} where the rate is $\min(\sqrt n,\sqrt T)$.
Parts (b), (c), and (d) follow directly from part (a). In particular, part (c) coincides with the result by
\citet[Proposition 1.i]{bai2020simpler}.
The proof makes use of the Cauchy-Schwarz inequality, therefore, the implied rates are not the sharpest possible. Hence this result does not allow to derive asymptotic normality. However, besides providing an intuitive and quick proof of consistency of PC, which is often enough for many purposes, Proposition \ref{prop:L} is also needed for proving successive results, where the consistency rates are refined, provided additional moment conditions are imposed (see Section \ref{sec:cons2}).
Part (a) is based essentially on a result on the $n\times n$ sample covariance matrix $\widehat{\bm\Gamma}^x$ which in turn is determined by the serial dependence of the considered stochastic process, as explained in Appendix \ref{sec:hannan}.
Its proof is based on four main steps proved in Lemma \ref{lem:covarianze}. First, we prove that the rescaled sample covariance matrix $n^{-1}\widehat{\bm\Gamma}^x$ is a $\sqrt T$-consistent estimator of the rescaled population covariance matrix $n^{-1}\bm\Gamma^x$. This result holds under a minimal and mild set of assumptions, namely Assumptions \ref{ass:common} and \ref{ass:eval} jointly with contemporaneous orthogonality as in
\eqref{eq:ortho} plus either the conditions (A) (see Lemma \ref{lem:covarianze}(i-b)) or the conditions (B) (see Lemma \ref{lem:covarianze}(i-a)).
We stress that under either (A) or (B) Assumption \ref{ass:ind} of independence between factors and idiosyncratic components is not needed.
Second, we prove that the rescaled sample covariance $n^{-1}\widehat{\bm\Gamma}^x$ actually converges to the rescaled common component covariance $n^{-1}\bm\Gamma^{C}$ with rate $\min(n,\sqrt T)$ (see Lemma \ref{lem:covarianze}(ii)). This result follows from Assumption \ref{ass:idio} which implies that the idiosyncratic covariance matrix $\bm\Gamma^ e$ has bounded eigenvalues for all $n\in\mathbb N$ (see \eqref{eq:evalidio}). It is clear that, in fact, Assumption \ref{ass:idio}(b) is not strictly needed since it is enough to make the weaker assumption of bounded idiosyncratic eigenvalues.
Third, we prove that the matrices of sorted eigenvalues, $\widehat{\mathbf M}^x$, and corresponding normalized eigenvectors, $\widehat{\mathbf V}^x$, of $\widehat{\bm\Gamma}^x$ converge to the matrices of sorted eigenvalues, ${\mathbf M^{C}}$, and corresponding normalized eigenvectors, ${\mathbf V^{C}}$, of $\bm\Gamma^{C}$. Specifically, from Weyl's inequality we get $n^{-1}\Vert \widehat{\mathbf M}^x-\mathbf M^{C}\Vert=O_{\mathrm P}(\max(n^{-1},T^{-1/2}))$ (see Lemma \ref{lem:covarianze}(iii) and \citealp[Theorem 1]{MK04}), while from Davis Kahan theorem we get $\Vert \widehat{\mathbf V}^x-\mathbf V^{C}\widehat{\mathbf J}\Vert =O_{\mathrm P}(\max(n^{-1},T^{-1/2}))$, where $\mathbf J$ is an $r\times r$ diagonal matrix with entries $\pm 1$ depending on $n$ and $T$, and accounting for the fact that a sample and a population eigenvector identify asymptotically the same one-dimensional subspaces, but might point in opposite directions (see Lemma \ref{lem:covarianze}(iv) and \citealp[Corollary 1]{yu15}).
Fourth, since $\bm\Gamma^{C} =\bm\Lambda\bm\Gamma^F\bm\Lambda^\prime = \mathbf V^{C}\mathbf M^{C}\mathbf V^{{C}\prime}$, the columns of the true loadings matrix $\bm\Lambda$ must span the same space as the normalized eigenvectors of the common component covariance matrix. In other words, it must be that $\bm\Lambda(\bm\Gamma^F)^{1/2}\mathbf K=\mathbf V^{C}(\mathbf M^{C})^{1/2}$ for some invertible $r\times r$ matrix $\mathbf K$. In particular, we prove that the columns of $\mathbf K$ are the normalized eigenvectors of $(\bm\Gamma^F)^{1/2} (n^{-1}\bm\Lambda^\prime\bm\Lambda) (\bm\Gamma^F)^{1/2}$ (see Lemma \ref{lem:KO1}(iii))), i.e., $\mathbf K$ is a rotation.
Summing up, given that the PC estimator of the loadings is $\widehat{\bm\Lambda}=\widehat{\mathbf V}^x(\widehat {\mathbf M}^x)^{1/2}$ (see \eqref{eq:estL}), part (a) follows, with
$\bm{\mathcal H}=(\bm\Gamma^F)^{1/2} \mathbf K{\mathbf J}$ being a finite and invertible linear transformation. So the true loadings are recovered up to: (i) a scale given by $(\bm\Gamma^F)^{1/2}$; (ii) a rotation $\mathbf K$, and (iii) a diagonal matrix of signs ${\mathbf J}$. Note that both $(\bm\Gamma^F)^{1/2}$ and $\mathbf K$ do not depend on the sample size $T$, but depend only on population quantities as well as $n$. The only dependence on $T$ is in the sign matrix ${\mathbf J}$, but such dependence can be easily fixed (see Section \ref{sec:II}).
The following result gives a series of more intuitive asymptotic expressions for $\bm{\mathcal H}$ (see Appendix \ref{corol:ovvioproof} for a proof).
\begin{prop}\label{corol:ovvio}
Under Assumptions \ref{ass:common} through \ref{ass:eval}, contemporaneous orthogonality as in
\eqref{eq:ortho} and either of the following:
\begin{inparaenum}
\item [(A)] Assumptions \ref{ass:Wold} or \ref{ass:Hannan} or \ref{ass:Wu} in Appendix \ref{sec:hannan} on serial dependence of the processes;
\item [(B)] the moment conditions in \eqref{eq:oslo2} and Assumption \ref{ass:common}(c-ii) holding with the rate $\sqrt T$;
\end{inparaenum}
it follows that, as $n,T\to\infty$,
\begin{compactenum}
\item [(a)] $\min(n,\sqrt T)\left\Vert\bm{\mathcal H} - (\bm\Lambda^\prime\bm\Lambda)^{-1}\bm\Lambda^\prime\widehat{\bm\Lambda} \right\Vert = O_{\mathrm P}(1)$;
\item [(b)] $\min(\sqrt n,\sqrt T)\left\Vert\bm{\mathcal H}^{-1} - \widehat{\bm F}^\prime \bm F (\bm F^\prime\bm F)^{-1} \right\Vert = O_{\mathrm P}(1)$;
\item [(c)] $\min(n,\sqrt T)\left\Vert\bm{\mathcal H}^{-1} - (\widehat{\bm\Lambda}^\prime\widehat{\bm\Lambda})^{-1}\widehat{\bm\Lambda}^\prime{\bm\Lambda} \right\Vert = O_{\mathrm P}(1)$, where we recall that $\widehat{\bm\Lambda}^\prime\widehat{\bm\Lambda}=\widehat{\mathbf M}^x$;
\item [(d)] $\min(\sqrt n,\sqrt T)\left\Vert\bm{\mathcal H} - {\bm F}^\prime \widehat{\bm F} (\widehat{\bm F}^\prime\widehat{\bm F})^{-1} \right\Vert = O_{\mathrm P}(1)$, where we recall that $\widehat{\bm F}^\prime \widehat{\bm F} = T\, \mathbf I_r$.
\end{compactenum}
\end{prop}
The proof is an immediate consequence of Proposition \ref{prop:L}(a) and \ref{prop:L}(c) and since the assumptions made are the same as those made for Proposition \ref{prop:L}, these results are directly applicable to the consistency results proved therein.
Proposition \ref{corol:ovvio} shows that asymptotically the matrix $\bm{\mathcal H}$ is obtained by linear projection of the estimated loadings onto the true loadings (part (a)) or of the true factors onto the estimated ones (part (d)). Similarly, asymptotically $\bm{\mathcal H}^{-1}$ is obtained by linear projection of the estimated factors onto the true factors (part (b)) or by linear projection of the true loadings onto the estimated ones (part (c)). The derived rates are not the sharpest possible for two reasons. First, the rates in part (a) and (b) are obtained by using the bound $n^{-1}\Vert\bm\Lambda^\prime(\widehat{\bm\Lambda}-\bm\Lambda\bm{\mathcal H})\Vert\le n^{-1/2}\Vert \bm\Lambda\Vert\, n^{-1/2}\Vert\widehat{\bm\Lambda}-\bm\Lambda\bm{\mathcal H} \Vert$, which, therefore, inherits the bound obtained for the loadings in Proposition \ref{prop:L}(a). Second, the rates in parts (b) and (d) are slower than those in parts (a) and (c) since the rates in Proposition \ref{prop:L}(c) are not the sharpest possible. As shown, in Proposition \ref{prop:LLFF} below, it is possible to define a different transformation matrix, which is asymptotically equivalent to $\bm{\mathcal H}$, but which converges to the projection matrices in Proposition \ref{corol:ovvio} at a faster rate.
Finally, consider the following spectral decomposition: \begin{equation}\label{eq:U0V0U0}
(\bm\Gamma^F)^{1/2}\bm\Sigma_\Lambda(\bm\Gamma^F)^{1/2} =: \bm\Upsilon_0\bm V_0\bm\Upsilon_0^\prime,
\end{equation}
where $\bm\Upsilon_0$ is the $r\times r$ matrix having as columns the normalized eigenvectors of $(\bm\Gamma^F)^{1/2}\bm\Sigma_\Lambda(\bm\Gamma^F)^{1/2}$, and $\bm V_0$ is the $r\times r$ matrix of corresponding eigenvalues sorted in decreasing order.
Then, we prove a result, which is the analogous of the result in \citet[Proposition 1]{Bai03}, and is needed for the next section. It gives the limit of the space spanned by the estimated loadings (see Appendix \ref{app:KKKc1} for a proof).
\begin{prop}\label{prop:KKK}
Under Assumptions \ref{ass:common} through \ref{ass:eval}, contemporaneous orthogonality as in
\eqref{eq:ortho} and either of the following:
\begin{inparaenum}
\item [(A)] Assumptions \ref{ass:Wold} or \ref{ass:Hannan} or \ref{ass:Wu} in Appendix \ref{sec:hannan} on serial dependence of the processes;
\item [(B)] the moment conditions in \eqref{eq:oslo2};
\end{inparaenum}
it follows that, as $n,T\to\infty$,
$
\left\Vert \frac{\widehat{\bm\Lambda}^\prime\bm\Lambda}{n}-\bm V_0\bm{\mathcal J}_0\bm\Upsilon_0^\prime (\bm\Gamma^F)^{-1/2}\right\Vert = o_{\mathrm P}(1),
$
where the columns of $\bm\Upsilon_0$ are the normalized eigenvectors of $(\bm\Gamma^F)^{1/2}\bm\Sigma_\Lambda(\bm\Gamma^F)^{1/2}$, $\bm V_0$ is the $r\times r$ diagonal matrix containing the corresponding eigenvalues, and $\bm{\mathcal J}_0$ is an $r\times r$ diagonal matrix with entries $\pm 1$ independent of $n$ and $T$.
If Assumption \ref{ass:common}(a) holds with rate $\sqrt n$ and Assumption \ref{ass:common}(c-ii) holds with rate $\sqrt T$, then the above holds with rate $\min(\sqrt n,\sqrt T)$. Furthermore, $\bm V_0\bm{\mathcal J}_0\bm\Upsilon_0^\prime (\bm\Gamma^F)^{-1/2}$ is finite and positive definite.
\end{prop}
\section{Consistent estimation of the loadings and factors spaces - part 2}\label{sec:cons2}
Let us start from the estimated loadings. By definition of $\widehat{\bm\Lambda}$ in \eqref{eq:estL} $\widehat{\mathbf V}^x=\widehat{\bm\Lambda}(\widehat{\mathbf M}^x)^{-1/2}$.
Thus, from \eqref{eq:evecXX2}, we obtain
\begin{equation}\label{eq:start}
\frac{\bm X^\prime\bm X}{nT}\widehat{\bm\Lambda}=\widehat{\bm\Lambda}\frac{\widehat{\mathbf M}^x}{n}.
\end{equation}
Then, substituting $\bm X^\prime\bm X=(\bm\Lambda\bm F^\prime+\bm E^\prime)^\prime(\bm F\bm\Lambda^\prime+\bm E)$ into \eqref{eq:start},
and letting
\begin{equation}\label{eq:acca}
\widehat{\mathbf H}:=
\frac{\bm F^\prime\bm F}{T}
\frac{\bm\Lambda^\prime\widehat{\bm\Lambda}}{n}
\left(\frac{\widehat{\mathbf M}^x}{n}\right)^{-1},
\end{equation}
we obtain
\begin{align}
\widehat{\bm\Lambda}-\bm\Lambda\widehat{\mathbf H}=&\, \left(\frac{\bm\Lambda\bm F^\prime\bm E\widehat{\bm\Lambda}}{nT}
+\frac{\bm E^\prime\bm F\bm\Lambda^\prime\widehat{\bm\Lambda}}{nT}
+\frac{\bm E^\prime\bm E\widehat{\bm\Lambda}}{nT}\right)\left(\frac{\widehat{\mathbf M}^x}{n}\right)^{-1}
.\label{eq:start4}
\end{align}
Notice that, as $n,T\to\infty$, $\widehat{\mathbf H}$ is well defined because of Lemma \ref{lem:HO1}(i). Taking the $i$th row of \eqref{eq:start4}, we get
\begin{align}
\widehat{\bm\lambda}_i^\prime-{\bm\lambda}_i^\prime\widehat{\mathbf H}=&\,
\left(
\underbrace{\frac 1{nT}{\bm\lambda}_i^\prime\sum_{t=1}^T\sum_{j=1}^n\mathbf F_t e_{jt}\widehat{\bm\lambda}_j^\prime}_{\text{(1.a)}}
+\underbrace{\frac 1{nT} \sum_{t=1}^T e_{it}\mathbf F_t^\prime\sum_{j=1}^n\bm\lambda_j\widehat{\bm\lambda}_j^\prime}_{\text{(1.b)}}
+\underbrace{\frac 1{nT} \sum_{t=1}^T\sum_{j=1}^n e_{it} e_{jt} \widehat{\bm\lambda}_j^\prime}_{\text{(1.c)}}
\right) \left(\frac{\widehat{\mathbf M}^x}{n}\right)^{-1}.\label{eq:sviluppoLambda}
\end{align}
The following bounds hold for the terms in \eqref{eq:sviluppoLambda} (see Appendix \ref{prop:loadproof} for a proof).
\begin{lem}\label{prop:load}
Under Assumptions \ref{ass:common} through \ref{ass:ind}, as $n,T\to\infty$,
\begin{inparaenum}[(a)]
\item $\sqrt {nT}\left\Vert \text{\upshape (1.a)}\right\Vert = O_{\mathrm P}(1)$;
\item $\sqrt T \left\Vert \text{\upshape (1.b)}\right\Vert = O_{\mathrm P}(1)$;
\item $\min(n,\sqrt{nT})\left\Vert \text{\upshape (1.c)}\right\Vert = O_{\mathrm P}(1)$;
\end{inparaenum}
uniformly in $i$.
\end{lem}
Consistency follows immediately (see Appendix \ref{prop:L2proof} for a proof).
\begin{prop}\label{prop:L2}
Under Assumptions \ref{ass:common} through \ref{ass:ind}, as $n,T\to\infty$,
\begin{compactenum}[(a)]
\item $\min(n,\sqrt {T})\left\Vert \widehat{\bm\lambda}_i-\widehat{\mathbf H}^\prime{\bm\lambda}_i\right\Vert = O_{\mathrm P} (1)$,
uniformly in $i$;
\item $\min(n,\sqrt {T}) \left\Vert \frac {\widehat{\bm\Lambda}-{\bm\Lambda}\widehat{\mathbf H}}{\sqrt n}\right\Vert=O_{\mathrm P}(1)$.
\end{compactenum}
\end{prop}
Part (a) refines the result in Proposition \ref{prop:L}(b). Part (b) has the same rate as in Proposition \ref{prop:L}(a). As shown in Proposition \ref{prop:LLFF}(a) below, the matrix $\widehat{\mathbf H}$ is asymptotically equivalent to the projection matrix of the estimated loadings onto the true ones.
Turning to the estimated factors, from \eqref{eq:estL} and \eqref{eq:estF}, we have
\begin{align}\label{eq:FFF}
\widehat{\bm F}&=\bm X\widehat{\bm\Lambda}(\widehat{\bm\Lambda}^\prime \widehat{\bm\Lambda})^{-1}=\bm F\bm\Lambda^\prime \widehat{\bm\Lambda}(\widehat{\bm\Lambda}^\prime \widehat{\bm\Lambda})^{-1}+\bm E\widehat{\bm\Lambda}(\widehat{\bm\Lambda}^\prime \widehat{\bm\Lambda})^{-1}= \frac{\bm F\bm\Lambda^\prime \widehat{\bm\Lambda}}n\left(\frac {\widehat{\mathbf M}^x}n\right)^{-1} +
\frac{\bm E\widehat{\bm\Lambda}}{n}\left(\frac {\widehat{\mathbf M}^x}n\right)^{-1}.
\end{align}
Then, substituting on the right-hand-side of \eqref{eq:FFF}
$\bm{\Lambda}$ with
$(\bm\Lambda-\widehat{\bm\Lambda}\widehat{\mathbf H}^{-1})+\widehat{\bm\Lambda}\widehat{\mathbf H}^{-1},$ in the first term and
$\widehat{\bm\Lambda}$ with
$
\widehat{\bm\Lambda}-\bm\Lambda\widehat{\mathbf H}+\bm\Lambda\widehat{\mathbf H}
$
in the second term,
we obtain
\begin{align}
\widehat{\bm F}-\bm F (\widehat{\mathbf H}^{-1})^\prime&=\left(\frac{\bm F(\bm\Lambda-\widehat{\bm\Lambda}\widehat{\mathbf H}^{-1})^\prime\widehat{\bm\Lambda}}{n}+
\frac{\bm E(\widehat{\bm\Lambda}-\bm\Lambda \widehat{\mathbf H})}{n}+
\frac{\bm E\bm\Lambda\widehat{\mathbf H}}{n}\right)\left(\frac {\widehat{\mathbf M}^x}n\right)^{-1}.\label{eq:FFF2}
\end{align}
Notice that, as $n,T\to\infty$, $\widehat{\mathbf H}^{-1}$ is well defined because of Lemma \ref{lem:HO1}(ii).
Taking the $t$th row of \eqref{eq:FFF2}, we get
\begin{align}
\widehat{\mathbf F}_t^\prime -\mathbf F_t^\prime (\widehat{\mathbf H}^{-1})^\prime&=\left(
\underbrace{\mathbf F_t^\prime\frac{(\bm\Lambda-\widehat{\bm\Lambda}\widehat{\mathbf H}^{-1})^\prime\widehat{\bm\Lambda}}{n}}_{\text{ (2.a)}}+
\underbrace{\bm e_t^\prime \frac{(\widehat{\bm\Lambda}-\bm\Lambda \widehat{\mathbf H})}{n}}_{\text{(2.b)}}+
\underbrace{\bm e_t^\prime \frac{\bm\Lambda\widehat{\mathbf H}}{n}}_{\text{(2.c)}}
\right)\left(\frac {\widehat{\mathbf M}^x}n\right)^{-1}.\label{eq:sviluppoFactor}
\end{align}
The following bounds hold for the terms in \eqref{eq:sviluppoFactor} (see Appendix \ref{prop:factorproof} for a proof).
\begin{lem}\label{prop:factor}
Under Assumptions \ref{ass:common} through \ref{ass:ind}, as $n,T\to\infty$,
\begin{inparaenum}[(a)]
\item $\min(n,\sqrt{nT},T)\Vert \text{2.a}\Vert=O_{\mathrm {P}}(1)$;\linebreak
\item $\min(n,\sqrt{nT},T)\Vert \text{2.b}\Vert=O_{\mathrm {P}}(1)$;
\item $\sqrt n\Vert \text{2.c}\Vert=O_{\mathrm {P}}(1)$;
\end{inparaenum}
uniformly in $t$.
\end{lem}
Consistency follows immediately (see Appendix \ref{prop:F2proof} for a proof).
\begin{prop}\label{prop:F2}
Under Assumptions \ref{ass:common} through \ref{ass:ind}, as $n,T\to\infty$
\begin{compactenum}[(a)]
\item $\min(\sqrt n,{T})\left\Vert \widehat{\mathbf F}_t-\widehat{\mathbf H}^{-1}{\mathbf F}_t \right\Vert = O_{\mathrm P} (1)$,
uniformly in $t$;
\item $\min(\sqrt n, {T}) \left\Vert \frac{\widehat{\bm F}-{\bm F}(\widehat{\mathbf H}^{-1})^\prime}{\sqrt T}\right\Vert=O_{\mathrm P}(1)$.
\end{compactenum}
\end{prop}
Parts (a) and (b) refine the results in Proposition \ref{prop:L}(c)-\ref{prop:L}(d). As shown in Proposition \ref{prop:LLFF}(b) below, the matrix $\widehat{\mathbf H}^{-1}$ is asymptotically equivalent to the projection matrix of the estimated loadings onto the true ones.
Finally, we derive a result analogous to the result in \citet[Theorem 1]{baing02} and \citet[Lemma A.1]{Bai03} (see Appendix \ref{prop:sumFproof} for a proof).
\begin{prop}\label{prop:sumF}
Under Assumptions \ref{ass:common} through \ref{ass:ind}, as $n,T\to\infty$,
$$
\min(n,T)\left\{\frac 1T\sum_{t=1}^T \left\Vert \widehat{\mathbf F}_t-\widehat{\mathbf H}^{-1}{\mathbf F}_t \right\Vert^2\right\}=O_{\mathrm P}(1).
$$
\end{prop}
This result is often needed when using the estimated quantities for further analysis, e.g., in factor augmented regressions \citep{baing06}. It is also needed to prove consistency for the selection of the number of factors based on Information Criteria \citep{baing02,ABC10}.
\section{The role of $\widehat{\mathbf H}$ and its relation with $\pmb{\mathcal H}$}\label{sec:Hhat}
To understand the meaning of $\widehat{\mathbf H}$ in Propositions \ref{prop:L2} and \ref{prop:F2}, we can derive a series of more intuitive asymptotic expressions (see Appendix \ref{prop:LLFFproof} for a proof).
\begin{prop}\label{prop:LLFF}
Under Assumptions \ref{ass:common} through \ref{ass:ind}, as $n,T\to\infty$,
\begin{compactenum}[(a)]
\item $\min(n,\sqrt{nT}, {T}) \left\Vert \widehat{\mathbf H} - (\bm\Lambda^\prime\bm\Lambda)^{-1}
{\bm\Lambda}^\prime\widehat{\bm\Lambda}
\right\Vert=O_{\mathrm P}(1)$;
\item $\min( n,\sqrt{nT}, {T}) \left\Vert\widehat{\mathbf H}^{-1}-
\widehat{\bm F}^\prime{\bm F}(\bm F^\prime\bm F)^{-1}
\right\Vert=O_{\mathrm P}(1)$;
\item $
\min(n,\sqrt{nT}, {T})\left\Vert \widehat{\mathbf H}^{-1} - (\widehat{\bm\Lambda}^\prime\widehat{\bm\Lambda})^{-1}
\widehat{\bm\Lambda}^\prime{\bm\Lambda}
\right\Vert=O_{\mathrm P}(1)$, where we recall that $\widehat{\bm\Lambda}^\prime\widehat{\bm\Lambda}=\widehat{\mathbf M}^x$;
\item $\min(n,\sqrt{nT}, {T}) \left\Vert \widehat{\mathbf H} - {\bm F}^\prime\widehat{\bm F}(\widehat{\bm F}^\prime\widehat{\bm F})^{-1}\right\Vert= O_{\mathrm P}(1)$, where we recall that $\widehat{\bm F}^\prime \widehat{\bm F} = T\, \mathbf I_r$;
\item $\min(n,\sqrt{nT}, {T})\left\Vert \widehat{\mathbf H}^{-1} -({\bm\Lambda}^\prime\widehat{\bm\Lambda})^{-1} \bm\Lambda^\prime\bm\Lambda
\right\Vert=O_{\mathrm P}(1)$;
\item $\min( n,\sqrt{nT}, {T}) \left\Vert\widehat{\mathbf H}-
\bm F^\prime\bm F (\widehat{\bm F}^\prime{\bm F})^{-1}
\right\Vert=O_{\mathrm P}(1)$;
\item $\min( n,\sqrt{nT}, {T}) \left\Vert \frac{(\widehat{\mathbf H}^{-1})^\prime\widehat{\bm\Lambda}^\prime\widehat{\bm\Lambda}\widehat{\mathbf H}^{-1}}n-\frac{\bm\Lambda^\prime\bm\Lambda}n\right\Vert=O_{\mathrm P}(1)$, where we recall that $\widehat{\bm\Lambda}^\prime\widehat{\bm\Lambda}=\widehat{\mathbf M}^x$;
\item $\min( n,\sqrt{nT}, {T}) \left\Vert \frac{\widehat{\mathbf H} \widehat{\bm F}^\prime\widehat{\bm F}\widehat{\mathbf H}^\prime}T- \frac{{\bm F}^\prime{\bm F}}T\right\Vert=O_{\mathrm P}(1)$, where we recall that $\widehat{\bm F}^\prime \widehat{\bm F} = T\, \mathbf I_r$.
\end{compactenum}
\end{prop}
This result, which is similar to what proved by \citet[Lemma 3.i]{bai2020simpler}, shows that asymptotically the matrix $\widehat{\mathbf H}$ is obtained by linear projection of the estimated loadings onto the true loadings (part (a)) or of the true factors onto the estimated ones (part (d)). Similarly, asymptotically $\widehat{\mathbf H}^{-1}$ is obtained by linear projection of the estimated factors onto the true factors (part (b)) or by linear projection of the true loadings onto the estimated ones (part (c)). Part (e)-(h) are less intuitive, but still useful for proving further results.
Moreover, $\widehat{\mathbf H}$ has also the following asymptotic expansion (see Appendix \ref{prop:HHATproof} for a proof).
\begin{prop}\label{prop:HHAT}
Under Assumptions \ref{ass:common} through \ref{ass:ind}, as $n,T\to\infty$,
\begin{compactenum}
\item [(a)] $\min(n,\sqrt{T})\left\Vert \widehat{\mathbf H} - \left(\frac{\bm F^\prime\bm F}{T}\right)^{1/2}\widehat{\mathbf Q}\right\Vert = O_{\mathrm P}(1)$;
\item [(b)] $\min(n,\sqrt{T})\left\Vert\widehat{\mathbf H}^{-1} - \widehat{\mathbf Q}^\prime \left(\frac{\bm F^\prime\bm F}T\right)^{-1/2}\right\Vert =O_{\mathrm P}(1)$;
\end{compactenum}
where the columns of $\widehat{\mathbf Q}$ are the normalized eigenvectors of $(T^{-1}\bm F^\prime\bm F)^{1/2}(n^{-1}\bm \Lambda^\prime\bm \Lambda) (T^{-1}\bm F^\prime\bm F)^{1/2}$.
\end{prop}
According to this result and Propositions \ref{prop:L2} and \ref{prop:F2}, the loadings and the factors can be consistently estimated up to: (i) a scale $(T^{-1}{\bm F^\prime\bm F})^{1/2}$ and (ii) a rotation $\widehat{\mathbf Q}$. The proof of this result rests on four intermediate relevant results, which we summarize here. First and second, from Proposition \ref{prop:LLFF}, it follows that (see also \eqref{eq:warwick3} and \eqref{eq:warwick4} in the proof of Proposition \ref{prop:HHAT})
\begin{align}
&\left\Vert\widehat{\mathbf H} - \frac{\bm F^\prime\bm F}{T} (\widehat{\mathbf H}^{-1})^{\prime}\right\Vert =O_{\mathrm P}\left(\max\left(\frac 1n,\frac{1}{\sqrt{nT}},\frac 1T\right)\right),\label{eq:starop1}\\
&\left\Vert\frac{\widehat{\mathbf M}^x}{n} - \widehat{\mathbf H}^{-1}\frac{\bm F^\prime\bm F}{T} \frac{\bm\Lambda^\prime\bm\Lambda}{n} \widehat{\mathbf H}\right\Vert=O_{\mathrm P}\left(\max\left(\frac 1n,\frac{1}{\sqrt{nT}},\frac 1T\right)\right).\label{eq:starop2}
\end{align}
Third, \eqref{eq:starop1} and \eqref{eq:starop2} imply (see also \eqref{eq:madre} in the proof of Proposition \ref{prop:HHAT})
\begin{align}
\left\Vert
\frac{\widehat{\mathbf M}^x}{n} - \widehat{\mathbf H}^{-1} \frac{\bm F^\prime\bm F}{T} \frac{\bm\Lambda^\prime\bm\Lambda}{n}\frac{\bm F^\prime\bm F}{T} (\widehat{\mathbf H}^{-1})^{\prime}
\right\Vert=O_{\mathrm P}\left(\max\left(\frac 1n,\frac{1}{\sqrt{nT}},\frac 1T\right)\right),\label{eq:starop3}
\end{align}
Last, the $r$ non-zero eigenvalues of
$n^{-1} \bm\Lambda(T^{-1}\bm F^\prime\bm F) \bm\Lambda^\prime$, collected into the diagonal $r\times r$ matrix $n^{-1}\widehat{\mathbf M}^C$ coincide with the eigenvalues of $(T^{-1}\bm F^\prime\bm F)^{1/2}(n^{-1}\bm \Lambda^\prime\bm \Lambda) (T^{-1}\bm F^\prime\bm F)^{1/2}$, and are such that (see also \eqref{eq:mistake} in the proof of Proposition \ref{prop:HHAT})
\begin{equation}\label{eq:starop3bis}
\left\Vert \frac{\widehat{\mathbf M}^C}n-\frac{\widehat{\mathbf M}^x}n\right\Vert = O_{\mathrm P}\left(\max\left(\frac 1n,\frac{1}{\sqrt{T}}\right)\right).
\end{equation}
Then, part (b) follows directly from \eqref{eq:starop3} and \eqref{eq:starop3bis}.
We then prove a result which is the counterpart of Propositions \ref{prop:KKK} and it gives the limit of the space spanned by the estimated factors. It follows essentially from the definition of $\widehat{\mathbf H}$ in \eqref{eq:acca} and Propositions \ref{prop:KKK} and \ref{prop:LLFF}(f) (see Appendix \ref{prop:KKKbisproof} for a proof).
\begin{prop}\label{prop:KKKbis}
Under Assumptions \ref{ass:common} through \ref{ass:ind}, as $n,T\to\infty$,
$
\left\Vert \frac{\widehat{\bm F}^\prime\bm F}{T}-\bm{\mathcal J}_0\bm\Upsilon_0^\prime (\bm\Gamma^F)^{1/2}\right\Vert = o_{\mathrm P}(1),
$
where the columns of $\bm\Upsilon_0$ are the normalized eigenvectors of $(\bm\Gamma^F)^{1/2}\bm\Sigma_\Lambda(\bm\Gamma^F)^{1/2}$, and $\bm{\mathcal J}_0$ is an $r\times r$ diagonal matrix with entries $\pm 1$ independent of $n$ and $T$. If Assumption \ref{ass:common}(a) holds with rate $\sqrt n$ and Assumption \ref{ass:common}(c-ii) holds with rate $\sqrt T$, then the above holds with rate $\min(\sqrt n,\sqrt T)$. Furthermore, $\bm{\mathcal J}_0\bm\Upsilon_0^\prime (\bm\Gamma^F)^{1/2}$ is finite and positive definite.
\end{prop}
Finally, we can derive useful limits for $\widehat{\mathbf H}$, which depends only on population quantities (see Appendix \ref{cor:sempliceproof} for a proof).
\begin{prop}\label{cor:semplice}
Under Assumptions \ref{ass:common} through \ref{ass:ind}, as $n,T\to\infty$,
\begin{inparaenum}
\item [(a)] $\left\Vert\widehat{\mathbf H}- (\bm\Gamma^F)^{1/2}\bm\Upsilon_0 \bm{\mathcal J}_0 \right\Vert =o_{\mathrm P}(1)$;
\item [(b)] $\left\Vert\widehat{\mathbf H}^{-1} - \bm{\mathcal J}_0\bm\Upsilon_0^\prime (\bm\Gamma^F)^{-1/2} \right\Vert =o_{\mathrm P}(1)$;
\end{inparaenum}
where the columns of $\bm\Upsilon_0$ are the normalized eigenvectors of $(\bm\Gamma^F)^{1/2}\bm\Sigma_\Lambda(\bm\Gamma^F)^{1/2}$, and $\bm{\mathcal J}_0$ is an $r\times r$ diagonal matrix with entries $\pm 1$ independent of $n$ and $T$. If Assumption \ref{ass:common}(a) holds with rate $\sqrt n$ and Assumption \ref{ass:common}(c-ii) holds with rate $\sqrt T$, then the above holds with rate $\min(\sqrt n,\sqrt T)$.
\end{prop}
Part (a) is analogous to the result in \citet[Lemma 3.ii]{bai2020simpler}. According to this result and Propositions \ref{prop:L2} and \ref{prop:F2}, the loadings and the factors are consistently estimated up to: (i) $(\bm\Gamma^F)^{1/2}$, (ii) a rotation, $\bm{\Upsilon}_0$, and (iii) a diagonal matrix of signs $\bm{\mathcal J}_0$.
Proposition \ref{cor:semplice} can be proved either by directly combining Propositions \ref{prop:LLFF}(d) and \ref{prop:KKKbis}, or from Proposition \ref{prop:HHAT}, by noticing that Assumptions \ref{ass:common}(a) and \ref{ass:common}(c-ii) imply $\Vert(T^{-1}\bm F^\prime\bm F)^{1/2}(n^{-1}\bm \Lambda^\prime\bm \Lambda) (T^{-1}\bm F^\prime\bm F)^{1/2}-(\bm\Gamma^F)^{1/2} \bm\Sigma_\Lambda(\bm\Gamma^F)^{1/2}\Vert =o_{\mathrm P}(1)$, and, therefore, by Davis Kahan theorem \citep[Corollary 1]{yu15}, the corresponding normalized eigenvectors satisfy:
$\Vert\widehat{\mathbf Q}-\bm\Upsilon_0\bm{\mathcal J}_0\Vert=o_{\mathrm P}(1)$.
Proposition \ref{prop:HHAT} and \ref{cor:semplice} convey the same message. The difference is in the way we approximate $\widehat{\mathbf H}$. In Proposition \ref{prop:HHAT} the rate is $\min(n,\sqrt{T})$, however, the limiting quantity is still random and depends on $n$ and $T$. In Proposition \ref{cor:semplice} we obtain a deterministic limiting quantity independent of $n$ and $T$, but the rate is slower and based on the rates at which Assumptions \ref{ass:common}(a) and \ref{ass:common}(c-ii) are satisfied, which are, obviously, $\sqrt n$ and $\sqrt T$, respectively.
We can now compare $\bm{\mathcal H}$ and $\widehat{\mathbf H}$ by means of the results in Propositions \ref{corol:ovvio} and \ref{prop:LLFF}.
The main difference is that while $\bm{\mathcal H}$ depends only on population quantities (but for an irrelevant sign indeterminacy), $\widehat{\mathbf H}$ depends both on population and estimated quantities.
Such difference has a series of important consequences: (i) $\widehat{\mathbf H}$ is positive definite and finite only asymptotically as $n,T\to\infty$ (see Lemma \ref{lem:HO1}), while $\bm{\mathcal H}$ is always positive definite and finite (see Lemma \ref{lem:HO1bis}); (ii) $\widehat{\mathbf H}$ can be interpreted only asymptotically (see Propositions \ref{prop:LLFF}, \ref{prop:HHAT}, and \ref{cor:semplice}), while $\bm{\mathcal H}$ has a well defined expression for all $n$ and $T$ as given in Proposition \ref{prop:L}; (iii) any identification restriction imposed on the loadings and/or the factors will constrain $\widehat{\mathbf H}$ only asymptotically, while we can derive exact expressions for $\bm{\mathcal H}$ (see Section \ref{sec:II}).
Still, $\bm{\mathcal H}$ and $\widehat{\mathbf H}$ are asymptotically equivalent. This can be seen in two ways. First, by comparing Proposition \ref{corol:ovvio}(a) and Proposition \ref{prop:LLFF}(a),
it follows that
\begin{align}
\left\Vert \bm{\mathcal H}-\widehat{\mathbf H}\right\Vert &\le
\left\Vert\bm{\mathcal H} - (\bm\Lambda^\prime\bm\Lambda)^{-1}\bm\Lambda^\prime\widehat{\bm\Lambda} \right\Vert
+
\left\Vert \left(\bm\Lambda^\prime\bm\Lambda\right)^{-1}
{\bm\Lambda}^\prime\widehat{\bm\Lambda}- \widehat{\mathbf H}
\right\Vert
\nonumber\\
&= O_{\mathrm P}\left(\max\left(\frac 1n,\frac 1{\sqrt T}\right)\right)+O_{\mathrm P}\left(\max\left(\frac 1n,\frac 1{\sqrt{n T}},\frac 1T\right)\right) = O_{\mathrm P}\left(\max\left(\frac 1n,\frac 1{\sqrt T}\right)\right).\label{eq:accheeq}
\end{align}
It is important to stress that $\bm{\mathcal H}$ approximates the projection matrix of the estimated loadings onto the true loadings at a slower rate than $\widehat{\mathbf H}$. Indeed, while to prove Proposition \ref{prop:LLFF}(a) we are able to use a tighter bound than the one obtained for the loadings in Proposition \ref{prop:L2}(b), by bounding $ n^{-1}\Vert\bm\Lambda^\prime(\widehat{\bm\Lambda}-\bm\Lambda\widehat{\mathbf H})\Vert$, we already noticed that to prove Proposition \ref{corol:ovvio}(a), we can only use a looser bound, which therefore inherits the weaker bound obtained for the loadings in Proposition \ref{prop:L}(a). This result has important implications when it comes to discussing the behaviour of $\bm{\mathcal H}$ and $\widehat{\mathbf H}$ under various identification conditions (see Section \ref{sec:II}).
Second, since the columns of $\mathbf K$ in the definition of $\bm{\mathcal H}$ (see Proposition \ref{prop:L}) are the normalized eigenvectors of $(\bm\Gamma^F)^{1/2} (n^{-1}\bm\Lambda^\prime\bm\Lambda) (\bm\Gamma^F)^{1/2}$, then, by Assumption \ref{ass:common}(a) and Davis Kahan theorem \citep[Corollary 1]{yu15}
it must hold that $\lim_{n\to\infty}\Vert \mathbf K-\bm\Upsilon_0\bm{\mathcal J}^*\Vert =0$ for some $r\times r$ diagonal matrix $\bm{\mathcal J}^*$ with entries $\pm 1$ and independent of $n$, and letting $\bm{\mathcal J}_0$ be such that $\Vert\bm{\mathcal J}^*\mathbf J -\bm{\mathcal J}_0\Vert=o_{\mathrm P}(1)$ as $n,T\to\infty$, we get $\Vert \mathbf K\mathbf J-\bm\Upsilon_0\bm{\mathcal J}_0\Vert=o_{\mathrm P}(1)$ (see \eqref{eq:ups000} and \eqref{eq:ups} in the proof of Proposition \ref{prop:KKK}).
As a consequence of this result, of the definition of $\bm{\mathcal H}$ in Proposition \ref{prop:L}, and of Proposition \ref{cor:semplice}(a), it immediately follows that
\begin{align}
\left\Vert \bm{\mathcal H}-\widehat{\mathbf H}\right\Vert &\le \left\Vert \bm{\mathcal H}-(\bm\Gamma^F)^{1/2}\bm\Upsilon_0\bm{\mathcal J}_0\right\Vert +\left\Vert (\bm\Gamma^F)^{1/2}\bm\Upsilon_0\bm{\mathcal J}_0-\widehat{\mathbf H}\right\Vert = o_{\mathrm P}(1),\label{eq:accheeq2}
\end{align}
thus, providing a proof of the asymptotic equivalence of $\bm{\mathcal H}$ and $\widehat{\mathbf H}$ alternative to the proof in \eqref{eq:accheeq}. This proof, however, relies on a looser bound than the one used in \eqref{eq:accheeq}, indeed, it would deliver the slower rate $\min(\sqrt n,\sqrt T)$ inherited from Proposition \ref{cor:semplice}.
\section{Asymptotic normality}\label{sec:AN}
In this section we prove asymptotic normality of the PC estimators. In particular, Theorems \ref{th:CLTL} and \ref{th:CLTF} proved below directly show the relationship with the unfeasible OLS estimators.
The analogous results given in \citet[Theorem 2 and 1, respectively]{Bai03} are not so intuitive, but more intuitive asymptotic expansions, equivalent to those derived in this section, can be easily derived (see Section \ref{sec:cmppca} for details).
It is crucial to stress that, in order to prove asymptotic normality without need of imposing too restrictive bounds on the rates of divergence of $n$ and $T$, we have to make use of the rates derived in Propositions \ref{prop:L2} and \ref{prop:F2}, which are sharper than those derived in Proposition \ref{prop:L}. This, however, requires the use of the matrix $\widehat{\mathbf H}$ rather than $\bm{\mathcal H}$.
Asymptotic normality of the PC estimator of the loadings follows from Lemma \ref{prop:load}.
\begin{theorem}\label{th:CLTL}
Under Assumptions \ref{ass:common} through \ref{ass:rates}, as $n,T\to\infty$, for any given $i=1,\ldots,n$,
$$
\sqrt T(\widehat{\bm\lambda}_i-\widehat{\mathbf H}^\prime{\bm\lambda}_i)
\to_d\mathcal N\left(\mathbf 0_r,
\bm\Upsilon_0^\prime(\bm\Gamma^F)^{1/2}\bm\Theta_i^{\text{\tiny \upshape OLS}} (\bm\Gamma^F)^{1/2} \bm\Upsilon_0
\right),
$$
with $\bm\Upsilon_0$ defined in \eqref{eq:U0V0U0} and $\bm\Theta_i^{\text{\tiny \upshape OLS}}:=(\bm\Gamma^F)^{-1}
\bm\Phi_i
(\bm\Gamma^F)^{-1}$ with $\bm\Phi_i$ defined in Assumption \ref{ass:CLT}(a).
\end{theorem}
\noindent
\textbf{Proof of Theorem \ref{th:CLTL}.}
From the transposed of \eqref{eq:sviluppoLambda} for any $i=1,\ldots, n$, we get
\begin{align}
\sqrt T(\widehat{\bm\lambda}_i-\widehat{\mathbf H}^\prime{\bm\lambda}_i )
&=\sqrt T \left(\frac{\widehat{\mathbf M}^x}{n}\right)^{-1} \bm{\mathcal H}^\prime \left\{\text{(1.a)+(1.b)+(1.c)}\right\}+ \sqrt T \left(\frac{\widehat{\mathbf M}^x}{n}\right)^{-1} \left\{\text{(1.d)+(1.e)+(1.f)}\right\}\nonumber\\
&=\sqrt T\left(\frac{\widehat{\mathbf M}^x}{n}\right)^{-1} \bm{\mathcal H}^\prime\text{(1.b)} + O_{\mathrm {P}}\left(\max\left(\frac {\sqrt T} n, \frac 1{\sqrt{n}},\frac 1{\sqrt T}\right)\right) \nonumber\\
&= \left(\frac{\widehat{\mathbf M}^x}{n}\right)^{-1}\frac{ \bm{\mathcal H}^\prime \bm\Lambda^\prime\bm\Lambda}{n} \left(\frac 1{\sqrt T}\sum_{t=1}^T \mathbf F_t e_{it}\right)
+ O_{\mathrm {P}}\left(\max\left(\frac {\sqrt T} n, \frac 1{\sqrt{n}},\frac 1{\sqrt T}\right)\right)\nonumber\\
&= \left(\frac{\widehat{\mathbf M}^x}{n}\right)^{-1}\frac{ \widehat{ \bm\Lambda}^\prime\bm\Lambda}{n} \left(\frac 1{\sqrt T}\sum_{t=1}^T \mathbf F_t e_{it}\right)
+ O_{\mathrm {P}}\left(\max\left(\frac {\sqrt T} n, \frac 1{\sqrt{n}},\frac 1{\sqrt T}\right)\right)\nonumber\\
&=\widehat{\mathbf H}^\prime \left(\frac{\bm F^\prime\bm F}{T}\right)^{-1} \left(\frac 1{\sqrt T}\sum_{t=1}^T \mathbf F_t e_{it}\right)
+ O_{\mathrm {P}}\left(\max\left(\frac {\sqrt T} n, \frac 1{\sqrt{n}},\frac 1{\sqrt T}\right)\right),\label{eq:finaleL}
\end{align}
because of Proposition \ref{corol:ovvio}(a) and Lemma \ref{prop:load}, the definition of $\widehat{\mathbf H}$ in \eqref{eq:acca}, and since
$\Vert n ({\widehat{\mathbf M}^x})^{-1}\Vert=O_{\mathrm P}(1)$
and
$\Vert\bm{\mathcal H}\Vert=O(1)$ by Lemmas \ref{lem:MO1}(iv) and \ref{lem:HO1bis}(i), respectively.
Define
\[
\widehat{\bm\lambda}_i^{\text{\tiny OLS}} := \left(\frac{\bm F^\prime\bm F}{T}\right)^{-1} \left(\frac 1{\sqrt T}\sum_{t=1}^T \mathbf F_t x_{it}\right),
\]
which is the unfeasible OLS estimator of $\bm\lambda_i$ when regressing $x_{it}$ onto $\mathbf F_t$.
Then, since by Assumption \ref{ass:rates} $\sqrt T/n\to 0$ as $n,T\to\infty$, from \eqref{eq:finaleL}, using Proposition \ref{cor:semplice}(a),
\begin{align}\label{eq:CLTL}
\sqrt T(\widehat{\bm\lambda}_i-\widehat{\mathbf H}^\prime{\bm\lambda}_i)
&= \left\{ \underset{n,T\to\infty}{\text{P-lim}} \widehat{\mathbf H}^\prime \right\} \left(\frac{\bm F^\prime\bm F}{T}\right)^{-1} \left(\frac 1{\sqrt T}\sum_{t=1}^T \mathbf F_t e_{it}\right)
+ o_{\mathrm {P}}(1)\nonumber\\
&=
\bm{\mathcal J}_0\bm\Upsilon_0^\prime (\bm\Gamma^F)^{1/2} \left(\frac{\bm F^\prime\bm F}{T}\right)^{-1} \left(\frac 1{\sqrt T}\sum_{t=1}^T \mathbf F_t e_{it}\right)+o_{\mathrm P}(1)\nonumber\\
&=
\bm{\mathcal J}_0\bm\Upsilon_0^\prime (\bm\Gamma^F)^{1/2} \sqrt T(\widehat{\bm\lambda}_i^{\text{\tiny OLS}}-{\bm\lambda}_i) +o_{\mathrm P}(1).
\end{align}
Moreover, as $T\to\infty$,
\begin{equation}\label{eq:thetaOLS}
\sqrt T(\widehat{\bm\lambda}_i^{\text{\tiny OLS}}-{\bm\lambda}_i) =(\bm\Gamma^F)^{-1} \left(\frac 1{\sqrt T}\sum_{t=1}^T \mathbf F_t e_{it}\right)+o_{\mathrm P}(1)
\to_d\mathcal N\left(\mathbf 0_r, (\bm\Gamma^F)^{-1}
\bm\Phi_i
(\bm\Gamma^F)^{-1}
\right),
\end{equation}
by Assumptions \ref{ass:common}(b), \ref{ass:common}(c-ii), and \ref{ass:CLT}(a), and Slutsky's theorem.
The proof follows by substituting \eqref{eq:thetaOLS} into \eqref{eq:CLTL}, using again Slutsky's theorem, and since $\bm{\mathcal J}_0$ being a diagonal sign matrix plays no role in the asymptotic covariance. $\Box$\\
Moving to the PC estimator of the factors, asymptotic normality follows from Lemma \ref{prop:factor}.
\begin{theorem}\label{th:CLTF}
Under Assumptions \ref{ass:common} through \ref{ass:rates}, as $n,T\to\infty$, for any given $t=1,\ldots,T$,
$$
\sqrt n(\widehat{\mathbf F}_t-\widehat{\mathbf H}^{-1}{\mathbf F}_t) \to_d\mathcal N\left(\mathbf 0_r, \bm\Upsilon_0^\prime(\bm\Gamma^F)^{-1/2}\bm\Pi_t^{\text{\tiny \upshape OLS}}(\bm\Gamma^F)^{-1/2}\bm\Upsilon_0\right),
$$
with $\bm\Upsilon_0$ defined in \eqref{eq:U0V0U0}
and $\bm\Pi_t^{\text{\tiny \upshape OLS}}:=(\bm\Sigma_\Lambda)^{-1}\bm\Gamma_t(\bm\Sigma_\Lambda)^{-1}$ with $\bm\Gamma_t$ defined in Assumption \ref{ass:CLT}(b).
\end{theorem}
\noindent
\textbf{Proof of Theorem \ref{th:CLTF}.}
From the transposed of \eqref{eq:sviluppoFactor}, for any $t=1,\ldots, T$, we get
\begin{align}
\sqrt n(\widehat{\mathbf F}_t-\widehat{\mathbf H}^{-1}{\mathbf F}_t )
&=\sqrt n \left(\frac{\widehat{\mathbf M}^x}{n}\right)^{-1} \left\{\text{(2.a)+(2.b)+(2.c)}\right\}\nonumber\\
&=\sqrt n \left(\frac{\widehat{\mathbf M}^x}{n}\right)^{-1}\text{(2.c)}+ O_{\mathrm {P}}\left(\max\left(\frac 1{\sqrt n},\frac {\sqrt n} T, \frac 1{\sqrt{T}}\right)\right)\nonumber\\
&= \left(\frac{\widehat{\mathbf M}^x}{n}\right)^{-1} \widehat{\mathbf H}^\prime \left(\frac 1{\sqrt n}\sum_{i=1}^n \bm \lambda_i e_{it}\right)
+ O_{\mathrm {P}}\left(\max\left(\frac 1{\sqrt n},\frac {\sqrt n} T, \frac 1{\sqrt{T}}\right)\right)\nonumber\\
&=\left(\frac{\widehat{\bm\Lambda}^\prime\widehat{\bm\Lambda}}{n}\right)^{-1} \widehat{\mathbf H}^\prime \left(\frac 1{\sqrt n}\sum_{i=1}^n \bm\lambda_i e_{it}\right)+ O_{\mathrm {P}}\left(\max\left(\frac 1{\sqrt n},\frac {\sqrt n} T, \frac 1{\sqrt{T}}\right)\right)\nonumber\\
&= \left(\frac{\widehat{\mathbf H}^\prime{\bm\Lambda}^\prime{\bm\Lambda}\widehat{\mathbf H}}{n}\right)^{-1} \widehat{\mathbf H}^\prime \left(\frac 1{\sqrt n}\sum_{i=1}^n \bm\lambda_i e_{it}\right)+ O_{\mathrm {P}}\left(\max\left(\frac 1{\sqrt n},\frac {\sqrt n} T, \frac 1{\sqrt{T}}\right)\right)\nonumber\\
& = \widehat{\mathbf H}^{-1} \left(\frac{\bm\Lambda^\prime\bm\Lambda}{n}\right)^{-1}\left(\frac 1{\sqrt n}\sum_{i=1}^n \bm\lambda_i e_{it}\right)+ O_{\mathrm {P}}\left(\max\left(\frac 1{\sqrt n},\frac {\sqrt n} T, \frac 1{\sqrt{T}}\right)\right),
\label{eq:finaleF}
\end{align}
because of Proposition \ref{prop:L2}(b) and Lemma \ref{prop:factor},
the definition of $\widehat{\bm\Lambda}$ in \eqref{eq:estL}, and since $\Vert n({\widehat{\mathbf M}^x})^{-1}\Vert=O_{\mathrm P}(1)$ by Lemma \ref{lem:MO1}(iv).
Define,
\[
\widehat{\mathbf F}_t^{\text{\tiny OLS}}:=\left(\frac{\bm\Lambda^\prime\bm\Lambda}{n}\right)^{-1}\left(\frac 1{\sqrt n}\sum_{i=1}^n \bm\lambda_i x_{it}\right),
\]
which is the unfeasible OLS estimator of $\mathbf F_t$ when regressing $x_{it}$ onto $\bm\lambda_i$. Then, since by Assumption \ref{ass:rates} $\sqrt n/T\to 0$ as $n,T\to\infty$, from \eqref{eq:finaleF}, using Proposition \ref{cor:semplice}(b),
\begin{align}\label{eq:CLTF}
\sqrt n(\widehat{\mathbf F}_t-\widehat{\mathbf H}^{-1}{\mathbf F}_t)
& = \left\{ \underset{n,T\to\infty}{\text{P-lim}} \widehat{\mathbf H}^{-1} \right\}
\left(\frac{\bm\Lambda^\prime\bm\Lambda}{n}\right)^{-1}\left(\frac 1{\sqrt n}\sum_{i=1}^n \bm\lambda_i e_{it}\right)+o_{\mathrm P}(1)\nonumber\\
&= \bm{\mathcal J}_0\bm\Upsilon_0^\prime (\bm\Gamma^F)^{-1/2} \left(\frac{\bm\Lambda^\prime\bm\Lambda}{n}\right)^{-1}\left(\frac 1{\sqrt n}\sum_{i=1}^n \bm\lambda_i e_{it}\right)+o_{\mathrm P}(1)\nonumber\\
&= \bm{\mathcal J}_0\bm\Upsilon_0^\prime (\bm\Gamma^F)^{-1/2} \sqrt n(\widehat{\mathbf F}_t^{\text{\tiny OLS}}-\mathbf F_t)+ o_{\mathrm P}(1).
\end{align}
Moreover, as $n\to\infty$,
\begin{equation}\label{eq:gammaOLS}
\sqrt n(\widehat{\mathbf F}_t^{\text{\tiny OLS}}-\mathbf F_t) = (\bm\Sigma_\Lambda)^{-1} \left(\frac 1{\sqrt n}\sum_{i=1}^n \bm\lambda_i e_{it}\right)
+o(1)
\to_d\mathcal N\left(\mathbf 0_r, (\bm\Sigma_\Lambda)^{-1}
\bm\Gamma_t
(\bm\Sigma_\Lambda)^{-1}
\right),
\end{equation}
by Assumptions \ref{ass:common}(a) and \ref{ass:CLT}(b), and Slutsky's theorem.
The proof follows by substituting \eqref{eq:gammaOLS} into \eqref{eq:CLTF}, using again Slutsky's theorem, and since $\bm{\mathcal J}_0$ being a diagonal sign matrix plays no role in the asymptotic covariance. $\Box$\\
It is important to stress that Theorems \ref{th:CLTL} and \ref{th:CLTF}, as well as the analogous theorems in \citet{Bai03}, cannot be used to make inference on the loadings or the factors. This is due to the presence of the unknown matrix $\widehat{\mathbf H}$, which depends on the unknown true factors and loadings. Notice that even when $r=1$, so that $\widehat{\mathbf H}$ is a scalar, still factors and loadings are in general consistently estimated only up to an unknown scale.
In particular, the presence of $\widehat{\mathbf H}$ has two effects. First, the location of the asymptotic distribution of $\widehat{\bm\lambda}_i$ and $\widehat{\mathbf F}_t$ are random.
Second, the asymptotic covariance matrix cannot be estimated consistently, since any such estimator would require consistent estimators of $\bm\Gamma^F$ and/or $\bm\Sigma_\Lambda$, but these cannot be estimated consistently. Indeed, the candidate estimators $T^{-1}\widehat{\bm F}^\prime \widehat{\bm F}$ and $n^{-1}\widehat{\bm \Lambda}^\prime \widehat{\bm \Lambda}$ have probability limits which still depend on $\widehat{\mathbf H}$ (see Proposition \ref{prop:LLFF}(g) and \ref{prop:LLFF}(h)).
\begin{sidewaystable}[htbp]
\caption{Asymptotic normality}\label{tab:H}
\centering
\scriptsize{
\begin{tabular}{l | l | c | c | c }
\hline
\hline
&&&&\\
\eqref{eq:CLTL} $\sqrt T(\widehat{\bm\lambda}_i-\widehat{\mathbf H}^\prime{\bm\lambda}_i)\to_d\mathcal N\left(\mathbf 0_r,\mathbf H_{0,A}^\prime \bm\Theta_i^{\text{OLS}}\mathbf H_{0,A}\right) $&&&&\\
\eqref{eq:ANLbis} $\sqrt T(\widehat{\bm\lambda}_i-\widehat{\mathbf H}^\prime{\bm\lambda}_i)\to_d\mathcal N\left(\mathbf 0_r,\mathbf H_{1,A}^\prime \bm\Theta_i^{\text{OLS}}\mathbf H_{1,A}\right) $&$\widehat{\mathbf H}$ & $\left(\frac{\bm F^\prime\bm F}{T}\right)^{1/2}\widehat{\mathbf Q}$ & $\mathbf H_{0,A}:=(\bm \Gamma^F)^{1/2} \bm\Upsilon_0 \bm{\mathcal J}_0$ & $\mathbf H_{1,A}:=(\bm\Sigma_\Lambda)^{-1/2}\bm\Upsilon_1 \bm{\mathcal J}_1\bm V_0^{1/2}$\\[3pt]
\eqref{eq:CLTF} $\sqrt n(\widehat{\mathbf F}_t-\widehat{\mathbf H}^{-1}{\mathbf F}_t)\to_d\mathcal N\left(\mathbf 0_r,\mathbf H_{0,A}^{-1} \bm\Pi_t^{\text{OLS}}(\mathbf H_{0,A}^{-1})^\prime\right)$& rate & $\min(n,\sqrt{T})$ & $\min(\sqrt n,\sqrt T)$ & $\min(\sqrt n,\sqrt T)$\\[3pt]
\eqref{eq:ANFbis} $\sqrt n(\widehat{\mathbf F}_t-\widehat{\mathbf H}^{-1}{\mathbf F}_t)\to_d\mathcal N\left(\mathbf 0_r,\mathbf H_{1,A}^{-1} \bm\Pi_t^{\text{OLS}}(\mathbf H_{1,A}^{-1})^\prime\right)$& &Proposition \ref{prop:HHAT} & Proposition \ref{cor:semplice} & \eqref{eq:ennesimaespansione}\\[3pt]
\hline
&&&&\\[-3pt]
\eqref{eq:ANLBAI1} $\sqrt T(\widetilde{\bm\lambda}_i-\widetilde{\mathbf H}^{-1}{\bm\lambda}_i)\to_d\mathcal N\left(\mathbf 0_r,\mathbf H_{0,B}^{-1}\bm\Theta_i^{\text{OLS}}(\mathbf H_{0,B}^{-1})^\prime\right)$&&&&\\
\eqref{eq:ANLBAI} $\sqrt T(\widetilde{\bm\lambda}_i-\widetilde{\mathbf H}^{-1}{\bm\lambda}_i)\to_d\mathcal N\left(\mathbf 0_r,\mathbf H_{1,B}^{-1}\bm\Theta_i^{\text{OLS}}(\mathbf H_{1,B}^{-1})^\prime\right)$&$\widetilde{\mathbf H}$ &$ \left(\frac{\bm F^\prime\bm F}{T}\right)^{-1/2} \widehat{\mathbf Q}$& $\mathbf H_{0,B}:=(\bm\Gamma^F)^{-1/2} \bm\Upsilon_0\bm{\mathcal J}_0$ & $\mathbf H_{1,B}:=(\bm\Sigma_\Lambda)^{1/2}\bm\Upsilon_1 \bm{\mathcal J}_1 \bm V_0^{-1/2}$\\[3pt]
\eqref{eq:ANFBAI1} $\sqrt n(\widetilde{\mathbf F}_t-\widetilde{\mathbf H}^{\prime}{\mathbf F}_t)\to_d\mathcal N\left(\mathbf 0_r,\mathbf H_{0,B}^\prime\bm\Pi_t^{\text{OLS}} \mathbf H_{0,B}\right)$& rate & $\min(n,\sqrt{T})$ & $\min(\sqrt n,\sqrt T)$ & $\min(\sqrt n,\sqrt T)$\\[3pt]
\eqref{eq:ANFBAI} $\sqrt n(\widetilde{\mathbf F}_t-\widetilde{\mathbf H}^{\prime}{\mathbf F}_t)\to_d\mathcal N\left(\mathbf 0_r,\mathbf H_{1,B}^\prime\bm\Pi_t^{\text{OLS}} \mathbf H_{1,B}\right)$&&\eqref{eq:sameQ}&\eqref{eq:Htildelim2}& \eqref{eq:Htildelim}\\[3pt]
\hline
&&&&\\[-3pt]
\eqref{eq:ANLSW1} $\sqrt T(\wideparen{\bm\lambda}_i-\wideparen{\mathbf H}^{\prime}{\bm\lambda}_i)\to_d\mathcal N\left(\mathbf 0_r,\mathbf H_{0,C}^{\prime}\bm\Theta_i^{\text{OLS}}\mathbf H_{0,C}\right)$&&&&\\
\eqref{eq:ANLSW2} $\sqrt T(\wideparen{\bm\lambda}_i-\wideparen{\mathbf H}^{\prime}{\bm\lambda}_i)\to_d\mathcal N\left(\mathbf 0_r,\mathbf H_{1,C}^{\prime}\bm\Theta_i^{\text{OLS}}\mathbf H_{1,C}\right)$&$\wideparen{\mathbf H}$ &$ \left(\frac{\bm\Lambda^\prime\bm\Lambda}{n}\right)^{-1/2} \wideparen {\mathbf Q}$& $\mathbf H_{0,C}:=(\bm\Gamma^F)^{1/2} \bm\Upsilon_0\bm{\mathcal J}_0\bm V_0^{-1/2}$ & $\mathbf H_{1,C}:=(\bm\Sigma_\Lambda)^{-1/2} \bm\Upsilon_1\bm{\mathcal J}_1$\\[3pt]
\eqref{eq:ANFSW1} $\sqrt n(\wideparen{\mathbf F}_t-\wideparen{\mathbf H}^{-1}{\mathbf F}_t ) \to_d\mathcal N\left(\mathbf 0_r, \mathbf H_{0,C}^{-1}\bm\Pi_t^{\text{\tiny \upshape OLS}}(\mathbf H_{0,C}^{-1})^\prime\right)$& rate & $\min(n,\sqrt{T})$ & $\min(\sqrt n,\sqrt T)$ & $\min(\sqrt n,\sqrt T)$\\[3pt]
\eqref{eq:ANFSW2} $\sqrt n(\wideparen{\mathbf F}_t-\wideparen{\mathbf H}^{-1}{\mathbf F}_t ) \to_d\mathcal N\left(\mathbf 0_r, \mathbf H_{1,C}^{-1}\bm\Pi_t^{\text{\tiny \upshape OLS}}(\mathbf H_{1,C}^{-1})^\prime\right)$&&\eqref{eq:sameQsw}&\eqref{eq:semplicesw2}& \eqref{eq:semplicesw}\\[3pt]
\hline
&&&&\\[-3pt]
\eqref{eq:ANLD1} $\sqrt T(\bar{\bm\lambda}_i-\bar{\mathbf H}^{-1}{\bm\lambda}_i)\to_d\mathcal N\left(\mathbf 0_r,\mathbf H_{0,D}^{-1}\bm\Theta_i^{\text{OLS}}(\mathbf H_{0,D}^{-1})^\prime\right)$
&&&&\\
\eqref{eq:ANLD2} $\sqrt T(\bar{\bm\lambda}_i-\bar{\mathbf H}^{\prime}{\bm\lambda}_i)\to_d\mathcal N\left(\mathbf 0_r,\mathbf H_{1,D}^{-1}\bm\Theta_i^{\text{OLS}}(\mathbf H_{1,D}^{-1})^\prime\right)$&$\bar{\mathbf H}$ &$ \left(\frac{\bm\Lambda^\prime\bm\Lambda}{n}\right)^{1/2} \wideparen {\mathbf Q}$& $\mathbf H_{0,D}:=(\bm\Gamma^F)^{-1/2} \bm\Upsilon_0\bm{\mathcal J}_0\bm V_0^{1/2}$ & $\mathbf H_{1,D}:=(\bm\Sigma_\Lambda)^{1/2} \bm\Upsilon_1\bm{\mathcal J}_1$\\[3pt]
\eqref{eq:ANFD1} $\sqrt n(\bar{\mathbf F}_t-\bar{\mathbf H}^{\prime}{\mathbf F}_t ) \to_d\mathcal N\left(\mathbf 0_r, \mathbf H_{0,D}^{\prime}\bm\Pi_t^{\text{\tiny \upshape OLS}}\mathbf H_{0,D}\right)$& rate & $\min(n,\sqrt{T})$ & $\min(\sqrt n,\sqrt T)$ & $\min(\sqrt n,\sqrt T)$\\[3pt]
\eqref{eq:ANFD2} $\sqrt n(\bar{\mathbf F}_t-\bar{\mathbf H}^{\prime}{\mathbf F}_t ) \to_d\mathcal N\left(\mathbf 0_r, \mathbf H_{1,D}^{\prime}\bm\Pi_t^{\text{\tiny \upshape OLS}}\mathbf H_{1,D}\right)$
&&\eqref{eq:sameQswD}&\eqref{eq:sempliceD2}& \eqref{eq:sempliceD}\\
&&&&\\
\hline
\hline
\end{tabular}
\begin{tabular}{p{.8\textwidth}}
$\widehat{\mathbf Q}$ are normalized eigenvectors of $(T^{-1}\bm F^\prime\bm F)^{1/2}(n^{-1}\bm \Lambda^\prime\bm \Lambda) (T^{-1}\bm F^\prime\bm F)^{1/2}$; \\
$\wideparen {\mathbf Q}$ are the normalized eigenvectors of $(n^{-1}\bm \Lambda^\prime\bm \Lambda)^{1/2} (T^{-1}\bm F^\prime\bm F) (n^{-1}\bm \Lambda^\prime\bm \Lambda)^{1/2}$; \\
$\bm\Upsilon_0$ are normalized eigenvectors of $(\bm\Gamma^F)^{1/2}\bm\Sigma_\Lambda(\bm\Gamma^F)^{1/2}$; $\bm\Upsilon_1$ are normalized eigenvectors of $(\bm\Sigma_\Lambda)^{1/2}\bm\Gamma^F(\bm\Sigma_\Lambda)^{1/2}$;\\
$\bm V_0=\lim_{n\to\infty} n^{-1}\mathbf M^{C}$ are eigenvalues of $(\bm\Gamma^F)^{1/2}\bm\Sigma_\Lambda(\bm\Gamma^F)^{1/2}$ and $(\bm\Sigma_\Lambda)^{1/2}\bm\Gamma^F(\bm\Sigma_\Lambda)^{1/2}$;\\
$\bm{\mathcal J}_0$ and $\bm{\mathcal J}_1$ are diagonal with entries $\pm 1$.
\end{tabular}
}
\end{sidewaystable}
\section{Equivalence with approach B1} \label{sec:cmppca}
The PC estimators studied so far coincide with those studied by \citet{Bai03}, indeed, by \eqref{eq:stessi}, $\widetilde{\bm\Lambda}=\widehat{\bm\Lambda}$ and $\widetilde{\bm F}=\widehat{\bm F}$. The purpose of this section is then to show that, by following the proofs of \citet{Bai03} we can derive the same asymptotic expansions as in Theorems \ref{th:CLTL} and \ref{th:CLTF}, showing the relation between the PC and the unfeasible OLS estimators. These are more interpretable than the original results.
Notice that, by virtue of Table \ref{tab:ass}, the results in \citet{Bai03} can be proved under the same assumptions made in this paper so there is no need to prove them again here.\footnote{Although the estimators are identical, in this section we keep using the notation $\widetilde{\bm \Lambda}$ and $\widetilde{\bm F}$ to highlight that their properties are derived using a different approach with respect to the one used so far.}
\paragraph{Consistency.}
First of all notice that Proposition \ref{prop:L} still holds, and its proof using $\widetilde{\bm\Lambda}$ and $\widetilde{\bm F}$ is a special case of the proof by \citet[Theorem 4]{FLM13}, when no uniform bounds are computed.\footnote{Note that the proof by \citet{FLM13} follows the same steps as in \citet{Bai03}, but it derives slower rates since fewer cross-moment assumptions are imposed.}
Now, by definition of eigenvectors, it holds that
\begin{equation}\label{eq:EVECBAI}
\frac{\bm X\bm X^\prime}{nT}\widetilde{\bm F}=\widetilde{\bm F}\frac{\widetilde{\mathbf M}^x}{T}.
\end{equation}
By using $\bm X=\bm F\bm\Lambda^\prime+\bm E$ and taking the $t$th row of \eqref{eq:EVECBAI}, for any $t=1,\ldots,T$,
\begin{align}
\widetilde{\mathbf F}_t^\prime - \mathbf F_t^\prime \widetilde{\mathbf H} &= \left(
\underbrace{\frac 1{nT} {\mathbf F}_t^\prime \sum_{i=1}^n \sum_{s=1}^T \bm\lambda_i e_{is} \widetilde{\mathbf F}_s^\prime}_{\text{(1.a)}}
+
\underbrace{\frac 1{nT} \sum_{i=1}^n e_{it}\bm\lambda_i^\prime \sum_{s=1}^T {\mathbf F}_s \widetilde{\mathbf F}_s^\prime}_{\text{(1.b)}}
+\underbrace{\frac 1{nT} \sum_{i=1}^n \sum_{s=1}^T e_{is} e_{it}\widetilde{\mathbf F}_s^\prime}_{\text{(1.c)}}
\right) \left(\frac{\widetilde {\mathbf M}^x}{T}\right)^{-1},\label{eq:sviluppoLambdaBAI}
\end{align}
which is the expansion given in \citet[Equation (A.1)]{Bai03} and where
\begin{equation}\label{eq:Htilde}
\widetilde{\mathbf H}:=\frac{\bm\Lambda^\prime\bm\Lambda}{n}\frac{\bm F^\prime\widetilde{\bm F}}{T} \left(\frac{\widetilde {\mathbf M}^x}{T}\right)^{-1}.
\end{equation}
Notice that $\Vert\widetilde{\mathbf H}\Vert= O_{\mathrm P}(1)$ and $\Vert\widetilde{\mathbf H}^{-1}\Vert= O_{\mathrm P}(1)$ because, as $n,T\to\infty$, all its terms tend to finite and positive definite quantities (see Assumption \ref{ass:common}(a),
Proposition \ref{prop:KKKbis}, and
Lemma \ref{lem:covarianze}(iii) jointly with Lemmas \ref{lem:Vzero}(i) and \ref{lem:Vzero}(iii)).
Then, from \citet[Lemma A.2]{Bai03},
$\Vert \text{(1.a)}\Vert= O_{\mathrm P}(\max(n^{-1},(nT)^{-1/2}) )$,
$\Vert \text{(1.b)}\Vert= O_{\mathrm P}( n^{-1/2})$, and
$\Vert \text{(1.c)}\Vert= O_{\mathrm P}(\max(n^{-1},(nT)^{-1/2}, T^{-1}) )$. Therefore, the estimated factors are consistent:
$$
\left\Vert \widehat{\mathbf F}_t-\widetilde{\mathbf H}^{\prime}{\mathbf F}_t \right\Vert = O_{\mathrm P} \left(\max\left(\frac 1{\sqrt n},\frac 1T\right)\right),
$$
which coincides with Proposition \ref{prop:F2}, but when stated using $\widetilde{\mathbf H}$ in place of $\widehat{\mathbf H}$.
From \citet[Equation before (B.2)]{Bai03}, for any $i=1,\ldots, n$,
\begin{align}
\widetilde{\bm\lambda}_i^\prime-{\bm\lambda}_i^\prime (\widetilde{\mathbf H}^{-1})^\prime &=
\underbrace{\bm\lambda_i^\prime\frac{(\bm F-\widetilde{\bm F} \widetilde{\mathbf H}^{-1})^\prime\widetilde{\bm F} }{T}}_{\text{(2.a)}}
+
\underbrace{\bm\varepsilon_i^\prime\frac{(\widetilde{\bm F}- \bm F\widetilde{\mathbf H} )}{T}}_{\text{(2.b)}}
+
\underbrace{\frac{\bm\varepsilon_i^\prime\bm F\widetilde{\mathbf H}}{T}}_{\text{(2.c)}},\label{eq:sviluppoFactorBAI}
\end{align}
which is the obtained from the linear projection of $\bm X^\prime$ onto the estimated factors $\widetilde{\bm F}$ in a similar way as we obtained the linear projection of $\bm X$ onto the estimated loadings $\widehat{\bm \Lambda}$ giving \eqref{eq:sviluppoFactor}. Then, from \citet[Lemmas B.3 and B.1]{Bai03},
$\Vert \text{(2.a)}\Vert= O_{\mathrm P}(\max(n^{-1},(nT)^{-1/2},T^{-1}) )$,
$\Vert \text{(2.b)}\Vert= O_{\mathrm P}(\max(n^{-1},(nT)^{-1/2},T^{-1}) )$,
and $\Vert \text{(2.c)}\Vert= O_{\mathrm P}(T^{-1/2} )$. Therefore, the estimated loadings are consistent:
$$
\left\Vert \widetilde{\bm\lambda}_i-\widetilde{\mathbf H}^{-1}{\bm\lambda}_i\right\Vert = O_{\mathrm P} \left(\max\left(\frac 1{ n},\frac 1{\sqrt T}\right)\right),
$$
which coincides with Proposition \ref{prop:L2}, but when stated using $\widetilde{\mathbf H}$ in place of $\widehat{\mathbf H}$.
\paragraph{The role of $ \widetilde{\mathbf H}$.} First, from \citet[Lemma B.3]{Bai03},
\begin{align}
&\left\Vert \widetilde{\mathbf H}- ({\bm F}^\prime {\bm F})^{-1}{\bm F}^\prime\widetilde{\bm F} \right\Vert = O_{\mathrm P}\left(\max\left(\frac 1n,\frac 1{\sqrt{nT}},\frac 1T\right)\right),\label{eq:prop9Htilde}
\end{align}
which is the analogous of Proposition \ref{prop:LLFF}(b). This immediately implies also the analogous of Propositions \ref{prop:LLFF}(d), \ref{prop:LLFF}(f), and \ref{prop:LLFF}(h).
Second, following the same reasoning used to prove Proposition \ref{prop:HHAT}, but this time using \eqref{eq:prop9Htilde}, it follows that
\begin{align}
&\left\Vert \frac{\widetilde{\mathbf M}^x}{T} - \widetilde{\mathbf H}^{-1} \frac{\bm\Lambda^\prime\bm\Lambda}{n} (\widetilde{\mathbf H}^{-1})^{\prime}\right\Vert = O_{\mathrm P}\left(\max\left(\frac 1n,\frac 1{\sqrt{nT}},\frac 1T\right)\right),\label{eq:starp1BAI}
\end{align}
from which it follows that we must have
\begin{align}
&\left\Vert
\widetilde{\mathbf H} - \left(\frac{\bm F^\prime\bm F}{T}\right)^{-1/2} \widehat{\mathbf Q}
\right\Vert= O_{\mathrm P}\left(\max\left(\frac 1n,\frac 1{\sqrt{T}}\right)\right),\label{eq:sameQ}
\\
&\left\Vert
\widetilde{\mathbf H}^{-1} -\widehat{\mathbf Q}^\prime \left(\frac{\bm F^\prime\bm F}{T}\right)^{1/2}
\right\Vert= O_{\mathrm P}\left(\max\left(\frac 1n,\frac 1{\sqrt{T}}\right)\right),\label{eq:sameQinv}
\end{align}
where $\widehat {\mathbf Q}$ are the normalized eigenvectors of $(T^{-1}\bm F^\prime\bm F)^{1/2}(n^{-1}\bm \Lambda^\prime\bm \Lambda) (T^{-1}\bm F^\prime\bm F)^{1/2}$, which has $n^{-1}\widehat{\mathbf M}^C$ as eigenvalues, and these, in turn, are such that $n^{-1}\Vert \widehat{\mathbf M}^C-\widehat{\mathbf M}^x\Vert = O_{\mathrm P}(\max(n^{-1},T^{-1/2}))$ with
$T^{-1}\widetilde{\mathbf M}^x=n^{-1}\widehat{\mathbf M}^x$. Results \eqref{eq:sameQ} and \eqref{eq:sameQinv} are the analogous of Proposition \ref{prop:HHAT}. Once again, we see that the loadings and the factors can be consistently estimated up to: (i) a scale $(T^{-1}{\bm F^\prime\bm F})^{-1/2}$ and (ii) a rotation $\widehat{\mathbf Q}$.
We can then derive a limit for $\widetilde{\mathbf H}$ which depends only on population quantities. This can be done in two ways.
First, by Assumptions \ref{ass:common}(a) and \ref{ass:common}(c-ii), we also have $\Vert(T^{-1}\bm F^\prime\bm F)^{1/2}(n^{-1}\bm \Lambda^\prime\bm \Lambda) (T^{-1}\bm F^\prime\bm F)^{1/2}-(\bm\Gamma^F)^{1/2} \bm\Sigma_\Lambda(\bm\Gamma^F)^{1/2}\Vert =o_{\mathrm P}(1)$. Therefore, by Davis Kahan theorem \citep[Corollary 1]{yu15}, the corresponding normalized eigenvectors satisfy:
$\Vert\widehat{\mathbf Q}-\bm\Upsilon_0\bm{\mathcal J}_0\Vert=o_{\mathrm P}(1)$, which, jointly with \eqref{eq:sameQ} and Assumption \ref{ass:common}(c-ii), implies
\begin{equation}
\left\Vert \widetilde{\mathbf H}- (\bm\Gamma^F)^{-1/2} \bm\Upsilon_0\bm{\mathcal J}_0\right\Vert = o_{\mathrm P}(1)
\;\text{ and }\;
\left\Vert \widetilde{\mathbf H}^{-1}- \bm{\mathcal J}_0 \bm\Upsilon_0^\prime (\bm\Gamma^F)^{1/2}\right\Vert = o_{\mathrm P}(1), \label{eq:Htildelim2}
\end{equation}
which is the analogous of Proposition \ref{cor:semplice}.
Second, consider the following spectral decomposition:
\begin{equation}\label{eq:U1V0U1}
(\bm\Sigma_\Lambda)^{1/2}\bm\Gamma^F(\bm\Sigma_\Lambda)^{1/2} =: \bm\Upsilon_1\bm V_0\bm\Upsilon_1^\prime,
\end{equation}
where $\bm\Upsilon_1$ is the $r\times r$ matrix having as columns the normalized eigenvectors of $(\bm\Sigma_\Lambda)^{1/2}\bm\Gamma^F(\bm\Sigma_\Lambda)^{1/2}$, and $\bm V_0$ is the $r\times r$ matrix of corresponding eigenvalues sorted in decreasing order, which coincide with those of
$(\bm\Gamma^F)^{1/2}\bm\Sigma_\Lambda ( \bm\Gamma^F)^{1/2}$, and are given by $\bm V_0=\lim_{n\to\infty} n^{-1}\mathbf M^{C}$.
Then, from \citet[Proposition 1]{Bai03}
\begin{equation}\label{eq:KKKBai}
\left\Vert \frac{\widetilde{\bm F}^\prime\bm F}{T}-\bm V_0^{1/2} \bm{\mathcal J}_1\bm\Upsilon_1^\prime (\bm\Sigma_\Lambda)^{-1/2}\right\Vert = o_{\mathrm P}(1),
\end{equation}
where $\bm{\mathcal J}_1$ is an $r \times r$ diagonal matrix with entries $\pm 1$.\footnote{Notice that the sign indeterminacy due to $\bm{\mathcal J}_1$ is not present in the original proof since the sign of the eigenvectors is implicitly fixed.} Notice that the limiting quantity in \eqref{eq:KKKBai} is finite and positive definite because of Assumption \ref{ass:common}(a) and Lemmas \ref{lem:Vzero}(i) and \ref{lem:Vzero}(iii).
From \eqref{eq:Htilde}, \eqref{eq:KKKBai}, Assumption \ref{ass:common}(a), and Lemma \ref{lem:Vzero}(iv),
\begin{equation}\label{eq:Htildelim}
\left\Vert \widetilde{\mathbf H}-(\bm\Sigma_\Lambda)^{1/2} \bm\Upsilon_1 \bm{\mathcal J}_1 \bm V_0^{-1/2} \right\Vert = o_{\mathrm P}(1) \;\text{ and }\; \left\Vert \widetilde{\mathbf H}^{-1}- \bm V_0^{1/2}\bm{\mathcal J}_1\bm\Upsilon_1^\prime (\bm\Sigma_\Lambda)^{-1/2} \right\Vert = o_{\mathrm P}(1).
\end{equation}
If Assumption \ref{ass:common}(a) holds with rate $\sqrt n$ and Assumption \ref{ass:common}(c-ii) holds with rate $\sqrt T$, then \eqref{eq:Htildelim2}, \eqref{eq:KKKBai}, and \eqref{eq:Htildelim} hold with rate $\min(\sqrt n,\sqrt T)$.
\paragraph{A new limit for $\widehat{\mathbf H}$.} From \eqref{eq:Htildelim2} and \eqref{eq:Htildelim} and uniqueness of the limits,
\begin{align}
&
(\bm\Sigma_\Lambda)^{1/2} \bm\Upsilon_1 \bm{\mathcal J}_1 \bm V_0^{-1/2}
=(\bm\Gamma^F)^{-1/2} \bm\Upsilon_0\bm{\mathcal J}_0
,\label{eq:riunioni2}\\
&
\bm V_0^{1/2}
\bm{\mathcal J}_1
\bm\Upsilon_1^\prime
(\bm\Sigma_\Lambda)^{-1/2}
=
\bm{\mathcal J}_0\bm\Upsilon_0^\prime(\bm\Gamma^F)^{1/2}
.\label{eq:riunioni}
\end{align}
And, from Proposition \ref{prop:KKK}(b) and \eqref{eq:riunioni2}, it follows that we also have
\begin{equation}\label{eq:KKKBaiLL}
\left\Vert \frac{\widehat{\bm \Lambda}^\prime\bm \Lambda}{n}-
\bm V_0^{1/2}\bm{\mathcal J}_1\bm\Upsilon_1^\prime (\bm\Sigma_\Lambda)^{1/2}\right\Vert = o_{\mathrm P}(1).
\end{equation}
Therefore, from \eqref{eq:acca}, \eqref{eq:KKKBaiLL}, Assumption \ref{ass:common}(c-ii), and Lemma \ref{lem:Vzero}(iv),
\begin{equation}
\left\Vert\widehat{\mathbf H}- \bm\Gamma^F(\bm\Sigma_\Lambda)^{1/2}\bm\Upsilon_1 \bm{\mathcal J}_1\bm V_0^{-1/2}\right\Vert= o_{\mathrm P}(1),
\label{eq:ennesimaespansioneAA}
\end{equation}
and, since from \eqref{eq:U1V0U1}, $\bm\Gamma^F(\bm\Sigma_\Lambda)^{1/2}=(\bm\Sigma_\Lambda)^{-1/2}\bm\Upsilon_1\bm V_0\bm\Upsilon_1^\prime$, from \eqref{eq:ennesimaespansioneAA} we obtain
\begin{equation}
\left\Vert\widehat{\mathbf H}- (\bm\Sigma_\Lambda)^{-1/2}\bm\Upsilon_1 \bm{\mathcal J}_1\bm V_0^{1/2}\right\Vert= o_{\mathrm P}(1)
\;\text{ and }\;
\left\Vert\widehat{\mathbf H}^{-1}- \bm V_0^{-1/2} \bm{\mathcal J}_1\bm\Upsilon_1^\prime(\bm\Sigma_\Lambda)^{1/2}\right\Vert= o_{\mathrm P}(1).
\label{eq:ennesimaespansione}
\end{equation}
Once again, if Assumption \ref{ass:common}(a) holds with rate $\sqrt n$ and Assumption \ref{ass:common}(c-ii) holds with rate $\sqrt T$, the rate in \eqref{eq:KKKBaiLL}-\eqref{eq:ennesimaespansione} is $\min(\sqrt n,\sqrt T)$.
\paragraph{The relation between $\widetilde{\mathbf H}$ and $\widehat{\mathbf H}$.}
Since $\widetilde{\bm F}=\widehat{\bm F}$ from Proposition \ref{prop:LLFF}(b) and \eqref{eq:prop9Htilde}, we see that we must have
\begin{align}
&\left\Vert\widetilde{\mathbf H}-(\widehat{\mathbf H}^{-1})^\prime\right\Vert = O_{\mathrm P}\left(\max\left(\frac 1n,\frac 1{\sqrt{nT}},\frac 1T\right)\right)\;\text{and}\;
&\left\Vert\widetilde{\mathbf H}^{-1}-\widehat{\mathbf H}^\prime \right\Vert = O_{\mathrm P}\left(\max\left(\frac 1n,\frac 1{\sqrt{nT}},\frac 1T\right)\right).\label{eq:sameH}
\end{align}
It follows that, by substituting \eqref{eq:sameH} into Proposition \ref{prop:LLFF}(a), we obtain its analogous:
\begin{align}
&\left\Vert \widetilde{\mathbf H}^{-1} - \widetilde{\bm\Lambda}^\prime{\bm\Lambda}(\bm\Lambda^\prime\bm\Lambda)^{-1}\right\Vert= O_{\mathrm P}\left(\max\left(\frac 1n,\frac 1{\sqrt{nT}},\frac 1T\right)\right),\label{eq:prop9HtildeL}
\end{align}
which immediately implies also the analogous of Propositions \ref{prop:LLFF}(c), \ref{prop:LLFF}(e), and \ref{prop:LLFF}(g).
\paragraph{Asymptotic normality.} From \eqref{eq:sviluppoLambdaBAI} and \citet[Theorem 1]{Bai03}, if $\sqrt n/T\to 0$, as $n,T\to\infty$,
\begin{align}
\sqrt n(\widetilde{\mathbf F}_t&-\widetilde{\mathbf H}^{\prime}{\mathbf F}_t) =\left(\frac{\widetilde {\mathbf M}^x}{T} \right)^{-1} \frac{\widetilde{\bm F}^\prime \bm F}{T}\left(\frac 1{\sqrt n}\sum_{i=1}^n {\bm\lambda}_i e_{it}\right)+o_{\mathrm P}(1)= \widetilde{\mathbf H}^\prime \left(\frac{\bm\Lambda^\prime\bm\Lambda}{n}\right)^{-1}\left(\frac 1{\sqrt n}\sum_{i=1}^n {\bm\lambda}_i e_{it}\right)+o_{\mathrm P}(1)\nonumber\\
& = \left\{ \underset{n,T\to\infty}{\text{P-lim}} \widetilde{\mathbf H}^\prime \right\}
\left(\frac{\bm\Lambda^\prime\bm\Lambda}{n}\right)^{-1}\left(\frac 1{\sqrt n}\sum_{i=1}^n \bm\lambda_i e_{it}\right)+o_{\mathrm P}(1)= \bm V_0^{-1/2}\bm{\mathcal J}_1\bm\Upsilon_1^\prime (\bm\Sigma_\Lambda)^{1/2} \sqrt n(\widehat{\mathbf F}_t^{\text{\tiny OLS}}-\mathbf F_t)+ o_{\mathrm P}(1)\nonumber\\
&= \bm{\mathcal J}_0\bm\Upsilon_0^\prime (\bm\Gamma^F)^{-1/2} \sqrt n(\widehat{\mathbf F}_t^{\text{\tiny OLS}}-\mathbf F_t)+ o_{\mathrm P}(1).
\label{eq:CLTFBAI}
\end{align}
where the first line is the only asymptotic expansion reported by \citet[Theorem 1]{Bai03}, and for the subsequent equalities we used also \eqref{eq:Htilde}, \eqref{eq:Htildelim2}, and \eqref{eq:Htildelim}, or \eqref{eq:riunioni2}. The last expression in \eqref{eq:CLTFBAI} coincides with \eqref{eq:CLTF} in the proof of Theorem \ref{th:CLTF}, hence, by Slutsky's theorem,
\begin{equation}\label{eq:ANFBAI1}
\sqrt n(\widetilde{\mathbf F}_t-\widetilde{\mathbf H}^{\prime}{\mathbf F}_t) \to_d\mathcal N\left(\mathbf 0_r, \bm\Upsilon_0^\prime (\bm\Gamma^F)^{-1/2}
\bm\Pi_t^{\text{\tiny \upshape OLS}}
(\bm\Gamma^F)^{-1/2}\bm\Upsilon_0
\right),
\end{equation}
Moreover, using the second last line of \eqref{eq:CLTFBAI}, by Slutsky's theorem, we also have
\begin{equation}\label{eq:ANFBAI}
\sqrt n(\widetilde{\mathbf F}_t-\widetilde{\mathbf H}^{\prime}{\mathbf F}_t) \to_d\mathcal N\left(\mathbf 0_r, \bm V_0^{-1/2}\bm\Upsilon_1^\prime (\bm\Sigma_\Lambda)^{1/2}\bm\Pi_t^{\text{\tiny \upshape OLS}}(\bm\Sigma_\Lambda)^{1/2}\bm\Upsilon_1 \bm V_0^{-1/2}
\right),
\end{equation}
where $\bm\Pi_t^{\text{\tiny \upshape OLS}}=(\bm\Sigma_\Lambda)^{-1}\bm\Gamma_t(\bm\Sigma_\Lambda)^{-1}$ with $\bm\Gamma_t$ is defined in Assumption \ref{ass:CLT}(b). Furthermore, letting $\bm{\mathcal Q}:=\bm V_0^{1/2}\bm\Upsilon_1^\prime(\bm\Sigma_\Lambda)^{-1/2}$ the asymptotic covariance matrix of $\widetilde{\mathbf F}_t$ can also be written as $(\bm V_0)^{-1}\bm {\mathcal Q}\bm\Gamma_t\bm {\mathcal Q}^\prime (\bm V_0)^{-1}$, which is the expression given in \citet[Theorem 1]{Bai03}. Notice that from \eqref{eq:ANFBAI} it follows that we can state Thoreom \ref{th:CLTF} also as
\begin{equation}\label{eq:ANFbis}
\sqrt n(\widehat{\mathbf F}_t-\widehat{\mathbf H}^{-1}{\mathbf F}_t) \to_d\mathcal N\left(\mathbf 0_r, \bm V_0^{-1/2}\bm\Upsilon_1^\prime (\bm\Sigma_\Lambda)^{1/2}\bm\Pi_t^{\text{\tiny \upshape OLS}}(\bm\Sigma_\Lambda)^{1/2}\bm\Upsilon_1 \bm V_0^{-1/2}
\right).
\end{equation}
Similarly, from \eqref{eq:sviluppoFactorBAI} and \citet[Theorem 2]{Bai03}, if $\sqrt T/n\to 0$, as $n,T\to\infty$,
\begin{align}
\sqrt T(\widetilde{\bm\lambda}_i&-\widetilde{\mathbf H}^{-1}{\bm\lambda}_i)
= \widetilde{\mathbf H}^\prime \left(\frac 1{\sqrt T}\sum_{t=1}^T \mathbf F_t e_{it}\right)+o_{\mathrm P}(1)=\left(\frac{\widetilde{\bm F}^\prime\widetilde{\bm F}}{T}\right)^{-1} \widetilde{\mathbf H}^\prime \left(\frac 1{\sqrt T}\sum_{t=1}^T \mathbf F_t e_{it}\right)+o_{\mathrm P}(1)\nonumber\\
&=\left(\frac{\widetilde{\mathbf H}^{\prime}{\bm F}^\prime{\bm F}\widetilde{\mathbf H}}{T}\right)^{-1} \widetilde{\mathbf H}^\prime \left(\frac 1{\sqrt T}\sum_{t=1}^T \mathbf F_t e_{it}\right)+o_{\mathrm P}(1)= \widetilde{\mathbf H}^{-1}\left(\frac{{\bm F}^\prime{\bm F}}{T}\right)^{-1}\left(\frac 1{\sqrt T}\sum_{t=1}^T \mathbf F_t e_{it}\right)+o_{\mathrm P}(1)\nonumber\\
&=\left\{ \underset{n,T\to\infty}{\text{P-lim}} \widetilde{\mathbf H}^{-1} \right\}\left(\frac{{\bm F}^\prime{\bm F}}{T}\right)^{-1}\left(\frac 1{\sqrt T}\sum_{t=1}^T \mathbf F_t e_{it}\right)+o_{\mathrm P}(1)= \bm V_0^{1/2}\bm{\mathcal J}_1\bm\Upsilon_1^\prime (\bm\Sigma_\Lambda)^{-1/2} \sqrt T(\widehat{\bm\lambda}_i^{\text{\tiny OLS}}-{\bm\lambda}_i) +o_{\mathrm P}(1)\nonumber\\
&= \bm{\mathcal J}_0\bm\Upsilon_0^\prime(\bm\Gamma^F)^{1/2}\sqrt T(\widehat{\bm\lambda}_i^{\text{\tiny OLS}}-{\bm\lambda}_i) +o_{\mathrm P}(1)
,\label{eq:CLTLBAI}
\end{align}
where the first line is the only asymptotic expansion reported by \citet[proof of Theorem 2]{Bai03}, and for the subsequent equalities we used also
the fact that $T^{-1}\widetilde{\mathbf F}^\prime\widetilde{\mathbf F}=\mathbf I_r$ by definition,
\eqref{eq:prop9Htilde} which implies $\Vert T^{-1}\widetilde{\bm F}^\prime\widetilde{\bm F}-T^{-1} \widetilde{\mathbf H}^{\prime}{\bm F}^\prime{\bm F}\widetilde{\mathbf H} \Vert=o_{\mathrm P}(1)$,
\eqref{eq:Htildelim2}, and \eqref{eq:Htildelim}, or \eqref{eq:riunioni}.
The last expression in \eqref{eq:CLTLBAI} coincides with \eqref{eq:CLTL} in the proof of Theorem \ref{th:CLTL}, hence, by Slutsky's theorem,
\begin{equation}\label{eq:ANLBAI1}
\sqrt T(\widetilde{\bm\lambda}_i-\widetilde{\mathbf H}^{-1}{\bm\lambda}_i) \to_d\mathcal N\left(\mathbf 0_r, \bm\Upsilon_0^\prime(\bm\Gamma^F)^{1/2} \bm\Theta_i^{\text{\tiny \upshape OLS}}(\bm\Gamma^F)^{1/2}\bm\Upsilon_0
\right).
\end{equation}
Moreover, using the second last line of \eqref{eq:CLTLBAI}, by Slutsky's theorem, we also have
\begin{equation}\label{eq:ANLBAI}
\sqrt T(\widetilde{\bm\lambda}_i-\widetilde{\mathbf H}^{-1}{\bm\lambda}_i) \to_d\mathcal N\left(\mathbf 0_r, \bm V_0^{1/2}\bm\Upsilon_1^\prime (\bm\Sigma_\Lambda)^{-1/2} \bm\Theta_i^{\text{\tiny \upshape OLS}}(\bm\Sigma_\Lambda)^{-1/2}\bm\Upsilon_1\bm V_0^{1/2}
\right),
\end{equation}
where $\bm\Theta_i^{\text{\tiny \upshape OLS}}=(\bm\Gamma^F)^{-1}\bm\Phi_i(\bm\Gamma^F)^{-1}$ with $\bm\Phi_i$ is defined in Assumption \ref{ass:CLT}(a). Furthermore, from the second line of \eqref{eq:CLTLBAI}, using \eqref{eq:Htildelim}, we also have that the asymptotic covariance matrix of $\widetilde{\bm\lambda}_i$ can also be written as $\bm V_0^{-1/2}\bm\Upsilon_1^\prime (\bm\Sigma_\Lambda)^{1/2}
\bm\Phi_i
(\bm\Sigma_\Lambda)^{1/2}\bm\Upsilon_1\bm V_0^{-1/2}$
and by letting $\bm {\mathcal Q}:=(\bm V_0)^{1/2}\bm\Upsilon_1^\prime(\bm\Sigma_\Lambda)^{-1/2}$ this is equivalent to
$(\bm {\mathcal Q}^{\prime})^{-1} \bm\Phi_i\bm {\mathcal Q}^{-1}$, which is the expression given in \citet[Theorem 2]{Bai03}. Notice that from \eqref{eq:ANLBAI} it follows that we can state Thoreom \ref{th:CLTL} also as
\begin{equation}\label{eq:ANLbis}
\sqrt T(\widehat{\bm\lambda}_i-\widehat{\mathbf H}^{\prime}{\bm\lambda}_i) \to_d\mathcal N\left(\mathbf 0_r, \bm V_0^{1/2}\bm\Upsilon_1^\prime (\bm\Sigma_\Lambda)^{-1/2} \bm\Theta_i^{\text{\tiny \upshape OLS}}(\bm\Sigma_\Lambda)^{-1/2}\bm\Upsilon_1\bm V_0^{1/2}
\right).
\end{equation}
\section{Identification and inference}\label{sec:II}
\paragraph{Constraints in exploratory factor analysis.} In this section, we consider a series of identifying restriction where either the factors or the loadings are assumed to be orthonormal. These are statistical restrictions in the sense that the identified factors have no direct interpretation for the considered data. In other words, we are considering only exploratory factor analysis, but not confirmatory factor analysis. Let us first summarize the identifying conditions that are typically found in the literature.\footnote{
Other identifying conditions related to confirmatory factor analysis exist in the literature. In particular, an appealing one is to assume $\bm\Lambda^\prime=[\mathbf I_r | \bm\Lambda^{\prime}_1]$ with $\bm\Lambda_1$ being $N-r\times r$ and full and leave $\bm\Gamma^F$ and, especially, $T^{-1}\bm F^\prime \bm F$ unrestricted. This is tantamount as assuming the the first $r$ elements of the observables are noisy measures of the latent factors, i.e., $x_{jt}= F_{jt}+ e_{jt}$, $j=1,\ldots, r$ \citep{PF86,AFP87}.
}
\begin{compactenum}
\item [(A)] $n^{-1}\bm\Lambda^\prime\bm\Lambda$ diagonal for all $n\in\mathbb N$ (\citealp[condition (c) page 122]{AR56}; \citealp[PC1]{baing13}).
\item [(B)] $n^{-1}\bm\Lambda^\prime(\text{diag}(\bm\Gamma^ e))^{-1}\bm\Lambda$ diagonal for all $n\in\mathbb N$ (\citealp[condition (d) page 122]{AR56}; \citealp[Chapter 2]{lawleymaxwell71};
\citealp[Chapter 9.2]{mardia1979multivariate}; \citealp[IC3]{baili12,baili16}).
\item [(C)] $\bm\Gamma^F=\mathbf I_r$ (\citealp[Equation (2)]{hotelling1933analysis};
\citealp[Section 4]{AR56};
\citealp[Chapter 2]{lawleymaxwell71};
\citealp[Chapter 9.2]{mardia1979multivariate}; \citealp[Cases 1 and 2]{RT82}; \citealp[Chapter 7.1]{jolliffe2002principal}; \citealp[Chapter 14.2]{anderson2003introduction}).
\item [(D)] $T^{-1}\bm F^\prime\bm F=\mathbf I_r$ for all $T\in\mathbb N$ (\citealp[Chapter 14.2]{anderson2003introduction}; \citealp[IC3]{baili12,baili16}; \citealp[Assumption 1]{onatski2012asymptotics}; \citealp[PC1]{baing13};
\citealp[Assumption 1]{Freyaldenhoven2022}).
\end{compactenum}
Two comments follow from inspection of the above works. First, when considering maximum likelihood estimation of an exact factor model, for the loadings it is usually imposed (B) rather than (A). Indeed, the maximum likelihood estimates of loadings and idiosyncratic variances depend on each other so they are usually jointly constrained. Typically, (A) is instead imposed when considering PC estimation, since in that case we are not interested in estimating the idiosyncratic covariance. Hence we will focus only on (A). Classical works consider only the fixed $n$ case, so it is natural to assume that (A) holds for all $n\in\mathbb N$. However, when we allow $n\to\infty$, (A) might seem too restrictive, and it is interesting to consider also its limiting case, i.e., $\bm\Sigma_\Lambda=\mathbf I_r$.
The choice between (C) and (D) depends on the nature of the factors. The distinction between random and deterministic factors can be found only in the context of maximum likelihood estimation (\citealp[Chapter 14.2]{anderson2003introduction}; \citealp{AFP87}; \citealp{baili12}), where it is shown that, regardless of such distinction, asymptotic normality of the maximum likelihood estimator of the loadings can be derived provided we impose (A) or, more often, (B) and either (C) if factors are random or (D) if factors are deterministic. So in maximum likelihood estimation the distinction is irrelevant. However, this distinction is extremely relevant in PC estimation since, if the factors are random, (D) is not realistic and (C) should be imposed instead. Nevertheless, we will also consider (D) in the following to clearly show its implications.
\paragraph{Identification of the true loadings and factors.}
The next results show the implications of various identifying constraints for the true loadings and factors (see Appendix \ref{corol:K00proof} for a proof).
\begin{prop}\label{corol:K00}
Under Assumptions \ref{ass:common} and \ref{ass:eval},
\begin{compactenum}
\item [(I.a)] if $\bm\Gamma^F=\mathbf I_r$ and
$\bm\Lambda$ is unrestricted,
then:
\begin{compactenum}
\item [(i.1)] $\bm\lambda_i^\prime=\mathbf v_i^{C\prime}(\mathbf M^{C})^{1/2}\mathbf K^\prime$ for all $i=1,\ldots, n$ and all $n\in\mathbb N$;
\item [(ii.1)] $\mathbf F_t=\mathbf K(\mathbf M^{C})^{-1/2} \mathbf V^{{C}\prime}\bm{C}_t$ for all $t\in\mathbb Z$ and all $n\in\mathbb N$;
\item [(i.2)] $\lim_{n\to\infty}\Vert \bm\lambda_i^\prime-\bm p_{i0}^\prime (\bm V_0)^{1/2} \bm\Upsilon_0^\prime \Vert=0$ for all $i\in\mathbb N$;
\item [(ii.2)] $\text{\upshape m.s.-}\lim_{n\to\infty}\Vert \mathbf F_t-\bm\Upsilon_0(\bm V_0)^{-1/2} \bm W_{\infty}^\prime\bm C_{t,\infty}\Vert=0$ for all $t\in\mathbb Z$;
\end{compactenum}
\item [(II.a)] if $\bm\Gamma^F=\mathbf I_r$ and
$n^{-1}\bm\Lambda^\prime\bm \Lambda$ is diagonal for all $n\in\mathbb N$,
then:
\begin{compactenum}
\item [(i.1)] $\bm\lambda_i^\prime=\mathbf v_i^{C\prime}(\mathbf M^{C})^{1/2}\bm S$ for all $i=1,\ldots, n$ and all $n\in\mathbb N$;
\item [(ii.1)] $\mathbf F_t=\bm S(\mathbf M^{C})^{-1/2}\mathbf V^{{C}\prime}\bm{C}_t$ for all $t\in\mathbb Z$ and all $n\in\mathbb N$;
\item [(i.2)] $\lim_{n\to\infty}\Vert \bm\lambda_i^\prime-\bm p_{i0}^\prime (\bm V_0)^{1/2} \bm {\mathcal S}_0 \Vert=0$ for all $i\in\mathbb N$;
\item [(ii.2)] $\text{\upshape m.s.-}\lim_{n\to\infty}\Vert\mathbf F_t-\bm{\mathcal S}_0(\bm V_0)^{-1/2} \bm W_{\infty}^{\prime}\bm{C}_{t,\infty}\Vert=0$ for all $t\in\mathbb Z$;
\end{compactenum}
\item [(III.a)] if $\bm\Gamma^F=\mathbf I_r$ and
$\bm\Sigma_\Lambda$ is diagonal,
then:
\begin{compactenum}
\item [(i)] $\lim_{n\to\infty}\Vert \bm\lambda_i^\prime-\bm p_{i0}^{\prime}(\bm V_0)^{1/2}\bm{\mathcal S}_0\Vert=0$ for all $i\in\mathbb N$;
\item [(ii)] $\text{\upshape m.s.-}\lim_{n\to\infty}\Vert\mathbf F_t-\bm{\mathcal S}_0(\bm V_0)^{-1/2} \bm W_{\infty}^{\prime}\bm{C}_{t,\infty}\Vert=0$ for all $t\in\mathbb Z$;
\end{compactenum}
\item [(IV.a)] if $T^{-1} \bm F^\prime\bm F=\mathbf I_r$ for all $T\in\mathbb N$ and
$\bm\Lambda$ is unrestricted,
then:
\begin{compactenum}
\item [(i.1)] $\bm\lambda_i^\prime=\widehat{\mathbf v}_i^{C\prime}(\widehat{\mathbf M}^{C})^{1/2}\widehat{\mathbf Q}^\prime$ for all $i=1,\ldots, n$ and all $n,T\in\mathbb N$;
\item [(ii.1)] $\mathbf F_t=\widehat{\mathbf Q}(\widehat{\mathbf M}^{C})^{-1/2} \widehat{\mathbf V}^{{C}\prime}\bm{C}_t=0$ for all $t=1,\ldots, T$ and all $n,T\in\mathbb N$;
\item [(i.2)] $\lim_{T\to\infty}\Vert \bm\lambda_i^\prime-{\mathbf v}_i^{C\prime}({\mathbf M}^{C})^{1/2}{\mathbf K}^\prime\Vert =0$ for all $i=1,\ldots, n$ and all $n\in\mathbb N$;
\item [(ii.2)] $\text{\upshape m.s.-}\lim_{T\to\infty}\Vert\mathbf F_t-{\mathbf K}(\widehat{\mathbf M}^{C})^{-1/2} {\mathbf V}^{{C}\prime}\bm{C}_t\Vert=0$ for all $t\in\mathbb N$ and all $n\in\mathbb N$;
\item [(i.3)] $\lim_{n,T\to\infty}\Vert \bm\lambda_i^\prime-\bm p_{i0}^\prime (\bm V_0)^{1/2} \bm\Upsilon_0^\prime \Vert=0$ for all $i\in\mathbb N$;
\item [(ii.3)] $\text{\upshape m.s.-}\lim_{n,T\to\infty}\Vert \mathbf F_t- \bm\Upsilon_0(\bm V_0)^{-1/2} \bm W_{\infty}^\prime\bm C_{t,\infty}\Vert=0$ for all $t\in\mathbb N$;
\end{compactenum}
\item [(V.a)] if $T^{-1} \bm F^\prime\bm F=\mathbf I_r$ and
$n^{-1}\bm\Lambda^\prime\bm \Lambda$ is diagonal for all $n,T\in\mathbb N$,
then:
\begin{compactenum}
\item [(i.1)] $\bm\lambda_i^\prime=\widehat{\mathbf v}_i^{C\prime}(\widehat{\mathbf M}^{C})^{1/2}\widehat{\bm S}$ for all $i=1,\ldots, n$ and all $n,T\in\mathbb N$;
\item [(ii.1)] $\mathbf F_t=\widehat{\bm S}(\widehat{\mathbf M}^{C})^{-1/2} \widehat{\mathbf V}^{{C}\prime}\bm{C}_t=0$ for all $t=1,\ldots, T$ and all $n,T\in\mathbb N$;
\item [(i.2)] $\lim_{T\to\infty}\Vert \bm\lambda_i^\prime-{\mathbf v}_i^{C\prime}({\mathbf M}^{C})^{1/2}{\bm S}^\prime\Vert =0$ for all $i=1,\ldots, n$ and all $n\in\mathbb N$;
\item [(ii.2)] $\text{\upshape m.s.-}\lim_{T\to\infty}\Vert\mathbf F_t-{\bm S}(\widehat{\mathbf M}^{C})^{-1/2} {\mathbf V}^{{C}\prime}\bm{C}_t\Vert=0$ for all $t\in\mathbb N$ and all $n\in\mathbb N$;
\item [(i.3)] $\lim_{n,T\to\infty}\Vert \bm\lambda_i^\prime-\bm p_{i0}^\prime (\bm V_0)^{1/2} \bm{\mathcal S}_0 \Vert=0$ for all $i\in\mathbb N$;
\item [(ii.3)] $\text{\upshape m.s.-}\lim_{n,T\to\infty}\Vert \mathbf F_t- \bm{\mathcal S}_0(\bm V_0)^{-1/2} \bm W_{\infty}^\prime\bm C_{t,\infty}\Vert=0$ for all $t\in\mathbb N$;
\end{compactenum}
\item [(VI.a)] if $T^{-1} \bm F^\prime\bm F=\mathbf I_r$ for all $T\in\mathbb N$ and
$\bm\Sigma_\Lambda$ is diagonal,
then:
\begin{compactenum}
\item [(i)] $\lim_{n,T\to\infty}\Vert \bm\lambda_i^\prime-\bm p_{i0}^{\prime}(\bm V_0)^{1/2}\bm{\mathcal S}_0\Vert=0$ for all $i\in\mathbb N$;
\item [(ii)] $\text{\upshape m.s.-}\lim_{n,T\to\infty}\Vert\mathbf F_t-\bm{\mathcal S}_0(\bm V_0)^{-1/2} \bm W_{\infty}^\prime\bm C_{t,\infty}\Vert=0$ for all $t\in\mathbb N$;
\end{compactenum}
\item [(I.b)] if $\bm\Sigma_\Lambda=\mathbf I_r$ and
$\bm F$ is unrestricted,
then:
\begin{compactenum}
\item [(i)] $\lim_{n\to\infty} \left\Vert\bm\lambda_i^\prime - \bm p_{i0}^{\prime} \bm\Upsilon_1^\prime \right \Vert=0$ for all $i\in\mathbb N$;
\item [(ii)] $\text{\upshape m.s.-}\lim_{n\to\infty}\Vert\mathbf F_t-\bm \Upsilon_1 \bm W_{\infty}^\prime\bm C_{t,\infty}\Vert =0$ for all $t\in\mathbb Z$;
\end{compactenum}
\item[(II.b)] if $\bm\Sigma_\Lambda=\mathbf I_r$ and
$T^{-1} \bm F^\prime\bm F$ is diagonal for all $T\in\mathbb N$,
then:
\begin{compactenum}
\item [(i)] $\text{\upshape P-}\lim_{n,T\to\infty}\Vert\bm\lambda_i^\prime- \bm p_{i0}^\prime \bm{\mathcal S}_0\Vert=0$ for all $i\in\mathbb N$;
\item [(ii)] $\text{\upshape P-}\lim_{n,T\to\infty}\Vert\mathbf F_t-\bm {\mathcal S}_0 \bm W_{\infty}^\prime\bm C_{t,\infty}\Vert =0$ for all $t\in\mathbb N$;
\end{compactenum}
\item[(III.b)] if $\bm\Sigma_\Lambda=\mathbf I_r$ and
$\bm\Gamma^F$ is diagonal,
then:
\begin{compactenum}
\item [(i)] $\lim_{n\to\infty}\Vert\bm\lambda_i^\prime- \bm p_{i0}^\prime \bm{\mathcal S}_0\Vert=0$ for all $i\in\mathbb N$;
\item [(ii)] $\text{\upshape m.s.-}\lim_{n\to\infty}\Vert\mathbf F_t-\bm {\mathcal S}_0 \bm W_{\infty}^\prime\bm C_{t,\infty}\Vert =0$ for all $t\in\mathbb Z$;
\end{compactenum}
\item [(IV.b)] if $n^{-1}\bm\Lambda^\prime\bm \Lambda=\mathbf I_r$ for all $n\in\mathbb N$ and
$\bm F$ is unrestricted,
then:
\begin{compactenum}
\item [(i.1)] $\bm\lambda_i^\prime= n \mathbf v_i^{C\prime} (\mathbf M^{C})^{-1/2}\mathbf K^\prime(\bm\Gamma^F)^{1/2}$ for all $i=1,\ldots, n$ and all $n\in\mathbb N$;
\item [(ii.1)] $\mathbf F_t= (\bm\Gamma^F)^{1/2}\mathbf K(\mathbf M^{C})^{-1/2}\mathbf V^{C\prime}\bm C_t$ for all $t\in\mathbb Z$ and all $n\in\mathbb N$;
\item [(i.2)] $\lim_{n\to\infty} \left\Vert\bm\lambda_i^\prime - \bm p_{i0}^{\prime} \bm\Upsilon_1^\prime \right \Vert=0$ for all $i\in\mathbb N$;
\item [(ii.2)] $\text{\upshape m.s.-}\lim_{n\to\infty}\Vert\mathbf F_t-\bm \Upsilon_1 \bm W_{\infty}^\prime\bm C_{t,\infty}\Vert =0$ for all $t\in\mathbb Z$;
\end{compactenum}
\item [(V.b)] if $n^{-1}\bm\Lambda^\prime\bm \Lambda=\mathbf I_r$ and
$T^{-1} \bm F^\prime\bm F$ is diagonal for all $n,T\in\mathbb N$,
then:
\begin{compactenum}
\item [(i.1)] $\bm\lambda_i^\prime =\sqrt n \widehat{\mathbf v}_i^{C\prime} \widehat{\bm S}$ for all $i=1,\ldots, n$ and all $n,T\in\mathbb N$;
\item [(ii.1)] $\mathbf F_t=n^{-1/2}\widehat{\bm S} \widehat{\mathbf V}^{C\prime}\bm C_t$ for all $t=1,\ldots,T$ and all $n,T\in\mathbb N$;
\item [(i.2)] $\text{\upshape P-}\lim_{T\to\infty}\Vert \bm\lambda_i^\prime -\sqrt n {\mathbf v}_i^{C\prime} {\bm S}\Vert=0$ for all $i=1,\ldots, n$ and all $n\in\mathbb N$;
\item [(ii.2)] $\text{\upshape P-}\lim_{T\to\infty}\Vert\mathbf F_t-n^{-1/2}{\bm S} {\mathbf V}^{C\prime}\bm C_t\Vert=0$ for all $t\in\mathbb N$ and all $n\in\mathbb N$;
\item [(i.3)] $\text{\upshape P-}\lim_{n,T\to\infty}\Vert \bm\lambda_i^\prime-\bm p_{i0}^\prime \bm{\mathcal S}_0\Vert=0$ for all $i\in\mathbb N$;
\item [(ii.3)]$\text{\upshape P-}\lim_{n,T\to\infty}\Vert \mathbf F_t-\bm {\mathcal S}_0 \bm W_{\infty}^\prime\bm C_{t,\infty}\Vert=0$ for all $t\in\mathbb N$;
\end{compactenum}
\item[(VI.b)] if $n^{-1}\bm\Lambda^\prime\bm \Lambda=\mathbf I_r$ for all $n\in\mathbb N$ and
$\bm\Gamma^F$ is diagonal,
then:
\begin{compactenum}
\item [(i.1)] $\bm\lambda_i^\prime = \sqrt n \mathbf v_i^{C\prime}\bm S$ for all $i=1,\ldots, n$ and all $n\in\mathbb N$;
\item [(ii.1)] $\mathbf F_t=n^{-1/2}\bm S \mathbf V^{C\prime}\bm C_t$ for all $t\in\mathbb Z$ and all $n\in\mathbb N$;
\item [(i.2)] $\lim_{n\to\infty}\Vert \bm\lambda_i^\prime-\bm p_{i0}^\prime \bm{\mathcal S}_0\Vert=0$ for all $i\in\mathbb N$;
\item [(ii.2)] $\text{\upshape m.s.}\lim_{n\to\infty}\Vert \mathbf F_t-\bm {\mathcal S}_0 \bm W_{\infty}^\prime\bm C_{t,\infty}\Vert=0$ for all $t\in\mathbb Z$;
\end{compactenum}
\end{compactenum}
and if we replace Assumption \ref{ass:common}(c-ii) with Assumption \ref{ass:Wold}, or \ref{ass:Hannan}, or \ref{ass:Wu} in Appendix \ref{sec:hannan},
the limits in (II.b) and (V.b) hold also in mean-square. We used the following notation:
the columns of $\mathbf V^C$ and $\widehat{\mathbf V}^C$ are the normalized eigenvectors of $\bm\Lambda\bm\Gamma^F\bm\Lambda$ (with eigenvalues $\mathbf M^C$) and of
$\bm\Lambda(T^{-1}\bm F^\prime\bm F)\bm\Lambda$ (with eigenvalues $\widehat{\mathbf M}^C$), respectively,
the columns of ${\mathbf K}$ and $\widehat{\mathbf Q}$ are the normalized eigenvectors of $(\bm\Gamma^F)^{1/2}(n^{-1}\bm \Lambda^\prime\bm \Lambda) (\bm\Gamma^F)^{1/2}$ (with eigenvalues $n^{-1}\mathbf M^C$)
and of $(T^{-1}\bm F^\prime\bm F)^{1/2}(n^{-1}\bm \Lambda^\prime\bm \Lambda) (T^{-1}\bm F^\prime\bm F)^{1/2}$ (with eigenvalues $n^{-1}\widehat{\mathbf M}^C$), respectively,
the columns of $\bm\Upsilon_0$ and $\bm\Upsilon_1$ are the normalized eigenvectors of $(\bm\Gamma^F)^{1/2}\bm \Sigma_\Lambda (\bm\Gamma^F)^{1/2}$ (with eigenvalues $\bm V_0$) and of
$(\bm \Sigma_\Lambda)^{1/2}\bm\Gamma^F(\bm \Sigma_\Lambda)^{1/2}$ (with eigenvalues $\bm V_0$), respectively,
$\bm S$ is an $r\times r$ diagonal matrix with entries $\pm 1$ depending only on $n$,
$\widehat{\bm S}$ is an $r\times r$ diagonal matrix with entries $\pm 1$ depending on $n$ and $T$,
$\bm{\mathcal S}_0$ is an $r\times r$ diagonal matrix with entries $\pm 1$ independent of $n$ and $T$; $\bm p_{i0}:=\lim_{n\to\infty}\sqrt n \mathbf v_i^C$, $i\in\mathbb N$,
$\bm W_{\infty}:=\lim_{n\to\infty}n^{-1/2}\bm P_{n,\infty}$, with
$\bm P_{n,\infty}:=(\bm p_{10}\cdots \bm p_{n0})^\prime$, and
$\bm{C}_{t,\infty}:=\text{\upshape m.s.-}\lim_{n\to\infty} n^{-1/2}\bm C_t$, $t\in\mathbb Z$, such that, as $n\to\infty$, the following hold:
$\Vert \bm p_{i0}\Vert = O(1)$,
$n^{-1/2}\Vert\bm P_{n,\infty}\Vert = O(1)$, $\Vert\bm W_{\infty}\Vert=O(1)$, and $\Vert \bm C_{t,\infty}\Vert = O_{\mathrm{m.s.}}(1)$.
\end{prop}
First note that, unless we constrain $T^{-1}\bm F^\prime\bm F$, $T$ plays no role, since in population the factors are a stochastic process so $t\in\mathbb Z$ and then we consider only cases with either $n$ fixed or $n\to\infty$ (parts (I.a), (II.a), (III.a), (I.b), (III.b), (IV.b), (VI.b)). If instead we constrain also $T^{-1}\bm F^\prime\bm F$, then we either consider the case of $n$ and $T$ fixed or $n,T\to\infty$ (parts (IV.a), (V.a), (VI.a), (II.b), (V.b)). Moreover, in three cases we can derive limits for fixed $n$ and $T\to\infty$ (parts (IV.a), (V.a), (V.b)).
Overall it is clear that $n$ and $T$ have no symmetric roles. This is apparent by looking in detail at the implications of the various conditions considered, which are also summarized in Table \ref{tab:ident}.
\begin{compactenum}
\item[(i)] From parts (I.a), (IV.a), (I.b), and (IV.b) it is obvious that it is not enough to impose orthonormal factors or loadings to achieve identification. This is because the imposed conditions are only $r(r+1)/2$ so identification can be achieved only up to a rotation accounting for the remaining $r(r-1)/2$ degrees-of-freedom. Note that part (I.b) is the weakest since identification requires also $n\to\infty$.
\item[(ii)] Parts (II.a) and (VI.b) allow to identify the loadings and the factors up to a sign for any given fixed $n$ and regardless of $T$. In particular, the factors are identified as the population PCs of the common component. These are the classical identification conditions for time series.
\item[(iii)] Parts (V.a) and (V.b) allow to identify the loadings and the factors up to a sign for any given fixed $n$ and $T$. Provided we let $n,T\to\infty$, these have the same implications as parts (II.a) and (VI.b), respectively. But, given that they impose stronger conditions on the factors, they also imply that, for any fixed $n$ and $T$, the true loadings are random while the true factors are the sample PCs of the common component. However, these restrictions are not credible in time series. Indeed, if $\bm F$ is stochastic, then $\mathrm P(T^{-1}\bm F^\prime\bm F=\mathbf I_r)=0$, so for the restrictions to hold the factors have to be deterministic, implying that the common components, and so the loadings, are deterministic too. For these reasons parts (V.a) and (V.b), should never be considered in a time series context.
\item[(iv)] Parts (VI.a) and (II.b) not only require deterministic factors to hold, but do not even allow for identification unless we also require $n,T\to\infty$, so also these constraints should never be considered in a time series context.
\item [(v)] Finally, parts (III.a) and (III.b) require $n\to\infty$ for achieving identification and do not depend on $T$. As shown below, in this case the factors are identified as the PCs of the infinite dimensional process of common components. They are the most realistic conditions which should be employed consistently with the fact that the approximate factor model is identified only when $n\to\infty$.\smallskip
\end{compactenum}
Following \citet{gersing_id}, who firstly derived it, part (III.a) can be reformulated as follows. First, let $\bm\Lambda_{n,\infty}$ be the $n\times r$ matrix with rows $\bm\lambda_{i,\infty}^\prime:=\lim_{n\to\infty} \mathbf v_i^{C\prime}(\mathbf M^C)^{1/2}$. Then, by definition of $\bm p_{i0}$ and Lemma \ref{lem:Vzero}(i),
\begin{equation}\label{eq:phil1}
\bm\lambda_{i,\infty}^\prime=\bm p_{i0}^\prime (\bm V_0)^{1/2},
\end{equation}
and part (III.a.i) is equivalent to $\lim_{n\to\infty}\Vert \bm\lambda_i^\prime -\bm\lambda_{i,\infty}^\prime \bm{\mathcal S}_0 \Vert=0.$ Note that
the sequence $\{\bm\Lambda_{n,\infty},\, n\in\mathbb N\}$ is nested as $n$ grows since the rows of $\bm\Lambda_{n,\infty}$ do not depend on $n$.
Second, by definition of $\bm W_\infty$ and $\bm p_{i0}$,
\begin{align}
\bm W_{\infty}^\prime \bm W_{\infty}=\lim_{n\to\infty} \frac 1n \bm P_{n,\infty}^\prime \bm P_{n,\infty}=\lim_{n\to\infty} \frac 1n\sum_{i,j=1}^n \bm p_{0i}\bm p_{0j}^\prime
=\lim_{n\to\infty} \sum_{i,j=1}^n \mathbf v_{i}^C\mathbf v_{j}^{C\prime} =\lim_{n\to\infty} \mathbf V^{C\prime}\mathbf V^C= \mathbf I_r.\label{eq:Pinfnrom}
\end{align}
Now let $\mathbf F_{t,\infty}:=\text{m.s.-}\lim_{n\to\infty}(\bm\Lambda_{n,\infty}^\prime \bm\Lambda_{n,\infty})^{-1}\bm\Lambda_{n,\infty}^\prime \bm C_t$. Then, since $\bm\Lambda_{n,\infty} = \bm P_{n,\infty} (\bm V_0)^{1/2}$ for all $n\in\mathbb N$, from \eqref{eq:Pinfnrom} it follows that $n^{-1}\bm\Lambda_{n,\infty}^\prime \bm\Lambda_{n,\infty}=\bm V_0$ for all $n\in\mathbb N$, and also
\begin{equation}\label{eq:phil2}
\mathbf F_{t,\infty}=(\bm V_0)^{-1/2} \text{m.s.-}\lim_{n\to\infty}\frac{\bm P_{n,\infty}^\prime\bm C_t}{n}=(\bm V_0)^{-1/2}\bm W_{\infty}^\prime\bm C_{t,\infty}.
\end{equation}
So part (III.a.ii) is equivalent to $\text{m.s.-}\lim_{n\to\infty}\Vert \mathbf F_t-\bm{\mathcal S}_0\mathbf F_{t,\infty}\Vert=0$. Note that $\mathbf F_{t,\infty}$ does not depend on $n$ and it is obtained as a weighted average of infinitely many common components. Specifically, the elements of $\mathbf F_{t,\infty}$ are the normalized PCs of the infinite dimensional vector $\bm C_{t,\infty}$ which is the mean-squared limit of the $n$-dimensional vector $n^{-1/2}\mathbf C_t$. Indeed, by definition of $\bm W_{\infty}$ and $\bm C_{t,\infty}$, and by Lemma \ref{lem:Vzero}(i),
\begin{align}
\bm W_{\infty}^\prime \mathbb{E}[\bm C_{t,\infty}\bm C_{t,\infty}^\prime]\bm W_{\infty}&=
\lim_{n\to\infty} \left\{\frac{ \bm P_{n,\infty}^\prime}{\sqrt n} \frac{\bm\Gamma^C}{n}\frac{\bm P_{n,\infty}}{\sqrt n} \right\}\nonumber\\
&=\lim_{n\to\infty} \left\{ \frac 1{n^2} \sum_{i,j=1}^n \bm p_{0i} [\bm\Gamma^C]_{ij} \bm p_{j0}^{\prime} \right\}=\lim_{n\to\infty} \left\{ \frac 1n \sum_{i,j=1}^n \mathbf v_{i}^C [\bm\Gamma^C]_{ij} \mathbf v_{j}^{C\prime} \right\}\nonumber\\
&=\lim_{n\to\infty} \left\{\mathbf V^{C\prime} \frac{\bm\Gamma^C}{n} \mathbf V^C\right\}=\lim_{n\to\infty} \left\{ \frac{\mathbf M^C}{n}\right\}
= \bm V_0,\label{eq:Pinfevec}
\end{align}
and, because of \eqref{eq:Pinfnrom} and \eqref{eq:Pinfevec}, the columns of $\bm W_{\infty}$ are the normalized eigenvectors of the infinite dimensional matrix $\mathbb{E}[\bm C_{t,\infty}\bm C_{t,\infty}^\prime]=\lim_{n\to\infty} n^{-1} \bm\Gamma^C$ which is the covariance matrix of the infinite dimensional process $\bm C_{t,\infty}$. An analogous reasoning could be done for part (III.b), but in that case the factors would be identified with the non-normalized PCs of $\bm C_{t,\infty}$, having covariance matrix $\bm V_0$.
Finally, since the columns of $\mathbf V^C$ and $\bm\Lambda$ span the same space, then
$\bm\Lambda=\mathbf V^{C}\mathbf V^{C\prime}\bm\Lambda$, so that
$\bm\lambda_i^\prime=\mathbf v_i^{C\prime}\mathbf V^{C\prime}\bm\Lambda$ and $C_{it} = \bm\lambda_i^\prime \mathbf F_t= \mathbf v_i^{C\prime}\mathbf V^{C\prime} \bm\Lambda \mathbf F_t= \mathbf v_i^{C\prime}\mathbf V^{C\prime}\bm C_t$ for all $i=1,\ldots, n$ and all $n\in\mathbb N$. Therefore, since $C_{it}$ does not depend on $n$ and, by definition, $\bm P_{n,\infty}=\sqrt n\mathbf V^{C}$ for all $n\in\mathbb N$, from \eqref{eq:phil1} and \eqref{eq:phil2}, we obtain
\begin{align}
C_{it} =
\text{m.s.-}\lim_{n\to\infty} C_{it}&= \text{m.s.-}\lim_{n\to\infty} \mathbf v_i^{C\prime} {\mathbf V^{C\prime}\bm C_t }= \text{m.s.-}\lim_{n\to\infty} \sqrt n \mathbf v_i^{C\prime} \frac{\mathbf V^{C\prime}\bm C_t }{\sqrt n}= \text{m.s.-}\lim_{n\to\infty} \sqrt n \mathbf v_i^{C\prime}\frac{\bm P_{n,\infty}^{\prime}\bm C_t }{n}\nonumber\\
& =\bm p_{i0}^\prime \bm W_{\infty}^\prime\bm C_{t,\infty}=\bm p_{i0}^\prime (\bm V_0)^{1/2} (\bm V_0)^{-1/2} \bm W_{\infty}^\prime\bm C_{t,\infty} = \bm\lambda_{i,\infty}^\prime\mathbf F_{t,\infty}.\nonumber
\end{align}
Hence, in general, for all $i\in\mathbb N$, there exists an invertible $r\times r$ matrix $\bm{\mathcal H}_{\infty}$, independent of $n$ and $T$, such that
\begin{equation}\label{eq:phil3}
C_{it}=\bm\lambda_i^\prime \mathbf F_t=\bm\lambda_i^\prime\bm{\mathcal H}_{\infty}\bm{\mathcal H}_{\infty}^{-1} \mathbf F_t= \bm\lambda_{i,\infty}^\prime\mathbf F_{t,\infty}.
\end{equation}
From part (I.a), we see that, if we just impose orthonormal factors, then $\bm{\mathcal H}_{\infty}= \lim_{n\to\infty}\mathbf K=\bm\Upsilon_0$, where $\bm\Upsilon_0$ is a rotation, while from parts (II.a) or (III.a), we see that, if we also impose orthogonal loadings, then $\bm{\mathcal H}_{\infty}=\bm{\mathcal S}_0$. For more details we refer to \citet{gersing_id,gersing_lag}.
\begin{sidewaystable}[htbp]
\caption{Identifying constraints and their implications}\label{tab:ident}
\centering
\scriptsize{
\begin{tabular}{l | ll | l | l | l | l}
\hline
\hline
&&&&&&\\[-8pt]
& \multicolumn{2}{c|}{Constraints}&$n,T$&$\bm\lambda_i^\prime$ &$\mathbf F_t$ & $\widehat{\mathbf H}$ \\[1pt]
\hline
&&&&&&\\[-8pt]
I.a& $\bm\Gamma^F=\mathbf I_r$
& $\bm\Lambda$ unrestricted
& $n$ fixed
& $\mathbf v_i^{C\prime}(\mathbf M^{C})^{1/2}\mathbf K^\prime$
& $\mathbf K(\mathbf M^{C})^{-1/2} \mathbf V^{{C}\prime}\bm{C}_t$
& $ \widehat{\mathbf H}^\prime\widehat{\mathbf H}=\mathbf I_r+ O_{\mathrm P}\left(\frac 1n,\frac 1{\sqrt T}\right)$\\[3pt]
&&& $n\to\infty$
& $ \bm p_{i0}^\prime (\bm V_0)^{1/2}\bm\Upsilon_0^\prime$
& ${\text{\tiny m.s.}} \bm\Upsilon_0(\bm V_0)^{-1/2}\bm W_{\infty}^\prime \bm C_{t,\infty}$
&\\[3pt]
\hline
&&&&&&\\[-8pt]
II.a& $\bm\Gamma^F=\mathbf I_r$
& $\frac{\bm\Lambda^\prime\bm \Lambda}n$ diagonal
& $n$ fixed
& $\mathbf v_i^{C\prime}(\mathbf M^{C})^{1/2}\bm S$
& $\bm S(\mathbf M^{C})^{-1/2} \mathbf V^{{C}\prime}\bm{C}_t$
& $\widehat{\mathbf H}=\mathbf J+ O_{\mathrm P}\left(\frac 1n,\frac 1{\sqrt T}\right)$\\[3pt]
&&& $n\to\infty$
& $ \bm p_{i0}^\prime (\bm V_0)^{1/2}\bm{\mathcal S}_0$
& ${\text{\tiny m.s.}} \bm{\mathcal S}_0(\bm V_0)^{-1/2}\bm W_{\infty}^\prime \bm C_{t,\infty}$
&\\[3pt]
\hline
&&&&&&\\[-8pt]
III.a&$\bm\Gamma^F=\mathbf I_r$
& $\bm\Sigma_\Lambda$ diagonal
& $n\to\infty$
& $ \bm p_{i0}^\prime (\bm V_0)^{1/2}\bm{\mathcal S}_0$
& ${\text{\tiny m.s.}} \bm{\mathcal S}_0(\bm V_0)^{-1/2}\bm W_{\infty}^\prime \bm C_{t,\infty}$
& $\widehat{\mathbf H}=\mathbf J+ O_{\mathrm P}\left(\frac 1{\sqrt n},\frac 1{\sqrt T}\right)$\\[3pt]
\hline
&&&&&&\\[-8pt]
IV.a& $\frac{ \bm F^\prime\bm F}T=\mathbf I_r$
& $\bm\Lambda$ unrestricted
& $n,T$ fixed
& $\widehat{\mathbf v}_i^{C\prime}(\widehat{\mathbf M}^C)^{1/2} \widehat{\mathbf Q}^\prime$
& $\widehat{\mathbf Q}(\widehat{\mathbf M}^{C})^{-1/2} \widehat{\mathbf V}^{{C}\prime}\bm{C}_t$
& $\widehat{\mathbf H}^\prime\widehat{\mathbf H}=\mathbf I_r+ O_{\mathrm P}\left(\frac 1n,\frac 1{ T}\right)$\\[3pt]
&&& $T\to\infty$
& $\mathbf v_i^{C\prime}(\mathbf M^{C})^{1/2}\mathbf K^\prime$
& $\text{\tiny m.s.} \mathbf K(\mathbf M^{C})^{-1/2} \mathbf V^{{C}\prime}\bm{C}_t$
&\\[3pt]
&&& $n,T\to\infty$
& $\bm p_{i0}^\prime (\bm V_0)^{1/2}\bm\Upsilon_0^\prime$
& ${\text{\tiny m.s.}} \bm\Upsilon_0(\bm V_0)^{-1/2}\bm W_{\infty}^\prime \bm C_{t,\infty}$
&\\[3pt]
\hline
&&&&&&\\[-8pt]
V.a& $\frac{ \bm F^\prime\bm F}T=\mathbf I_r$
& $\frac{\bm\Lambda^\prime\bm \Lambda}n$ diagonal
& $n,T$ fixed
& $\widehat{\mathbf v}_i^{C\prime}(\widehat{\mathbf M}^C)^{1/2} \widehat{\bm S}$
& $\widehat{\bm S}(\widehat{\mathbf M}^{C})^{-1/2} \widehat{\mathbf V}^{{C}\prime}\bm{C}_t$
& $\widehat{\mathbf H}=\mathbf J+ O_{\mathrm P}\left(\frac 1n,\frac 1{ T}\right)$\\[3pt]
&&&$T\to\infty$
& $\mathbf v_i^{C\prime}(\mathbf M^{C})^{1/2}\bm S$
& $\text{\tiny m.s.} \bm S(\mathbf M^{C})^{-1/2} \mathbf V^{{C}\prime}\bm{C}_t$
&\\[3pt]
&&& $n,T\to\infty$
& $ \bm p_{i0}^\prime (\bm V_0)^{1/2}\bm{\mathcal S}_0$
& ${\text{\tiny m.s.}} \bm{\mathcal S}_0(\bm V_0)^{-1/2}\bm W_{\infty}^\prime \bm C_{t,\infty}$
&\\[3pt]
\hline
&&&&&&\\[-8pt]
VI.a& $\frac{ \bm F^\prime\bm F}T=\mathbf I_r$
& $\bm\Sigma_\Lambda$ diagonal
& $n,T\to\infty$
& $\bm p_{i0}^\prime (\bm V_0)^{1/2}\bm{\mathcal S}_0$
& ${\text{\tiny m.s.}}\bm{\mathcal S}_0(\bm V_0)^{-1/2}\bm W_{\infty}^\prime \bm C_{t,\infty}$
& $\widehat{\mathbf H}=\mathbf J+ O_{\mathrm P}\left(\frac 1{\sqrt n},\frac 1{ T}\right)$\\[3pt]
\hline
\hline
&&&&&&\\[-8pt]
I.b& $\bm\Sigma_\Lambda=\mathbf I_r$
& $\bm F$ unrestricted
& $n\to\infty$
& $\bm p_{i0}^{\prime} \bm\Upsilon_1^\prime$
& ${\text{\tiny m.s.}} \bm \Upsilon_1 \bm W_{\infty}^\prime\bm C_{t,\infty}$
& $\widehat{\mathbf H}^\prime\widehat{\mathbf H}=\frac{\widehat{\mathbf M}^x}{n}+ O_{\mathrm P}\left(\frac 1{\sqrt n},\frac 1{T}\right)$\\[3pt]
\hline
&&&&&\\[-8pt]
II.b& $\bm\Sigma_\Lambda=\mathbf I_r$
& $\frac{ \bm F^\prime\bm F}T$ diagonal
& $n,T\to\infty$
& $\text{\tiny m.s.} \bm p_{i0}^\prime \bm{\mathcal S}_0$
& $\text{\tiny m.s.} \bm {\mathcal S}_0 \bm W_{\infty}^\prime\bm C_{t,\infty}$
&$\widehat{\mathbf H}=\mathbf J\left(\frac{\widehat{\mathbf M}^x}n\right)^{1/2}+ O_{\mathrm P}\left(\frac 1{\sqrt n},\frac 1{T}\right)$\\[3pt]
\hline
&&&&&\\[-8pt]
III.b& $\bm\Sigma_\Lambda=\mathbf I_r$
& $\bm\Gamma^F$ diagonal
& $n\to\infty$
& $\bm p_{i0}^\prime \bm{\mathcal S}_0$
& $\text{\tiny m.s.} \bm {\mathcal S}_0 \bm W_{\infty}^\prime\bm C_{t,\infty}$
&$\widehat{\mathbf H}=\mathbf J\left(\frac{\widehat{\mathbf M}^x}n\right)^{1/2}+ O_{\mathrm P}\left(\frac 1{\sqrt n},\frac 1{\sqrt T}\right)$\\[3pt]
\hline
&&&&&\\[-8pt]
IV.b& $\frac{\bm\Lambda^\prime\bm \Lambda}n=\mathbf I_r$
& $\bm F$ unrestricted
& $n$ fixed
& $n \mathbf v_i^{C\prime} (\mathbf M^{C})^{-1/2}\mathbf K^\prime(\bm\Gamma^F)^{1/2}$
& $ (\bm\Gamma^F)^{1/2}\mathbf K(\mathbf M^{C})^{-1/2}\mathbf V^{C\prime}\bm C_t$
& $\widehat{\mathbf H}^\prime\widehat{\mathbf H}=\frac{\widehat{\mathbf M}^x}{n}+ O_{\mathrm P}\left(\frac 1n,\frac 1{T}\right)$ \\[3pt]
&&& $n\to\infty$
& $\bm p_{i0}^{\prime} \bm\Upsilon_1^\prime$
& $\text{\tiny m.s.}\bm \Upsilon_1 \bm W_{\infty}^\prime\bm C_{t,\infty}$
& \\[3pt]
\hline
&&&&&\\[-8pt]
V.b& $\frac{\bm\Lambda^\prime\bm \Lambda}n=\mathbf I_r$
& $\frac{ \bm F^\prime\bm F}T$ diagonal
& $n,T$ fixed
& $\sqrt n \widehat{\mathbf v}_i^{C\prime} \widehat{\bm S}$
& $\frac 1{\sqrt n}\widehat{\bm S} \widehat{\mathbf V}^{C\prime}\bm C_t$
&$\widehat{\mathbf H}=\mathbf J\left(\frac{\widehat{\mathbf M}^x}n\right)^{1/2}+ O_{\mathrm P}\left(\frac 1{ n},\frac 1{ T}\right)$\\[3pt]
&&& $T\to\infty$
& $\text{\tiny m.s.} \sqrt n {\mathbf v}_i^{C\prime} {\bm S}$
& $\text{\tiny m.s.} \frac 1{\sqrt n}{\bm S} {\mathbf V}^{C\prime}\bm C_t$
& \\[3pt]
&&& $n,T\to\infty$
& $\text{\tiny m.s.} \bm p_{i0}^\prime \bm{\mathcal S}_0$
& $\text{\tiny m.s.} \bm {\mathcal S}_0 \bm W_{\infty}^\prime\bm C_{t,\infty}$
& \\[3pt]
\hline
&&&&&\\[-8pt]
VI.b& $\frac{\bm\Lambda^\prime\bm \Lambda}n=\mathbf I_r$
& $\bm\Gamma^F$ diagonal
& $n$ fixed
& $\sqrt n \mathbf v_i^{C\prime}\bm S$
& $\frac 1{\sqrt n}\bm S \mathbf V^{C\prime}\bm C_t$
&$\widehat{\mathbf H}=\mathbf J\left(\frac{\widehat{\mathbf M}^x}n\right)^{1/2}+ O_{\mathrm P}\left(\frac 1{ n},\frac 1{\sqrt T}\right)$\\[3pt]
&&& $n\to\infty$
& $ \bm p_{i0}^\prime \bm{\mathcal S}_0$
& $\text{\tiny m.s.} \bm {\mathcal S}_0 \bm W_{\infty}^\prime\bm C_{t,\infty}$
& \\[3pt]
\hline
\hline
\end{tabular}
}
\begin{tabular}{p{.77\textwidth}}
$\mathbf V^C$ normalized eigenvectors of $\bm\Lambda \bm\Gamma^F \bm\Lambda ^\prime$ with eigenvalues $\mathbf M^C$;
$\mathbf K$ normalized eigenvectors of $(\bm\Gamma^F)^{1/2}(n^{-1}\bm \Lambda^\prime\bm \Lambda) (\bm\Gamma^F)^{1/2}$ with eigenvalues $n^{-1}\mathbf M^C$;\\
$\bm \Upsilon_0$ normalized eigenvectors of $(\bm\Gamma^F)^{1/2}\bm \Sigma_\Lambda (\bm\Gamma^F)^{1/2}$ with eigenvalues $\bm V_0$;
$\bm\Upsilon_1$ normalized eigenvectors of $(\bm \Sigma_\Lambda)^{1/2}\bm\Gamma^F(\bm \Sigma_\Lambda)^{1/2}$ with eigenvalues $\bm V_0$;\\
$\widehat{\mathbf V}^C$ normalized eigenvectors of $\bm\Lambda(T^{-1}\bm F^\prime\bm F) \bm\Lambda ^\prime$ with eigenvalues $\widehat{\mathbf M}^C$; $\widehat{\mathbf Q}$ normalized eigenvectors of $(T^{-1}\bm F^\prime\bm F)^{1/2}(n^{-1}\bm \Lambda^\prime\bm \Lambda) (T^{-1}\bm F^\prime\bm F)^{1/2}$ with eigenvalues $\widehat{\mathbf M}^C$;\\
$\bm S$ is an $r\times r$ diagonal matrix with entries $\pm 1$ depending only on $n$; $\bm{\mathcal S}_0$ is an $r\times r$ diagonal matrix with entries $\pm 1$ independent of $n$ and $T$; $\widehat{\bm S}$ is an $r\times r$ diagonal matrix with entries $\pm 1$ depending on $n$ and $T$;\\
$\bm p_{i0}:=\lim_{n\to\infty}\sqrt n \mathbf v_i^C$, $i\in\mathbb N$; $\bm W_{\infty}:=\lim_{n\to\infty}n^{-1/2}\bm P_{n,\infty}$, with $\bm P_{n,\infty}:=(\bm p_{10}\cdots \bm p_{n0})^\prime$; $\bm{C}_{t,\infty}:=\text{\upshape m.s.-}\lim_{n\to\infty} n^{-1/2}\bm C_t$.
\end{tabular}
\end{sidewaystable}
\paragraph{Identification of $\widehat{\mathbf H}$ and $\widetilde{\mathbf H}$.}
The following result states the behavior of $\widehat{\mathbf H}$ defined in \eqref{eq:acca} and $\widetilde{\mathbf H}$ defined in \eqref{eq:Htilde}, when imposing further identifying constraints (see Appendix \ref{corol:H00proof} for a proof).
\begin{prop}\label{corol:H00}
Under Assumptions \ref{ass:common} through \ref{ass:ind},
with Assumption \ref{ass:common}(a) holding with rate $\sqrt n$
and Assumption \ref{ass:common}(c-ii) holding with rate $\sqrt T$, as $n,T\to\infty$,
\begin{compactenum}
\item [(I.a)]
if $\bm\Gamma^F=\mathbf I_r$ and
$\bm\Lambda$ is unrestricted,
$\min(n,\sqrt T) \Vert \widehat{\mathbf H}^\prime\widehat{\mathbf H}-\mathbf I_r\Vert=O_{\mathrm P}(1)$;
\item [(II.a)]
if $\bm\Gamma^F=\mathbf I_r$ and
$n^{-1}\bm\Lambda^\prime\bm \Lambda$ is diagonal for all $n\in\mathbb N$,
$\min(n,\sqrt T)\Vert\widehat{\mathbf H}-\mathbf J\Vert=O_{\mathrm P}(1)$;
\item [(III.a)] if
$\bm\Gamma^F=\mathbf I_r$ and
$\bm\Sigma_\Lambda$ is diagonal,
$\min(\sqrt n,\sqrt T)\Vert\widehat{\mathbf H}-\mathbf J\Vert=O_{\mathrm P}(1)$;
\item [(IV.a)]
if $T^{-1}\bm F^\prime\bm F=\mathbf I_r$ for all $T\in\mathbb N$ and
$\bm\Lambda$ is unrestricted,
$\min(n, T)\Vert \widehat{\mathbf H}^\prime\widehat{\mathbf H} -\mathbf I_r \Vert = O_{\mathrm P}(1)$;
\item [(V.a)]
if $T^{-1} \bm F^\prime \bm F=\mathbf I_r$ and
$n^{-1}\bm\Lambda^\prime\bm \Lambda$ is diagonal for all $n,T\in\mathbb N$,
$\min(n, T)\Vert\widehat{\mathbf H}-\mathbf J\Vert=O_{\mathrm P}(1)$;
\item [(VI.a)]
if $T^{-1} \bm F^\prime \bm F=\mathbf I_r$ for all $T\in\mathbb N$ and
$\bm\Sigma_\Lambda$ is diagonal,
$\min(\sqrt n, T)\Vert\widehat{\mathbf H}-\mathbf J\Vert=O_{\mathrm P}(1)$;
\item [(I.b)]
if $\bm\Sigma_\Lambda=\mathbf I_r$
and $\bm F$ is unrestricted,
$\min(\sqrt n, T)\Vert \widehat{\mathbf H}^{^\prime}\widehat{\mathbf H} - n^{-1}\widehat{\mathbf M}^x \Vert = O_{\mathrm P}(1)$;
\item [(II.b)]
if $\bm\Sigma_\Lambda=\mathbf I_r$ and
$T^{-1}\bm F^\prime \bm F$ is diagonal for all $T\in\mathbb N$,
$\min(\sqrt n, T)\Vert\widehat{\mathbf H}-\mathbf J(n^{-1}\widehat{\mathbf M}^x)^{1/2}\Vert=O_{\mathrm P}(1)$;
\item [(III.b)] if
$\bm\Sigma_\Lambda=\mathbf I_r$ and
$\bm\Gamma^F$ is diagonal,
$\min(\sqrt n, \sqrt T)\Vert\widehat{\mathbf H}-\mathbf J(n^{-1}\widehat{\mathbf M}^x)^{1/2}\Vert=O_{\mathrm P}(1)$;
\item [(IV.b)] if
$n^{-1}\bm\Lambda^\prime\bm \Lambda=\mathbf I_r$ for all $n\in\mathbb N$ and
$\bm F$ is unrestricted,
$\min(n,T)\Vert \widehat{\mathbf H}^{^\prime}\widehat{\mathbf H} - n^{-1}\widehat{\mathbf M}^x \Vert = O_{\mathrm P}(1)$;
\item [(V.b)] if
$n^{-1}\bm\Lambda^\prime\bm \Lambda=\mathbf I_r$ and
$T^{-1} \bm F^\prime \bm F$ is diagonal for all $n,T\!\in\mathbb N$,
$\min(n,T)\Vert\widehat{\mathbf H}-\mathbf J(n^{-1}\widehat{\mathbf M}^x)^{1/2}\Vert=O_{\mathrm P}(1)$;
\item [(VI.b)] if
$n^{-1}\bm\Lambda^\prime\bm \Lambda=\mathbf I_r$ for all $n\in\mathbb N$ and
$\bm\Gamma^F$ is diagonal,
$\min(n,\sqrt T)\Vert\widehat{\mathbf H}-\mathbf J(n^{-1}\widehat{\mathbf M}^x)^{1/2}\Vert=O_{\mathrm P}(1)$;
\end{compactenum}
where $\mathbf J$ is an $r\times r$ diagonal matrix with entries $\pm 1$, depending on $n$ and $T$.
Moreover, all statements hold also for $\widetilde{\mathbf H}$.
\end{prop}
As an immediate consequence of Proposition \ref{corol:H00} notice that,
by means of \eqref{eq:sameHsw}, \eqref{eq:sameHD}, and \eqref{eq:sameHD3}, we can derive also the properties of $\wideparen {\mathbf H}$ and $\bar{\mathbf H}$, which are the matrices determining identification under approaches A.2 and B.2. In particular, the results for $\widehat {\mathbf H}$ and $\widetilde{\mathbf H}$ under the constraints (a) in Proposition \ref{corol:H00} hold for $\wideparen {\mathbf H}$ and $\bar{\mathbf H}$ but under the constraints (b).
Although in some cases we impose non-asymptotic identifying conditions, we can derive only asymptotic results. Indeed, $\widehat{\mathbf H}$ depends on estimated quantities so its properties hold only asymptotically. The different convergence rates depend on the columns of the loadings being exactly or just asymptotically orthonormal.
Under parts (II.a), (III.a), (V.a) and (VI.a), we see that $\widehat{\mathbf H}$ can be reduced to an asymptotically diagonal matrix of signs $\mathbf J$. The sign indeterminacy, given by $\mathbf J$, can always be fixed, for example, as follows. Let $\iota_j$ be the first index of the $j$th column of $\bm\Lambda$ such that $\lambda_{\iota_j j}\ne 0$, $j=1,\ldots, r$. Then, we assume $\lambda_{\iota_j j}> 0$, for $j=1,\ldots,r$. This condition, which is also imposed, although implicitly, by \citet[proof of Proposition 1]{Bai03}, implies that $\mathbf J=\mathbf I_r$ in Proposition \ref{corol:H00}.
Given that the assumptions made are the same as those under which Propositions \ref{prop:L2} and \ref{prop:F2} hold, the results in Proposition \ref{corol:H00} are directly applicable to the consistency results proved therein. Specifically, under all (a) cases the PC estimators are consistent for the true loadings and factors as defined in Proposition \ref{corol:K00}. Indeed, by Propositions \ref{prop:L2} and \ref{prop:F2},
\begin{align}
\Vert \widehat{\bm\lambda}_i-{\bm\lambda}_i \Vert&\le \Vert\widehat{\bm\lambda}_i-\widehat{\mathbf H}^\prime{\bm\lambda}_i \Vert + \Vert\widehat{\mathbf H}^\prime-\mathbf I_r\Vert\, \Vert{\bm\lambda}_i\Vert = O_{\mathrm P}\left(\max\left(\frac 1n,\frac 1{\sqrt T}\right)\right) + O_{\mathrm P}\left(\eta_{nT}\right),\label{eq:finita1}\\
\Vert\widehat{\mathbf F}_t- {\mathbf F}_t\Vert &= \Vert \widehat{\mathbf F}_t-\widehat{\mathbf H}^{-1}{\mathbf F}_t\Vert + \Vert \widehat{\mathbf H}^{-1}-\mathbf I_r \Vert\, \Vert\mathbf F_t \Vert
= O_{\mathrm P}\left(\max\left(\frac 1{\sqrt n},\frac 1{ T}\right)\right) + O_{\mathrm P}\left(\eta_{nT}\right).\label{eq:finita2}
\end{align}
However, the term $\eta_{nT}$ which is given in Proposition \ref{corol:K00} depends on the imposed constraints. This has important implications for inference.
Under part (II.a) $\eta_{nT}=\max(n^{-1},T^{-1/2})$ this means that $ \sqrt T(\widehat{\mathbf H}^\prime-\mathbf J) = O_{\mathrm P}(1)$ and the CLT for the loadings in Theorem \ref{th:CLTL} cannot hold, but $\sqrt n(\widehat{\mathbf H}^{-1}-\mathbf J)= O_{\mathrm P}(\sqrt {n/T})$, so the CLT for the factors in Theorem \ref{th:CLTF} can still hold provided $n/T\to 0$, as $n,T\to\infty$, which is a stronger constraint. Under part (VI.a) $\eta_{nT}=\max(n^{-1/2},T^{-1})$ which implies that Theorem \ref{th:CLTF} for the factors cannot hold but Theorem \ref{th:CLTL} for the loadings can still hold provided $T/n\to 0$, as $n,T\to\infty$. Under part (III.a) $\eta_{nT}=\max(n^{-1/2},T^{-1/2})$ and neither Theorem \ref{th:CLTL} nor Theorem \ref{th:CLTF} can hold and we conjecture that if any asymptotic normality can be proved the asymptotic covariance will be larger than in those theorems.
Still, under (II.a), (III.a) or (VI.a) we can test for hypothesis of the form $\text H_0: \mathbf R \bm\lambda_i=\mathbf 0_q$ with $\mathbf R$ being $q\times r$ and full-rank for some $q\le r$. Indeed, in these cases, under $\text H_0$, by following \eqref{eq:CLTL} we get:
\begin{align}
\sqrt T\, \mathbf R\widehat{\bm\lambda}_i\to_d\mathcal N\left(\mathbf 0_r, \mathbf R(\bm \Gamma^F)^{-1}\bm \Phi_i (\bm \Gamma^F)^{-1}\mathbf R^\prime \right),\nonumber
\end{align}
because ${\text{P-lim}}_{n,T\to\infty} \widehat{\mathbf H}^\prime=\mathbf I_r$, and we can immediately build a Wald statistics for $\text H_0$. Similarly, we can build pointwise $(1-\alpha)$-confidence intervals for the factors given by $\{\widehat{F}_{jt}\pm z_{1-\alpha/2}{T^{-1/2} [\widehat{\bm \Pi}_t^{\text{\tiny OLS}}]_{jj}^{1/2}} \}$, $j=1,\ldots, r$.
However, to make inference for general hypothesis, we have to resort to the constraints in part (V.a), which is also considered by \citet[Theorem 1]{baing13}. In this case, following again \eqref{eq:CLTL} and \eqref{eq:CLTF} and combining them with \eqref{eq:finita1} and \eqref{eq:finita2},
\begin{align}
\sqrt T(\widehat{\bm\lambda}_i-{\bm\lambda}_i )&=\sqrt T(\widehat{\bm\lambda}_i-\widehat{\mathbf H}^\prime{\bm\lambda}_i ) + o_{\mathrm P}(1)\to_d\mathcal N\left(\mathbf 0_r,
(\bm\Gamma^F)^{-1}
\bm\Phi_i
(\bm\Gamma^F)^{-1}
\right),\label{eq:CLTLesatto}\\
\sqrt n(\widehat{\mathbf F}_t- {\mathbf F}_t) &= \sqrt n(\widehat{\mathbf F}_t-\widehat{\mathbf H}^{-1}{\mathbf F}_t) + o_{\mathrm P}(1)\to_d\mathcal N\left(\mathbf 0_r, (\bm\Sigma_\Lambda)^{-1}\bm\Gamma_t(\bm\Sigma_\Lambda)^{-1}\right),\label{eq:CLTFesatto}
\end{align}
where used also the fact that $\Vert\bm\lambda_i\Vert = O(1)$, because of Assumption \ref{ass:common}(a), and $\Vert \mathbf F_t\Vert = O_{\mathrm P}(1)$, because of Lemma \ref{lem:FTLN}(ii).
It is then clear that under condition (V.a), neither the location nor the covariance of the asymptotic distribution depend on $\widehat{\mathbf H}$, or its limit, anymore.
Still, the price to be paid is not irrelevant. Indeed, condition (V.a) is stated for a sample quantity, and, thus, in order for it to hold, we should either consider the factors as a deterministic sequence or we should interpret the CLTs in \eqref{eq:CLTLesatto} and \eqref{eq:CLTFesatto} as conditional on a realization of the factors $\mathbf F_1,\ldots, \mathbf F_T$ (see also the comments made above and in \citealp[page~439]{baili12}, in the case of maximum likelihood estimation). Under this interpretation, we can then impose condition (V.a) and use \eqref{eq:CLTLesatto} and \eqref{eq:CLTFesatto} to make inference not only on the loadings, but also on the factors since now $\mathbf F_t$ is no more random.
\paragraph{Identification of $\bm{\mathcal H}$.}
To conclude, recall from Proposition \ref{prop:L} that $\bm{\mathcal H}$ depends on $T$ only through a sign matrix $\mathbf J$. If we let $\widehat{\mathbf V}^{x}_j$ and $\mathbf V_j^{C}$, $j=1,\ldots, r$, be the columns of $\widehat{\mathbf V}^{x}$ and $\mathbf V^{C}$, respectively, we can assume, without loss of generality, that we always choose sample eigenvectors $\widehat{\mathbf V}^x_j$ such that $\widehat{\mathbf V}^{x\prime}_j \mathbf V_j^{C}\ge 0$ for all $j=1,\ldots, r$. Hereafter, we work under this condition and, therefore, $\mathbf J=\mathbf I_r$ and $\bm{\mathcal H}$ depends only on population quantities. We then derive the implications for $\bm{\mathcal H}$ as defined in Proposition \ref{prop:L} (see Appendix \ref{corol:K00Hproof} for a proof).\footnote{If we did not set $\mathbf J=\mathbf I_r$, then $\bm S$ in Proposition \ref{corol:K00H}(II.a) would depend also on $T$
while convergence in Proposition \ref{corol:K00H}(III.a) would be in probability only, thus requiring also $T\to\infty$.
Moreover, $\bm S$ and $\bm{\mathcal S}_0$ in Proposition \ref{corol:K00H} would still be diagonal matrices with entries $\pm 1$, but, in general, they would be different from $\bm S$ and $\bm{\mathcal S}_0$ in Proposition \ref{corol:K00}.}
\begin{prop}\label{corol:K00H}
Under Assumptions \ref{ass:common} and \ref{ass:eval}, and if $\mathbf J=\mathbf I_r$,
\begin{compactenum}
\item [(I.a)] if $\bm\Gamma^F=\mathbf I_r$ and
$\bm\Lambda$ is unrestricted, then,
for all $n\in\mathbb N$, $\bm{\mathcal H}^\prime \bm{\mathcal H}=\mathbf I_r$;
\item [(II.a)] if $\bm\Gamma^F=\mathbf I_r$ and $n^{-1}\bm\Lambda^\prime\bm \Lambda$ is diagonal for all $n\in\mathbb N$, then, for all $n\in\mathbb N$, $\bm{\mathcal H}=\bm S$;
\item [(III.a)] if $\bm\Gamma^F=\mathbf I_r$ and
$\bm\Sigma_\Lambda$ is diagonal, then, $\lim_{n\to\infty}\Vert\bm{\mathcal H}-\bm{\mathcal S}_0\Vert=0$;
\end{compactenum}
where $\bm S$ is an $r\times r$ diagonal matrix with entries $\pm 1$ depending only on $n$, and $\bm{\mathcal S}_0$ is an $r\times r$ diagonal matrix with entries $\pm 1$ independent of $n$ and $T$.
\end{prop}
Here, we consider only the case of orthonormal factors and orthogonal loadings, which is consistent with the PC estimators considered under approach A.1 or, equivalently, B.1. The case of orthonormal loadings and orthogonal factors, consistent with approaches A.2 and B.2, can be easily derived from the proof or Proposition \ref{corol:K00}, hence it is omitted. Furthermore, for simplicity we also do not cover the cases in which we constraint $T^{-1}\bm F^\prime\bm F$.
Part (I.a) shows that if we restrict only the factors to be orthonormal, then $\bm{\mathcal H}$ is an orthogonal matrix. This is straightforward from its definition in Proposition \ref{prop:L}.
Under parts (II.a) and (III.a) we see that $\bm{\mathcal H}$ can be reduced exactly or, at least, asymptotically to a diagonal matrix of signs. By comparing Propositions \ref{corol:K00} and \ref{corol:K00H}, from \eqref{eq:phil3}, we also see that, $\bm{\mathcal H}_{\infty}=\lim_{n\to\infty}\bm{\mathcal H}$.
\singlespacing
{{
\setlength{\bibsep}{.2cm}
\bibliographystyle{chicago}
\bibliography{BL_biblio}
}}
\setcounter{section}{0}
\setcounter{subsection}{0}
\setcounter{equation}{0}
\gdef\thesection{ \Alph{section}}
\gdef\thesubsection{\Alph{section}.\arabic{subsection}}
\gdef\thefigure{\Alph{section}\arabic{figure}}
\gdef\theequation{\Alph{section}\arabic{equation}}
\gdef\varphible{\Alph{section}\arabic{table}}
\gdef\arabic{footnote}{\Alph{section}\arabic{footnote}}