Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.
Asymptotic equivalence of Principal Components and Quasi Maximum Likelihood estimators in Large Approximate Factor Models
center[center omitted — 86 chars of source]
abstractThis paper investigates the properties of Quasi Maximum Likelihood estimation of an approximate factor model for an $n$-dimensional vector of stationary time series.
We prove that the factor loadings estimated by Quasi Maximum Likelihood are asymptotically equivalent, as $n\to\infty$, to those estimated via Principal Components. Both estimators are, in turn, also asymptotically equivalent, as $n\to\infty$, to the unfeasible Ordinary Least Squares estimator we would have if the factors were observed.
We also show that the usual sandwich form of the asymptotic covariance matrix of the Quasi Maximum Likelihood estimator is asymptotically equivalent to the simpler asymptotic covariance matrix of the unfeasible Ordinary Least Squares. All these results hold in the general case in which the idiosyncratic components are cross-sectionally heteroskedastic, as well as serially and cross-sectionally weakly correlated. The intuition behind these results is that as $n\to\infty$ the factors can be considered as observed, thus showing that factor models enjoy a blessing of dimensionality.
Keywords:
Approximate Dynamic Factor Model; Principal Component Analysis; Quasi Maximum Likelihood.
\thispagestyle{empty}
\footnotetext{ Universit\`a di Bologna, [email removed]}
Introduction
Factor models are one of the major dimension reduction techniques used to analyze large panels of time series. Some of their most successful applications are, among many others, in finance (Chamberlain and Rothschild, 1983, connor2006common, ait2017using, kim2019factor), and macroeconomics (Stock and Watson, 2002b,
BBE05, FHLR05, de2008forecasting, Nowcasting).
Let us assume to observe an $n$-dimensional zero-mean stochastic process over $T$ periods: $\{x_{it},\, i=1,\ldots , n,\, t=1,\ldots, T\}$, such that
align[align omitted — 123 chars of source]
where $\bm\lambda_i=(\lambda_{i1}\cdots\lambda_{ir})^\prime$ and $\mathbf F_t=(F_{1t}\cdots F_{rt})^\prime$ are $r$-dimensional unobserved vectors, called loadings and factors, respectively, and with $r<\min(n,T)$. We also call $\xi_{it}$ the idiosyncratic component and
$\chi_{it}=\bm \lambda_i^\prime \mathbf F_t$ the common component of the $i$th observed variable at time $t$. Both the factors and the idiosyncratic components are allowed to be serially correlated. Furthermore, the idiosyncratic components are also allowed to be cross-sectionally correlated. We call such model an approximate factor model. This is the class of factor models we consider in this paper and studied, e.g., by Bai03. It is a restricted version of the generalized dynamic factor model originally proposed by FHLR00, where the factors are loaded with lags (the loadings are linear filters) and not just contemporaneously. Another popular alternative, not considered in this paper, is the factor model studied, e.g., by lam2012factor, where the idiosyncratic components are instead assumed to be serially uncorrelated.
There are two main ways to estimate the factor loadings in (ref). First, by Principal Component (PC) analysis stockwatson02JASA,Bai03,FLM13, and, second, by Quasi Maximum Likelihood (QML) estimation baili12,baili16. In both cases, we can first estimate the loadings and then estimate the factors by linear projection, possibly weighted by the idiosyncratic variances, of the observables onto the estimated loadings. The PC estimator of the loadings is easily implementable since it just requires to compute the $r$ leading eigenvectors and eigenvalues of the sample covariance matrix of the observables. The QML estimator of the loadings does not have a closed form solution, so numerical maximization is required.
It is typically defined as the maximizer of a mis-specified log-likelihood where the idiosyncratic components are treated as if they were uncorrelated even if the true ones are correlated.
Although QML is the classical way to estimate a factor model, dating back more than fifty years ago lawleymaxwell71,MBQML, in the recent years, PC analysis has gained popularity, given its non-parametric nature and ease of implementation. Nevertheless, QML estimation still retains an important role in factor analysis
since for many reasons. For example, it allows to easily impose constraints on the loadings CGM16,DCGF2021, and it fully addresses idiosyncratic cross-sectional heteroskedasticity, while the PC method does not baili12,baili16.
In this paper, we compare the PC and QML estimators of the loadings and we show that, under a minimal unified set of standard assumptions and identifying constraints, the two estimators are asymptotically equivalent as $n\to\infty$ (see Theorem (ref)).
It is well known that both estimators of the loadings are $\min(n,\sqrt T)$-consistent as $n,T\to\infty$ and are also asymptotically normal if we assume $n^{-1}\sqrt T\to 0$. Standard references for these results are, among others, Bai03 for PC, and baili16 for QML. Proving the asymptotic properties of the PC estimator is long but straightforward. However, things are much more complicated for the QML estimator, essentially because no closed form expression exists for this estimator. Moreover, the existing proofs for the two estimators do not make use of the same assumptions nor of the same identification constraints needed to uniquely identify the loadings.
Our main result has some important implications. First, if $n^{-1}\sqrt T\to 0$, as $n,T\to\infty$, then the QML estimator has the same asymptotic distribution as the unfeasible OLS estimator we would obtain if the factors were observed. Therefore, we are able to prove asymptotic normality of the QML estimator in an indirect and easy way (see Theorem (ref)). Our approach provides a much simpler alternative to the approach used by baili16, and it does not require to study also the properties of the QML estimators of the idiosyncratic variances. Second, it nests the case of spherical idiosyncratic components tippingbishop99, in which case the equivalence of PC and QML estimator is often quoted, but, to the best of our knowledge, it has never been formally proved, at least under the standard minimal set of assumptions used in this paper. Related to these two points, we stress that our result is, in fact, more general since it holds without the need of considering QML estimation based on a mis-specified log-likelihood with a diagonal idiosyncratic covariance. Third, and last, our result paves the way towards studying the asymptotic properties of the QML estimator of a factor model where we also explicitly model the dynamics of the factors, and for which only partial results exist (DGRqml, baili16).
The theoretical analysis of approximate factor models requires a double asymptotic framework. Indeed, only if $n,T\to\infty$ we can consistently estimate the eigenvectors of the covariance matrix of the data which are needed for PC estimation, and we can control for the mis-specifications introduced in the log-likelihood when considering QML estimation. This means that QML estimation of approximate factor models does not fall into the framework of classical QML estimation, where $n$ is fixed.
In a second contribution, we reconcile the asymptotic distribution of the QML estimator of the loadings, derived in this paper, with the results of classical likelihood-based inference. In particular, we show that, when evaluated in the true value of the parameters or in the value of the QML estimator, the first and second derivatives of the factor model log-likelihood we maximize are asymptotically, as $n,T\to\infty$, equivalent to the first and second derivatives of the log-likelihood we would have if the factors were observed (see Theorem (ref)). This result also allows us to compute simple estimators of the asymptotic covariance matrix of the estimated loadings without the need of computing the derivatives of the log-likelihood which have rather long and complex expressions.
The paper is organized as follows. In Section (ref) we present the model and all assumptions. In Section (ref) we reconsider the PC estimator of the loadings and in Theorem (ref) we derive its asymptotic properties using a new approach equivalent to the one proposed by Bai03 but which is more convenient for the present work. In Section (ref) we review the existing results on the QML estimator of the loadings by baili16. In Section (ref) we present our first contribution in Theorem (ref) where we prove that PC and QML estimators of the loadings are asymptotically equivalent.
In Section (ref) we present our second contribution in Theorem (ref), where we prove the asymptotically equivalence of the first and second derivatives of the factor model log-likelihood and the log-likelihood of a model with observed factors.
In Section (ref) we briefly discuss estimation of factors.
In Section (ref) we provide simulation results confirming our theoretical results. In Section (ref) we conclude.
In the Supplementary Material we prove all other main and auxiliary theoretical results.
Model and Assumptions
Given the $n$-dimensional vector $\mathbf x_t=(x_{1t}\cdots x_{nt})^\prime$, we can also write model (ref) in vector notation:
\[
\mathbf x_t = \bm\Lambda\mathbf F_t +\bm \xi_t,\quad t=1,\ldots, T,
\]
where $\bm\xi_{t}=(\xi_{1t}\cdots\xi_{nt})^\prime$ is the $n$-dimensional vector of idiosyncratic components and $\bm\Lambda=(\bm\lambda_1\cdots\bm\lambda_n)^\prime$ is the $n\times r$ matrix of factor loadings. We call $\bm\chi_t=\bm\Lambda\mathbf F_t$ the vector of common components.
Moreover, we can collected all observations into the $T\times n$ matrix $\bm X=(\mathbf x_1\cdots\mathbf x_T)^\prime$, and we can write model (ref) also in matrix notation:
equation[equation omitted — 74 chars of source]
where $\bm \Xi=(\bm\xi_1\cdots\bm \xi_T)^\prime$ is a $T\times n$ matrix of idiosyncratic components, and $\bm F=(\mathbf F_1\cdots\mathbf F_T)^\prime$ is the $T\times r$ matrix of factors.
The following assumptions are similar to those made by Bai03 for PC estimation while slightly differ from those in baili16 for QML estimation. Throughout, we highlight the main differences or similarities.
We start by characterizing the common component by means of the following assumption.
ass[common component] $\,$
\begin{compactenum}[(a)]
• $\lim_{n\to\infty}\Vert n^{-1}\bm\Lambda^\prime\bm\Lambda-\bm\Sigma_{\Lambda}\Vert=0$, where $\bm\Sigma_{\Lambda}$ is $r\times r$ positive definite, and, for all $i\in\mathbb N$,
$\Vert\bm\lambda_i\Vert\le M_\Lambda$ for some finite positive real $M_\Lambda$ independent of $i$.
• For all $t\in\mathbb Z$, $\mathbb{E}[\mathbf F_{t}]=\mathbf 0_r$ and $\bm\Gamma^F=\mathbb{E}[\mathbf F_t\mathbf F_t^\prime]$ is $r\times r$ positive definite and $\Vert\bm\Gamma^F\Vert\le M_F$ for some finite positive real $M_F$ independent of $t$.
• \begin{inparaenum}
• For all $t\in\mathbb Z$, $\mathbb{E}[\Vert \mathbf F_t\Vert^4]\le K_F$ for some finite positive real $K_F$ independent of $t$; \\
• for all $i,j=1,\ldots, r$, all $s=1,\ldots, T$, and all $T\in\mathbb N$,
\[
\mathbb{E}\left[\left\vert\frac 1{\sqrt{T}}\sum_{t=1}^T\left\{F_{is}F_{jt}-\mathbb{E}[F_{is}F_{jt}]\right\} \right\vert^2\right]
\le C_F
\]
for some finite positive real $C_F$ independent of $i$, $j$, $s$, and $T$.
\end{inparaenum}
• There exists an integer $N$ such that for all $n> N$, $r$ is a finite positive integer, independent of $n$.
\end{compactenum}
Part (a) is standard Bai03. It implies that, asymptotically, as $n\to\infty$, the loadings matrix has asymptotically maximum column rank $r$, and that
for any given $n\in\mathbb N$, each factor has a finite contribution to each of the $n$ observed series (upper bound on $\Vert\bm\lambda_i\Vert$). A similar requirement is in baili16.
We only consider non-random factor loadings for simplicity and in agreement with classical factor analysis where the loadings are the parameters of the model (lawleymaxwell71).
Part (b) assumes that the factors have zero mean and have a finite full-rank covariance matrix $\bm\Gamma^F$, so they are non-degenerate. In part (c-i) we assume finite 4th order moments of the factors. These are standard requirements Bai03.
Part (c-ii) is very general, it implies that the sample covariance matrix of the factors is a $\sqrt T$-consistent estimator of its population counterpart $\bm\Gamma^F$, which has full rank because of part (b). It is immediate to see that it is equivalent to asking for 4th order summable cross-cumulants which a necessary and sufficient condition for consistent estimation (hannan).
This approach is high-level in that it does not make any specific assumption on the dynamics of $\{\mathbf F_t\}$. Obviously this implies the usual assumption of convergence in probability made in this literature: $\mathrm P\text{-}\lim_{T\to\infty}\Vert T^{-1}{\bm F^\prime\bm F}-\bm\Gamma^F\Vert=0$ Bai03. Nothing would change in our proofs if we directly assumed this latter condition instead of part (c-ii). Notice that baili16 treat the factors as being deterministic, but essentially make the same assumption as our parts (b) and (c).
Part (d) implies the existence of a finite number of factors. In particular, the number of common factors, $r$, is identified only as $n\to\infty$. Here $N$ is the minimum number of series we need to be able to identify $r$ so that $r\le N$. Hereafter, when we say “for all $n\in\mathbb N$” we always mean that $n>N$ so that $r$ can be identified. In practice, we must always work with $n$ such that $r<n$. Moreover, because PC estimation is based on eigenvalues of an $n\times n$ matrix estimated using $T$ observations, then we must also have samples of size $T$ such that $r<T$. Therefore, sometimes it is directly assumed that $r<\min(n,T)$.
To characterize the idiosyncratic component, we make the following assumptions.
ass[idiosyncratic component]
$\,$
\begin{compactenum}[(a)]
• For all $i\in\mathbb N$ and all $t\in\mathbb Z$, $\mathbb{E}[\xi_{it}]= 0$ and $\sigma_i^2=\mathbb{E}[\xi_{it}^2]$ is such that $C_\xi\le \sigma_i^2\le C_\xi^\prime$ for some finite positive reals $C_\xi$ and $C_\xi^\prime$ independent of $i$ and $t$.
• For all $i,j\in\mathbb N$, all $t\in\mathbb Z$, and all $k\in\mathbb Z$, $\vert \mathbb{E}[\xi_{it}\xi_{j,t-k}]\vert\le \rho^{\vert k\vert} M_{ij}$, where $\rho$ and $M_{ij}$ are finite positive reals independent of $t$ such that $0\le \rho <1$, $\sum_{j=1,j\ne i}^n M_{ij}\le M_\xi$, and $\sum_{i=1, i\ne j}^n M_{ij}\le M_\xi$ for some finite positive real $M_{\xi}$ independent of $i$, $j$, and $n$.
• \begin{inparaenum}
• For all $i=1,\ldots, n$, all $t=1,\ldots, T$, and all $n,T\in\mathbb N$, $\mathbb{E}[\xi_{it}^4]\le Q_\xi$ for some finite positive real $Q_\xi$ independent of $i$ and $t$;\\
• for all $j=1,\ldots, n$, all $s=1,\ldots, T$, and all $n,T\in\mathbb N$,
\[
\mathbb{E}\left[\left\vert\frac 1{\sqrt{nT}} \sum_{i=1}^n\sum_{t=1}^T\left\{\xi_{is}\xi_{jt}-\mathbb{E}[\xi_{is}\xi_{jt}]\right\} \right\vert^2\right]
\le K_\xi
\]
for some finite positive real $K_\xi$ independent of $j$, $s$, $n$, and $T$.
\end{inparaenum}
\end{compactenum}
By part (a), we have that the idiosyncratic components have zero mean. This, jointly with Assumption (ref)(b) by which $\mathbb{E}[\mathbf F_t]=\mathbf 0_r$, implies that we are implicitly assuming that each observed series has zero mean, i.e., $\mathbb{E}[ x_{it}]=0$ (this is without loss of generality), and strictly positive variance, i.e., $\mathbb V\text{ar}(x_{it})>0$, for all $i\in\mathbb N$ (baili16). For the case of non-zero mean see Remark (ref).
Part (b) has a twofold purpose. First, it limits the degree of serial correlation of the idiosyncratic components by imposing geometric decay of their autocovariances. Second, it also limits the degree of cross-sectional correlation between idiosyncratic components, which is usually assumed in approximate factor models.
In particular, by setting $k=0$, it follows also that for all $i\in\mathbb N$, $\sigma_i^2 \le M_\xi$. Thus, all idiosyncratic components have finite variance. This jointly with Assumptions (ref)(a) and (ref)(b) implies that
each observed time series has finite variance, i.e., $\mathbb V\text{ar}(x_{it})<\infty$, for all $i\in\mathbb N$. In Lemma (ref) we show that part (b) implies all usual conditions on second order moments typically found in the literature (Bai03, and baili16).
Part (c-i) assumes finite 4th order moments of the idiosyncratic components. Part (c-ii) gives summability conditions across the cross-section and time dimensions for the 4th order cumulants of $\{\xi_{it}\}$. Jointly with Assumptions (ref)(a) and (ref)(a) it nests the requirement in baili16. It implies that the sample (auto)covariances between $\{\xi_{it}\}$ and $\{\xi_{jt}\}$ are $\sqrt T$-consistent estimators of their population counterparts. In particular, by choosing $s=t$ we see that we can consistently estimate the $(i,j)$th entry of the idiosyncratic covariance matrix $\bm\Gamma^\xi$. Notice that, contrary to the existing literature (Bai03, and baili16), there is no need to ask for finite 8th order moments and cumulants.
We then make a series of identifying assumptions.
ass[independence]
The processes
$\{\xi_{it},\, i\in\mathbb N,\, t\in\mathbb Z\}$ and $\{F_{jt},\, j=1,\ldots, r,\, t\in\mathbb Z\}$ are mutually independent.
This assumption obviously implies that the factors and, therefore, the common components are independent of the idiosyncratic components at all leads and lags and across all units. This is compatible for example with a macroeconomic interpretation of factor models, according to which the factors driving the common component are independent of the idiosyncratic components representing measurement errors or local dynamics. This assumption is made for simplicity and could be easily relaxed Bai03.
Let $\bm\Gamma^\chi=\mathbb{E}[\bm\chi_t\bm\chi_t^\prime]=\bm\Lambda\bm\Gamma^F \bm\Lambda^\prime$ with $r$ largest eigenvalues $\mu_j^\chi$, $j=1,\ldots, r$, sorted in decreasing order. In Lemma (ref)(iv) we prove that,
for all $j=1,\ldots,r$,
equation[equation omitted — 161 chars of source]
where $\underline C_j$ and $ \overline C_j$ are finite positive reals. This means we consider only strong factors, i.e., fully pervasive. Furthermore in Lemma (ref)(v) we prove that the largest eigenvalue of $\bm\Gamma^\xi$ is such that
equation[equation omitted — 93 chars of source]
where $C_\xi^\prime$ and $M_\xi$ are defined in Assumptions (ref)(a) and (ref)(b), respectively.
Because of Weyl's inequality (see, e.g., MK04, Theorem 1), conditions (ref)-(ref) and Assumption (ref) imply the eigengap in the eigenvalues $\mu_j^x$, $j=1,\ldots,n$, of $\bm\Gamma^x=\mathbb{E}[\mathbf x_t\mathbf x_t^\prime]=\bm\Gamma^\chi+\bm\Gamma^\xi$ (see Lemma (ref)(vi)):
align[align omitted — 217 chars of source]
This property allows us to identify, asymptotically, as $n\to\infty$, the number of factors $r$ (see, e.g., baing02, onatski10, trapani2018randomized, among many others). And, therefore, as $n\to\infty$ we also identify the common and idiosyncratic components.
ass[distinct eigenvalues]
For all $n\in\mathbb N$ and all $i=1,\ldots, r-1$, $\mu_i^\chi > \mu_{i+1}^\chi$.
This assumption is needed in order to identify the eigenvectors in PC estimation.
Note that it implies that $\overline C_j<\underline C_{j-1}$, $j=2,\ldots, r$, in (ref) and that the $r$ eigenvalues of $\bm\Sigma_\Lambda\bm\Gamma^F$ are distinct Bai03. Indeed, these coincide with the non-zero eigenvalues of $\lim_{n\to\infty}n^{-1}{\bm\Gamma^\chi}$ which are given by $\lim_{n\to\infty}n^{-1}{\mathbf M^\chi}=(\bm\Sigma_\Lambda)^{1/2}\bm\Gamma^F(\bm\Sigma_\Lambda)^{1/2}$.
The factors and the loadings can be identified by means of the following assumption.
ass[Identification]
$\,$
\begin{inparaenum}[(a)]
• For all $n\in\mathbb N$, $n^{-1}{\bm\Lambda^\prime\bm\Lambda}$ diagonal.
• For all $T\in\mathbb N$, $T^{-1}{\bm F^\prime\bm F}=\mathbf I_r$.
\end{inparaenum}
This assumption a standard requirement in PC based exploratory factor analysis (see, e.g, baing13). It has some important and usueful implications in terms of identification. In particular, in Proposition (ref) we prove that, for all $n\in\mathbb N$, $n^{-1}{\bm\Lambda^\prime\bm\Lambda} = n^{-1}{\mathbf M^\chi}$
and
$\bm\Lambda=\mathbf V^\chi(\mathbf M^\chi)^{1/2}$, where $\mathbf M^\chi$ is the $r\times r$ diagonal matrix of the nonzero eigenvalues of $\bm\Gamma^\chi$ (sorted in decreasing order) and $\mathbf V^\chi$ is the $n\times r$ matrix having as columns the corresponding normalized eigenvectors.
Notice that part (a) differs from baili16, where it is assumed that $n^{-1}{\bm\Lambda^\prime(\bm\Sigma^\xi)^{-1}\bm\Lambda}$ is diagonal, for all $n\in\mathbb N$, with $\bm\Sigma^\xi$ being the diagonal matrix with entries the idiosyncratic variances $\sigma_i^2$, $i=1,\ldots, n$. This is a common requirement in QML based exploratory factor analysis. One of the aims of this paper is to reconsider QML estimation under Assumption (ref)(a).
Part (b) is instead also assumed in baili16 but for deterministic factors.
Clearly since Assumption (ref) concerns only the product $\bm\Lambda^\prime\bm\Lambda$, it allows us to identify the loadings only up to right multiplication by a diagonal matrix with entries $\pm 1$, i.e., the columns of $\bm\Lambda$ are identified only up to a sign. We can fix such sign by means of the following assumption.
ass[Global identification]
For all $j=1,\ldots, r$, one of the two following conditions holds:
\begin{inparaenum}[(a)]
• $\lambda_{1j}> 0$; or
• $F_{j1}> 0$.
\end{inparaenum}
Since loadings are considered as deterministic and are estimated as eigenvectors, part (a) is more natural and can be easily imposed just by setting the sign of the eigenvectors of the sample covariance matrix $\widehat{\bm\Gamma}^x$ accordingly. As a consequence of Assumption (ref) both the loadings and the factors are globally identified baing13.
Finally, in order to derive the asymptotic distribution of the considered estimators of the loadings it is common to assume the following assumption (Bai03, and baili16).
ass[Central limit theorem]
For all $i\in\mathbb N$, as $T\to\infty$,
\[
\frac 1{\sqrt T}\sum_{t=1}^T
\mathbf F_t \xi_{it}
\to_d
\mathcal N\left(\mathbf 0_r, \bm\Phi_i
\right),
\]
where $\bm\Phi_i=\lim_{T\to\infty}T^{-1}\sum_{t,s=1}^T
\mathbb{E}_{}[\mathbf F_t\mathbf F_s^\prime\xi_{it}\xi_{is}]$.
There are many ways to derive this condition from more primitive assumptions, for example, by assuming strong mixing factors and idiosyncratic components both with finite $4+\epsilon$ moments.
rem{
As discussed above, all assumptions for PC estimation by Bai03 are implied or equivalent to ours. Regarding QML estimation and the assumptions in baili16, we do not need their Assumption D, which requires the QML estimators of the idiosyncratic variances, $\sigma_i^2$, $i=1,\ldots, n$, to be finite and strictly positive, and we do not require the moment condition E.3, which is needed only for estimation of $\sigma_i^2$ and it is not necessary to prove our results. Assumptions E.4, E.5, E.6, F.2, and F.3 in baili16 are needed only for studying the behavior of the estimated factors so are not needed here. All other assumptions in baili16 are equivalent or nested into ours, with the crucial exception of their identifying condition on the loadings which, as noted above, differs from ours.}
rem{
Allowing for data with non-zero mean simply amounts to adding a constant term, $\alpha_i\ne 0$, to (ref), so that $x_{it}=\alpha_i+\bm\lambda_i^\prime\mathbf F_t+\xi_{it}$. All estimators described in the following retain the same properties even in this case. Indeed, estimation of such model by PCs simply requires to center data first, i.e., to work with $x_{it}-\bar {x}_i$ where $\bar x_i=T^{-1}\sum_{t=1}^T x_{it}$ is clearly a consistent estimator of $\alpha_i$. As for QML estimation, it is straightforward to see that $\bar x_i$ is precisely the QML estimator of $\alpha_i$ and thus it is enough to work with the likelihood of the centered data baili12.
}
Principal Component Analysis
The PC estimators of $\bm\Lambda$ and $\bm F$ are the solutions of the following minimization:
align[align omitted — 495 chars of source]
where $\underline{\bm\Lambda}$ and $\underline{\bm F}$ indicate generic values of the loadings and the factors, respectively, and it is intended that they also satisfy Assumptions (ref), (ref), and (ref). Given the identifying constraints in Assumption (ref), the solution to (ref) can be found in two steps. There are two equivalent ways to do that:
inparaenum• based on the $n\times n$ matrix $\bm X^\prime\bm X$ solve first for loadings and then get the factors by projecting $\bm X$ onto the estimated loadings;
• based on the $T\times T$ matrix $\bm X\bm X^\prime$ solve first for the factors and then get the loadings by projecting $\bm X$ onto the estimated factors.
Although the majority of the literature on factor models considers approach B, thus estimating the factors as normalized eigenvectors, and it derives the theory accordingly Bai03, in the rest of the paper we follow approach A. The main reason for this choice is that approach A does not require to estimate the factors first, which is convenient given that our focus is on estimation of the loadings. Notice that approach A is also the classical one (see, e.g., lawleymaxwell71, mardia1979multivariate, jolliffe2002principal).
Nevertheless, which approach to choose is just a matter of taste and it has no theoretical or practical implications. Indeed, numerically both approaches give the same results and all the following theory can be equivalently derived under approach B. Incidentally, by choosing approach A, we also contribute to PC literature with new proofs, alternative to those in Bai03.
More in detail, consider the $n\times n$ sample covariance matrix (recall that $\mathbb{E}[\mathbf x_t]=\mathbf 0_n$ by assumption) $\widehat{\bm\Gamma}^x= T^{-1}{\bm X^\prime\bm X}$, having its $r$ largest eigenvalues collected in the $r\times r$ diagonal matrix $\widehat{\mathbf M}^x$ (sorted in descending order) with corresponding normalized eigenvectors as columns of the $n\times r$ matrix $\widehat{\mathbf V}^x$.
For any given $\underline{\bm\Lambda}$ the solution of (ref) for $\underline{\bm F}$ is just the linear projection $\underline{\bm F}=\bm X\underline{\bm\Lambda}(\underline{\bm\Lambda}^\prime\underline{\bm\Lambda})^{-1}$. By substituting this expression in (ref) we have
align[align omitted — 712 chars of source]
Now since by construction each column of $ \underline{\bm\Lambda}
(\underline{\bm\Lambda}^\prime\underline{\bm\Lambda})^{-1/2}$ is normalized (since we assumed ${\underline{\bm\Lambda}^\prime\underline{\bm\Lambda}}$ to be diagonal), then the above maximizaton, once solved, should return the $r$ largest eigenvalues of $\widehat{\bm\Gamma}^x$ divided by $n$, i.e., it must give $n^{-1}{\widehat{\mathbf M}^x}$. In other words, our estimator $\widehat{\bm\Lambda}$ must be such that $\widehat{\bm\Lambda}
(\widehat{\bm\Lambda}^\prime\widehat{\bm\Lambda})^{-1/2}$ is the matrix of normalized eigenvectors corresponding the $r$ largest eigenvalues of $(nT)^{-1}{\bm X^\prime\bm X}$, i.e., such that:
equation[equation omitted — 257 chars of source]
but also, by definition of eigenvectors,
equation[equation omitted — 148 chars of source]
Therefore, since we must have $\bm\Lambda^\prime\bm\Lambda$ diagonal with distinct entries by Assumption (ref)(a), from (ref) and (ref),
equation[equation omitted — 103 chars of source]
which is such that $n^{-1}{\widehat{\bm\Lambda}^\prime \widehat{\bm\Lambda}}=n^{-1}{\widehat{\mathbf M}^x}$ is diagonal. The factors are then estimated as the linear projections: $\widehat{\bm F}=\bm X\widehat{\bm\Lambda}(\widehat{\bm\Lambda}^\prime\widehat{\bm\Lambda})^{-1}=\bm X\widehat{\mathbf V}^x(\widehat{\mathbf M}^x)^{-1/2}$, which are the normalized PCs of $\bm X$, such that $T^{-1}{\widehat{\bm F}^\prime\widehat{\bm F}}=\mathbf I_r$. The latter, however, are not needed in the following.
Letting $\widehat{\bm\lambda}_i^\prime$ be the $i$th row of the PC estimator of the loadings in (ref), we have the following asymptotic results.
theorem[]
Under Assumptions (ref) through (ref), as $n,T\to\infty$,
\begin{compactenum}[(a)]
• $\min(n,\sqrt T)\Vert\widehat{\bm\lambda}_i-{\bm\lambda}_i\Vert=O_{\mathrm P}(1)$, for any given $i=1,\ldots,n$, and, if $n^{-1}\sqrt T\to0$,
$
\sqrt T(\widehat{\bm\lambda}_i-{\bm\lambda}_i)
\to_d\mathcal N\left(\mathbf 0_r, \bm\Phi_i\right),
$
where $\bm\Phi_i$ is defined in Assumption (ref);
• $\min(n,\sqrt T)\Vert n^{-1/2}({\widehat{\bm\Lambda}-\bm\Lambda})\Vert=O_{\mathrm P}(1)$.
\end{compactenum}
An important consequence of Theorem (ref) is that the PC estimator is asymptotically equivalent to the unfeasible OLS estimator, which we would obtain if the factors were observed, and denoted as $\bm\Lambda^{\text{\tiny OLS}}$, with $i$th row given by ${\bm\lambda}_i^{\text{\tiny \upshape OLS}\prime}$.
cor[]
Under Assumptions (ref) through (ref), as $n,T\to\infty$,
\begin{inparaenum}[(a)]
•
$
\min(n,\sqrt{nT})\Vert\widehat{\bm\lambda}_i - {\bm\lambda}_i^{\text{\tiny \upshape OLS}}\Vert = O_{\mathrm P}\left(1\right)
$, for any given $i=1,\ldots,n$;
• $\min(n,\sqrt {nT})\Vert n^{-1/2}({\widehat{\bm\Lambda}-{\bm\Lambda}^{\text{\tiny \upshape OLS}}})\Vert=O_{\mathrm P}(1)$.
\end{inparaenum}
Quasi Maximum Likelihood
Define $\bm {\mathcal X}=\text{vec}(\bm X^\prime)=(\mathbf x_{1}^\prime\cdots \mathbf x_{T}^\prime)^\prime$ and $\bm{\mathcal Z}=\text{vec}(\bm \Xi^\prime)=(\bm\xi_{1}^\prime\cdots \bm\xi_{T}^\prime)^\prime$ as
the $nT$-dimensional vectors of observations and idiosyncratic components,
$\bm {\mathcal F}=\text{vec}(\bm F^\prime)=(\mathbf F_{1}^\prime\cdots \mathbf F_T^\prime)^\prime$ as the $rT$-dimensional vector of factors, and $\bm {\mathfrak L}=\mathbf I_T\otimes \bm\Lambda$ as the $nT\times rT$ matrix containing all factor loadings replicated $T$ times. Then, by vectorizing the transposed of (ref) we have:
align[align omitted — 103 chars of source]
Let $\bm\Omega^x=\mathbb{E}_{}[\bm {\mathcal X}\bm {\mathcal X}^\prime]$ and $\bm\Omega^\xi=\mathbb{E}_{}[\bm{\mathcal Z}\bm{\mathcal Z}^\prime]$, be the $nT\times nT$ covariance matrices of the $nT$-dimensional vectors of data and idiosyncratic components, respectively. Let also $\bm\Omega^F= \mathbb{E}_{}[\bm {\mathcal F}\bm {\mathcal F}^{\prime}]$ be the $rT\times rT$ covariance matrix of the $rT$-dimensional factor vector. Then, because of Assumption (ref),
\[
\bm\Omega^x = \bm {\mathfrak L}\,\bm\Omega^F\bm {\mathfrak L}^\prime + \bm\Omega^\xi.
\]
In principle, the parameters that need to be estimated are then given by the vector
$$
\bm\varphi=(\mathrm{vec}(\bm\Lambda)^\prime, \mathrm{vech}(\bm\Omega^{\xi})^\prime,\mathrm{vech}(\bm\Omega^{F})^\prime)^\prime,
$$ and
the Gaussian quasi-log-likelihood computed in a generic value of the parameters, denoted as $\underline{\bm{\varphi}}$, is given by (omitting the constant term for simplicity)
align[align omitted — 226 chars of source]
In general, maximization of (ref) is an unfeasible task since the parameters vector ${\bm{\varphi}}$ to be estimated has $\simeq (nT)^2$ elements. The common practice is then to consider simpler mis-specified log-likelihoods which depend on fewer parameters baili12,baili16.
First of all, consistently with the fact that we do not assume any parametric model describing the dynamics of the factors, we can consider a simpler log-likelihood with $\bm\Omega^F=\mathbf I_{nr}$, where we imposed also the identification constraint $\bm\Gamma^F=\mathbf I_r$ implied by Assumption (ref)(b).
A second simplification consists in considering a mis-specified log-likelihood of an approximate factor model where also the idiosyncratic components are treated as serially uncorrelated. As a result of these two mis-specifications the log-likelihood (ref) is reduced to:
equation[equation omitted — 315 chars of source]
The parameters to be estimated are then reduced:
$\bm\varphi=(\mathrm{vec}(\bm\Lambda)^\prime, \mathrm{vech}(\bm\Gamma^\xi)^\prime)^\prime$.
Nevertheless, estimation of $\bm\varphi$ by means of maximization of (ref) seems still hopeless since in general ${\bm\Gamma}^\xi$ contains $n(n+1)/2$ distinct elements.
Some further mis-specification of the log-likelihood (ref), based on regularizing $\bm\Gamma^\xi$, is then usually introduced in order to reduce the number of parameters to be estimated. A possibility in this sense is explored, for example, by bailiao16 who propose to maximize (ref) subject to an $\ell_1$ penalty imposed on the off-diagonal entries of $\bm\Gamma^\xi$. This approach forces sparsity, thus reducing the number of parameters to be estimated, but at the same time it makes estimation dependent on the chosen penalization level, which affects also the rate of consistency.
An even simpler approach consists in estimating only the diagonal entries of $\bm\Gamma^\xi$. Specifically, letting $\bm\Sigma^\xi=\text{dg}(\sigma_1^2\cdots\sigma_n^2)$ be the diagonal matrix with entries the diagonal entries of $\bm\Gamma^\xi$, we can focus on maximization of the further mis-specified log-likelihood:
equation[equation omitted — 333 chars of source]
The parameters to be estimated are then reduced to: $\bm\varphi=(\mathrm{vec}(\bm\Lambda)^\prime, \sigma_1^2,\cdots, \sigma_n^2)^\prime$, which are just $nr+n$. This is a feasible task given that we have $nT$ data points.
The log-likelihood (ref) is the one considered in classical factor analysis, where, however, $n$ is assumed to be fixed and small, and, moreover, the true idiosyncratic covariance matrix is assumed to be diagonal, i.e., the likelihood is not mis-specified (see, e.g., lawleymaxwell71, and RT82). In the high-dimensional case, i.e., when we allow $n\to\infty$, maximization of (ref) has been studied by baili12,baili16 under a variety of possible identifying constraints. In particular, while baili12 consider the case of no idiosyncratic serial or cross-correlation, i.e., they assume $\bm\Omega^\xi=\mathbf I_T\otimes \bm\Sigma^\xi$, which is diagonal, thus considering (ref) as a correctly specified log-likelihood, baili16 allow instead for idiosyncratic serial and cross-sectional correlations, as we do in this paper, thus treating (ref) as a mis-specified log-likelihood.
In the rest of this section we briefly review the properties of the QML estimator of the loadings. We refer to MBQML for a full review of QML estimation of factor models.
We start with the simplest case in which we consider a further mis-specification of the log-likelihood (ref), where we treat the idiosyncratic components as if they were homoskedastic, thus in the log-likelihood we replace $\bm\Sigma^\xi$ with an even simpler covariance matrix ${\sigma}^2\mathbf I_n$, with ${\sigma}^2>0$ and finite.
In this case tippingbishop99 prove that the QML estimator of the loadings matrix $\bm\Lambda$, has a closed form given by
equation[equation omitted — 328 chars of source]
with $\widehat{\sigma}^{2\text{\tiny QML,E}_0}$ being the QML estimator of $\sigma^2$. Intuitively, since under our assumptions we should have $\widehat{\mathbf M}^x=O_{\mathrm P}(n)$, as $n\to\infty$, then $\widehat{\bm\Lambda}^{\text{\tiny QML,E}_0}$ seems to coincide asymptotically with the PC estimator given in (ref). This is a well known fact and it is often quoted in the literature (see, e.g., DGRqml): in the case of spherical idiosyncratic components the PC and QML estimators are asymptotically equivalent. However, to the best of our knowledge no formal proof exists, at least under the present set of assumptions. In fact, the proof would essentially require to prove that the estimated $(r+1)$th largest eigenvalue is such that $\widehat{\mu}^x_{r+1}=O_{\mathrm P}(1)$, which is not an easy task in a high-dimensional setting, because, although we know that under our assumptions ${\mu}^x_{r+1}=O(1)$ (see Lemma (ref)(v)), in general $\widehat{\mu}^x_{r+1}$ is not a consistent estimator of ${\mu}^x_{r+1}$ (see, e.g., trapani2018randomized).
Things become more complicated if we allow for heteroskedasticity. Indeed, no closed form solution exists for the QML estimator maximizing the mis-specified log-likelihood (ref) and numerical maximization is required instead (baili12, for a proposed algorithm). Still it is possible to derive its asymptotic properties.
Let us denote as $\widehat{\bm\lambda}_i^{\text{\tiny \upshape QML,E}}$, $i=1,\ldots, n$, the QML estimator of the $i$th row of $\bm\Lambda$. In the classical fixed $n$ case, such estimator retains the classical propertied of the QML estimators, so it is $\sqrt T$-consistent and asymptotically normal. However, since the first and second derivatives of (ref) are very complex, the asymptotic covariance matrix has also a very complicated form, which in turn makes its estimation not at all easy (see, e.g., AR56, AFP87, and AA88). If, instead, we study the properties of the QML estimator when allowing also for $n\to\infty$, things become, perhaps surprisingly, simpler. This is shown in the following theorem proved by baili16.
theorem[]
Assume $n^{-1}{\bm\Lambda^\prime(\bm\Sigma^\xi)^{-1}\bm\Lambda}$ to be diagonal for all $n\in\mathbb N$ and $T^{-1}{\bm F^\prime\bm F}=\mathbf I_r$ for all $T\in\mathbb N$.
Then, under Assumptions (ref), (ref), (ref), and (ref), if $n^{-1}\sqrt T\to0$, as $n,T\to\infty$, for any given $i=1,\ldots,n$,
$
\sqrt T(\widehat{\bm\lambda}_i^{\text{\tiny \upshape QML,E}} -\bm\lambda_i)\to_d\mathcal N\left(\mathbf 0_r, \bm\Phi_i\right),
$
where $\bm\Phi_i$ is defined in Assumption (ref).
The proof of this result is based on asymptotic expansions of a set of conditions derived from first order conditions computed for the log-likelihood (ref). It is a very long proof, aimed at showing that $\widehat{\bm\lambda}_i^{\text{\tiny \upshape QML,E}}$ is asymptotically equivalent to the unfeasible OLS we would obtain if we knew the factors, which in turn has a known asymptotic distribution baili16. The proof requires a series of technical assumptions different from, but nested into, ours (see Remark (ref)). However, as noticed already in Section (ref), it is important to stress that Theorem (ref) holds under an identification constraint for the loadings which differs from our Assumption (ref)(a). At the end of the next section we prove that this theorem holds also under our assumptions (see Theorem (ref) and Corollary (ref)).
rem{
Notice that the fact that we consider a mis-specified log-likelihood with serially and cross-sectionally uncorrelated idiosyncratic components, does not affect consistency of the estimated loadings, but only their asymptotic covariance. And in particular, notice that what matters for the asymptotic covariance is just the idiosyncratic serial correlation. Indeed, if the idiosyncratic components were cross-sectionally uncorrelated but serially correlated the asymptotic covariance in Theorem (ref) would be the same, while if they were serially uncorrelated the asymptotic covariance would be $ \sigma_i^2 \mathbf I_r$ (due to the identification $\bm\Gamma^F=\mathbf I_r$), regardless of the presence or not of cross-sectional correlation. This result, which is a special case of Theorem (ref), does not require any constraint between the rates of divergence of $n$ and $T$ and it is proved by baili12. }
rem{
The case of QML estimation when $\bm\Omega^F$ depends explicitly on additional parameters capturing the autocorrelations in the factors is considered in DGRqml. However, in that case QML estimation requires the use of the EM algorithm jointly with the Kalman smoother and the results of this paper do not directly apply unless we first prove convergence of the EM to the same QML estimator considered here.
}
The PC and QML estimators are asymptotically equivalent
Given the discussion in the previous section, we might argue that the PC and QML estimators have the same asymptotic properties. However, Theorems (ref) and (ref) are derived under a similar but different sets of identifying assumptions and the two results cannot be directly compared. Even in the spherical case the proof seems to be not so easy due to the unknown properties of the smallest $N-r$ sample eigenvalues.
It is then natural to ask the following question. Can we prove in a simple way that the PC and QML estimators are asymptotically equivalent under
the same set of assumptions given in this paper, which are the standard PC assumptions? Moreover, if we prove such equivalence and
since the PC estimator of the loadings does not depend on the idiosyncratic covariance matrix, it is natural to ask also the following additional question. Can we prove the asymptotic equivalence of the two estimators when considering QML based on the log-likelihood (ref) of an approximate factor model, i.e., without constraining the idiosyncratic covariance to be diagonal or even homoskedastic? In this section we answer both questions.
Denote as $\widehat{\bm\Lambda}^{\text{\tiny QML}}$ the QML estimator of the loadings matrix maximizing the mis-specified log-likleihood (ref), where the idiosyncratic components are treated as serially uncorrelated, but their covariance matrix is unrestricted so it is correctly specified.
Let also $\widehat{\bm\lambda}_i^{\text{\tiny \upshape QML}\prime}$, $i=1,\ldots, n$, be the $i$th row of $\widehat{\bm\Lambda}^{\text{\tiny QML}}$. Then, we state our main result.
theoremUnder Assumptions (ref) through (ref) and assuming also that $\bm\Gamma^\xi$ is positive definite, as $n,T\to\infty$,
\begin{inparaenum}[(a)]
• $n\Vert
n^{-1/2}({\widehat{\bm\Lambda}^{\text{\tiny \upshape QML}}-\widehat{\bm\Lambda}})\Vert= O_{\mathrm {P}}(1)$;
•
$n \Vert\widehat{\bm\lambda}_i^{\text{\tiny \upshape QML}}-\widehat{\bm\lambda}_i \Vert = O_{\mathrm {P}}(1)$, for any given $i=1,\ldots,n$.
\end{inparaenum}
Consistency and asymptotic normality of the QML estimator of the loadings maximizing the log-likelihood (ref) immediately follow.
theoremUnder Assumptions (ref) through (ref) and assuming also that $\bm\Gamma^\xi$ is positive definite, as $n,T\to\infty$:
\begin{compactenum}[(a)]
• $\min(n,\sqrt T)\Vert\widehat{\bm\lambda}_i^{\text{\tiny \upshape QML}}-{\bm\lambda}_i\Vert = O_{\mathrm {P}}(1)$, for any given $i=1,\ldots,n$,, and, if $n^{-1}\sqrt T\to 0$
$
\sqrt T(\widehat{\bm\lambda}_i^{\text{\tiny \upshape QML}}-\bm\lambda_i)\to_d\mathcal N\left(\mathbf 0_r, \bm\Phi_i\right),
$
where $\bm\Phi_i$ is defined in Assumption (ref);
• $\min(n,\sqrt T)\Vert n^{-1/2}({\widehat{\bm\Lambda}^{\text{\tiny \upshape QML}}-\bm\Lambda})\Vert=O_{\mathrm P}(1)$.
\end{compactenum}
The results in Theorems (ref) and (ref) hold when considering the log-likelihood (ref) which depends on any generic idiosyncratic covariance matrix provided that it is positive definite and Assumption (ref) is satisfied. Therefore, from Theorem (ref) it follows that if we replace in the log-likelihood the full idiosyncratic covariance with a diagonal covariance matrix as in baili12,baili16 or if we also impose homoskedasticity as in tippingbishop99, we still get QML estimators that are asymptotically equivalent to the PC estimator.
corUnder Assumptions (ref) through (ref), as $n,T\to\infty$,
\begin{inparaenum}[(a)]
• $n\Vert \widehat{\bm\lambda}_i - \widehat{\bm\lambda}_i^{\text{\tiny \upshape QML,E}}\Vert=O_{\mathrm P}(1)$;
• $n\Vert \widehat{\bm\lambda}_i - \widehat{\bm\lambda}_i^{\text{\tiny \upshape QML,E}_0}\Vert=O_{\mathrm P}(1)$.
\end{inparaenum}
From these results and Theorem (ref), we directly have a proof of Theorem (ref) by baili16, which now holds under our identifying constraint in Assumption (ref)(a), and we also prove the often quoted statement that the PC and QML estimators are asymptotically equivalent under sphericity of idiosyncratic components.
rem{
The appealing feature of the log-likelihood (ref), is that it does not depend on the factors, and, under the identifying constraint $T^{-1}{\bm F'\bm F}=\mathbf I_r$ for all $T\in\mathbb N$, it does not depend on the second moments of the factors either. Still, if we follow the classical approach to consider the vector of factors $\bm{\mathcal F}$ as an $rT$-dimensional sequence of deterministic constants lawleymaxwell71, then a log-likelihood alternative to (ref) can be considered, namely:
\begin{equation}
\ell_{\tiny E}(\bm{\mathcal X};\bm\varphi,\bm{\mathcal F})= -\frac T2\log\det(\bm\Gamma^\xi)-\frac 12\sum_{t=1}^T
(\mathbf x_t-\bm\Lambda\,\mathbf F_t)^\prime
(\underline{\bm\Gamma}^\xi)^{-1}
(\mathbf x_t-\underline{\bm\Lambda}\,\underline{\mathbf F}_t).
\end{equation}
It is straightforward to see that, for given factors, the loadings maximizing (ref) are given by their OLS estimator, while, for given loadings, the factors maximizing (ref) are given by their GLS estimator. Hence, full maximization of (ref) requires knowing, or estimating, the factors too. Although seemingly simpler this approach presents at least two major drawbacks with respect to the approach followed in the proof of Theorem (ref) (see also the comments in
baili12, and AR56).
First, by iterating, between OLS and GLS we can think of finding a solution. This approach is similar to the one recently considered by pz23, but the convergence properties of such algorithm are not proved nor discussed, so it is unclear how this estimator is related to the maximizer of the full likelihood (ref). Moreover, this strategy would require a positive definite estimator of $\bm\Gamma^\xi$ in order to compute the GLS estimator of the factors, a hard task when $n$ is large. This, in general, requires again mis-specifying or regularizing the log-likelihood (ref), e.g., by replacing $\bm\Gamma^\xi$ with the diagonal $\bm\Sigma^\xi$. The asymptotic properties of the estimator of the loadings will then explicitly depend on the properties of an estimator of the idiosyncratic covariance, or at least of its diagonal elements. As shown in Theorem (ref), this is not the case for the QML estimator maximizing the log-likelihood (ref).
Second, if the factors are treated as random variables, as, e.g., in the popular Factor Augmented VAR models BBE05, then they cannot be considered as constant parameters, which means that (ref) is not the full log-likelihood of the data $\bm{\mathcal X}$ but it is just the conditional log-likelihood of $\bm{\mathcal X}$ given the factors, so, in principle, not all information is used to estimate the loadings. Indeed, if the factors are random variables then the log-likelihood (ref) is decomposed as
\begin{equation}
\ell_{\text{\tiny E}}(\bm{\mathcal X};\underline{\bm\varphi})=\ell_{\text{\tiny E}}(\bm{\mathcal X}|\bm{\mathcal F};\underline{\bm\varphi}) +\ell_{\text{\tiny E}}(\bm{\mathcal F};\underline{\bm\varphi})-\ell_{\text{\tiny E}}(\bm{\mathcal F}|\bm{\mathcal X};\underline{\bm\varphi}),
\end{equation}
where $\ell_{\text{\tiny E}}(\bm{\mathcal X}|\bm{\mathcal F};\underline{\bm\varphi})$ coincides with (ref) but it does not coincide with $\ell_{\text{\tiny E}}(\bm{\mathcal X};\underline{\bm\varphi})$ anymore.
This last point has both a theoretical and an applied implication. From the theory point of view, to directly show from (ref) that the QML estimator of the loadings maximizing $\ell_{\text{\tiny E}}(\bm{\mathcal X};\underline{\bm\varphi})$ is asymptotically equivalent to the unfeasible OLS maximizing $\ell_{\text{\tiny E}}(\bm{\mathcal X}|\bm{\mathcal F};\underline{\bm\varphi})$ would require showing that $\ell_{\text{\tiny E}}(\bm{\mathcal F};\underline{\bm\varphi})$ and $\ell_{\text{\tiny E}}(\bm{\mathcal F}|\bm{\mathcal X};\underline{\bm\varphi})$ are asymptotically negligible. This is the argument sketched by BT11, but to make it formal is not an easy task. Consider the simplest case in which the factors are treated as serially uncorrelated, then, while $\ell_{\text{\tiny E}}(\bm{\mathcal F};\underline{\bm\varphi})$ does not depend on any parameter and can be discarded, the expression of $\ell_{\text{\tiny E}}(\bm{\mathcal F}|\bm{\mathcal X};\underline{\bm\varphi})$ will depend on the conditional moments (mean and covariance) of the factors given $\bm{\mathcal X}$. These in turn have simple expressions only if we are willing to assume Gaussianity, in which case the conditional mean is just a linear projection. Otherwise computation of those moments is not straightforward. The proof of Theorem (ref) relies instead only on $\ell_{\text{\tiny E}}(\bm{\mathcal X};\underline{\bm\varphi})$, so, as noticed above, it does not require knowing the factors or their conditional moments.
Finally, from a practical point of view, we could use the right-hand-side of (ref) to compute the QML estimator by means of the Expectation Maximization (EM) algorithm, which simplifies our task since it allows us to discard $\ell_{\text{\tiny E}}(\bm{\mathcal F}|\bm{\mathcal X};\underline{\bm\varphi})$ wu83. However, once again we would still have to compute the conditional moments of the factors when in the E-step we need compute the conditional expectation of $\ell_{\text{\tiny E}}(\bm{\mathcal X}|\bm{\mathcal F};\underline{\bm\varphi}) +\ell_{\text{\tiny E}}(\bm{\mathcal F};\underline{\bm\varphi})$ given $\bm{\mathcal X}$.
This introduces a correction in the estimation of the loadings and the resulting estimator is not given by a simple OLS anymore.
}
Asymptotic covariance matrices of the QML and PC estimators
In this section, we focus on the mis-specified log-likelihood (ref), which is commonly used in empirical work baili12,baili16. For such log-likelihood, denote the Fisher information and the population Hessian matrices for $\bm\lambda_i$, $i=1,\ldots, n$, as:
align[align omitted — 848 chars of source]
From QML theory the asymptotic covariance of the QML estimator should be given by:
equation[equation omitted — 269 chars of source]
This would be the matrix to estimate if we were to conduct QML based inference on the loadings.
In general, estimation of (ref) is very difficult given the complex expressions of the Hessian and Fisher information matrices (see also (A.15) and (A.42) in the Supplementary Material). Moreover, from Theorem (ref) we know that, in fact, if $n^{-1}\sqrt T\to 0$, as $n,T\to\infty$, then the asymptotic covariance of the QML estimator is
equation[equation omitted — 238 chars of source]
where we used the definition of $\bm\Phi_i$ in Assumption (ref). So what is the relation between the asymptotic covariances (ref) and (ref)?
First, notice that (ref) coincides with the asymptotic covariance of the unfeasible OLS estimator, which, in turn, is the QML estimator maximizing the log-likelihood for an exact factor model conditional on observing the factors, i.e.,
equation[equation omitted — 312 chars of source]
It is easy to compute the Fisher information and the population Hessian matrices of (ref) for $\bm\lambda_i$:
(see also (A.37) and (A.47) in the Supplementary Material)
align[align omitted — 990 chars of source]
where we imposed orthonormality of the factors as required by Assumption (ref)(b). Clearly,
equation[equation omitted — 276 chars of source]
It follows that for the asymptotic covariances (ref) and (ref) to be equivalent it must be that the Fisher information and Hessian matrices of the full log-likelihood $\ell_{\text{\tiny E}}(\bm{\mathcal X};\underline{\bm\varphi})$ in (ref) and of the conditional log-likelihood $\ell_{\text{\tiny E}}(\bm{\mathcal X}|\bm{\mathcal F};\underline{\bm\varphi})$ in (ref) are asymptotically equivalent when computed in the true values of the parameters $\bm\varphi$. This is proved in the following theorem.
theoremUnder Assumptions (ref) through (ref), as $n,T\to\infty$, for any given $i=1,\ldots,n$,
\begin{compactenum}[(a)]
• \[
\frac 1{\sqrt T}\left\Vert \left.\frac{\partial \ell_{\text{\tiny \upshape E}}(\bm{\mathcal X};\underline{\bm\varphi})}{\partial \underline{\bm\lambda}_i^\prime}\right\vert_{\underline{\bm\varphi}={\bm\varphi}}
-\left.\frac{\partial \ell_{\text{\tiny \upshape E}}(\bm{\mathcal X}|\bm{\mathcal F};\underline{\bm\varphi})}{\partial \underline{\bm\lambda}_i^\prime}\right\vert_{\underline{\bm\varphi}={\bm\varphi}}
\right\Vert=O_{\mathrm P}\left(\max\left(\frac 1{\sqrt n},\frac {\sqrt T}{n}\right)\right);
\]
• \[
\frac 1T\left\Vert \left.\frac{\partial^2 \ell_{\text{\tiny \upshape E}}(\bm{\mathcal X};\underline{\bm\varphi})}{\partial \underline{\bm\lambda}_i^\prime\partial \underline{\bm\lambda}_i}\right\vert_{\underline{\bm\varphi}={\bm\varphi}}
-\left.\frac{\partial^2 \ell_{\text{\tiny \upshape E}}(\bm{\mathcal X}|\bm{\mathcal F};\underline{\bm\varphi})}{\partial \underline{\bm\lambda}_i^\prime\partial \underline{\bm\lambda}_i}\right\vert_{\underline{\bm\varphi}={\bm\varphi}}
\right\Vert=O_{\mathrm P}\left(\max\left(\frac 1n,\frac 1{\sqrt {nT}}\right)\right);
\]
• \[
\frac 1{\sqrt T}\left\Vert \left.\frac{\partial \ell_{\text{\tiny \upshape E}}(\bm{\mathcal X};\underline{\bm\varphi})}{\partial \underline{\bm\lambda}_i^\prime}\right\vert_{\underline{\bm\varphi}=\widehat{\bm\varphi}^{\text{\tiny \upshape QML,E}}}
-\left.\frac{\partial \ell_{\text{\tiny \upshape E}}(\bm{\mathcal X}|\bm{\mathcal F};\underline{\bm\varphi})}{\partial \underline{\bm\lambda}_i^\prime}\right\vert_{\underline{\bm\varphi}=\widehat{\bm\varphi}^{\text{\tiny \upshape QML,E}}}
\right\Vert=O_{\mathrm P}\left(\max\left(\frac 1{\sqrt n},\frac {\sqrt T}{n}\right)\right);
\]
• \[
\frac 1T\left\Vert \left.\frac{\partial^2 \ell_{\text{\tiny \upshape E}}(\bm{\mathcal X};\underline{\bm\varphi})}{\partial \underline{\bm\lambda}_i^\prime\partial \underline{\bm\lambda}_i}\right\vert_{\underline{\bm\varphi}=\widehat{\bm\varphi}^{\text{\tiny \upshape QML,E}}}
-\left.\frac{\partial^2 \ell_{\text{\tiny \upshape E}}(\bm{\mathcal X}|\bm{\mathcal F};\underline{\bm\varphi})}{\partial \underline{\bm\lambda}_i^\prime\partial \underline{\bm\lambda}_i}\right\vert_{\underline{\bm\varphi}=\widehat{\bm\varphi}^{\text{\tiny \upshape QML,E}}}
\right\Vert=O_{\mathrm P}\left(\max\left(\frac 1n,\frac 1{\sqrt {nT}}\right)\right).
\]
\end{compactenum}
From Theorem (ref) it is clear that if $n^{-1}\sqrt T\to 0$, as $n,T\to\infty$, then we can consistently estimate $\text{A$\!\mathbb V$ar}_0(\sqrt T \widehat{\bm\lambda}_i^{\text{\tiny QML,E}})$ in (ref) by means of any consistent estimator of $\text{A$\!\mathbb V$ar}_1(\sqrt T \widehat{\bm\lambda}_i^{\text{\tiny QML,E}})=\bm\Phi_i$, as the classical HAC estimator:
align[align omitted — 538 chars of source]
where $\widehat{\mathbf F}_t=(\widehat{\mathbf M}^x)^{-1/2}\widehat{\mathbf V}^{x\prime} \mathbf x_t$ and $\widehat{\xi}_{it}=x_{it}-\widehat{\bm\lambda}_i^\prime\widehat{\mathbf F}_t$ are the PC estimators of the factors and idiosyncratic components, respectively. Consistency of this estimator is proved by Bai03.
rem{
The sandwich form of the asymptotic covariances (ref) and (ref) comes from the fact that in the log-likelihood we treat the idiosyncratic components as serially uncorrelated, while, in fact, they might be autocorrelated. Indeed, if we assumed $\mathbb{E}[\xi_{it}\xi_{is}]=0$ for all $s,t=1,\ldots, T$ with $s\ne t$ and all $i=1,\ldots, T$, then $\bm \Phi_i = \sigma_i^2\mathbf I_r$ (see also Remark (ref)). It follows that: $\bm{\mathcal I}_i(\bm{\mathcal X}|\bm{\mathcal F};\bm\varphi)={\sigma_i^{-2}}\mathbf I_r=-\bm{\mathcal H}_{i}(\bm{\mathcal X}|\bm{\mathcal F};\bm\varphi)$.
And, by virtue of Theorem (ref), $\bm{\mathcal I}_i(\bm{\mathcal X};\bm\varphi)$ must coincide asymptotically with $-\bm{\mathcal H}_{i}(\bm{\mathcal X};\bm\varphi)$, thus, from (ref), we have
$\text{A$\!\mathbb V$ar}_0(\sqrt T \widehat{\bm\lambda}_i^{\text{\tiny QML,E}})=\{\bm{\mathcal I}_i(\bm{\mathcal X};\bm\varphi)\}^{-1}$.
}
Estimation of the factors
Up to this point we did not say anything about estimating the factors. This is because, first, the factors are typically estimated once we have estimated the loadings, and, second, there is not a clear definition of what a QML estimator of the factors is.
If we treat the factors as random variables, then they are not parameters and they do not have a QML estimator. We know that their optimal (in mean-squared sense) estimator is their conditional mean given $\bm{\mathcal X}$, which, under Gaussianity can be estimated as the linear projection. If, consistently with the discussion in Section (ref) about QML estimation, we mis-specify the second order structure of the factors, by replacing $\bm\Omega^F$ with $\mathbf I_{rT}$, and of the idiosyncratic components, by replacing $\bm\Omega^\xi$ with $\bm\Sigma^\xi$,
then such linear projection is given by:
equation[equation omitted — 298 chars of source]
where we used Woodbury formula. Alternatively, it is common to use the OLS or GLS estimators:
equation[equation omitted — 318 chars of source]
If instead we treat the factors as constant parameters then, as discussed in Remark (ref), we can see the GLS estimator as the QML estimator maximizing the joint log-likelihood (ref) of factors and data.
The three estimators defined above are unfeasible unless we first compute estimates of the parameters. By virtue of our result in Theorem (ref) there is asymptotically no difference if we use $\widehat{\bm\Lambda}$ or $\widehat{\bm\Lambda}^{\text{\tiny QML}}$, and an estimator of $\bm\Sigma^\xi$ can easily be computed either from the residuals of PC estimation or by QML as suggested by baili12. Once we use these estimated parameters in (ref) and (ref), we have the estimators $\widehat {\mathbf F}_t^{\text{\tiny LP}}$, $\widehat{\mathbf F}_t^{\text{\tiny OLS}}$, and $\widehat{\mathbf F}_t^{\text{\tiny GLS}}$. Notice that, as shown in Section (ref), $\widehat{\mathbf F}_t^{\text{\tiny OLS}}$ is nothing else but the PC estimator of the factors. The GLS has also been studied by BT11 and choi12 when using different estimators of $\bm\Sigma^\xi$.
By construction, the OLS and GLS in (ref) are both less efficient than the linear projection in (ref). Moreover, the GLS is always more efficient than the OLS if we could use an estimator of the full idiosyncratic covariance matrix, but, since this is in general unfeasible and we typically estiamate only its diagonal as in (ref), then we can just conjecture that the more sparse is the true covariance matrix the more likely the GLS is to be more efficient.
Finally, these estimators are all $\min(\sqrt n, T)$-consistent. For the OLS we refer to Bai03. For the GLS we refer to baili16. Moreover, it is straightforward to see that, by Lemmas (ref)(v) and (ref)(i), we have $\Vert \widehat{\mathbf F}_t^{\text{\tiny GLS}}-\widehat{\mathbf F}_t^{\text{\tiny LP}} \Vert=O_{\mathrm P}(n^{-1})$.
rem{
If we explicitly model the dynamics of $\mathbf F_t$ then the expression of $\widehat{\mathbf F}_t^{\text{\tiny LP}}$ in (ref) is replaced by the Kalman smoother DGRfilter. This, in fact, can be shown to be asymptotically equivalent, as $n\to\infty$, to the GLS estimator in (ref) baili16,PR22.
}
rem{
In the case of deterministic factors, we could also write the mis-specified log-likelihood (ref) of an exact factor model as:
\begin{equation}
\ell_{\tiny E}(\bm{\mathcal X};\bm\varphi)= -\frac 1 2\sum_{i=1}^n \log\det (\bm F\,\bm \lambda_i\bm \lambda_i^\prime\bm F^\prime+\underline{\sigma}_i^2\mathbf I_T)-\frac 12\sum_{i=1}^n \bm x_i^\prime(\underline{\bm F}\,\underline{\bm \lambda}_i\underline{\bm \lambda}_i^\prime\underline{\bm F}^\prime+\underline{\sigma}_i^2\mathbf I_T)^{-1}\bm x_i,
\end{equation}
where $\bm x_i=(x_{i1}\cdots x_{iT})'$, thereby exchanging the role of $n$ and $T$.
Then, we can conjecture that the QML estimator of the factors maximizing (ref) will be asymptotically equivalent, this time as $T\to\infty$, to their PC estimator.
This is approach is also considered in fortin2023latent. However, since, as noted above, the PC estimator of the factors is asymptotically equivalent to the OLS, it is not the most efficient estimator because it does not account for possible cross-sectional heteroskedasticity of idiosyncratic components.
}
Monte Carlo study
Throughout, we let $n\in\{20, 50, 100,200\}$, $T=100$, and $r=2$, and, for all $i=1,\ldots, n$ and $t=1,\ldots, T$, we simulate the data according to
align[align omitted — 175 chars of source]
where $\bm\ell_i$ and $\bm f_t$ are $r$-dimensional vectors. Specifically,
inparaenum• $\bm\ell_i$ has entries ${\ell}_{ij}{\sim}iid\,\mathcal{N}(1,1)$, $i=1,\ldots, N$, $j=1,\ldots, r$;
• ${\bm A}=0.9 \check{\bm A} \Vert{\bm A}\Vert^{-1}$, where $\check{\bm A}$ is $r\times r$ with entries $\check{a}_{jj}{\sim} iid\,\mathcal U[0.5,0.8]$ for all $j$, and $\check{ a}_{jk}{\sim} iid\, \mathcal U[0,0.3]$ if $j\ne k$;
• $u_{jt}{\sim}(0,1)$, $j=1,\ldots, r$, $t=1,\ldots, t$, with either a Gaussian or an Asymmetric Laplace distribution and with $\mathbb{C}\mathrm{ov}(u_{it},u_{jt})=0$ if $i\ne j$,
and $\mathbb{C}\mathrm{ov}(u_{it},u_{jt-k})=0$ for all $i,j$ and if $k\ne 0$;
• $e_{it}{\sim}(0,\sigma_{ei}^2)$, with either a Gaussian or an Asymmetric Laplace distribution and with
$\sigma_{ei}^2\sim \mathcal U[0.5, 1.5]$ for all $i$,
$\mathbb{C}\mathrm{ov}(e_{it},e_{jt})=\tau^{\vert i-j\vert }\mathbb I(\vert i-j\vert \le 10)$ with $\tau\in\{0,0.5\}$ if $i\ne j$,
and $\mathbb{C}\mathrm{ov}(e_{it},e_{jt-k})=0$ for all $i,j$ and if $k\ne 0$;
• $\delta_i{\sim}iid\,\mathcal{U}(0,\delta)$, and $\delta\in\{0,0.5\}$;
• $\phi_i=\{{\theta_i (\sum_{t=1}^T \chi_{it}^2)/(\sum_{t=1}^T \xi_{it}^2)}\}^{1/2}$,
and
$\theta_i{\sim}iid\,\mathcal{U}(0.25,0.5)$.
In the case of the Asymmetric Laplace distribution, all the innovations have location 0, asymmetry index $\kappa$, with $\kappa \sim \mathcal{U}(.9,1.1)$ and scale index $\lambda=\sqrt{{1+\kappa^4}/{\kappa^2}}$, so that the variance is 1. The parameters $\tau$ and $\delta$ control the degrees of
cross-sectional and serial idiosyncratic correlation in the idiosyncratic components. The the noise-to-signal ratio for series $i$ is given by $\theta_i$.
Finally, in order to satisfy Assumptions (ref)(a) and (ref)(b) we proceed as follows. Given the common components is generated as $\chi_{it}={\bm\ell}_{i}^\prime {\bm f}_t$, let $\bm\chi_t=(\chi_{1t}\cdots \chi_{nt})^\prime$ and compute $\check{\bm\Gamma}^\chi=T^{-1}\sum_{t=1}^T \bm\chi_t\bm\chi_t^\prime$. Collect its $r$ non-zero eigenvalues into the $r\times r$ diagonal matrix $\check{\mathbf M}^{\chi}$ and the corresponding normalized eigenvectors as the columns of the $n\times r$ matrix $\check{\mathbf V}^{\chi}$, with sign fixed such that it has non-negative entries in the first row. The loadings are then simulated as $\bm \Lambda=\check{\mathbf V}^{\chi} (\check{\mathbf M}^{\chi})^{1/2}$ and the factors as $\mathbf F_t = (\check{\mathbf M}^{\chi})^{-1/2}\check{\mathbf V}^{\chi\prime}\bm\chi_t$.
We simulate the model described above $B=500$ times, and at each replication we estimate the loadings via PC and QML, where the latter is defined as the maximizer of the log-likelihood (ref) and it is computed numerically in the way proposed by baili12,baili16. In this way, we obtain two estimates of the loadings matrix $\widehat{\bm\Lambda}^{(b)}$ and $\widehat{\bm\Lambda}^{\text{\tiny QML,E}(b)}$, respectively, $b=1,\ldots, B$. We also compute the unfeasible OLS estimator ${\bm\Lambda}^{\text{\tiny OLS}(b)}$ by regressing $\mathbf x_t$ onto the simulated factors.
table[table omitted — 4,032 chars of source]
table[table omitted — 3,242 chars of source]
In Table (ref) we report the Mean-Squared-Error (MSE) for each column, $j=1,\ldots, r$, of the considered estimators, averaged over the $B$ replications (with standard deviations in parenthesis):
align[align omitted — 531 chars of source]
Results show that QML and PC estimator have similar MSEs and both improve as $n$ increases to the point that when $n=200$ their MSEs are comparable with the one of the unfeasible OLS, i.e., the estimators behave as if the factors were observed. In most cases the QML estimator has a smaller MSE even when the true distribution is not Gaussian.
Finally, in Table (ref) for each column of the loadings, $j=1,\ldots, r$, we report the distance between the QML and PC estimators measured as (with standard deviations in parenthesis):
align[align omitted — 189 chars of source]
And we also report the relative MSE of the PC estimator with respect to the MSE of the QML estimator: $\text{MSE}_j^{\text{\tiny REL}}={\text{MSE}_j^{\text{\tiny PC}}}/{\text{MSE}_j^{\text{\tiny QML}}}$. Results clearly show that as $n$ grows the PC and QML estimators become almost indistinguishable.
Concluding remarks
To compute in practice the QML estimator of the loadings there are at least two main issues. First, in finite samples the QML estimator of the loadings has no closed form and depends also on the estimator of the idiosyncratic covariance for which no closed form exists either. This is true even if we use the the log-likelihood (ref) of an exact factor model. Second, the convergence properties of the various available EM algorithms used to compute the QML estimator RT82,baili12,baili16,pz23 have never been fully investigated. On the one hand, it is easy to prove, that at each iteration of an EM algorithm the log-likelihood evaluated in the parameters estimated at that iteration is larger than at the previous iteration wu83, but, on the other hand, no formal proof exists of convergence of those algorithms to a global maximum of the likelihood, at least to the best of our knowledge.
The results of this paper offer a possible solution by showing that, if we are just interested in the factor loadings and we do not need to estimate the idiosyncratic variances, then we can simply use the PC estimator of the loadings and of its asymptotic covariance matrix to approximate the corresponding QML estimator and its asymptotic covariance matrix. Once this is done, the factors can be estimated via OLS as in PC analysis.
As a consequence, we might think that there is no apparent advantage in directly computing the QML estimator of the loadings and of the idiosyncratic variances.
Nevertheless, QML estimation has at least three advantages. First, it allows us to easily impose restrictions on the parameters of the model. Second, having also the QML estimator of the idiosyncratic variances allows us to compute estimators of the factors as the GLS, which are possibly more efficient. Third, QML estimation, as presented in this paper, is a first step towards estimating a model where we explicitly model the dynamics of the factors, something we cannot do with PC analysis. This last point, already briefly discussed in Remarks (ref) and (ref), is the subject of our ongoing research BLqml.
appendix\numberwithin{equation}{section}
\numberwithin{prop}{section}
\section{Proof of main results}
\subsection{Proof of Theorem (ref)}
From (ref) and (ref), and since by (ref) we have $\widehat{\bm\Lambda}^\prime\widehat{\bm\Lambda}=\widehat{\mathbf M}^x$ and $\widehat{\mathbf V}^x=\widehat{\bm\Lambda}(\widehat{\mathbf M}^x)^{-1/2}$ which is well defined because of Lemma (ref)(iv), we get
\begin{equation}
\frac{\bm X^\prime\bm X}{nT}\widehat{\bm\Lambda}=\widehat{\bm\Lambda}\frac{\widehat{\mathbf M}^x}{n}.
\end{equation}
Then, substituting $\bm X^\prime\bm X=(\bm\Lambda\bm F^\prime+\bm\Xi^\prime)^\prime(\bm F\bm\Lambda^\prime+\bm\Xi)$ into (ref)
\begin{align}
\frac{\bm\Lambda\bm F^\prime\bm F\bm\Lambda^\prime\widehat{\bm\Lambda}}{nT}
+\frac{\bm\Lambda\bm F^\prime\bm \Xi\widehat{\bm\Lambda}}{nT}
+\frac{\bm\Xi^\prime\bm F\bm\Lambda^\prime\widehat{\bm\Lambda}}{nT}
+\frac{\bm\Xi^\prime\bm \Xi\widehat{\bm\Lambda}}{nT}=\widehat{\bm\Lambda}\frac{\widehat{\mathbf M}^x}{n}.
\end{align}
Define
\begin{equation}
\widehat{\mathbf H}=
\left(\frac{\bm F^\prime\bm F}{T}\right)
\left(\frac{\bm\Lambda^\prime\widehat{\bm\Lambda}}{n}\right)
\left(\frac{\widehat{\mathbf M}^x}{n}\right)^{-1}=
\left(\frac{\bm\Lambda^\prime\widehat{\bm\Lambda}}{n}\right)
\left(\frac{\widehat{\mathbf M}^x}{n}\right)^{-1},
\end{equation}
by Assumption (ref)(b). Notice that, as $n,T\to\infty$, $\widehat{\mathbf H}$ is well defined because $\left(\frac{\widehat{\mathbf M}^x}{n}\right)^{-1}$ is well defined because of Lemma (ref)(iv). From (ref) and (ref)
\begin{align}
\widehat{\bm\Lambda}&-\bm\Lambda\widehat{\mathbf H}= \left(\frac{\bm\Lambda\bm F^\prime\bm \Xi\widehat{\bm\Lambda}}{nT}
+\frac{\bm\Xi^\prime\bm F\bm\Lambda^\prime\widehat{\bm\Lambda}}{nT}
+\frac{\bm\Xi^\prime\bm \Xi\widehat{\bm\Lambda}}{nT}\right)\left(\frac{\widehat{\mathbf M}^x}{n}\right)^{-1}\nonumber\\
=&\,
\left(\frac{\bm\Lambda\bm F^\prime\bm \Xi{\bm\Lambda}}{nT}
+\frac{\bm\Xi^\prime\bm F\bm\Lambda^\prime{\bm\Lambda}}{nT}
+\frac{\bm\Xi^\prime\bm \Xi{\bm\Lambda}}{nT}\right)\bm{\mathcal H} \left(\frac{\widehat{\mathbf M}^x}{n}\right)^{-1}\!\!\!\!\!+
\left(\frac{\bm\Lambda\bm F^\prime\bm \Xi}{nT}
+\frac{\bm\Xi^\prime\bm F\bm\Lambda^\prime}{nT}
+\frac{\bm\Xi^\prime\bm \Xi}{nT}\right)(\widehat{\bm\Lambda}-\bm\Lambda\bm{\mathcal H}) \left(\frac{\widehat{\mathbf M}^x}{n}\right)^{-1}.
\end{align}
Taking the $i$th row of (ref)
\begin{align}
\widehat{\bm\lambda}_i^\prime-{\bm\lambda}_i^\prime\widehat{\mathbf H}=&\,
\left(
\underbrace{\frac 1{nT}{\bm\lambda}_i^\prime\sum_{t=1}^T\sum_{j=1}^n\mathbf F_t\xi_{jt}{\bm\lambda}_j^\prime}_{(1.a)}
+\underbrace{\frac 1{nT} \sum_{t=1}^T \xi_{it}\mathbf F_t^\prime\sum_{j=1}^n\bm\lambda_j\bm\lambda_j^\prime}_{(1.b)}
+\underbrace{\frac 1{nT} \sum_{t=1}^T\sum_{j=1}^n\xi_{it}\xi_{jt} \bm\lambda_j^\prime}_{(1.c)}
\right)\bm{\mathcal H} \left(\frac{\widehat{\mathbf M}^x}{n}\right)^{-1}\nonumber\\
&+
\left(
\underbrace{\frac 1{nT}{\bm\lambda}_i^\prime\sum_{t=1}^T\sum_{j=1}^n\mathbf F_t\xi_{jt}(\widehat{\bm\lambda}_j^\prime-\bm\lambda_j^\prime\bm{\mathcal H})}_{(1.d)}
+\underbrace{\frac 1{nT} \sum_{t=1}^T \xi_{it}\mathbf F_t^\prime\sum_{j=1}^n\bm\lambda_j(\widehat{\bm\lambda}_j^\prime-\bm\lambda_j^\prime\bm{\mathcal H})}_{(1.e)}\right.\nonumber\\
&\;\;\left.+\underbrace{\frac 1{nT} \sum_{t=1}^T\sum_{j=1}^n\xi_{it}\xi_{jt} (\widehat{\bm\lambda}_j^\prime-\bm\lambda_j^\prime\bm{\mathcal H})}_{(1.f)}\right) \left(\frac{\widehat{\mathbf M}^x}{n}\right)^{-1}.
\end{align}
From Proposition (ref) we see that, under Assumptions (ref) through (ref), the terms $\text{\upshape (1.a)}$, $\text{\upshape (1.c)}$, $\text{\upshape (1.d)}$, $\text{\upshape (1.e)}$, and $\text{\upshape (1.f)}$ are all $o_{\mathrm P}\left(\frac 1{\sqrt T}\right)$. In particular, from (ref), for any $i=1,\ldots, n$, we get
\begin{align}
\widehat{\bm\lambda}_i-\widehat{\mathbf H}^\prime{\bm\lambda}_i = \left(\frac{\widehat{\mathbf M}^x}{n}\right)^{-1} \bm{\mathcal H}^\prime\left(\frac{\bm\Lambda^\prime\bm\Lambda}{n}\right) \left(\frac 1T\sum_{t=1}^T \mathbf F_t\xi_{it}\right)
+ O_{\mathrm {P}}\left(\max\left(\frac 1 n, \frac 1{\sqrt{nT}}\right)\right).
\end{align}
Consistency follows immediately since the first term in (ref) is $O_{\mathrm P}\left(\frac 1{\sqrt T}\right)$ because of Assumption (ref) (see also (ref) in the proof of Theorem (ref)) and since $\left\Vert \left(\frac{\widehat{\mathbf M}^x}{n}\right)^{-1} \bm{\mathcal H}^\prime \right\Vert = O_{\mathrm {P}}(1)$ because of Lemma (ref)(iv) and (ref).
Then, from (ref),
by Proposition (ref)(b):
\begin{align}
\sqrt T(\widehat{\bm\lambda}_i-\widehat{\mathbf H}^\prime{\bm\lambda}_i) &= \left(\frac{\widehat{\mathbf M}^x}{n}\right)^{-1} \bm{\mathcal H}^\prime\left(\frac{\bm\Lambda^\prime\bm\Lambda}{n}\right) \left(\frac 1{\sqrt T}\sum_{t=1}^T \mathbf F_t\xi_{it}\right)
+ o_{\mathrm {P}}(1)\nonumber\\
&= \left(\frac{\widehat{\mathbf M}^x}{n}\right)^{-1}\left(\frac{\widehat{\bm\Lambda}^\prime\bm\Lambda}{n}\right) \left(\frac 1{\sqrt T}\sum_{t=1}^T \mathbf F_t\xi_{it}\right)
+ o_{\mathrm {P}}(1).
\end{align}
Now, from Proposition (ref) we have that, as $n,T\to\infty$,
$\min(\sqrt{n},\sqrt T)\Vert \widehat{\mathbf H}-\bm J\Vert = o_{\mathrm P}(1)$, and, by using the definition of $\widehat{\mathbf H}$ in (ref), it follows that (ref) is equivalent to (note that $\Vert\bm\lambda_i\Vert=O(1)$ by Assumption (ref)(a))
\begin{align}
\sqrt T(\widehat{\bm\lambda}_i-\bm J\bm\lambda_i)&=\sqrt T(\widehat{\bm\lambda}_i-\widehat{\mathbf H}^\prime\bm\lambda_i)+\sqrt T(\widehat{\mathbf H}^\prime-\bm J)\bm\lambda_i+o_{\mathrm P}(1)\nonumber\\
&= \left(\frac{\widehat{\mathbf M}^x}{n}\right)^{-1} \left(\frac{\widehat{\bm\Lambda}^\prime\bm\Lambda}{n}\right) \left(\frac 1{\sqrt T}\sum_{t=1}^T \mathbf F_t\xi_{it}\right)+ o_{\mathrm {P}}(1) \nonumber\\
&=\widehat{\mathbf H}^\prime \left(\frac 1{\sqrt T}\sum_{t=1}^T \mathbf F_t\xi_{it}\right)+o_{\mathrm P}(1)\nonumber\\
&= \bm J\left(\frac 1{\sqrt T}\sum_{t=1}^T \mathbf F_t\xi_{it}\right)+o_{\mathrm P}(1)\nonumber\\
&\to_d\mathcal N\left(\mathbf 0_r, \bm\Phi_i\right),
\end{align}
where we used Slutsky's theorem and Assumption (ref). Notice that, $\bm J$ plays no role in the covariance since it is diagonal and $\bm J^2=\mathbf I_r$. Because of Assumption (ref) the sign indeterminacy on the left-hand-side can be easily fixed so that $\bm J=\mathbf I_r$. By substituting $\mathbf I_r$ in place of $\widehat{\mathbf H}$ in (ref) and using Assumption (ref) it follows also that:
\[
\left\Vert \widehat{\bm\lambda}_i-{\bm\lambda}_i \right\Vert = O_{\mathrm {P}}\left(\max\left(\frac 1 n, \frac 1{\sqrt{T}}\right)\right).
\]
This completes the proof of part (a).
For part (b), from (ref)
\begin{align}
\left\Vert\frac{\widehat{\bm\Lambda}-\bm\Lambda\widehat{\mathbf H}}{\sqrt n}\right\Vert\le&\,
\left(\left\Vert\frac{\bm F^\prime\bm \Xi}{\sqrt nT}\right\Vert\,\left\Vert\frac{\bm\Lambda}{\sqrt n}\right\Vert^2
+\left\Vert \frac{\bm\Xi^\prime\bm F}{\sqrt {n}T}\right\Vert\,\left\Vert \frac{\bm\Lambda^\prime{\bm\Lambda}}{n}\right\Vert
+\left\Vert\frac{\bm\Xi^\prime\bm \Xi\bm\Lambda}{n^{3/2}T}\right\Vert\right)\left\Vert\bm{\mathcal H} \right\Vert \,\left\Vert\left(\frac{\widehat{\mathbf M}^x}{n}\right)^{-1}\right\Vert\nonumber\\
&+
\left(\left\Vert\frac{\bm F^\prime\bm \Xi}{\sqrt nT}\right\Vert\, \left\Vert\frac{\bm\Lambda}{\sqrt n}\right\Vert
+\left\Vert\frac{\bm\Xi^\prime\bm F}{\sqrt nT}\right\Vert\, \left\Vert\frac{\bm\Lambda}{\sqrt n}\right\Vert
+\left\Vert\frac{\bm\Xi^\prime\bm \Xi}{nT}\right\Vert\right)\left\Vert\frac{\widehat{\bm\Lambda}-\bm\Lambda\bm{\mathcal H}}{\sqrt n}\right\Vert \left\Vert\left(\frac{\widehat{\mathbf M}^x}{n}\right)^{-1}\right\Vert,\nonumber
\end{align}
and the proof of part (b) follows from
Proposition (ref) and Lemma
(ref)(i),
(ref),
(ref)(i),
(ref)(ii),
(ref)(iii),
(ref)(iv),
(ref). This proves part (b) and completes the proof.
$\Box$
\subsection{Proof of Corollary (ref)}
From (ref) and the second last line of (ref) in the proof of Theorem (ref) and by imposing the identification constraint of orthonormal factors in Assumption (ref)(b) and $\bm J=\mathbf I_r$ by Assumption (ref), we have
\begin{align}
\widehat{\bm\lambda}_i-\bm\lambda_i &= \frac 1{ T}\sum_{t=1}^T \mathbf F_t\xi_{it}+ O_{\mathrm P}\left(\max\left(\frac 1n,\frac 1{\sqrt {nT}}\right)\right),
\end{align}
where the rate of the last term comes from (ref) in the proof of Theorem (ref)(a). By definition of OLS, and again imposing Assumption (ref)(b), we have:
\begin{align}
{\bm\lambda}_i^{\text{\tiny OLS}}-\bm\lambda_i
= \frac 1{ T}\sum_{t=1}^T \mathbf F_t\xi_{it}.
\end{align}
By comparing (ref) and (ref) we complete the proof of part (a). Part (b) follows also directly from Theorem (ref)(b). This completes the proof.
$\Box$
\subsection{Proof of Theorem (ref)}
In principle, we could try to replicate the proofs by baili16 under our identifying constraints and using only our assumptions. However, there is a much simpler and intuitive way to proceed.
Consider the log-likelihood (ref). The parameters to be estimated are given by ${\bm\varphi}=(\mathrm{vec}({\bm\Lambda})^\prime, \mathrm{vech}({\bm\Gamma}^\xi)^\prime)^\prime$. Let also $\widehat{\bm\varphi}^{\text{\tiny \upshape QML}}=(\mathrm{vec}(\widehat{\bm\Lambda}^{\text{\tiny \upshape QML}})^\prime, \mathrm{vech}(\widehat{\bm\Gamma}^{\xi,\text{\tiny \upshape QML}})^\prime)^\prime$ denote the maximizer of (ref) and
$\underline{\bm\varphi}=(\mathrm{vec}(\underline{\bm\Lambda})^\prime, \mathrm{vech}(\underline{\bm\Gamma}^\xi)^\prime)^\prime$ denote a generic value of the parameters.
Whenever we consider $\underline{\bm\varphi}$ it is intended that its elements satisfy Assumptions (ref) through (ref).
Then, the elements of $\widehat{\bm\varphi}^{\text{\tiny \upshape QML}}$ are such that:
\begin{equation}
\widehat{\bm\Lambda}^{\text{\tiny QML}}\widehat{\bm\Lambda}^{\text{\tiny QML}\prime}+\widehat{\bm\Gamma}^{\xi,\text{\tiny QML}}=\widehat{\bm\Gamma}^x,
\end{equation}
where $\widehat{\bm\Gamma}^x=T^{-1} {\bm X^\prime\bm X}$. To see that (ref) defines the global maximum of the log-likelihood we proceed in two steps.
First, notice that the first order conditions derived from the log-likelihood (ref) are satisfied when (ref) holds:
\begin{align}
&\left.\frac{\partial \ell(\bm{\mathcal X};\underline{\bm\varphi})}{\partial\underline{\bm\Lambda}}\right\vert_{\underline{\bm\Lambda}=\widehat{\bm\Lambda}^{\text{\tiny QML}}}\nonumber\\
&= T
\left(\widehat{\bm\Lambda}^{\text{\tiny QML}}\widehat{\bm\Lambda}^{\text{\tiny QML}\prime}+\widehat{\bm\Gamma}^{\xi,\text{\tiny QML}}\right)^{-1}
\widehat{\bm\Gamma}^x
\left(\widehat{\bm\Lambda}^{\text{\tiny QML}}\widehat{\bm\Lambda}^{\text{\tiny QML}\prime}+\widehat{\bm\Gamma}^{\xi,\text{\tiny QML}}\right)^{-1}
\widehat{\bm\Lambda}^{\text{\tiny QML}}
-T\left(\widehat{\bm\Lambda}^{\text{\tiny QML}}\widehat{\bm\Lambda}^{\text{\tiny QML}\prime}+\widehat{\bm\Gamma}^{\xi,\text{\tiny QML}}\right)^{-1}
\widehat{\bm\Lambda}^{\text{\tiny QML}}\nonumber\\
&= T
\left(\widehat{\bm\Lambda}^{\text{\tiny QML}}\widehat{\bm\Lambda}^{\text{\tiny QML}\prime}+\widehat{\bm\Gamma}^{\xi,\text{\tiny QML}}\right)^{-1}
\widehat{\bm\Lambda}^{\text{\tiny QML}}
-T\left(\widehat{\bm\Lambda}^{\text{\tiny QML}}\widehat{\bm\Lambda}^{\text{\tiny QML}\prime}+\widehat{\bm\Gamma}^{\xi,\text{\tiny QML}}\right)^{-1}
\widehat{\bm\Lambda}^{\text{\tiny QML}}=\mathbf 0_{r\times n},\nonumber\\
&\left.\frac{\partial \ell(\bm{\mathcal X};\underline{\bm\varphi})}{\partial\underline{\bm\Gamma}^\xi}\right\vert_{\underline{\bm\Gamma}^\xi=\widehat{\bm\Gamma}^{\xi,\text{\tiny QML}}}\nonumber\\
&=
\frac T2 \left(\widehat{\bm\Lambda}^{\text{\tiny QML}}\widehat{\bm\Lambda}^{\text{\tiny QML}\prime}+\widehat{\bm\Gamma}^{\xi,\text{\tiny QML}}\right)^{-1} \widehat{\bm\Gamma}^x
\left(\widehat{\bm\Lambda}^{\text{\tiny QML}}\widehat{\bm\Lambda}^{\text{\tiny QML}\prime}+\widehat{\bm\Gamma}^{\xi,\text{\tiny QML}}\right)^{-1}
- \frac T2 \left(\widehat{\bm\Lambda}^{\text{\tiny QML}}\widehat{\bm\Lambda}^{\text{\tiny QML}\prime}+\widehat{\bm\Gamma}^{\xi,\text{\tiny QML}}\right)^{-1}\nonumber\\
&=
\frac T2
\left(\widehat{\bm\Lambda}^{\text{\tiny QML}}\widehat{\bm\Lambda}^{\text{\tiny QML}\prime}+\widehat{\bm\Gamma}^{\xi,\text{\tiny QML}}\right)^{-1}
- \frac T2 \left(\widehat{\bm\Lambda}^{\text{\tiny QML}}\widehat{\bm\Lambda}^{\text{\tiny QML}\prime}+\widehat{\bm\Gamma}^{\xi,\text{\tiny QML}}\right)^{-1}=\mathbf 0_{n\times n}.\nonumber
\end{align}
Notice also that the conditions given in baili12, which are derived from the first order conditions above, are also satisfied. Namely, it holds that:
\begin{align}
&\widehat{\bm\Lambda}^{\text{\tiny QML}\prime}\left(\widehat{\bm\Lambda}^{\text{\tiny QML}}\widehat{\bm\Lambda}^{\text{\tiny QML}\prime}+\widehat{\bm\Gamma}^{\xi,\text{\tiny QML}}\right)^{-1}\left\{\widehat{\bm\Gamma}^x-\widehat{\bm\Lambda}^{\text{\tiny QML}}\widehat{\bm\Lambda}^{\text{\tiny QML}\prime}-\widehat{\bm\Gamma}^{\xi,\text{\tiny QML}}\right\}=\mathbf 0_{r\times n},\nonumber\\
&\left(\widehat{\bm\Lambda}^{\text{\tiny QML}}\widehat{\bm\Lambda}^{\text{\tiny QML}\prime}+\widehat{\bm\Gamma}^{\xi,\text{\tiny QML}}\right)^{-1}=
\left(\widehat{\bm\Lambda}^{\text{\tiny QML}}\widehat{\bm\Lambda}^{\text{\tiny QML}\prime}+\widehat{\bm\Gamma}^{\xi,\text{\tiny QML}}\right)^{-1}
\widehat{\bm \Gamma}^x
\left(\widehat{\bm\Lambda}^{\text{\tiny QML}}\widehat{\bm\Lambda}^{\text{\tiny QML}\prime}+\widehat{\bm\Gamma}^{\xi,\text{\tiny QML}}\right)^{-1},\nonumber\\
&\widehat{\bm\Lambda}^{\text{\tiny QML}\prime}
\left(\widehat{\bm\Lambda}^{\text{\tiny QML}}\widehat{\bm\Lambda}^{\text{\tiny QML}\prime}+\widehat{\bm\Gamma}^{\xi,\text{\tiny QML}}\right)^{-1}
\widehat{\bm\Lambda}^{\text{\tiny QML}}=
\widehat{\bm\Lambda}^{\text{\tiny QML}\prime}
\left(\widehat{\bm\Lambda}^{\text{\tiny QML}}\widehat{\bm\Lambda}^{\text{\tiny QML}\prime}+\widehat{\bm\Gamma}^{\xi,\text{\tiny QML}}\right)^{-1}
\widehat{\bm \Gamma}^x
\left(\widehat{\bm\Lambda}^{\text{\tiny QML}}\widehat{\bm\Lambda}^{\text{\tiny QML}\prime}+\widehat{\bm\Gamma}^{\xi,\text{\tiny QML}}\right)^{-1}
\widehat{\bm\Lambda}^{\text{\tiny QML}}.\nonumber
\end{align}
Second, given the log-likelihood (ref), for any $\underline{\bm\varphi}$, we have:
\begin{align}
\ell(\bm{\mathcal X};\widehat{\bm\varphi}^{\text{\tiny QML}})&-\ell(\bm{\mathcal X};\underline{\bm\varphi})\nonumber\\
&=-\frac{T}2\log\frac{\det\left( \widehat{\bm\Lambda}^{\text{\tiny QML}}\widehat{\bm\Lambda}^{\text{\tiny QML}\prime}+\widehat{\bm\Gamma}^{\xi,\text{\tiny QML}}\right)}
{\det\left(\underline{\bm\Lambda}\,\underline{\bm\Lambda}^\prime+\underline{\bm\Gamma}^\xi \right)}-\frac{nT}{2}
+\frac T2
\text{tr}\left\{
\widehat{\bm\Gamma}^x
\left(\underline{\bm\Lambda}\,\underline{\bm\Lambda}^\prime+\underline{\bm\Gamma}^\xi \right)^{-1}
\right\}\\
&=-\frac{T}2\log\frac{\det\left( \widehat{\bm\Lambda}^{\text{\tiny QML}}\widehat{\bm\Lambda}^{\text{\tiny QML}\prime}+\widehat{\bm\Gamma}^{\xi,\text{\tiny QML}}\right)}
{\det\left(\underline{\bm\Lambda}\,\underline{\bm\Lambda}^\prime+\underline{\bm\Gamma}^\xi \right)}-\frac{nT}{2}
+\frac T2
\text{tr}\left\{
\left(
\widehat{\bm\Lambda}^{\text{\tiny QML}}\widehat{\bm\Lambda}^{\text{\tiny QML}\prime}+\widehat{\bm\Gamma}^{\xi,\text{\tiny QML}}
\right)
\left(\underline{\bm\Lambda}\,\underline{\bm\Lambda}^\prime+\underline{\bm\Gamma}^\xi \right)^{-1}
\right\},\nonumber
\end{align}
because of (ref). Now, denote as $\zeta_j$, $j=1,\ldots, n$, the $n$ roots of
\[
\det\left\{\left(\widehat{\bm\Lambda}^{\text{\tiny \upshape QML}}\widehat{\bm\Lambda}^{\text{\tiny \upshape QML}\prime}+\widehat{\bm\Gamma}^{\xi,\text{\tiny \upshape QML}}\right)-\zeta\left(
\underline{\bm\Lambda}\,\underline{\bm\Lambda}^\prime+\underline{\bm\Gamma}^\xi
\right) \right\}=0,
\]
which are all real since $\left(\widehat{\bm\Lambda}^{\text{\tiny \upshape QML}}\widehat{\bm\Lambda}^{\text{\tiny \upshape QML}\prime}+\widehat{\bm\Gamma}^{\xi,\text{\tiny \upshape QML}}\right)\left(
\underline{\bm\Lambda}\,\underline{\bm\Lambda}^\prime+\underline{\bm\Gamma}^\xi
\right)^{-1}
$ is a symmetric matrix. Then, (ref) reads:
\begin{equation}
\ell(\bm{\mathcal X};\widehat{\bm\varphi}^{\text{\tiny QML}})-\ell(\bm{\mathcal X};\underline{\bm\varphi}) = \frac T2\sum_{j=1}^n \left\{-\log\zeta_j-1+\zeta_j\right\}\ge 0,
\end{equation}
since $x\le e^{x-1}$ and so $-\log x-1+x\ge 0$. Therefore, from (ref) we see that (ref) defines indeed the global maximum of the log-likleihood (ref).
Now, consider the Singular Value Decomposition (SVD) of the true loadings:
\begin{equation}
\frac{{\bm\Lambda}}{\sqrt n}={\mathbf V}{\mathbf D}{\mathbf U},
\end{equation}
where ${\mathbf V}$ is $n\times r$ and such that ${\mathbf V}^\prime{\mathbf V}=\mathbf I_r$ for all $n\in\mathbb N$,
${\mathbf D}$ is $r\times r$ diagonal with strictly positive entries, and ${\mathbf U}$ is $r\times r$ such that
${\mathbf U}{\mathbf U}^\prime={\mathbf U}^{\prime}{\mathbf U}=\mathbf I_r$.
We know that under Assumption (ref) and (ref), $\bm\Lambda$ is globally identified. Let us show that ${\mathbf V}$, ${\mathbf D}$, and ${\mathbf U}$ in (ref) are also globally identified. First, notice that, given Assumption (ref)(a) which requires $n^{-1}{{\bm\Lambda}^\prime{\bm\Lambda}}$ to be diagonal, in order to estimate ${\bm\Lambda}$ we need to estimate $nr-{r(r-1)}/{2}$ parameters. Then, by looking at the
right-hand-side of (ref) we see that to estimate ${\mathbf V}$ we need to estimate $nr-\frac{r(r+1)}{2}$ parameters and to estimate ${\mathbf D}$ we need to estimate $r$ parameters, thus ${\mathbf V}{\mathbf D}$ depends on $nr-{r(r+1)}/{2}+r=nr-{r(r-1)}/{2}$ parameters as ${\bm\Lambda}$. However, in principle the
right-hand-side of (ref) depends also on ${\mathbf U}$ which in turn requires estimating ${r(r+1)}/2$ parameters more. But if we impose Assumption (ref)(a) also to the right-hand-side of (ref) we have that $n^{-1}{{\bm\Lambda}^\prime{\bm\Lambda}}={\mathbf U}^\prime{\mathbf D}^2{\mathbf U}$ has to be diagonal, and since ${\mathbf D}$ is diagonal, without loss of generality we can set ${\mathbf U}=\mathbf J$, a diagonal $r\times r$ matrix with entries $\pm 1$.
Furthermore, from Proposition (ref)(a) we also see that we must have
$n^{-1}{{\bm\Lambda}^\prime{\bm\Lambda}}=n^{-1}{\mathbf M^\chi}=\mathbf D^2$.
Hence,
\begin{equation}
\mathbf D = \left(\frac{\mathbf M^\chi}n\right)^{1/2},
\end{equation}
and by Lemma (ref)(iv) the entries of $\mathbf D$, denoted as $d_j$, $j=1,\ldots,r$, are such that
\begin{equation}
\sqrt{\underline C_j}\!\le \lim\inf_{n\to\infty} d_j \le\lim\sup_{n\to\infty} d_j \le\! \sqrt{\overline C_j},
\end{equation}
where $\underline C_j$ and $\overline C_j$ are finite positive reals. Last, from (ref) and Proposition (ref)(b), we must have
$n^{-1/2}{{\bm\Lambda}}={\mathbf V}{\mathbf D}\mathbf J = \mathbf V^\chi \left(n^{-1}{\mathbf M^\chi}\right)^{1/2}$, and by (ref) it follows that
\begin{equation}
{\mathbf V} = \mathbf V^\chi \mathbf J.
\end{equation}
This shows that under Assumption (ref)(a), the parameters in ${\mathbf D}$ are globally identified while the parameters in $\mathbf V$ are uniquely identified up to a right-multiplication by $\mathbf J$, which can be pinned down by means of Assumption (ref), thus achieving global identification of $\mathbf V$ as well. Given this discussion, hereafter, we can directly set $\mathbf U=\mathbf I_r$.
It is clear that the problem of QML estimation of the loadings can be rewritten as a problem of QML estimation of their SVD in (ref), namely of $\mathbf V$ and $\mathbf D$. To this end, we introduce also the SVDs of the QML estimator of the loadings and of a generic value of the loadings:
\begin{align}
&\frac{\widehat{\bm\Lambda}^{\text{\tiny QML}}}{\sqrt n}=\widehat{\mathbf V}^{\text{\tiny QML}}\widehat{\mathbf D}^{\text{\tiny QML}}\widehat{\mathbf U}^{\text{\tiny QML}},\qquad\frac{\underline{\bm\Lambda}}{\sqrt n}=\underline{\mathbf V}\,\underline{\mathbf D}\,\underline{\mathbf U},
\end{align}
where
$\widehat{\mathbf V}^{\text{\tiny \upshape QML}}$ and $\underline{\mathbf V}$ have the same properties as ${\mathbf V}$,
$\widehat{\mathbf D}^{\text{\tiny \upshape QML}}$ and $\underline{\mathbf D}$ have the same properties as ${\mathbf D}$,
and
$\widehat{\mathbf U}^{\text{\tiny \upshape QML}}$ and $\underline{\mathbf U}$ have the same properties as ${\mathbf U}$.
Because we set $\mathbf U=\mathbf I_r$, it follows that we can set $\widehat{\mathbf U}^{\text{\tiny \upshape QML}} = \mathbf I_r$, and we are left with the task of finding $\widehat{\mathbf V}^{\text{\tiny \upshape QML}}$ and $\widehat{\mathbf D}^{\text{\tiny \upshape QML}}$. From (ref) and since we must have $\widehat{\mathbf V}^{\text{\tiny \upshape QML}\prime}\widehat{\mathbf V}^{\text{\tiny \upshape QML}}=\mathbf I_r$, it follows that
\begin{equation}
\frac{\widehat{\bm\Gamma}^x}{n}-\widehat{\mathbf V}^{\text{\tiny QML}}\left(\widehat{\mathbf D}^{\text{\tiny QML}}\right)^2 \widehat{\mathbf V}^{\text{\tiny QML}\prime}-\frac{\widehat{\bm\Gamma}^{\xi,\text{\tiny QML}}}{n}=\mathbf 0_{n\times n},\nonumber
\end{equation}
which is equivalent to
\begin{equation}
\frac{\widehat{\mathbf V}^{\text{\tiny QML}\prime}\widehat{\bm\Gamma}^x\widehat{\mathbf V}^{\text{\tiny QML}}}{n}-\left(\widehat{\mathbf D}^{\text{\tiny QML}}\right)^2 -\frac{\widehat{\mathbf V}^{\text{\tiny QML}\prime}\widehat{\bm\Gamma}^{\xi,\text{\tiny QML}}\widehat{\mathbf V}^{\text{\tiny QML}}}{n}=\mathbf 0_{r\times r}.
\end{equation}
Now, let $\widehat{\bm v}_j^{\text{\tiny \upshape QML}}$, $j=1,\ldots, r$, be the $j$th column of $\widehat{\mathbf V}^{\text{\tiny \upshape QML}}$ and let
$\widehat{d}_j^{\,\text{\tiny \upshape QML}}$, $j=1,\ldots, r$, be the $j$th diagonal entry of $\widehat{\mathbf D}^{\text{\tiny \upshape QML}}$. Then, from (ref), for all $j=1,\ldots, r$, we have
\begin{equation}
\frac{\widehat{\bm v}_j^{\text{\tiny QML}\prime}\widehat{\bm\Gamma}^x\widehat{\bm v}_j^{\text{\tiny QML}}}{n}-\left(\widehat{ d}_j^{\,\text{\tiny QML}}\right)^2 -\frac{\widehat{\bm v}_j^{\text{\tiny QML}\prime}\widehat{\bm\Gamma}^{\xi,\text{\tiny QML}}\widehat{\bm v}_j^{\text{\tiny QML}}}{n}=0.
\end{equation}
Then, trivially, the QML estimators are such that:
\begin{equation}
\left(\widehat{\bm v}_j^{\text{\tiny QML}}, \widehat{d}_j^{\,\text{\tiny QML}},\widehat{\bm\Gamma}^{\xi,\text{\tiny QML}}\right)=\arg\!\!\!\!\min_{\underline{\bm v}_j,\underline{d}_j,\underline{\bm\Gamma}^\xi}
\left(\frac{\underline{\bm v}_j^{\prime}\widehat{\bm\Gamma}^x\underline{\bm v}_j}{n}-\underline{ d}_j^2 -
\frac{\underline{\bm v}_j^{\prime}\underline{\bm\Gamma}^{\xi}\underline{\bm v}_j}n
\right)^2=\arg\!\!\!\!\min_{\underline{\bm v}_j,\underline{d}_j,\underline{\bm\Gamma}^\xi}
\mathcal L_{0j}\left(\underline{\bm v}_j, \underline d_j,\underline{\bm\Gamma}^{\xi} \right), \;\text{say,}
\end{equation}
where $\underline{\bm v}_j$, $j=1,\ldots, r$, is the $j$th column of $\underline{\mathbf V}$ and
$\underline{d}_j$, $j=1,\ldots, r$, is the $j$th diagonal entry of $\underline{\mathbf D}$.
Define also the estimators $\widetilde{\bm v}_j$ and $\widetilde{d}_j$ such that:
\begin{equation}
\left(\widetilde{\bm v}_j, \widetilde{d}_j\right)=\arg\!\min_{\underline{\bm v}_j,\underline{d}_j}
\left(\frac{\underline{\bm v}_j^{\prime}\widehat{\bm\Gamma}^x\underline{\bm v}_j}{n}-\underline{ d}_j^2
\right)^2=\arg\!\min_{\underline{\bm v}_j,\underline{d}_j}\mathcal L_{1j}\left(\underline{\bm v}_j, \underline d_j\right), \;\text{say.}
\end{equation}
Consistently with (ref) and (ref), these define an estimator of $\bm\Lambda$ by means of its SVD:
\begin{equation}
\frac{\widetilde{\bm\Lambda}}{\sqrt n} = \widetilde{\mathbf V}\widetilde{\mathbf D}\widetilde{\mathbf U},
\end{equation}
where $\widetilde{\mathbf V}$ has columns $\widetilde{\bm v}_j$, $j=1,\ldots, r$, and it has the same properties as ${\mathbf V}$ (because $\underline{\mathbf V}$ does), and
$\widetilde{\mathbf D}$ has entries $\widetilde{ d}_j$, $j=1,\ldots, r$, and it has the same properties as ${\mathbf D}$ (because $\underline{\mathbf D}$ does). We set $\widetilde{\mathbf U}=\mathbf I_r$ consistently with the fact that $\mathbf U=\mathbf I_r$.
Then, it is easily seen that
\begin{equation}
\mathcal L_{0j}\left(\underline{\bm v}_j, \underline d_j,\underline{\bm\Gamma}^{\xi} \right)= \mathcal L_{1j}\left(\underline{\bm v}_j, \underline d_j\right)+\left(\frac{\underline{\bm v}_j^{\prime}\underline{\bm\Gamma}^{\xi}\underline{\bm v}_j}n\right)^2-2\left(\frac{\underline{\bm v}_j^{\prime}\underline{\bm\Gamma}^{\xi}\underline{\bm v}_j}n\right)\sqrt{\mathcal L_{1j}\left(\underline{\bm v}_j, \underline d_j\right)}.
\end{equation}
Now, we know that for any generic value of the idiosyncratic covariance matrix satisfying Assumption (ref)(b), it holds that
\begin{equation}
\frac{\underline{\bm v}_j^{\prime}\underline{\bm\Gamma}^{\xi}\underline{\bm v}_j}n\le \max_{\bm w\, :\, \bm w^\prime\bm w=1}\frac{\bm w'\underline{\bm\Gamma}^{\xi}\bm w}{n}=\frac{\underline {\mu}^\xi_1}{n}\le \frac{M_{2\xi}}n,
\end{equation}
since $\underline{\bm v}_j^{\prime}\underline{\bm v}_j=1$ and because of Lemma (ref)(v) which implies also that $M_{2\xi}$ is independent of $n$. Moreover, for any generic $\underline{\bm \Lambda}$ satisfying Assumption (ref)(a):
\begin{equation}
\sqrt{\mathcal L_{1j}\left(\underline{\bm v}_j, \underline d_j\right)} = \left\vert \frac{\underline{\bm v}_j^{\prime}\widehat{\bm\Gamma}^x\underline{\bm v}_j}{n}-\underline{ d}_j^2 \right\vert \le
\frac{\underline{\bm v}_j^{\prime}\widehat{\bm\Gamma}^x\underline{\bm v}_j}{n}+ \underline{ d}_j^2
\le \frac{\widehat{\mu}_1^x}{n} + \underline{ d}_j^2
= O_{\mathrm P}(1),
\end{equation}
since $\underline d_j= O(1)$ because $\Vert n^{-1/2}{\underline{\bm \Lambda}}\Vert= O(1)$ by Lemma (ref)(i), and $\widehat{\mu}_1^x = O_{\mathrm P}(n)$ because of Lemma (ref)(iii). Therefore, from (ref), (ref), and (ref), we have
\begin{equation}
\left\vert\mathcal L_{0j}\left(\underline{\bm v}_j, \underline d_j,\underline{\bm\Gamma}^{\xi} \right)-\mathcal L_{1j}\left(\underline{\bm v}_j, \underline d_j \right)\right\vert = O_{\mathrm P}\left(\frac 1n\right),
\end{equation}
which holds for any $\underline{\bm\Gamma}^{\xi}$ satisfying Assumption (ref)(b). By continuity of these loss functions, from (ref) it follows that their minima satisfy:
\begin{align}
\left\Vert \widehat{\bm v}_j^{\text{\tiny QML}}-\widetilde{\bm v}_j \right\Vert = O_{\mathrm P}\left(\frac 1n\right),\qquad \left\vert \widehat{d}_j^{\,\text{\tiny QML}}-\widetilde{d}_j \right\vert = O_{\mathrm P}\left(\frac 1n\right),
\end{align}
for all $j=1,\ldots, r$.
Let us now find $\widetilde{\bm v}_j $ and $\widetilde{d}_j $. From (ref), it is clear that the solutions must be such that:
\begin{equation}
\frac{\widetilde{\bm v}_j^{\prime}\widehat{\bm\Gamma}^x\widetilde{\bm v}_j}{n}=\widetilde{ d}_j^{\,2},
\end{equation}
which means that $\widetilde{ d}_j^{\,2}$ must be an eigenvalue of $n^{-1}{\widehat{\bm\Gamma}^x}$ and $\widetilde{\bm v}_j$ is the corresponding normalized eigenvector. Obviously, the solution in (ref) defines a global minimum of the loss $\mathcal L_1(\underline{\bm v}_j,\underline d_j)$ since $\mathcal L_1(\widetilde{\bm v}_j,\widetilde d_j)=0$ while $\mathcal L_1(\underline{\bm v}_j,\underline d_j)> 0$ for any other value $\underline{\bm v}_j\ne \widetilde{\bm v}_j$ and $\underline d_j\ne \widetilde d_j$.
Now, let us show that indeed it must be that $\widetilde{ d}_j^{\,2}=n^{-1}{\widehat{\mu}_j^x}$, $j=1,\ldots, r$, i.e., they have to be the $r$ largest eigenvalues of $n^{-1}{\widehat{\bm\Gamma}^x}$. First, by Lemma (ref)(i) and Weyl's inequality, for all $k=1,\ldots, r$, as $n,T\to\infty$,
\begin{equation}
\left\vert\frac{\widehat{\mu}_k^x}n-\frac{{\mu}_k^x}n\right\vert\le \left\Vert\frac{\widehat{\bm\Gamma}^x}n-\frac{{\bm\Gamma}^x}n\right\Vert= O_{\mathrm P}\left(\frac 1{\sqrt T}\right).
\end{equation}
Therefore, if for any given $j=1,\ldots,r $ we were to choose $\widetilde{ d}_j^{\,2}=n^{-1}{\widehat{\mu}_k^x}$ for, say, $k=r+1$, then, from (ref), we would have
\begin{align}
&\lim\inf_{n\to\infty} \widetilde{ d}_j^{\,2} =\lim\inf_{n\to\infty} \frac{\widehat{\mu}_{r+1}^x}n =\lim\inf_{n\to\infty} \frac{{\mu}_{r+1}^x}n+O_{\mathrm P}\left(\frac 1{\sqrt T}\right)
\ge \lim\inf_{n\to\infty} \frac{{\mu}_{n}^{\xi}}n+O_{\mathrm P}\left(\frac 1{\sqrt T}\right)= O_{\mathrm P}\left(\frac 1{\sqrt T}\right),\nonumber\\
&\lim\sup_{n\to\infty} \widetilde{ d}_j^{\,2} =\lim\sup_{n\to\infty} \frac{\widehat{\mu}_{r+1}^x}n =\lim\sup_{n\to\infty} \frac{{\mu}_{r+1}^x}n+O_{\mathrm P}\left(\frac 1{\sqrt T}\right)
\le \lim_{n\to\infty} \frac{M_\xi}n+O_{\mathrm P}\left(\frac 1{\sqrt T}\right)= O_{\mathrm P}\left(\frac 1{\sqrt T}\right),\nonumber
\end{align}
since we assumed $\bm\Gamma^{\xi}$ to be positive definite, so $\mu^\xi_n>0$ for all $n\in\mathbb N$, and by Lemma (ref)(vi). Therefore, as $n,T\to\infty$, this choice for $\widetilde{ d}_j$ cannot be a consistent estimator of $d_j$ (nor an approximation of $\widehat{d}_j^{\,\text{\tiny QML}}$), since while $\widetilde{ d}_j\to0$, as $n,T\to\infty$, it must be that $d_j>0$ for all $n\in\mathbb N$, as required in (ref).
So from (ref) and the above reasoning it follows that, for all $j=1,\ldots, r$,
\begin{equation}
\widetilde{ d}_j^{\,2}= \frac{\widehat{\mu}_j^x}{n},\qquad \widetilde{\bm v}_j = \widehat{\bm v}_j^x.
\end{equation}
where $ \widehat{\bm v}_j^x$ is the eigenvector of $n^{-1}{\widehat{\bm\Gamma}^x}$ corresponding to its $j$th largest eigenvalue.
Then, from (ref), first we have
\begin{align}
&\left\Vert \widehat{\mathbf V}^{\text{\tiny QML}}-\widehat{\mathbf V}^x \right\Vert=\left\Vert \widehat{\mathbf V}^{\text{\tiny QML}}-\widetilde{\mathbf V} \right\Vert = O_{\mathrm P}\left(\frac 1n\right),\\
&\left\Vert \widehat{\mathbf D}^{\,\text{\tiny QML}}-\left(\frac{\widehat{\mathbf M}^x}{n}\right)^{1/2} \right\Vert=
\left\Vert \widehat{\mathbf D}^{\,\text{\tiny QML}}-\widetilde{\mathbf D} \right\Vert = O_{\mathrm P}\left(\frac 1n\right),
\end{align}
because of (ref), and, second, we have
\begin{equation}
\frac{\widetilde{\bm\Lambda}}{\sqrt n} = \widehat{\mathbf V}^x\left(\frac{\widehat{\mathbf M}^x}{n}\right)^{1/2}=\frac{\widehat{\bm\Lambda}}{\sqrt n} ,
\end{equation}
which shows that the estimator $\widetilde{\bm\Lambda}$
is the PC estimator defined in (ref). Therefore, from (ref), (ref), and (ref), and by using the SVD of the QML estimator in (ref), it follows that:
\begin{align}
\left\Vert
\frac{\widehat{\bm\Lambda}^{\text{\tiny QML}}}{\sqrt n}-\frac{\widehat{\bm\Lambda}}{\sqrt n}
\right\Vert =&\,
\left\Vert
\widehat{\mathbf V}^{\text{\tiny QML}}\widehat{\mathbf D}^{\text{\tiny QML}}
- \widehat{\mathbf V}^x\left(\frac{\widehat{\mathbf M}^x}{n}\right)^{1/2}
\right\Vert\nonumber\\
\le&\, \left\Vert
\widehat{\mathbf V}^{\text{\tiny QML}}- \widehat{\mathbf V}^x
\right\Vert\,
\left\Vert
\left(\frac{\widehat{\mathbf M}^x}{n}\right)^{1/2}
\right\Vert
+
\left\Vert
\widehat{\mathbf D}^{\text{\tiny QML}}-\left(\frac{\widehat{\mathbf M}^x}{n}\right)^{1/2}
\right\Vert\,
\left\Vert
\widehat{\mathbf V}^x
\right\Vert\nonumber\\
&+\left\Vert
\widehat{\mathbf V}^{\text{\tiny QML}}- \widehat{\mathbf V}^x
\right\Vert\,
\left\Vert
\widehat{\mathbf D}^{\text{\tiny QML}}-\left(\frac{\widehat{\mathbf M}^x}{n}\right)^{1/2}
\right\Vert = O_{\mathrm P}\left(\frac 1n\right),\nonumber
\end{align}
because $ \Vert\widehat{\mathbf V}^x \Vert=1$ (it is a matrix of normalized eigenvectors) and $\Vert (n^{-1}{\widehat{\mathbf M}^x})^{1/2}\Vert=O_{\mathrm P}(1)$ by Lemma (ref)(iii). This proves part (a).
Finally, consider the $r$ dimensional $i$th rows, $i=1,\ldots, n$, of ${\mathbf V}^\chi$, $\widehat{\mathbf V}^x$, $\underline {\mathbf V}$, and $\widehat{\mathbf V}^{\text{\tiny \upshape QML}}$, denoted as ${\mathbf v}_i^{\chi\prime}$, $\widehat{\mathbf v}_i^{x\prime}$, $\underline{\mathbf v}_i^\prime$ , and $\widehat{\mathbf v}_i^{\text{\tiny \upshape QML}\prime}$, respectively. From Lemma (ref)(ii) and (ref)(iii), we know that $\sqrt n\Vert {\mathbf v}_i^{\chi\prime}\Vert = O(1)$ and $\sqrt n\Vert\widehat{\mathbf v}_i^{x\prime}\Vert=O_{\mathrm P}(1)$, which means that we must have $\sqrt n\Vert \underline {\mathbf v}_i^\prime\Vert= O(1)$ since any generic parameter we consider is assumed to satisfy the same assumptions as the corresponding true parameters (${\mathbf v}_i^{\chi\prime}$ in this case). Therefore, since the search of the elements of $\widehat{\mathbf V}^{\text{\tiny \upshape QML}}$ is made over all $\underline {\mathbf V}$ satisfying the same assumptions as ${\mathbf V}^\chi$, it follows that we must have also $\sqrt n\Vert \widehat{\mathbf v}_i^{\text{\tiny \upshape QML}\prime}\Vert=O_{\mathrm P}(1)$.
To see that this is the case, notice also that because of Assumption (ref)(a) the QML estimator must be such that $n^{-1}{\widehat{\bm \Lambda}^{\text{\tiny \upshape QML}\prime}\widehat{\bm \Lambda}^{\text{\tiny \upshape QML}}}$ is diagonal and positive definite. Thus, we must have $\Vert n^{-1/2}{\widehat{\bm \Lambda}^{\text{\tiny \upshape QML}}}\Vert = O_{\mathrm P}(1)$ (see also Lemma (ref)(i)). From the SVD in (ref) when $\widehat{\mathbf U}^{\text{\tiny \upshape QML}}=\mathbf I_r$ and since it must be that $\widehat{\mathbf V}^{\text{\tiny \upshape QML}\prime}\widehat{\mathbf V}^{\text{\tiny \upshape QML}}=\mathbf I_r$ for all $n\in\mathbb N$, then it follows that $\Vert \widehat{\mathbf V}^{\text{\tiny \upshape QML}}\Vert =O_{\mathrm P}(1)$ and $\Vert \widehat{\mathbf D}^{\text{\tiny \upshape QML}}\Vert =O_{\mathrm P}(\sqrt n)$. Moreover, since by Assumption (ref)(a) the $i$th row of
$\widehat{\bm \Lambda}^{\text{\tiny \upshape QML}}$ must be such that $\Vert\widehat{\bm \lambda}_i^{\text{\tiny \upshape QML}}\Vert=O_{\mathrm P}(1)$ and $\widehat{\bm \lambda}_i^{\text{\tiny \upshape QML}\prime}= \widehat{\mathbf v}_i^{\text{\tiny \upshape QML}\prime} \widehat{\mathbf D}^{\text{\tiny \upshape QML}}$, then it must be that $\sqrt n\Vert \widehat{\mathbf v}_i^{\text{\tiny \upshape QML}\prime}\Vert=O_{\mathrm P}(1)$.
From this reasoning and (ref) it follows that:
\begin{equation}
\sqrt n \left\Vert \widehat{\mathbf v}_i^{\text{\tiny QML}}-\widehat{\mathbf v}_i^x
\right\Vert = O_{\mathrm P}\left(\frac 1n\right).
\end{equation}
Now, by taking the $i$th row of the PC estimator in (ref) and the SVD of the QML estimator in (ref), for any $i=1,\ldots, n$, we get
\begin{align}
\left\Vert
\widehat{\bm\lambda}_i^{\text{\tiny QML}\prime}-\widehat{\bm\lambda}_i^{\prime}
\right\Vert =&\,
\left\Vert
\widehat{\mathbf v}_i^{\text{\tiny QML}\prime} \sqrt n \widehat{ \mathbf D}^{\text{\tiny QML}}-
\widehat{\mathbf v}_i^{x\prime}(\widehat{\mathbf M}^x)^{1/2}
\right\Vert\nonumber\\
\le&\,
\left\Vert
\widehat{\mathbf v}_i^{\text{\tiny QML}} - \widehat{\mathbf v}_i^x
\right\Vert \,
\left\Vert
\left({\widehat{\mathbf M}^x}\right)^{1/2}
\right\Vert +
\left\Vert
\sqrt n \widehat{ \mathbf D}^{\text{\tiny QML}}-\left({\widehat{\mathbf M}^x}\right)^{1/2}
\right\Vert\,
\left\Vert
\widehat{\mathbf v}_i^x
\right\Vert\nonumber\\
&+\left\Vert
\widehat{\mathbf v}_i^{\text{\tiny QML}} - \widehat{\mathbf v}_i^x
\right\Vert \,
\left\Vert
\sqrt n \widehat{ \mathbf D}^{\text{\tiny QML}}-\left({\widehat{\mathbf M}^x}\right)^{1/2}
\right\Vert= O_{\mathrm P}\left(\frac 1n\right),\nonumber
\end{align}
because of (ref), (ref), and since $\Vert
\widehat{\mathbf v}_i^x
\Vert=O_{\mathrm P}(n^{-1/2})$ by Lemma (ref)(ii) and (ref)(iii), and $\Vert ({\widehat{\mathbf M}^x})^{1/2}\Vert=O_{\mathrm P}(\sqrt n)$ by Lemma (ref)(iii). This proves part (b) and completes the proof. $\Box$\\
\subsection{Proof of Theorem (ref)}
From Theorem (ref)(b) and Corollary (ref) it follows that:
\begin{align}
\Vert\widehat{\bm\lambda}_i^{\text{\tiny QML}}-{\bm\lambda}_i\Vert&=\Vert\widehat{\bm\lambda}_i-{\bm\lambda}_i\Vert+O_{\mathrm {P}}\left(\frac 1{n}\right)\nonumber\\
&=\Vert\widehat{\bm\lambda}_i-{\bm\lambda}_i^{\text{\tiny OLS}}\Vert+\Vert{\bm\lambda}_i^{\text{\tiny OLS}}-{\bm\lambda}_i\Vert+O_{\mathrm {P}}\left(\frac 1{n}\right)\nonumber\\
&=\Vert{\bm\lambda}_i^{\text{\tiny OLS}}-{\bm\lambda}_i\Vert+O_{\mathrm {P}}\left(\max\left(\frac 1n,\frac 1{\sqrt {nT}}\right)\right)\nonumber\\
&=O_{\mathrm {P}}\left(\max\left(\frac 1n,\frac 1{\sqrt {nT}}, \frac 1{\sqrt {T}}\right)\right),
\end{align}
since the unfeasible OLS estimator is $\sqrt T$-consistent. Moreover, if $\sqrt T/n\to 0$ as $n,T\to\infty$,
by imposing the identification constraint of orthonormal factors in Assumption (ref)(b),
we have
\begin{align}
\sqrt T(\widehat{\bm\lambda}_i^{\text{\tiny QML}}-\bm\lambda_i)&=
\sqrt T(\widehat{\bm\lambda}_i^{\text{\tiny OLS}}-\bm\lambda_i)+ o_{\mathrm P}(1)
=\frac 1{\sqrt T}\sum_{t=1}^T \mathbf F_t\xi_{it}+ o_{\mathrm P}(1)\to_d\mathcal N\left(\mathbf 0_r, \bm\Phi_i\right).\nonumber
\end{align}
by Slutsky's theorem and Assumption (ref). This completes the proof of part (a). Part (b) follows similarly from Theorem (ref)(a). This completes the proof.
$\Box$
\subsection{Proof of Corollary (ref)}
The proof of part (a) follows by using the log-likelihood (ref) of an exact factor model in place of the log-likelihood (ref), then, by
noticing that $\bm{\Sigma}^{\xi}$ is positive definite by Assumption (ref)(a), and fnally by
replacing in the proof of Theorem (ref), $\widehat{\bm\Gamma}^{\xi,\text{\tiny QML}}$ and $\underline{\bm\Gamma}^{\xi}$ with $\widehat{\bm\Sigma}^{\xi,\text{\tiny QML}}$
and $\underline{\bm\Sigma}^{\xi}$ respectively. The proof of part (b) is the same but when using $\sigma^2\mathbf I_n$, with $\sigma^2>0$, in place of $\bm\Sigma^\xi$. $\Box$
\subsection{Proof of Theorem (ref)}
First of all, denote the log-likelihoods for one observation:
\begin{align}
&\ell_{t}(\mathbf x_t;\underline{\bm\varphi})=-\frac 12\log \det(\underline{\bm\Lambda}\,\underline{\bm\Lambda}^\prime+\underline{\bm\Sigma}^\xi)-\frac 12 \mathbf x_t^\prime (\underline{\bm\Lambda}\,\underline{\bm\Lambda}^\prime+\underline{\bm\Sigma}^\xi)^{-1}\mathbf x_t,\\
&\ell_{t}(\mathbf x_t|\mathbf F_t;\underline{\bm\varphi})=-\frac 12\log \det(\underline{\bm\Sigma}^\xi)-\frac 12 (\mathbf x_t-\underline{\bm\Lambda}\mathbf F_t)^\prime (\underline{\bm\Sigma}^\xi)^{-1}(\mathbf x_t-\underline{\bm\Lambda}\mathbf F_t),
\end{align}
Let us consider part (a). For any fixed value of the parameters, say $\widetilde{\bm\varphi}$, let
\begin{align}
\bm S(\bm{\mathcal X};\widetilde{\bm\varphi}) &= \sum_{t=1}^T\left.\frac{\partial \ell_t(\mathbf x_t;\underline{\bm\varphi})}{\partial \underline{\bm\Lambda}^\prime}\right\vert_{\underline{\bm\varphi}={\widetilde{\bm\varphi}}} = \left(\begin{array}{c}
\bm s_1^\prime(\bm{\mathcal X};\widetilde{\bm\varphi})\\
\vdots\\
\bm s_n^\prime(\bm{\mathcal X};\widetilde{\bm\varphi})\\
\end{array}
\right),\nonumber\\
\bm S(\bm{\mathcal X}|\bm{\mathcal F};\widetilde{\bm\varphi}) &= \sum_{t=1}^T\left.\frac{\partial \ell_t(\mathbf x_t|\mathbf F_t;\underline{\bm\varphi})}{\partial \underline{\bm\Lambda}^\prime}\right\vert_{\underline{\bm\varphi}={\widetilde{\bm\varphi}}} = \left(\begin{array}{c}
\bm s_1^\prime(\bm{\mathcal X}|\bm{\mathcal F};\widetilde{\bm\varphi})\\
\vdots\\
\bm s_n^\prime(\bm{\mathcal X}|\bm{\mathcal F};\widetilde{\bm\varphi})\\
\end{array}
\right),\nonumber
\end{align}
which are $n\times r$ matrices of first derivatives, and where, for any given $i=1,\ldots, n$,
\begin{equation}
\bm s_i(\bm{\mathcal X};\widetilde{\bm\varphi})=\sum_{t=1}^T\left.\frac{\partial \ell_t(\mathbf x_t;\underline{\bm\varphi})}{\partial \underline{\bm\lambda}_i^\prime}\right\vert_{\underline{\bm\varphi}={\widetilde{\bm\varphi}}}, \quad \bm s_i(\bm{\mathcal X}|\bm{\mathcal F};\widetilde{\bm\varphi})=\sum_{t=1}^T\left.\frac{\partial \ell_t(\mathbf x_t|\mathbf F_t;\underline{\bm\varphi})}{\partial \underline{\bm\lambda}_i^\prime}\right\vert_{\underline{\bm\varphi}={\widetilde{\bm\varphi}}},
\end{equation}
which are $r$-dimensional column vectors.
Then, recalling that $\widehat{\bm\Gamma}^x = \frac 1T\sum_{t=1}^T\mathbf x_t\mathbf x_t^\prime$, by computing the first derivatives of the log-likelihood (ref), we have
\begin{align}
\bm S(\bm{\mathcal X};\underline{\bm\varphi})=&\, -T (\underline{\bm\Lambda}\,\underline{\bm\Lambda}^\prime + \underline{\bm\Sigma}^\xi)^{-1} \underline{\bm\Lambda}+
(\underline{\bm\Lambda}\,\underline{\bm\Lambda}^\prime + \underline{\bm\Sigma}^\xi)^{-1}
\widehat{\bm\Gamma}^x
(\underline{\bm\Lambda}\,\underline{\bm\Lambda}^\prime + \underline{\bm\Sigma}^\xi)^{-1} \underline{\bm\Lambda}\nonumber\\
=&\, -T (\underline{\bm\Sigma}^\xi)^{-1}\underline{\bm\Lambda}\{\mathbf I_r+\underline{\bm\Lambda}^\prime(\underline{\bm\Sigma}^\xi)^{-1} \underline{\bm\Lambda}\}^{-1}
+ T (\underline{\bm\Lambda}\,\underline{\bm\Lambda}^\prime + \underline{\bm\Sigma}^\xi)^{-1} \widehat{\bm\Gamma}^x(\underline{\bm\Sigma}^\xi)^{-1}\underline{\bm\Lambda}\{\mathbf I_r+\underline{\bm\Lambda}^\prime(\underline{\bm\Sigma}^\xi)^{-1} \underline{\bm\Lambda}\}^{-1}\nonumber\\
=&\, -T (\underline{\bm\Sigma}^\xi)^{-1}\underline{\bm\Lambda}\{\mathbf I_r+\underline{\bm\Lambda}^\prime(\underline{\bm\Sigma}^\xi)^{-1} \underline{\bm\Lambda}\}^{-1}
+ T (\underline{\bm\Sigma}^\xi)^{-1} \widehat{\bm\Gamma}^x(\underline{\bm\Sigma}^\xi)^{-1}\underline{\bm\Lambda}\{\mathbf I_r+\underline{\bm\Lambda}^\prime(\underline{\bm\Sigma}^\xi)^{-1} \underline{\bm\Lambda}\}^{-1}\nonumber\\
&-T(\underline{\bm\Sigma}^\xi)^{-1}\underline{\bm\Lambda}
\{\mathbf I_r+\underline{\bm\Lambda}^\prime(\underline{\bm\Sigma}^\xi)^{-1} \underline{\bm\Lambda}\}^{-1}
\underline{\bm\Lambda}^\prime
(\underline{\bm\Sigma}^\xi)^{-1}
\widehat{\bm\Gamma}^x(\underline{\bm\Sigma}^\xi)^{-1}\underline{\bm\Lambda}\{\mathbf I_r+\underline{\bm\Lambda}^\prime(\underline{\bm\Sigma}^\xi)^{-1} \underline{\bm\Lambda}\}^{-1},
\end{align}
where we used the Woodbury identities
\begin{align}
&(\underline{\bm\Lambda}\,\underline{\bm\Lambda}^\prime + \underline{\bm\Sigma}^\xi)^{-1} \underline{\bm\Lambda} = (\underline{\bm\Sigma}^\xi)^{-1}\underline{\bm\Lambda}\{\mathbf I_r+\underline{\bm\Lambda}^\prime(\underline{\bm\Sigma}^\xi)^{-1} \underline{\bm\Lambda}\}^{-1},\\
&(\underline{\bm\Lambda}\,\underline{\bm\Lambda}^\prime + \underline{\bm\Sigma}^\xi)^{-1}= (\underline{\bm\Sigma}^\xi)^{-1}-(\underline{\bm\Sigma}^\xi)^{-1}
\underline{\bm\Lambda}
\{\mathbf I_r+\underline{\bm\Lambda}^\prime(\underline{\bm\Sigma}^\xi)^{-1} \underline{\bm\Lambda}\}^{-1}
\underline{\bm\Lambda}^\prime
(\underline{\bm\Sigma}^\xi)^{-1}.
\end{align}
In what follows, we make use of the following results. Let $\widehat{\bm\Gamma}^\xi=\frac 1T\sum_{t=1}^T \bm\xi_t\bm\xi_t^\prime$ and let $\widehat{\bm\Sigma}^\xi=\text{diag}(\widehat{\bm\Gamma}^\xi)$, the diagonal matrix having as diagonal entries the diagonal entries of $\widehat{\bm\Gamma}^\xi$. Then, by using twice Lemma (ref)(ii) and by Lemma (ref)(v), we have
\begin{align}
\frac 1n\left\Vert \bm\Sigma^\xi -\widehat{\bm\Gamma}^\xi\right\Vert &\le \frac 1n\left\Vert \bm\Sigma^\xi - \widehat{\bm\Sigma}^\xi\right\Vert +\frac 1n\left\Vert\widehat{\bm\Gamma}^\xi-\widehat{\bm\Sigma}^\xi \right\Vert \le O_{\mathrm P}\left(\frac 1{\sqrt T}\right) + \frac 1n\left\Vert\widehat{\bm\Gamma}^\xi \right\Vert\nonumber\\
&\le O_{\mathrm P}\left(\frac 1{\sqrt T}\right) + \frac 1n \left\Vert\bm\Gamma^\xi\right\Vert + O_{\mathrm P}\left(\frac 1{\sqrt T}\right) = O_{\mathrm P}\left(\frac 1{\sqrt T}\right) + O\left(\frac 1n\right).
\end{align}
which implies:
\begin{equation}
\left\Vert (\bm\Sigma^\xi)^{-1}\widehat{\bm\Gamma}^\xi - \mathbf I_n\right\Vert= O_{\mathrm P}\left(\max\left(\frac 1{n},\frac 1{\sqrt T}\right)\right).
\end{equation}
Moreover,
\begin{equation}
\left\Vert\{\mathbf I_r+{\bm\Lambda}^\prime({\bm\Sigma}^\xi)^{-1} {\bm\Lambda}\}^{-1}\left\{{\bm\Lambda}^\prime({\bm\Sigma}^\xi)^{-1}{\bm\Lambda}\right\}-\mathbf I_r\right\Vert= O\left(\frac 1n\right),
\end{equation}
because of Lemma (ref),
\begin{equation}
\left\Vert \left\{{\bm\Lambda}^\prime({\bm\Sigma}^\xi)^{-1}{\bm\Lambda}\right\}^{-1}\right\Vert
\le \frac 1{n\underline C_r \Vert({\bm\Sigma}^\xi)^{-1}\Vert}=\frac 1{n\underline C_r \min_{i=1,\ldots,n} \sigma_i^2}=\frac 1{n\underline C_r C_\xi}= O\left(\frac 1n\right),
\end{equation}
because of Lemma (ref)(iv)-(ref)(vi), Assumption (ref)(a), and MK04,
\begin{align}
\left\Vert\{\mathbf I_r+{\bm\Lambda}^\prime({\bm\Sigma}^\xi)^{-1} {\bm\Lambda}\}^{-1}-\left\{{\bm\Lambda}^\prime({\bm\Sigma}^\xi)^{-1}{\bm\Lambda}\right\}^{-1}\right\Vert= O\left(\frac 1{n^2}\right),\end{align}
because of (ref) and (ref),
\begin{align}
&\left\Vert \left\{{\bm\Lambda}^\prime({\bm\Sigma}^\xi)^{-1}{\bm\Lambda}\right\}^{-1}\bm\Lambda^\prime\right\Vert = O\left(\frac 1{\sqrt n}\right),\\
&\left\Vert\bm\Lambda\{\mathbf I_r+{\bm\Lambda}^\prime({\bm\Sigma}^\xi)^{-1} {\bm\Lambda}\}^{-1}-\bm\Lambda\left\{{\bm\Lambda}^\prime({\bm\Sigma}^\xi)^{-1}{\bm\Lambda}\right\}^{-1}\right\Vert= O\left(\frac 1{n^{3/2}}\right),
\end{align}
because of (ref), (ref) and since $\Vert {\bm\Lambda}\Vert = O(\sqrt n)$ by Lemma (ref)(i), and, last,
\begin{align}
&\left\Vert
{\bm\Lambda}{\bm\Lambda}^\prime
(\bm\Sigma^\xi)^{-1}\bm\Lambda\left\{\mathbf I_r+{\bm\Lambda}^\prime({\bm\Sigma}^\xi)^{-1}{\bm\Lambda}\right\}^{-1}-\bm\Lambda\right\Vert= O\left(\frac 1{\sqrt n}\right),
\end{align}
because of (ref) and since $\Vert {\bm\Lambda}\Vert = O(\sqrt n)$ by Lemma (ref)(i).
Now, denote $\widehat{\bm\Gamma}^{\xi F}=\frac 1T\sum_{t=1}^T \bm\xi_t\mathbf F_t^\prime$ and $\widehat{\bm\Gamma}^{F\xi}=\widehat{\bm\Gamma}^{\xi F\prime}$, so that we can write
\begin{equation}
\widehat{\bm\Gamma}^x
= \bm\Lambda\left(\frac 1T\sum_{t=1}^T\mathbf F_t\mathbf F_t^\prime\right)\bm\Lambda^\prime + \widehat{\bm\Gamma}^\xi+ \bm\Lambda\widehat{\bm\Gamma}^{F\xi}+\widehat{\bm\Gamma}^{\xi F}\bm\Lambda^\prime
= \bm\Lambda\bm\Lambda^\prime + \widehat{\bm\Gamma}^\xi+ \bm\Lambda\widehat{\bm\Gamma}^{F\xi}+\widehat{\bm\Gamma}^{\xi F}\bm\Lambda^\prime,
\end{equation}
because of Assumption (ref)(b). Let us consider (ref) when computed in the true value of the parameters. By means of (ref)-(ref) and (ref) we have
\begin{align}
\bm S(\bm{\mathcal X};{\bm\varphi})=&\, T(\bm\Sigma^\xi)^{-1}
\left\{
\bm\Lambda\bm\Lambda^\prime
+ \widehat{\bm\Gamma}^\xi
+ \bm\Lambda\widehat{\bm\Gamma}^{F\xi}
+\widehat{\bm\Gamma}^{\xi F}\bm\Lambda^\prime
\right\}
(\bm\Sigma^\xi)^{-1}\bm\Lambda\left\{{\bm\Lambda}^\prime({\bm\Sigma}^\xi)^{-1}{\bm\Lambda}\right\}^{-1}\nonumber\\
&- T(\bm\Sigma^\xi)^{-1}
\bm\Lambda\left\{{\bm\Lambda}^\prime({\bm\Sigma}^\xi)^{-1}{\bm\Lambda}\right\}^{-1}\bm\Lambda^\prime(\bm\Sigma^\xi)^{-1}
\left\{
\bm\Lambda\bm\Lambda^\prime + \widehat{\bm\Gamma}^\xi+ \bm\Lambda\widehat{\bm\Gamma}^{F\xi}+\widehat{\bm\Gamma}^{\xi F}\bm\Lambda^\prime
\right\}
(\bm\Sigma^\xi)^{-1}\bm\Lambda\left\{{\bm\Lambda}^\prime({\bm\Sigma}^\xi)^{-1}{\bm\Lambda}\right\}^{-1}\nonumber\\
&-T(\bm\Sigma^\xi)^{-1}\bm\Lambda\left\{{\bm\Lambda}^\prime({\bm\Sigma}^\xi)^{-1}{\bm\Lambda}\right\}^{-1}+ O\left(\frac T{\sqrt n}\right)\nonumber\\
=&\, T(\bm\Sigma^\xi)^{-1} \bm\Lambda\bm\Lambda^\prime (\bm\Sigma^\xi)^{-1}\bm\Lambda\left\{{\bm\Lambda}^\prime({\bm\Sigma}^\xi)^{-1}{\bm\Lambda}\right\}^{-1}+T(\bm\Sigma^\xi)^{-1} \widehat{\bm\Gamma}^\xi(\bm\Sigma^\xi)^{-1}\bm\Lambda\left\{{\bm\Lambda}^\prime({\bm\Sigma}^\xi)^{-1}{\bm\Lambda}\right\}^{-1}\nonumber\\
&+T(\bm\Sigma^\xi)^{-1} \bm\Lambda\widehat{\bm\Gamma}^{F\xi}(\bm\Sigma^\xi)^{-1}\bm\Lambda\left\{{\bm\Lambda}^\prime({\bm\Sigma}^\xi)^{-1}{\bm\Lambda}\right\}^{-1}+T(\bm\Sigma^\xi)^{-1}\widehat{\bm\Gamma}^{\xi F}\bm\Lambda^\prime
(\bm\Sigma^\xi)^{-1}\bm\Lambda\left\{{\bm\Lambda}^\prime({\bm\Sigma}^\xi)^{-1}{\bm\Lambda}\right\}^{-1}\nonumber\\
&-T(\bm\Sigma^\xi)^{-1}
\bm\Lambda\left\{{\bm\Lambda}^\prime({\bm\Sigma}^\xi)^{-1}{\bm\Lambda}\right\}^{-1}\bm\Lambda^\prime(\bm\Sigma^\xi)^{-1}\bm\Lambda\bm\Lambda^\prime(\bm\Sigma^\xi)^{-1}\bm\Lambda\left\{{\bm\Lambda}^\prime({\bm\Sigma}^\xi)^{-1}{\bm\Lambda}\right\}^{-1}\nonumber\\
&-T(\bm\Sigma^\xi)^{-1}
\bm\Lambda\left\{{\bm\Lambda}^\prime({\bm\Sigma}^\xi)^{-1}{\bm\Lambda}\right\}^{-1}\bm\Lambda^\prime(\bm\Sigma^\xi)^{-1}\widehat{\bm\Gamma}^\xi(\bm\Sigma^\xi)^{-1}\bm\Lambda\left\{{\bm\Lambda}^\prime({\bm\Sigma}^\xi)^{-1}{\bm\Lambda}\right\}^{-1}\nonumber\\
&-T(\bm\Sigma^\xi)^{-1}
\bm\Lambda\left\{{\bm\Lambda}^\prime({\bm\Sigma}^\xi)^{-1}{\bm\Lambda}\right\}^{-1}\bm\Lambda^\prime(\bm\Sigma^\xi)^{-1}\bm\Lambda\widehat{\bm\Gamma}^{F\xi}(\bm\Sigma^\xi)^{-1}\bm\Lambda\left\{{\bm\Lambda}^\prime({\bm\Sigma}^\xi)^{-1}{\bm\Lambda}\right\}^{-1}\nonumber\\
&-T(\bm\Sigma^\xi)^{-1}
\bm\Lambda\left\{{\bm\Lambda}^\prime({\bm\Sigma}^\xi)^{-1}{\bm\Lambda}\right\}^{-1}\bm\Lambda^\prime(\bm\Sigma^\xi)^{-1}\widehat{\bm\Gamma}^{\xi F}\bm\Lambda^\prime(\bm\Sigma^\xi)^{-1}\bm\Lambda\left\{{\bm\Lambda}^\prime({\bm\Sigma}^\xi)^{-1}{\bm\Lambda}\right\}^{-1}\nonumber\\
&-T(\bm\Sigma^\xi)^{-1}\bm\Lambda\left\{{\bm\Lambda}^\prime({\bm\Sigma}^\xi)^{-1}{\bm\Lambda}\right\}^{-1}+ O\left(\frac T{\sqrt n}\right)\nonumber\\
=&\,T(\bm\Sigma^\xi)^{-1} \bm\Lambda+T(\bm\Sigma^\xi)^{-1} \bm\Lambda\left\{{\bm\Lambda}^\prime({\bm\Sigma}^\xi)^{-1}{\bm\Lambda}\right\}^{-1}+T(\bm\Sigma^\xi)^{-1} \bm\Lambda\widehat{\bm\Gamma}^{F\xi}(\bm\Sigma^\xi)^{-1}\bm\Lambda\left\{{\bm\Lambda}^\prime({\bm\Sigma}^\xi)^{-1}{\bm\Lambda}\right\}^{-1}\nonumber\\
&+T(\bm\Sigma^\xi)^{-1}\widehat{\bm\Gamma}^{\xi F}-T(\bm\Sigma^\xi)^{-1}\bm\Lambda-T(\bm\Sigma^\xi)^{-1}\bm\Lambda\left\{{\bm\Lambda}^\prime({\bm\Sigma}^\xi)^{-1}{\bm\Lambda}\right\}^{-1}-T(\bm\Sigma^\xi)^{-1}\widehat{\bm\Gamma}^{F\xi}(\bm\Sigma^\xi)^{-1}\bm\Lambda\left\{{\bm\Lambda}^\prime({\bm\Sigma}^\xi)^{-1}{\bm\Lambda}\right\}^{-1}\nonumber\\
&-T(\bm\Sigma^\xi)^{-1}
\bm\Lambda\left\{{\bm\Lambda}^\prime({\bm\Sigma}^\xi)^{-1}{\bm\Lambda}\right\}^{-1}\bm\Lambda^\prime(\bm\Sigma^\xi)^{-1}\widehat{\bm\Gamma}^{\xi F}-T(\bm\Sigma^\xi)^{-1}\bm\Lambda\left\{{\bm\Lambda}^\prime({\bm\Sigma}^\xi)^{-1}{\bm\Lambda}\right\}^{-1}\nonumber\\
&+ O\left(\frac T{\sqrt n}\right)+ O_{\mathrm P}\left(\frac {\sqrt T}{\sqrt n}\right)+ O\left(\frac T{n^{3/2}}\right)\nonumber\\
=&\,T(\bm\Sigma^\xi)^{-1}\widehat{\bm\Gamma}^{\xi F}-T(\bm\Sigma^\xi)^{-1}
\bm\Lambda\left\{{\bm\Lambda}^\prime({\bm\Sigma}^\xi)^{-1}{\bm\Lambda}\right\}^{-1}\bm\Lambda^\prime(\bm\Sigma^\xi)^{-1}\widehat{\bm\Gamma}^{\xi F}\nonumber\\
&-T(\bm\Sigma^\xi)^{-1}\bm\Lambda\left\{{\bm\Lambda}^\prime({\bm\Sigma}^\xi)^{-1}{\bm\Lambda}\right\}^{-1}+ O\left(\frac T{\sqrt n}\right)+ O_{\mathrm P}\left(\frac {\sqrt T}{\sqrt n}\right)+ O\left(\frac T{n^{3/2}}\right).
\end{align}
Then, notice that $[(\bm\Sigma^\xi)^{-1}]_{i\cdot}\bm\xi_t= [(\bm\Sigma^\xi)^{-1}]_{ii} \xi_{it}= \frac{ \xi_{it}}{\sigma_i^{2}}$, and
$[(\bm\Sigma^\xi)^{-1}]_{i\cdot}\bm\Lambda = \frac{\bm\lambda_i^\prime}{\sigma_i^{2} }$, $i=1,\ldots, n$.
The following holds
\begin{align}
& \frac 1{\sigma_i^2}{\bm\lambda}_i^\prime \left\{{\bm\Lambda}^\prime({\bm\Sigma}^\xi)^{-1}{\bm\Lambda}\right\}^{-1}= O\left(\frac 1n\right),\\
&T\left[(\bm\Sigma^\xi)^{-1}\widehat{\bm\Gamma}^\xi(\bm\Sigma^\xi)^{-1}\bm\Lambda\left\{{\bm\Lambda}^\prime({\bm\Sigma}^\xi)^{-1}{\bm\Lambda}\right\}^{-1}
\right]_{i\cdot} = \frac T{\sigma_i^2}{\bm\lambda}_i^\prime \left\{{\bm\Lambda}^\prime({\bm\Sigma}^\xi)^{-1}{\bm\Lambda}\right\}^{-1}+ O_{\mathrm P}\left(\frac{\sqrt T}{n} \right)+ O\left(\frac{ T}{n^2} \right),\\
&T\left[
(\bm\Sigma^\xi)^{-1}\bm\Lambda \left\{{\bm\Lambda}^\prime({\bm\Sigma}^\xi)^{-1}{\bm\Lambda}\right\}^{-1}\bm\Lambda^\prime(\bm\Sigma^\xi)^{-1}\widehat{\bm\Gamma}^\xi(\bm\Sigma^\xi)^{-1}\bm\Lambda\left\{{\bm\Lambda}^\prime({\bm\Sigma}^\xi)^{-1}{\bm\Lambda}\right\}^{-1}
\right]_{i\cdot}=\frac T{\sigma_i^2}{\bm\lambda}_i^\prime \left\{{\bm\Lambda}^\prime({\bm\Sigma}^\xi)^{-1}{\bm\Lambda}\right\}^{-1}\nonumber\\
&+ O_{\mathrm P}\left(\frac{\sqrt T}{n} \right)+ O\left(\frac{ T}{n^2} \right),\\
&T\bm\lambda_i^\prime\{\mathbf I_r+{\bm\Lambda}^\prime({\bm\Sigma}^\xi)^{-1} {\bm\Lambda}\}^{-1}=T\bm\lambda_i^\prime\left\{{\bm\Lambda}^\prime({\bm\Sigma}^\xi)^{-1}{\bm\Lambda}\right\}^{-1}+ O\left(\frac T{n^{2}}\right),\\
&T{\bm\lambda}_i^\prime{\bm\Lambda}^\prime
(\bm\Sigma^\xi)^{-1}\bm\Lambda\left\{\mathbf I_r+{\bm\Lambda}^\prime({\bm\Sigma}^\xi)^{-1}{\bm\Lambda}\right\}^{-1}=T\bm\lambda_i^\prime + O\left(\frac T{n}\right),
\end{align}
where we used Assumption (ref)(a) and then (ref) follows from (ref), (ref) and (ref) follow from (ref) and (ref), (ref) follows from (ref), and (ref) follows from and (ref).
Therefore, by using (ref)-(ref) in (ref) we have that the $i$th row of $\bm S(\bm{\mathcal X};\bm{\varphi})$ is such that
\begin{align}
\bm s_i^\prime(\bm{\mathcal X};\bm{\varphi})=&\,\frac 1{\sigma_i^2}\sum_{t=1}^T {\xi}_{it}\mathbf F_t^\prime-\frac 1{\sigma_i^2}\bm\lambda_i^\prime
\left\{{\bm\Lambda}^\prime({\bm\Sigma}^\xi)^{-1}{\bm\Lambda}\right\}^{-1}{\bm\Lambda}^\prime(\bm\Sigma^\xi)^{-1}\sum_{t=1}^T {\bm\xi}_{t}\mathbf F_t^\prime\nonumber\\
&-\frac T{\sigma_i^2}{\bm\lambda}_i^\prime\left\{{\bm\Lambda}^\prime({\bm\Sigma}^\xi)^{-1}{\bm\Lambda}\right\}^{-1}+ O\left(\frac T{ n}\right)+ O_{\mathrm P}\left(\frac {\sqrt T}{n}\right)+ O\left(\frac T{ n^2}\right).
\end{align}
Hence, $\Vert\bm s_i^\prime(\bm{\mathcal X};\bm{\varphi})\Vert=O_{\mathrm P}(\sqrt T)$, because of Assumption (ref)(a), (ref) and since
\begin{align}
\mathbb{E}\left[\left\Vert \sum_{t=1}^T \frac{\mathbf F_t\bm\xi_t^\prime \bm\Lambda}{ nT}\right\Vert^2\right] = O\left(\frac 1{nT}\right) \;\text{ and }\; \mathbb{E}\left[\left\Vert \sum_{t=1}^T \frac{\mathbf F_t\bm\xi_t^\prime }{\sqrt nT}\right\Vert^2\right] = O\left(\frac 1{\sqrt T}\right)
,
\end{align}
because of (ref) in the proof of Proposition (ref)(a) and by Lemma (ref).
Moreover, from (ref) and by noticing that $(\bm\Sigma^\xi)^{-1}{\bm\Lambda}$ has the same properties as ${\bm\Lambda}$ since $\Vert (\bm\Sigma^\xi)^{-1}\Vert = O(1)$ by Assumption (ref)(a), we get
\begin{align}
\frac 1{\sqrt T}\left\Vert
\bm s_i^\prime(\bm{\mathcal X};\bm{\varphi})-\frac 1{\sigma_i^2}\sum_{t=1}^T {\xi}_{it}\mathbf F_t^\prime
\right\Vert\le &\,
\frac 1{C_\xi}\left\Vert
{\bm\lambda}_i^\prime
\left\{{\bm\Lambda}^\prime({\bm\Sigma}^\xi)^{-1}{\bm\Lambda}\right\}^{-1}
\right\Vert\,
\left\Vert\frac 1{\sqrt T} {\bm\Lambda}^\prime(\bm\Sigma^\xi)^{-1}\sum_{t=1}^T {\bm\xi}_{t}\mathbf F_t^\prime
\right\Vert\nonumber\\
&+
\frac{\sqrt T}{C_\xi} \left\Vert
{\bm\lambda}_i^\prime\left\{{\bm\Lambda}^\prime({\bm\Sigma}^\xi)^{-1}{\bm\Lambda}\right\}^{-1}
\right\Vert+ O\left(\frac {\sqrt T}{n}\right)+ O_{\mathrm P}\left(\frac {1}{n}\right)+ O\left(\frac {\sqrt T}{n^2}\right)\nonumber\\
=&\, O\left(\frac 1n\right) O_{\mathrm P}\left(\sqrt {n}\right)+ O\left(\frac {\sqrt T}{n}\right)+O_{\mathrm P}\left(\frac 1{n}\right)+ O\left(\frac {\sqrt T}{n^2}\right)\nonumber\\
=&\, O_{\mathrm P}\left(\max\left( \frac 1{\sqrt n},\frac{\sqrt T} {n}\right)\right),
\end{align}
.
Finally, by computing the first derivatives of the log-likelihood (ref), we have
\begin{align}
\bm S(\bm{\mathcal X}|\bm{\mathcal F};\underline{\bm\varphi})&= (\underline{\bm\Sigma}^\xi)^{-1}\sum_{t=1}^T (\mathbf x_t-\underline{\bm\Lambda}\mathbf F_t)\mathbf F_t^\prime.
\end{align}
And (ref) when computed in the true value of the parameters is
\begin{align}
\bm S(\bm{\mathcal X}|\bm{\mathcal F};{\bm\varphi})&= ({\bm\Sigma}^\xi)^{-1}\sum_{t=1}^T \bm\xi_t\mathbf F_t^\prime.
\end{align}
From (ref), the $i$th row of $\bm S(\bm{\mathcal X}|\bm{\mathcal F};{\bm\varphi})$ is then:
\begin{equation}
\bm s_i^\prime(\bm{\mathcal X}|\bm{\mathcal F};{\bm\varphi})=\frac 1{\sigma_i^2}\sum_{t=1}^T {\xi}_{it}\mathbf F_t^\prime.
\end{equation}
Hence, $\Vert \bm s_i^\prime(\bm{\mathcal X}|\bm{\mathcal F};{\bm\varphi})\Vert = O_{\mathrm P}(\sqrt T)$ because of (ref), and, by using (ref) in (ref), for any given $i=1,\ldots, n$ we have:
\[
\frac 1{\sqrt T}\left\Vert
\bm s_i^\prime(\bm{\mathcal X};\bm{\varphi})-\bm s_i^\prime(\bm{\mathcal X}|\bm{\mathcal F};{\bm\varphi})
\right\Vert = O_{\mathrm P}\left(\max\left( \frac 1{\sqrt n},\frac{\sqrt T} {n}\right)\right),
\]
which proves part (a).
Turning to part (b). For any specific value of the parameters, say $\widetilde{\bm\varphi}$, let
\begin{align}
\bm H(\bm{\mathcal X};\widetilde{\bm\varphi}) &= \sum_{t=1}^T\left.\frac{\partial^2 \ell_t(\mathbf x_t;\underline{\bm\varphi})}{\partial \text{vec}(\underline{\bm\Lambda})^\prime\partial \text{vec}(\underline{\bm\Lambda})}\right\vert_{\underline{\bm\varphi}={\widetilde{\bm\varphi}}} =
\left.
\frac{\partial \text{vec}(\bm S(\bm{\mathcal X};\underline{\bm\varphi}))^\prime}{\partial \text{vec}(\underline{\bm\Lambda})^\prime}
\right\vert_{\underline{\bm\varphi}={\widetilde{\bm\varphi}}}
= \left(\begin{array}{cccc}
\bm h_{11}(\bm{\mathcal X};\widetilde{\bm\varphi})&\ldots& \bm h_{1n}(\bm{\mathcal X};\widetilde{\bm\varphi})\\
\vdots&\ddots&\vdots\\
\bm h_{n1}(\bm{\mathcal X};\widetilde{\bm\varphi})&\ldots& \bm h_{nn}(\bm{\mathcal X};\widetilde{\bm\varphi})\\
\end{array}
\right),\nonumber\\
\bm H(\bm{\mathcal X}|\bm{\mathcal F};\widetilde{\bm\varphi}) &= \sum_{t=1}^T\left.\frac{\partial^2 \ell_t(\mathbf x_t|\mathbf F_t;\underline{\bm\varphi})}{\partial \text{vec}(\underline{\bm\Lambda})^\prime\partial \text{vec}(\underline{\bm\Lambda})}\right\vert_{\underline{\bm\varphi}={\widetilde{\bm\varphi}}} =
\left.
\frac{\partial \text{vec}(\bm S(\bm{\mathcal X}|\bm{\mathcal F};\underline{\bm\varphi}))^\prime}{\partial \text{vec}(\underline{\bm\Lambda})^\prime}
\right\vert_{\underline{\bm\varphi}={\widetilde{\bm\varphi}}}
= \left(\begin{array}{cccc}
\bm h_{11}(\bm{\mathcal X}|\bm{\mathcal F};\widetilde{\bm\varphi})&\ldots& \bm h_{1n}(\bm{\mathcal X}|\bm{\mathcal F};\widetilde{\bm\varphi})\\
\vdots&\ddots&\vdots\\
\bm h_{n1}(\bm{\mathcal X}|\bm{\mathcal F};\widetilde{\bm\varphi})&\ldots& \bm h_{nn}(\bm{\mathcal X}|\bm{\mathcal F};\widetilde{\bm\varphi})\\
\end{array}
\right),\nonumber
\end{align}
which are $nr\times nr$ matrices obtained from the matricization of the 4th order tensor of second derivatives, and where, for any given $ i=1,\ldots, n$,
\begin{align}
\bm h_{ii}(\bm{\mathcal X};\widetilde{\bm\varphi})&= \sum_{t=1}^T\left.\frac{\partial^2 \ell_t(\mathbf x_t;\underline{\bm\varphi})}{\partial \underline{\bm\lambda}_i^\prime\partial\underline{\bm\lambda}_i}\right\vert_{\underline{\bm\varphi}={\widetilde{\bm\varphi}}} =
\left.
\frac{\partial \bm s_i^\prime(\bm{\mathcal X};\underline{\bm\varphi})}{\partial \underline{\bm\lambda}_i^\prime}
\right\vert_{\underline{\bm\varphi}={\widetilde{\bm\varphi}}},\\
\bm h_{ii}(\bm{\mathcal X}|\bm{\mathcal F};\widetilde{\bm\varphi})&= \sum_{t=1}^T\left.\frac{\partial^2 \ell_t(\mathbf x_t|\mathbf F_t;\underline{\bm\varphi})}{\partial \underline{\bm\lambda}_i^\prime\partial \underline{\bm\lambda}_i}\right\vert_{\underline{\bm\varphi}={\widetilde{\bm\varphi}}} =
\left.
\frac{\partial \bm s_i^\prime(\bm{\mathcal X}|\bm{\mathcal F};\underline{\bm\varphi})}{\partial \underline{\bm\lambda}_i^\prime}
\right\vert_{\underline{\bm\varphi}={\widetilde{\bm\varphi}}},\nonumber
\end{align}
which are $r\times r$ matrices.
We now use the following relation:
\begin{align}
\text{vec}\left(\mathrm d \bm S^\prime(\bm{\mathcal X};\underline{\bm\varphi})\right) &=\left(
\frac{\partial \text{vec}(\bm S(\bm{\mathcal X};\underline{\bm\varphi}))^\prime}{\partial \text{vec}(\underline{\bm\Lambda})^\prime}\right)^\prime \text{vec}(\mathrm d \underline{\bm \Lambda}^\prime)=
\bm H(\bm{\mathcal X};\underline{\bm\varphi})
\text{vec}(\mathrm d \underline{\bm \Lambda}^\prime),
\end{align}
Then, by denoting $\underline{\bm P}=\left\{\mathbf I_r+\underline{\bm\Lambda}^\prime(\underline{\bm\Sigma}^\xi)^{-1}\underline{\bm\Lambda}\right\}$, from (ref) we have
\begin{align}
\text{vec}\left(\mathrm d \bm S^\prime(\bm{\mathcal X};\underline{\bm\varphi})\right)=&\,
-T
\left\{
(\underline{\bm\Sigma}^\xi)^{-1}\underline{\bm \Lambda}\otimes \mathbf I_r
\right\}
\text{vec}\left(\mathrm d \underline{\bm P}^{-1}\right)\nonumber\\
&-T
\left\{
(\underline{\bm\Sigma}^\xi)^{-1}\otimes \underline{\bm P}^{-1}
\right\}
\text{vec}(\mathrm d \underline{\bm \Lambda}^\prime)\nonumber\\
&+T
\left\{
(\underline{\bm\Sigma}^\xi)^{-1}\widehat{\bm\Gamma}^x (\underline{\bm\Sigma}^\xi)^{-1}\underline{\bm \Lambda}\otimes \mathbf I_r
\right\}
\text{vec}\left(\mathrm d \underline{\bm P}^{-1}\right)\nonumber\\
&+T
\left\{
(\underline{\bm\Sigma}^\xi)^{-1}\widehat{\bm\Gamma}^x (\underline{\bm\Sigma}^\xi)^{-1}\otimes \underline{\bm P}^{-1}
\right\}
\text{vec}(\mathrm d \underline{\bm \Lambda}^\prime)\nonumber\\
&-T
\left\{
(\underline{\bm\Sigma}^\xi)^{-1}\underline{\bm \Lambda}\,\underline{\bm P}^{-1}\underline{\bm \Lambda}^\prime(\underline{\bm\Sigma}^\xi)^{-1}\widehat{\bm\Gamma}^x (\underline{\bm\Sigma}^\xi)^{-1}\underline{\bm \Lambda}
\otimes
\mathbf I_r
\right\}
\text{vec}\left(\mathrm d \underline{\bm P}^{-1}\right)\nonumber\\
&-T
\left\{
(\underline{\bm\Sigma}^\xi)^{-1}\underline{\bm \Lambda}\,\underline{\bm P}^{-1}\underline{\bm \Lambda}^\prime(\underline{\bm\Sigma}^\xi)^{-1}\widehat{\bm\Gamma}^x (\underline{\bm\Sigma}^\xi)^{-1}
\otimes
\underline{\bm P}^{-1}
\right\}
\text{vec}(\mathrm d \underline{\bm \Lambda}^\prime)\nonumber\\
&-T
\left\{
(\underline{\bm\Sigma}^\xi)^{-1}\underline{\bm \Lambda}\,\underline{\bm P}^{-1}
\otimes
\underline{\bm P}^{-1}\underline{\bm \Lambda}^\prime
(\underline{\bm\Sigma}^\xi)^{-1}\widehat{\bm\Gamma}^x (\underline{\bm\Sigma}^\xi)^{-1}
\right\}
\bm C_{n,r}
\text{vec}(\mathrm d \underline{\bm \Lambda}^\prime)\nonumber\\
&-T
\left\{
(\underline{\bm\Sigma}^\xi)^{-1}\underline{\bm \Lambda}
\otimes
\underline{\bm P}^{-1}\underline{\bm \Lambda}^\prime
(\underline{\bm\Sigma}^\xi)^{-1}\widehat{\bm\Gamma}^x (\underline{\bm\Sigma}^\xi)^{-1}\underline{\bm \Lambda}
\right\}
\text{vec}\left(\mathrm d \underline{\bm P}^{-1}\right)\nonumber\\
&-T
\left\{(\underline{\bm\Sigma}^\xi)^{-1}
\otimes
\underline{\bm P}^{-1}\underline{\bm \Lambda}^\prime
(\underline{\bm\Sigma}^\xi)^{-1}\widehat{\bm\Gamma}^x (\underline{\bm\Sigma}^\xi)^{-1}\underline{\bm \Lambda}\underline{\bm P}^{-1}
\right\}
\text{vec}(\mathrm d \underline{\bm \Lambda}^\prime).
\end{align}
Moreover,
\begin{align}
\text{vec}\left(\mathrm d \underline{\bm P}^{-1}\right) &= -\left( \underline{\bm P}^{-1}\otimes \underline{\bm P}^{-1} \right)
\left\{
\left[
\underline{\bm \Lambda}^\prime(\underline{\bm\Sigma}^\xi)^{-1}\otimes \mathbf I_r
\right]
+
\left[
\mathbf I_r \otimes \underline{\bm \Lambda}^\prime(\underline{\bm\Sigma}^\xi)^{-1}
\right]\bm C_{n,r}
\right\}
\text{vec}(\mathrm d \underline{\bm \Lambda}^\prime)\nonumber\\
&= - \left\{
\left[
\underline{\bm P}^{-1}\underline{\bm \Lambda}^\prime(\underline{\bm\Sigma}^\xi)^{-1}\otimes \underline{\bm P}^{-1}
\right]
+\left[
\underline{\bm P}^{-1}\otimes\underline{\bm P}^{-1}\underline{\bm \Lambda}^\prime(\underline{\bm\Sigma}^\xi)^{-1}
\right]\bm C_{n,r}
\right\}
\text{vec}(\mathrm d \underline{\bm \Lambda}^\prime),
\end{align}
where $\bm C_{n,r}$ is the $nr\times nr$ commutation matrix such that $\text{vec}(\mathrm d \underline{\bm \Lambda})=\bm C_{n,r}\text{vec}(\mathrm d \underline{\bm \Lambda}^\prime)$.
Therefore, from (ref), (ref), and (ref) we get
\begin{align}
\bm H(\bm{\mathcal X};\underline{\bm\varphi}) =&\,
T\left[
(\underline{\bm\Sigma}^\xi)^{-1}\underline{\bm \Lambda}\,\underline{\bm P}^{-1} \underline{\bm \Lambda}^\prime (\underline{\bm\Sigma}^\xi)^{-1}
\otimes
\underline{\bm P}^{-1}
\right]\quad \text{\it A.1}\nonumber\\
&-T\left[
(\underline{\bm\Sigma}^\xi)^{-1}\widehat{\bm\Gamma}^x (\underline{\bm\Sigma}^\xi)^{-1}\underline{\bm \Lambda}\,\underline{\bm P}^{-1} \underline{\bm \Lambda}^\prime (\underline{\bm\Sigma}^\xi)^{-1}
\otimes
\underline{\bm P}^{-1}
\right]\quad \text{\it A.2}\nonumber\\
&+T\left[
(\underline{\bm\Sigma}^\xi)^{-1}\underline{\bm \Lambda}\,\underline{\bm P}^{-1} \underline{\bm \Lambda}^\prime (\underline{\bm\Sigma}^\xi)^{-1}\widehat{\bm\Gamma}^x (\underline{\bm\Sigma}^\xi)^{-1}\underline{\bm \Lambda}\,\underline{\bm P}^{-1} \underline{\bm \Lambda}^\prime (\underline{\bm\Sigma}^\xi)^{-1}
\otimes
\underline{\bm P}^{-1}
\right]\quad \text{\it A.3}\nonumber\\
&+T\left[
(\underline{\bm\Sigma}^\xi)^{-1}\underline{\bm \Lambda}\,\underline{\bm P}^{-1} \underline{\bm \Lambda}^\prime (\underline{\bm\Sigma}^\xi)^{-1}
\otimes
\underline{\bm P}^{-1}\underline{\bm \Lambda}^\prime(\underline{\bm\Sigma}^\xi)^{-1}\widehat{\bm\Gamma}^x (\underline{\bm\Sigma}^\xi)^{-1}\underline{\bm \Lambda}\,\underline{\bm P}^{-1}
\right]\quad \text{\it A.4}\nonumber\\
&+T\left[
(\underline{\bm\Sigma}^\xi)^{-1}\underline{\bm \Lambda}\,\underline{\bm P}^{-1}
\otimes
\underline{\bm P}^{-1}\underline{\bm \Lambda}^\prime(\underline{\bm\Sigma}^\xi)^{-1}
\right]
\bm C_{n,r}\quad \text{\it A.5}\nonumber\\
&-T\left[
(\underline{\bm\Sigma}^\xi)^{-1}\widehat{\bm\Gamma}^x (\underline{\bm\Sigma}^\xi)^{-1}\underline{\bm \Lambda}\,\underline{\bm P}^{-1}
\otimes
\underline{\bm P}^{-1}\underline{\bm \Lambda}^\prime(\underline{\bm\Sigma}^\xi)^{-1}
\right]
\bm C_{n,r}\quad \text{\it A.6}\nonumber\\
&+T\left[
(\underline{\bm\Sigma}^\xi)^{-1}\underline{\bm \Lambda}\,\underline{\bm P}^{-1}\underline{\bm \Lambda}^\prime(\underline{\bm\Sigma}^\xi)^{-1}\widehat{\bm\Gamma}^x (\underline{\bm\Sigma}^\xi)^{-1}\underline{\bm \Lambda}\,\underline{\bm P}^{-1}
\otimes
\underline{\bm P}^{-1}\underline{\bm \Lambda}^\prime(\underline{\bm\Sigma}^\xi)^{-1}
\right]
\bm C_{n,r}\quad \text{\it A.7}\nonumber\\
&+T\left[
(\underline{\bm\Sigma}^\xi)^{-1}\underline{\bm \Lambda}\,\underline{\bm P}^{-1}
\otimes
\underline{\bm P}^{-1}\underline{\bm \Lambda}^\prime(\underline{\bm\Sigma}^\xi)^{-1}\widehat{\bm\Gamma}^x (\underline{\bm\Sigma}^\xi)^{-1} \underline{\bm \Lambda}\,\underline{\bm P}^{-1} \underline{\bm \Lambda}^\prime(\underline{\bm\Sigma}^\xi)^{-1}
\right]
\bm C_{n,r}\quad \text{\it A.8}\nonumber\\
&-T\left[
(\underline{\bm\Sigma}^\xi)^{-1}
\otimes
\underline{\bm P}^{-1}
\right]\quad \text{\it B.1}\nonumber\\
&+T\left[
(\underline{\bm\Sigma}^\xi)^{-1}\widehat{\bm\Gamma}^x (\underline{\bm\Sigma}^\xi)^{-1}
\otimes
\underline{\bm P}^{-1}
\right]\quad \text{\it B.2}\nonumber\\
&-T\left[
(\underline{\bm\Sigma}^\xi)^{-1} \underline{\bm \Lambda}\,\underline{\bm P}^{-1} \underline{\bm \Lambda}^\prime(\underline{\bm\Sigma}^\xi)^{-1}\widehat{\bm\Gamma}^x (\underline{\bm\Sigma}^\xi)^{-1}
\otimes
\underline{\bm P}^{-1}
\right]\quad \text{\it B.3}\nonumber\\
&-T\left[
(\underline{\bm\Sigma}^\xi)^{-1} \underline{\bm \Lambda}\,\underline{\bm P}^{-1}
\otimes
\underline{\bm P}^{-1} \underline{\bm \Lambda}^\prime(\underline{\bm\Sigma}^\xi)^{-1}\widehat{\bm\Gamma}^x (\underline{\bm\Sigma}^\xi)^{-1}
\right]\bm C_{n,r}\quad \text{\it B.4}\nonumber\\
&-T\left[
(\underline{\bm\Sigma}^\xi)^{-1}
\otimes
\underline{\bm P}^{-1} \underline{\bm \Lambda}^\prime(\underline{\bm\Sigma}^\xi)^{-1}\widehat{\bm\Gamma}^x (\underline{\bm\Sigma}^\xi)^{-1}\underline{\bm \Lambda}\,\underline{\bm P}^{-1}
\right].\quad \text{\it B.5}
\end{align}
Let us consider (ref) when computed in the true value of the parameters. By means of (ref) we have:
{
\begin{align}
\bm H(\bm{\mathcal X};\bm\varphi)=&\,
T\left[(\bm\Sigma^\xi)^{-1}\bm\Lambda \bm P^{-1}\bm\Lambda^\prime(\bm\Sigma^\xi)^{-1}
\otimes
\bm P^{-1}
\right]\quad \text{\it A.1}
\nonumber\\
&-T\left[
(\bm\Sigma^\xi)^{-1}\bm\Lambda\bm\Lambda^\prime(\bm\Sigma^\xi)^{-1}\bm\Lambda\bm P^{-1}\bm\Lambda^\prime(\bm\Sigma^\xi)^{-1}
\otimes
\bm P^{-1}
\right]\quad \text{\it A.2.1}
\nonumber\\
&-T\left[
(\bm\Sigma^\xi)^{-1}\widehat{\bm\Gamma}^\xi(\bm\Sigma^\xi)^{-1}\bm\Lambda\bm P^{-1}\bm\Lambda^\prime(\bm\Sigma^\xi)^{-1}
\otimes
\bm P^{-1}
\right]\quad \text{\it A.2.2}
\nonumber\\
&-T\left[
(\bm\Sigma^\xi)^{-1}\bm\Lambda\widehat{\bm\Gamma}^{F\xi}(\bm\Sigma^\xi)^{-1}\bm\Lambda\bm P^{-1}\bm\Lambda^\prime(\bm\Sigma^\xi)^{-1}
\otimes
\bm P^{-1}
\right]\quad \text{\it A.2.3}
\nonumber\\
&-T\left[
(\bm\Sigma^\xi)^{-1}\widehat{\bm\Gamma}^{\xi F}\bm\Lambda^\prime(\bm\Sigma^\xi)^{-1}\bm\Lambda\bm P^{-1}\bm\Lambda^\prime(\bm\Sigma^\xi)^{-1}
\otimes
\bm P^{-1}
\right]\quad \text{\it A.2.4}
\nonumber\\
&+T\left[
(\bm\Sigma^\xi)^{-1} \bm\Lambda\bm P^{-1}\bm\Lambda^\prime(\bm\Sigma^\xi)^{-1} \bm\Lambda\bm\Lambda^\prime(\bm\Sigma^\xi)^{-1}\bm\Lambda\bm P^{-1}\bm\Lambda^\prime(\bm\Sigma^\xi)^{-1}
\otimes
\bm P^{-1}
\right]\quad \text{\it A.3.1}
\nonumber\\
&+T\left[
(\bm\Sigma^\xi)^{-1} \bm\Lambda\bm P^{-1}\bm\Lambda^\prime(\bm\Sigma^\xi)^{-1} \widehat{\bm\Gamma}^\xi(\bm\Sigma^\xi)^{-1}\bm\Lambda\bm P^{-1}\bm\Lambda^\prime(\bm\Sigma^\xi)^{-1}
\otimes
\bm P^{-1}
\right]\quad \text{\it A.3.2}
\nonumber\\
&+T\left[
(\bm\Sigma^\xi)^{-1} \bm\Lambda\bm P^{-1}\bm\Lambda^\prime(\bm\Sigma^\xi)^{-1} \bm\Lambda\widehat{\bm\Gamma}^{F\xi}(\bm\Sigma^\xi)^{-1}\bm\Lambda\bm P^{-1}\bm\Lambda^\prime(\bm\Sigma^\xi)^{-1}
\otimes
\bm P^{-1}
\right]\quad \text{\it A.3.3}
\nonumber\\
&+T\left[
(\bm\Sigma^\xi)^{-1} \bm\Lambda\bm P^{-1}\bm\Lambda^\prime(\bm\Sigma^\xi)^{-1} \widehat{\bm\Gamma}^{\xi F}\bm\Lambda^\prime(\bm\Sigma^\xi)^{-1}\bm\Lambda\bm P^{-1}\bm\Lambda^\prime(\bm\Sigma^\xi)^{-1}
\otimes
\bm P^{-1}
\right]\quad \text{\it A.3.4}
\nonumber\\
&+T\left[
(\bm\Sigma^\xi)^{-1} \bm\Lambda\bm P^{-1}\bm\Lambda^\prime(\bm\Sigma^\xi)^{-1}
\otimes
\bm P^{-1} \bm\Lambda^\prime(\bm\Sigma^\xi)^{-1}\bm\Lambda\bm\Lambda^\prime(\bm\Sigma^\xi)^{-1}\bm\Lambda\bm P^{-1}
\right]\quad \text{\it A.4.1}
\nonumber\\
&+T\left[
(\bm\Sigma^\xi)^{-1} \bm\Lambda\bm P^{-1}\bm\Lambda^\prime(\bm\Sigma^\xi)^{-1}
\otimes
\bm P^{-1} \bm\Lambda^\prime(\bm\Sigma^\xi)^{-1}\widehat{\bm\Gamma}^\xi(\bm\Sigma^\xi)^{-1}\bm\Lambda\bm P^{-1}
\right]\quad \text{\it A.4.2}
\nonumber\\
&+T\left[
(\bm\Sigma^\xi)^{-1} \bm\Lambda\bm P^{-1}\bm\Lambda^\prime(\bm\Sigma^\xi)^{-1}
\otimes
\bm P^{-1} \bm\Lambda^\prime(\bm\Sigma^\xi)^{-1}\bm\Lambda\widehat{\bm\Gamma}^{F\xi}(\bm\Sigma^\xi)^{-1}\bm\Lambda\bm P^{-1}
\right]\quad \text{\it A.4.3}
\nonumber\\
&+T\left[
(\bm\Sigma^\xi)^{-1} \bm\Lambda\bm P^{-1}\bm\Lambda^\prime(\bm\Sigma^\xi)^{-1}
\otimes
\bm P^{-1} \bm\Lambda^\prime(\bm\Sigma^\xi)^{-1}\widehat{\bm\Gamma}^{\xi F}\bm\Lambda^\prime(\bm\Sigma^\xi)^{-1}\bm\Lambda\bm P^{-1}
\right]\quad \text{\it A.4.4}
\nonumber\\
&+T\left[
(\bm\Sigma^\xi)^{-1}\bm\Lambda\bm P^{-1}
\otimes
\bm P^{-1} \bm\Lambda^\prime(\bm\Sigma^\xi)^{-1}
\right]\bm C_{n,r}\quad \text{\it A.5}
\nonumber\\
&-T\left[
(\bm\Sigma^\xi)^{-1}\bm\Lambda\bm\Lambda^\prime(\bm\Sigma^\xi)^{-1}\bm\Lambda\bm P^{-1}
\otimes
\bm P^{-1} \bm\Lambda^\prime (\bm\Sigma^\xi)^{-1}
\right]\bm C_{n,r}\quad \text{\it A.6.1}
\nonumber\\
&-T\left[
(\bm\Sigma^\xi)^{-1}\widehat{\bm\Gamma}^\xi(\bm\Sigma^\xi)^{-1}\bm\Lambda\bm P^{-1}
\otimes
\bm P^{-1} \bm\Lambda^\prime (\bm\Sigma^\xi)^{-1}
\right]\bm C_{n,r}\quad \text{\it A.6.2}
\nonumber\\
&-T\left[
(\bm\Sigma^\xi)^{-1}\bm\Lambda\widehat{\bm\Gamma}^{F\xi}(\bm\Sigma^\xi)^{-1}\bm\Lambda\bm P^{-1}
\otimes
\bm P^{-1} \bm\Lambda^\prime (\bm\Sigma^\xi)^{-1}
\right]\bm C_{n,r}\quad \text{\it A.6.3}
\nonumber\\
&-T\left[
(\bm\Sigma^\xi)^{-1}\widehat{\bm\Gamma}^{\xi F}\bm\Lambda^\prime(\bm\Sigma^\xi)^{-1}\bm\Lambda\bm P^{-1}
\otimes
\bm P^{-1} \bm\Lambda^\prime (\bm\Sigma^\xi)^{-1}
\right]\bm C_{n,r}\quad \text{\it A.6.4}
\nonumber\\
&+T\left[
(\bm\Sigma^\xi)^{-1}\bm\Lambda\bm P^{-1}\bm\Lambda^\prime(\bm\Sigma^\xi)^{-1}\bm\Lambda\bm\Lambda^\prime(\bm\Sigma^\xi)^{-1}\bm\Lambda\bm P^{-1}
\otimes
\bm P^{-1} \bm\Lambda^\prime (\bm\Sigma^\xi)^{-1}
\right]\bm C_{n,r}\quad \text{\it A.7.1}
\nonumber\\
&+T\left[
(\bm\Sigma^\xi)^{-1}\bm\Lambda\bm P^{-1}\bm\Lambda^\prime(\bm\Sigma^\xi)^{-1}\widehat{\bm\Gamma}^\xi(\bm\Sigma^\xi)^{-1}\bm\Lambda\bm P^{-1}
\otimes
\bm P^{-1} \bm\Lambda^\prime (\bm\Sigma^\xi)^{-1}
\right]\bm C_{n,r}\quad \text{\it A.7.2}
\nonumber\\
&+T\left[
(\bm\Sigma^\xi)^{-1}\bm\Lambda\bm P^{-1}\bm\Lambda^\prime(\bm\Sigma^\xi)^{-1}\bm\Lambda\widehat{\bm\Gamma}^{F\xi}(\bm\Sigma^\xi)^{-1}\bm\Lambda\bm P^{-1}
\otimes
\bm P^{-1} \bm\Lambda^\prime (\bm\Sigma^\xi)^{-1}
\right]\bm C_{n,r}\quad \text{\it A.7.3}
\nonumber\\
&+T\left[
(\bm\Sigma^\xi)^{-1}\bm\Lambda\bm P^{-1}\bm\Lambda^\prime(\bm\Sigma^\xi)^{-1}\widehat{\bm\Gamma}^{\xi F}\bm\Lambda^\prime(\bm\Sigma^\xi)^{-1}\bm\Lambda\bm P^{-1}
\otimes
\bm P^{-1} \bm\Lambda^\prime (\bm\Sigma^\xi)^{-1}
\right]\bm C_{n,r}\quad \text{\it A.7.4}
\nonumber\\
&+T\left[
(\bm\Sigma^\xi)^{-1}\bm\Lambda\bm P^{-1}
\otimes
\bm P^{-1} \bm\Lambda^\prime(\bm\Sigma^\xi)^{-1}\bm\Lambda\bm\Lambda^\prime(\bm\Sigma^\xi)^{-1}\bm\Lambda\bm P^{-1}\bm\Lambda^\prime(\bm\Sigma^\xi)^{-1}
\right]\bm C_{n,r}\quad \text{\it A.8.1}
\nonumber\\
&+T\left[
(\bm\Sigma^\xi)^{-1}\bm\Lambda\bm P^{-1}
\otimes
\bm P^{-1} \bm\Lambda^\prime(\bm\Sigma^\xi)^{-1}\widehat{\bm\Gamma}^\xi(\bm\Sigma^\xi)^{-1}\bm\Lambda\bm P^{-1}\bm\Lambda^\prime(\bm\Sigma^\xi)^{-1}
\right]\bm C_{n,r}\quad \text{\it A.8.2}
\nonumber\\
&+T\left[
(\bm\Sigma^\xi)^{-1}\bm\Lambda\bm P^{-1}
\otimes
\bm P^{-1} \bm\Lambda^\prime(\bm\Sigma^\xi)^{-1}\bm\Lambda\widehat{\bm\Gamma}^{F \xi}(\bm\Sigma^\xi)^{-1}\bm\Lambda\bm P^{-1}\bm\Lambda^\prime(\bm\Sigma^\xi)^{-1}
\right]\bm C_{n,r}\quad \text{\it A.8.3}
\nonumber\\
&+T\left[
(\bm\Sigma^\xi)^{-1}\bm\Lambda\bm P^{-1}
\otimes
\bm P^{-1} \bm\Lambda^\prime(\bm\Sigma^\xi)^{-1}\widehat{\bm\Gamma}^{\xi F}\bm\Lambda^\prime(\bm\Sigma^\xi)^{-1}\bm\Lambda\bm P^{-1}\bm\Lambda^\prime(\bm\Sigma^\xi)^{-1}
\right]\bm C_{n,r}\quad \text{\it A.8.4}
\nonumber\\
&-T\left[(\bm\Sigma^\xi)^{-1}
\otimes \bm P^{-1}
\right]\quad \text{\it B.1}
\nonumber\\
&+T\left[
(\bm\Sigma^\xi)^{-1}\bm\Lambda\bm\Lambda^\prime(\bm\Sigma^\xi)^{-1}
\otimes
\bm P^{-1}
\right]\quad \text{\it B.2.1}
\nonumber\\
&+T\left[
(\bm\Sigma^\xi)^{-1}\widehat{\bm\Gamma}^\xi(\bm\Sigma^\xi)^{-1}
\otimes
\bm P^{-1}
\right]\quad \text{\it B.2.2}
\nonumber\\
&+T\left[
(\bm\Sigma^\xi)^{-1}\bm\Lambda\widehat{\bm\Gamma}^{F\xi}(\bm\Sigma^\xi)^{-1}
\otimes
\bm P^{-1}
\right]\quad \text{\it B.2.3}
\nonumber\\
&+T\left[
(\bm\Sigma^\xi)^{-1}\widehat{\bm\Gamma}^{\xi F}\bm\Lambda^\prime(\bm\Sigma^\xi)^{-1}
\otimes
\bm P^{-1}
\right]\quad \text{\it B.2.4}
\nonumber\\
&[\ldots]\nonumber
\end{align}
\begin{align}
&[\ldots]\nonumber\\
&-T\left[
(\bm\Sigma^\xi)^{-1}\bm\Lambda\bm P^{-1}\bm\Lambda^\prime(\bm\Sigma^\xi)^{-1}\bm\Lambda\bm\Lambda^\prime(\bm\Sigma^\xi)^{-1}
\otimes
\bm P^{-1}
\right]\quad \text{\it B.3.1}
\nonumber\\
&-T\left[
(\bm\Sigma^\xi)^{-1}\bm\Lambda\bm P^{-1}\bm\Lambda^\prime(\bm\Sigma^\xi)^{-1}\widehat{\bm\Gamma}^\xi(\bm\Sigma^\xi)^{-1}
\otimes
\bm P^{-1}
\right]\quad \text{\it B.3.2}
\nonumber\\
&-T\left[
(\bm\Sigma^\xi)^{-1}\bm\Lambda\bm P^{-1}\bm\Lambda^\prime(\bm\Sigma^\xi)^{-1}\bm\Lambda\widehat{\bm\Gamma}^{F\xi}(\bm\Sigma^\xi)^{-1}
\otimes
\bm P^{-1}
\right]\quad \text{\it B.3.3}
\nonumber\\
&-T\left[
(\bm\Sigma^\xi)^{-1}\bm\Lambda\bm P^{-1}\bm\Lambda^\prime(\bm\Sigma^\xi)^{-1}\widehat{\bm\Gamma}^{\xi F}\bm\Lambda^\prime(\bm\Sigma^\xi)^{-1}
\otimes
\bm P^{-1}
\right]\quad \text{\it B.3.4}
\nonumber\\
&-T\left[
(\bm\Sigma^\xi)^{-1}\bm\Lambda\bm P^{-1}
\otimes
\bm P^{-1} \bm\Lambda^\prime(\bm\Sigma^\xi)^{-1}\bm\Lambda\bm\Lambda^\prime(\bm\Sigma^\xi)^{-1}
\right]\bm C_{n,r}\quad \text{\it B.4.1}
\nonumber\\
&-T\left[
(\bm\Sigma^\xi)^{-1}\bm\Lambda\bm P^{-1}
\otimes
\bm P^{-1} \bm\Lambda^\prime(\bm\Sigma^\xi)^{-1}\widehat{\bm\Gamma}^\xi(\bm\Sigma^\xi)^{-1}
\right]\bm C_{n,r}\quad \text{\it B.4.2}
\nonumber\\
&-T\left[
(\bm\Sigma^\xi)^{-1}\bm\Lambda\bm P^{-1}
\otimes
\bm P^{-1} \bm\Lambda^\prime(\bm\Sigma^\xi)^{-1}\bm\Lambda\widehat{\bm\Gamma}^{F\xi}(\bm\Sigma^\xi)^{-1}
\right]\bm C_{n,r}\quad \text{\it B.4.3}
\nonumber\\
&-T\left[
(\bm\Sigma^\xi)^{-1}\bm\Lambda\bm P^{-1}
\otimes
\bm P^{-1} \bm\Lambda^\prime(\bm\Sigma^\xi)^{-1}\widehat{\bm\Gamma}^{\xi F}\bm\Lambda^\prime(\bm\Sigma^\xi)^{-1}
\right]\bm C_{n,r}\quad \text{\it B.4.4}
\nonumber\\
&-T\left[
(\bm\Sigma^\xi)^{-1}
\otimes
\bm P^{-1} \bm\Lambda^\prime(\bm\Sigma^\xi)^{-1}\bm\Lambda \bm\Lambda^\prime(\bm\Sigma^\xi)^{-1}\bm\Lambda\bm P^{-1}
\right]\quad \text{\it B.5.1}
\nonumber\\
&-T\left[
(\bm\Sigma^\xi)^{-1}
\otimes
\bm P^{-1} \bm\Lambda^\prime(\bm\Sigma^\xi)^{-1}\widehat{\bm\Gamma}^\xi(\bm\Sigma^\xi)^{-1}\bm\Lambda\bm P^{-1}
\right]\quad \text{\it B.5.2}
\nonumber\\
&-T\left[
(\bm\Sigma^\xi)^{-1}
\otimes
\bm P^{-1} \bm\Lambda^\prime(\bm\Sigma^\xi)^{-1}\bm\Lambda\widehat{\bm\Gamma}^{F\xi}(\bm\Sigma^\xi)^{-1}\bm\Lambda\bm P^{-1}
\right]\quad \text{\it B.5.3}
\nonumber\\
&-T\left[
(\bm\Sigma^\xi)^{-1}
\otimes
\bm P^{-1} \bm\Lambda^\prime(\bm\Sigma^\xi)^{-1}\widehat{\bm\Gamma}^{\xi F}\bm\Lambda^\prime(\bm\Sigma^\xi)^{-1}\bm\Lambda\bm P^{-1}
\right].\quad \text{\it B.5.4}
\end{align}
}
By using (ref)-(ref) 30 terms of (ref) cancel out asymptotically, namely:
\begin{enumerate}
• $\Vert \text{\it A.1}-\text{\it A.2.2}\Vert = O_{\mathrm P}(\sqrt T n^{-1})+ O_{\mathrm P}(T n^{-2})$;
• $\Vert \text{\it A.2.1}-\text{\it A.3.1}\Vert = O_{\mathrm P}(T n^{-1})$;
• $\Vert \text{\it A.2.3}-\text{\it A.3.3}\Vert = O_{\mathrm P}(T n^{-1})$;
• $\Vert \text{\it A.5}-\text{\it A.6.2}\Vert = O_{\mathrm P}(\sqrt T n^{-1})+ O_{\mathrm P}(T n^{-2})$;
• $\Vert \text{\it A.6.1}-\text{\it A.7.1}\Vert = O_{\mathrm P}(T n^{-1})$;
• $\Vert \text{\it A.6.3}-\text{\it A.7.3}\Vert = O_{\mathrm P}(T n^{-1})$;
• $\Vert \text{\it A.2.4}-\text{\it B.2.4}\Vert = O_{\mathrm P}(T n^{-1})$;
• $\Vert \text{\it B.2.1}-\text{\it B.3.1}\Vert = O_{\mathrm P}(T n^{-1})$;
• $\Vert \text{\it A.3.2}-\text{\it B.3.2}\Vert = O_{\mathrm P}(\sqrt T n^{-1})+ O_{\mathrm P}(T n^{-2})$;
• $\Vert \text{\it B.2.3}-\text{\it B.3.3}\Vert = O_{\mathrm P}(T n^{-1})$;
• $\Vert \text{\it A.3.4}-\text{\it B.3.4}\Vert = O_{\mathrm P}(T n^{-1})$;
• $\Vert \text{\it A.8.1}-\text{\it B.4.1}\Vert = O_{\mathrm P}(T n^{-1})$;
• $\Vert \text{\it A.8.2}-\text{\it B.4.2}\Vert = O_{\mathrm P}(T n^{-1})+O_{\mathrm P}(\sqrt T n^{-1})+ O_{\mathrm P}(T n^{-2})$;
• $\Vert \text{\it A.8.4}-\text{\it B.4.4}\Vert = O_{\mathrm P}(T n^{-1})$;
• $\Vert \text{\it B.2.2}-\text{\it B.5.2}\Vert = O_{\mathrm P}(T n^{-1})+O_{\mathrm P}(\sqrt T n^{-1})+ O_{\mathrm P}(T n^{-2})$.
\end{enumerate}
and by using again (ref) we are left with the following 13 terms (ordered differently than in the previous expression):
\begin{align}
\bm H(\bm{\mathcal X};\bm\varphi)=&\,
-T \left[
(\bm\Sigma^{\xi})^{-1}
\otimes
\mathbf I_r
\right]\quad \text{\it B.5.1}\nonumber\\
&+T\left[
(\bm\Sigma^{\xi})^{-1}\bm\Lambda
\left\{\bm\Lambda^\prime(\bm\Sigma^{\xi})^{-1}\bm\Lambda \right\}^{-1}
\bm\Lambda^\prime(\bm\Sigma^{\xi})^{-1}
\otimes
\mathbf I_r
\right]\quad \text{\it A.4.1}\nonumber\\
&+T\left[
(\bm\Sigma^{\xi})^{-1}\bm\Lambda
\left\{\bm\Lambda^\prime(\bm\Sigma^{\xi})^{-1}\bm\Lambda \right\}^{-1}
\bm\Lambda^\prime(\bm\Sigma^{\xi})^{-1}
\otimes
\left\{\bm\Lambda^\prime(\bm\Sigma^{\xi})^{-1}\bm\Lambda \right\}^{-1}
\right]\quad \text{\it A.4.2}\nonumber\\
&+T\left[
(\bm\Sigma^{\xi})^{-1}\bm\Lambda
\left\{\bm\Lambda^\prime(\bm\Sigma^{\xi})^{-1}\bm\Lambda \right\}^{-1}
\bm\Lambda^\prime(\bm\Sigma^{\xi})^{-1}
\otimes
\widehat{\bm\Gamma}^{F\xi} (\bm\Sigma^{\xi})^{-1}\bm\Lambda\left\{\bm\Lambda^\prime(\bm\Sigma^{\xi})^{-1}\bm\Lambda \right\}^{-1}
\right]\quad \text{\it A.4.3}\nonumber\\
&+T\left[
(\bm\Sigma^{\xi})^{-1}\bm\Lambda
\left\{\bm\Lambda^\prime(\bm\Sigma^{\xi})^{-1}\bm\Lambda \right\}^{-1}
\bm\Lambda^\prime(\bm\Sigma^{\xi})^{-1}
\otimes
\left\{\bm\Lambda^\prime(\bm\Sigma^{\xi})^{-1}\bm\Lambda \right\}^{-1}
\bm\Lambda^\prime(\bm\Sigma^{\xi})^{-1} \widehat{\bm\Gamma}^{\xi F}
\right]\quad \text{\it A.4.4}\nonumber\\
&-T\left[
(\bm\Sigma^{\xi})^{-1}
\otimes
\left\{\bm\Lambda^\prime(\bm\Sigma^{\xi})^{-1}\bm\Lambda \right\}^{-1}
\right]\quad \text{\it B.1}\nonumber\\
&-T\left[
(\bm\Sigma^{\xi})^{-1}
\otimes
\widehat{\bm\Gamma}^{F\xi}(\bm\Sigma^{\xi})^{-1}\bm\Lambda\left\{\bm\Lambda^\prime(\bm\Sigma^{\xi})^{-1}\bm\Lambda \right\}^{-1}
\right]\quad \text{\it B.5.2}\nonumber\\
&-T\left[
(\bm\Sigma^{\xi})^{-1}
\otimes
\left\{\bm\Lambda^\prime(\bm\Sigma^{\xi})^{-1}\bm\Lambda \right\}^{-1} \bm\Lambda^\prime(\bm\Sigma^{\xi})^{-1}\widehat{\bm\Gamma}^{\xi F}
\right]\quad \text{\it B.5.3}\nonumber\\
&-T\left[
(\bm\Sigma^{\xi})^{-1}\bm\Lambda\left\{\bm\Lambda^\prime(\bm\Sigma^{\xi})^{-1}\bm\Lambda \right\}^{-1}
\otimes
\widehat{\bm\Gamma}^{F\xi}(\bm\Sigma^{\xi})^{-1}
\right]\bm C_{n,r}\quad \text{\it B.4.3}\nonumber\\
&-T\left[
(\bm\Sigma^{\xi})^{-1}\widehat{\bm\Gamma}^{\xi F}
\otimes
\left\{\bm\Lambda^\prime(\bm\Sigma^{\xi})^{-1}\bm\Lambda \right\}^{-1}\bm\Lambda^\prime (\bm\Sigma^{\xi})^{-1}
\right]\bm C_{n,r}\quad \text{\it A.6.4}\nonumber\\
&+T\left[
(\bm\Sigma^{\xi})^{-1}\bm\Lambda\left\{\bm\Lambda^\prime(\bm\Sigma^{\xi})^{-1}\bm\Lambda \right\}^{-1}
\otimes
\left\{\bm\Lambda^\prime(\bm\Sigma^{\xi})^{-1}\bm\Lambda \right\}^{-1} \bm\Lambda^\prime(\bm\Sigma^{\xi})^{-1}
\right]\bm C_{n,r}\quad \text{\it A.7.1}\nonumber\\
&+T\left[
(\bm\Sigma^{\xi})^{-1}\bm\Lambda\left\{\bm\Lambda^\prime(\bm\Sigma^{\xi})^{-1}\bm\Lambda \right\}^{-1} \bm\Lambda^\prime(\bm\Sigma^{\xi})^{-1}\widehat{\bm\Gamma}^{\xi F}
\otimes
\left\{\bm\Lambda^\prime(\bm\Sigma^{\xi})^{-1}\bm\Lambda \right\}^{-1} \bm\Lambda^\prime(\bm\Sigma^{\xi})^{-1}
\right]\bm C_{n,r}\quad \text{\it A.7.2}\nonumber\\
&+T\left[
(\bm\Sigma^{\xi})^{-1}\bm\Lambda\left\{\bm\Lambda^\prime(\bm\Sigma^{\xi})^{-1}\bm\Lambda \right\}^{-1}
\otimes
\widehat{\bm\Gamma}^{F\xi}(\bm\Sigma^{\xi})^{-1}\bm\Lambda\left\{\bm\Lambda^\prime(\bm\Sigma^{\xi})^{-1}\bm\Lambda \right\}^{-1} \bm\Lambda^\prime(\bm\Sigma^{\xi})^{-1}
\right]\bm C_{n,r}\quad \text{\it A.8.3}\nonumber\\
&+O\left(\frac {T}n\right)+O_{\mathrm P}\left(\frac {\sqrt T}n\right)+O\left(\frac {T}{n^2}\right).
\end{align}
Therefore,
by using again arguments as those in (ref)-(ref) in (ref) we have that the $i$th $r\times r$ sub-matrix of $\bm H(\bm{\mathcal X};\bm{\varphi})$ is such that
\begin{align}
\bm h_{ii}&(\bm{\mathcal X};\bm\varphi)= -\frac T{\sigma_i^2} \mathbf I_r\nonumber\\
& + \frac T{\sigma_i^4}{\bm\lambda}_i^\prime\left\{\bm\Lambda^\prime(\bm\Sigma^{\xi})^{-1}\bm\Lambda \right\}^{-1}{\bm\lambda}_i+\nonumber\\
& + \frac T{\sigma_i^4}{\bm\lambda}_i^\prime\left\{\bm\Lambda^\prime(\bm\Sigma^{\xi})^{-1}\bm\Lambda \right\}^{-1}{\bm\lambda}_i
\left\{
\left\{\bm\Lambda^\prime(\bm\Sigma^{\xi})^{-1}\bm\Lambda \right\}^{-1}+
\frac 1T\sum_{t=1}^T\mathbf F_t{\bm\xi}_t^\prime(\bm\Sigma^{\xi})^{-1}\bm\Lambda\left\{\bm\Lambda^\prime(\bm\Sigma^{\xi})^{-1}\bm\Lambda \right\}^{-1}\right.\nonumber\\
&\;\;+\left.\left\{\bm\Lambda^\prime(\bm\Sigma^{\xi})^{-1}\bm\Lambda \right\}^{-1}\bm\Lambda^\prime(\bm\Sigma^{\xi})^{-1}\frac 1T\sum_{t=1}^T{\bm\xi}_t\mathbf F_t^\prime
\right\}\nonumber\\
&- \frac T{\sigma_i^2}
\left\{
\left\{\bm\Lambda^\prime(\bm\Sigma^{\xi})^{-1}\bm\Lambda \right\}^{-1}+
\frac 1T\sum_{t=1}^T\mathbf F_t{\bm\xi}_t^\prime(\bm\Sigma^{\xi})^{-1}\bm\Lambda\left\{\bm\Lambda^\prime(\bm\Sigma^{\xi})^{-1}\bm\Lambda \right\}^{-1}+
\left\{\bm\Lambda^\prime(\bm\Sigma^{\xi})^{-1}\bm\Lambda \right\}^{-1}\bm\Lambda^\prime(\bm\Sigma^{\xi})^{-1}\frac 1T\sum_{t=1}^T{\bm\xi}_t\mathbf F_t^\prime
\right\}\nonumber\\
&- \frac T{\sigma_i^4}\left\{
\frac 1T \sum_{t=1}^T\mathbf F_t{\xi}_{it}
\otimes
{\bm\lambda}_i^\prime\left\{\bm\Lambda^\prime(\bm\Sigma^{\xi})^{-1}\bm\Lambda \right\}^{-1}
\right\}
- \frac T{\sigma_i^4}\left\{
\left\{\bm\Lambda^\prime(\bm\Sigma^{\xi})^{-1}\bm\Lambda \right\}^{-1}{\bm\lambda}_i
\otimes
\frac 1T \sum_{t=1}^T{\xi}_{it}\mathbf F_t^\prime
\right\}\nonumber\%&- \frac T{\sigma_i^4}\left\{
&+ \frac T{\sigma_i^4}
\Bigg\{
\left\{\bm\Lambda^\prime(\bm\Sigma^{\xi})^{-1}\bm\Lambda \right\}^{-1} {\bm\lambda}_i
\otimes
{\bm\lambda}_i^\prime\left\{\bm\Lambda^\prime(\bm\Sigma^{\xi})^{-1}\bm\Lambda \right\}^{-1}
\Bigg\}\nonumber\\
&+ \frac T{\sigma_i^4}\Bigg\{
\left\{\bm\Lambda^\prime(\bm\Sigma^{\xi})^{-1}\bm\Lambda \right\}^{-1}{\bm\lambda}_i
\otimes
{\bm\lambda}_i^\prime\left\{\bm\Lambda^\prime(\bm\Sigma^{\xi})^{-1}\bm\Lambda \right\}^{-1}\bm\Lambda^\prime(\bm\Sigma^\xi)^{-1}\frac 1T\sum_{t=1}^T{\bm\xi}_t\mathbf F_t^\prime
\Bigg\}\nonumber\\
&+ \frac T{\sigma_i^4}\Bigg\{
\frac 1T\sum_{t=1}^T\mathbf F_t{\bm\xi}_t^\prime(\bm\Sigma^\xi)^{-1}\bm\Lambda\left\{\bm\Lambda^\prime(\bm\Sigma^{\xi})^{-1}\bm\Lambda \right\}^{-1}{\bm\lambda}_i
\otimes
{\bm\lambda}_i^\prime\left\{\bm\Lambda^\prime(\bm\Sigma^{\xi})^{-1}\bm\Lambda \right\}^{-1}
\Bigg\}\nonumber\\
&+O\left(\frac{T}{n^2}\right)+O_{\mathrm P}\left(\frac{\sqrt T}{n^2}\right)+O\left(\frac{T}{n^4}\right).
\end{align}
Hence, $\Vert \bm h_{ii}(\bm{\mathcal X};\bm\varphi)\Vert=O_{\mathrm P}(T)$, because of Assumption (ref)(a), (ref), and (ref). Moreover, from (ref) and by noticing that $(\bm\Sigma^\xi)^{-1}{\bm\Lambda}$ has the same properties as ${\bm\Lambda}$ since $\Vert (\bm\Sigma^\xi)^{-1}\Vert = O(1)$ by Assumption (ref)(a), we get
\begin{align}
\frac 1T&\left\Vert\bm h_{ii}(\bm{\mathcal X};\bm\varphi)- \left(-\frac T{\sigma_i^2} \mathbf I_r\right)\right\Vert\le \frac{M_\Lambda^2}{C_\xi^2}
\left\Vert \left\{\bm\Lambda^\prime(\bm\Sigma^{\xi})^{-1}\bm\Lambda \right\}^{-1}\right\Vert\nonumber\\
& +
\frac{M_\Lambda^2}{C_\xi^2}
\left\Vert \left\{\bm\Lambda^\prime(\bm\Sigma^{\xi})^{-1}\bm\Lambda \right\}^{-1}\right\Vert^2
\Bigg\{
1 + 2\left\Vert\frac 1T\bm\Lambda^\prime(\bm\Sigma^\xi)^{-1}\sum_{t=1}^T{\bm\xi}_t\mathbf F_t^\prime \right\Vert
\Bigg\}\nonumber\\
&
\frac 1{C_\xi}
\left\Vert \left\{\bm\Lambda^\prime(\bm\Sigma^{\xi})^{-1}\bm\Lambda \right\}^{-1}\right\Vert
\Bigg\{
1 + 2\left\Vert\frac 1T\bm\Lambda^\prime(\bm\Sigma^\xi)^{-1}\sum_{t=1}^T{\bm\xi}_t\mathbf F_t^\prime \right\Vert
\Bigg\}\nonumber\\
&+ \frac{2M_\Lambda}{C_\xi^2} \left\Vert \left\{\bm\Lambda^\prime(\bm\Sigma^{\xi})^{-1}\bm\Lambda \right\}^{-1}\right\Vert\,\left\Vert\frac 1T\sum_{t=1}^T\mathbf F_t{\xi}_{it}\right\Vert+
\frac{M_\Lambda^2}{C_\xi^2}
\left\Vert \left\{\bm\Lambda^\prime(\bm\Sigma^{\xi})^{-1}\bm\Lambda \right\}^{-1}\right\Vert^2\Bigg\{
1 + 2\left\Vert\frac 1T\bm\Lambda^\prime(\bm\Sigma^\xi)^{-1}\sum_{t=1}^T{\bm\xi}_t\mathbf F_t^\prime \right\Vert
\Bigg\}\nonumber\\
&+O\left(\frac{1}{n^2}\right)+O_{\mathrm P}\left(\frac{1}{n^2\sqrt T}\right)+O\left(\frac{1}{n^4}\right)\nonumber\\
=&\,O\left(\frac{1}{n}\right) +O\left(\frac{1}{n^2}\right) \Bigg\{1+O_{\mathrm P}\left(\frac{\sqrt n}{\sqrt T}\right)\Bigg\}+O\left(\frac{1}{n}\right) \Bigg\{1+O_{\mathrm P}\left(\frac{\sqrt n}{\sqrt T}\right)\Bigg\}+O\left(\frac{1}{n}\right)O_{\mathrm P}\left(\frac{1}{\sqrt T}\right)\nonumber\\
&+O\left(\frac{1}{n^2}\right) \Bigg\{1+O_{\mathrm P}\left(\frac{\sqrt n}{\sqrt T}\right)\Bigg\}+O\left(\frac{1}{n^2}\right)+O_{\mathrm P}\left(\frac{1}{n^2\sqrt T}\right)+O\left(\frac{1}{n^4}\right)= O_{\mathrm P}\left(\max\left(\frac{1}{n},\frac 1{\sqrt{nT}}\right)\right).
\end{align}
Finally, from (ref) we have
\begin{align}
\mathrm d \bm S^\prime(\bm{\mathcal X}|\bm{\mathcal F};\underline{\bm\varphi})&=\mathrm d\left(\sum_{t=1}^T\mathbf F_t(\mathbf x_t-\underline{\bm\Lambda}\mathbf F_t)^\prime(\underline{\bm\Sigma}^\xi)^{-1} \right) = -\sum_{t=1}^T\mathbf F_t\mathbf F_t^\prime\left(\mathrm d\underline{\bm\Lambda}^\prime\right)(\underline{\bm\Sigma}^\xi)^{-1},\nonumber
\end{align}
so
\begin{equation}
\text{vec}(\mathrm d \bm S^\prime(\bm{\mathcal X}|\bm{\mathcal F};\underline{\bm\varphi})) = - (\underline{\bm\Sigma}^\xi)^{-1}\otimes \left(\sum_{t=1}^T\mathbf F_t\mathbf F_t^\prime\right)\text{vec}(\mathrm d \underline{\bm \Lambda}^\prime),\nonumber
\end{equation}
and
\begin{equation}
\bm H(\bm{\mathcal X}|\bm{\mathcal F};\underline{\bm\varphi}) = - (\underline{\bm\Sigma}^\xi)^{-1}\otimes \left(\sum_{t=1}^T\mathbf F_t\mathbf F_t^\prime\right)=-T (\underline{\bm\Sigma}^\xi)^{-1}\otimes \mathbf I_r,
\end{equation}
because of Assumption (ref)(b). And (ref) computed in the true value of the parameters is
\begin{equation}
\bm H(\bm{\mathcal X}|\bm{\mathcal F};{\bm\varphi}) = -T ({\bm\Sigma}^\xi)^{-1}\otimes \mathbf I_r.
\end{equation}
From (ref) the $i$th $r\times r$ sub-matrix of $\bm H(\bm{\mathcal X}|\bm{\mathcal F};\bm{\varphi})$ is then:
\begin{equation}
\bm h_{ii}(\bm{\mathcal X}|\bm{\mathcal F};{\bm\varphi}) = -\frac T{\sigma_i^2}\mathbf I_r.
\end{equation}
Hence, $\Vert\bm h_{ii}(\bm{\mathcal X}|\bm{\mathcal F};{\bm\varphi}) \Vert=O_{\mathrm P}(T)$, and, by using (ref) into (ref), for any given $i=1,\ldots, n$, we have:
\begin{equation}\nonumber
\frac 1T\left\Vert\bm h_{ii}(\bm{\mathcal X};\bm\varphi)- \bm h_{ii}(\bm{\mathcal X}|\bm{\mathcal F};{\bm\varphi})\right\Vert=O_{\mathrm P}\left(\max\left(\frac{1}{n},\frac 1{\sqrt{nT}}\right)\right),
\end{equation}
which proves part (b).
For part (c) we provide two different equivalent proofs. First,
\begin{align}
\left\Vert
\bm s_{i}(\bm{\mathcal X};\widehat{\bm{\varphi}}^{\text{\tiny QML,E}})-\bm s_{i}(\bm{\mathcal X}|\bm{\mathcal F};\widehat{\bm{\varphi}}^{\text{\tiny QML,E}})
\right\Vert
\le&\, \left\Vert
\bm s_{i}(\bm{\mathcal X};{\bm{\varphi}})-\bm s_{i}(\bm{\mathcal X}|\bm{\mathcal F};{\bm{\varphi}})
\right\Vert\nonumber\\
&+\left\Vert
\bm h_{ii}(\bm{\mathcal X};{\bm{\varphi}})-\bm h_{ii}(\bm{\mathcal X}|\bm{\mathcal F};{\bm{\varphi}})
\right\Vert\, \left\Vert \widehat{\bm\lambda}_i^{\text{\tiny QML,E}}-{\bm\lambda}_i \right\Vert+O_{\mathrm P}\left(\left\Vert \widehat{\bm\lambda}_i^{\text{\tiny QML,E}}-{\bm\lambda}_i \right\Vert^2\right)\nonumber\\
=&\,O_{\mathrm P}\left(\max\left(\frac {T}n,\frac {\sqrt T}{\sqrt{n}}\right)\right)+O_{\mathrm P}\left(\max\left(\frac {T}n,\frac {\sqrt T}{\sqrt{n}}\right)\right)O_{\mathrm P}\left(\max\left(\frac {1}n,\frac {1}{\sqrt{T}}\right)\right).\nonumber
\end{align}
because of parts (a) and (b), and (ref), which follows from Theorem (ref), which, in turn, holds even if we use the simpler log-likelihood (ref) in place of the log-likelihood (ref).
As a consequence,
\[
\frac 1{\sqrt T}\left\Vert
\bm s_{i}(\bm{\mathcal X};\widehat{\bm{\varphi}}^{\text{\tiny QML,E}})-\bm s_{i}(\bm{\mathcal X}|\bm{\mathcal F};\widehat{\bm{\varphi}}^{\text{\tiny QML,E}})
\right\Vert= O_{\mathrm P}\left(\max\left(\frac {\sqrt T}n,\frac 1{\sqrt{n}}\right)\right).
\]
Alternatively, let $\widehat{\bm\varphi}^{\text{\tiny OLS}}=(\mathrm{vec}(\bm\Lambda^{\text{\tiny OLS}})^\prime, \mathrm{vech}(\underline{\bm\Gamma}^\xi)^\prime)^\prime$ which is the vector of parameters made of the OLS estimator of the loadings and any generic value of the idiosyncratic covariance matrix satisfying Assumption (ref). Then, notice that $\bm s_{i}(\bm{\mathcal X};\widehat{\bm{\varphi}}^{\text{\tiny QML,E}})=\mathbf 0_r$ by definition of QML estimator, and $\bm s_{i}(\bm{\mathcal X}|\bm{\mathcal F};\widehat{\bm{\varphi}}^{\text{\tiny OLS}})=\mathbf 0_r$ because the OLS estimator is the QML estimator when maximizing the conditional log-likelihood, and the OLS estimator of the loadings does not depend on the estimator of the idiosyncratic covariance. Recall also that
\[
\left\Vert \widehat{\bm\lambda}_i^{\text{\tiny QML,E}}-{\bm\lambda}_i^{\text{\tiny OLS}} \right\Vert\le \left\Vert \widehat{\bm\lambda}_i^{\text{\tiny QML,E}}-\widehat{\bm\lambda}_i \right\Vert+\left\Vert \widehat{\bm\lambda}_i-{\bm\lambda}_i^{\text{\tiny OLS}} \right\Vert = O_{\mathrm P}\left(\max\left(\frac 1n,\frac 1{\sqrt{nT}}\right)\right),
\]
because of Theorem (ref) and (ref) which hold even if we use the log-likelihood (ref) in place of the log-likelihood (ref). Moreover, from (ref) it is straightforward to see that
$\Vert\bm h_{ii}(\bm{\mathcal X}|\bm{\mathcal F};\underline{\bm{\varphi}})\Vert=O_{\mathrm P}(T)$, for any $\underline{\bm{\varphi}}\in\mathcal O_n$. Therefore,
\begin{align}
\left\Vert
\bm s_{i}(\bm{\mathcal X};\widehat{\bm{\varphi}}^{\text{\tiny QML,E}})-\bm s_{i}(\bm{\mathcal X}|\bm{\mathcal F};\widehat{\bm{\varphi}}^{\text{\tiny QML,E}})
\right\Vert
=&\, \left\Vert
\bm s_{i}(\bm{\mathcal X}|\bm{\mathcal F};\widehat{\bm{\varphi}}^{\text{\tiny QML,E}})
\right\Vert\nonumber\\
\le&\, \left\Vert
\bm s_{i}(\bm{\mathcal X}|\bm{\mathcal F};\widehat{\bm{\varphi}}^{\text{\tiny OLS}})
\right\Vert+ \left\Vert
\bm h_{ii}(\bm{\mathcal X}|\bm{\mathcal F};\widehat{\bm{\varphi}}^{\text{\tiny OLS}})
\right\Vert\,
\left\Vert \widehat{\bm\lambda}_i^{\text{\tiny QML,E}}-{\bm\lambda}_i^{\text{\tiny OLS}} \right\Vert\nonumber\\
&+ O_{\mathrm P}\left(\left\Vert \widehat{\bm\lambda}_i^{\text{\tiny QML,E}}-{\bm\lambda}_i^{\text{\tiny OLS}} \right\Vert^2\right)\nonumber\\
\le &\, O_{\mathrm P}\left(T\right) O_{\mathrm P}\left(\max\left(\frac 1n,\frac 1{\sqrt{nT}}\right)\right)=O_{\mathrm P}\left(\max\left(\frac {T}n,\frac {\sqrt T}{\sqrt{n}}\right)\right).\nonumber
\end{align}
This proves part (c).
For part (d), for any specific value of the parameters, say $\widetilde{\bm\varphi}$, define the $n^2r^2\times nr$ matrices of third derivatives
\[
\bm T(\bm{\mathcal X};\widetilde{\bm\varphi}) =\left.\frac{\partial \text{vec}(\bm H(\bm{\mathcal X};\underline{\bm\varphi}))^\prime}{\partial \text{vec}(\underline{\bm\Lambda})^\prime}
\right|_{\underline{\bm\varphi}=\widetilde{\bm\varphi}},\quad \bm T(\bm{\mathcal X}|\bm{\mathcal F};\widetilde{\bm\varphi}) =\left.\frac{\partial \text{vec}(\bm H(\bm{\mathcal X}|\bm{\mathcal F};\underline{\bm\varphi}))^\prime}{\partial \text{vec}(\underline{\bm\Lambda})^\prime}
\right|_{\underline{\bm\varphi}=\widetilde{\bm\varphi}},
\]
with components
\[
\bm t_{iii}(\bm{\mathcal X};\widetilde{\bm\varphi}) = \left.\frac{\partial \text{vec}(\bm h_{ii}(\bm{\mathcal X};\underline{\bm\varphi}))^\prime}{\partial \underline{\bm\lambda}_i^\prime}
\right|_{\underline{\bm\varphi}=\widetilde{\bm\varphi}},\quad \bm t_{iii}(\bm{\mathcal X}|\bm{\mathcal F};\widetilde{\bm\varphi}) = \left.\frac{\partial \text{vec}(\bm h_{ii}(\bm{\mathcal X}|\bm{\mathcal F};\underline{\bm\varphi}))^\prime}{\partial \underline{\bm\lambda}_i^\prime}
\right|_{\underline{\bm\varphi}=\widetilde{\bm\varphi}},\quad i=1,\ldots, n,
\]
which are $r^2\times r$ matrices. Then, from (ref) the $i$th $r\times r$ sub-matrix of $\bm H(\bm{\mathcal X};\underline{\bm\varphi})$ is given by:
\begin{align}
\bm h_{ii}(\bm{\mathcal X};\underline{\bm\varphi})=&\,
\frac T{\underline{\sigma}_i^4}
\left[
\left\{\underline{\bm\lambda}_i^\prime \underline{\bm P}^{-1}\underline{\bm\lambda}_i\right\}
\underline{\bm P}^{-1}
\right]\nonumber\\
&-\frac T{\underline{\sigma}_i^4}
\left[
\left\{[\widehat{\bm\Gamma}^x]_{i\cdot}(\underline{\bm\Sigma}^\xi)^{-1}\underline{\bm\Lambda}\underline{\bm P}^{-1}\underline{\bm\lambda}_i
\right\}
\underline{\bm P}^{-1}
\right]\nonumber\\
&+\frac T{\underline{\sigma}_i^4}
\left[
\left\{\underline{\bm \lambda}_i\underline{\bm P}^{-1} \underline{\bm \Lambda}^\prime (\underline{\bm\Sigma}^\xi)^{-1}\widehat{\bm\Gamma}^x (\underline{\bm\Sigma}^\xi)^{-1}\underline{\bm \Lambda}\,\underline{\bm P}^{-1} \underline{\bm \lambda}_i
\right\}
\underline{\bm P}^{-1}
\right]\nonumber\\
&+\frac T{\underline{\sigma}_i^4}
\left[
\left\{\underline{\bm\lambda}_i^\prime \underline{\bm P}^{-1}\underline{\bm\lambda}_i
\right\}
\underline{\bm P}^{-1}\underline{\bm\Lambda}^\prime(\underline{\bm\Sigma}^\xi)^{-1}\widehat{\bm\Gamma}^\xi(\underline{\bm\Sigma}^\xi)^{-1}
\underline{\bm\Lambda}\,\underline{\bm P}^{-1}
\right]\nonumber\\
&+\frac T{\underline{\sigma}_i^4}
\left[
\underline{\bm P}^{-1}\underline{\bm\lambda}_i
\otimes
\underline{\bm\lambda}_i^\prime \underline{\bm P}^{-1}
\right]\nonumber\\
&-\frac T{\underline{\sigma}_i^4}
\left[
\underline{\bm P}^{-1}\underline{\bm\lambda}_i
\otimes
[\widehat{\bm\Gamma}^x]_{i\cdot}(\underline{\bm\Sigma}^\xi)^{-1}\underline{\bm\Lambda}\,\underline{\bm P}^{-1}
\right]\nonumber\\
&+\frac T{\underline{\sigma}_i^4}
\left[
\underline{\bm P}^{-1}\underline{\bm\lambda}_i
\otimes
\underline{\bm\lambda}_i^\prime \underline{\bm P}^{-1}\underline{\bm\Lambda}^\prime(\underline{\bm\Sigma}^\xi)^{-1}\widehat{\bm\Gamma}^\xi(\underline{\bm\Sigma}^\xi)^{-1}\underline{\bm\Lambda}\,\underline{\bm P}^{-1}
\right]\nonumber\\
&+\frac T{\underline{\sigma}_i^4}
\left[
\underline{\bm P}^{-1}\underline{\bm\Lambda}^\prime(\underline{\bm\Sigma}^\xi)^{-1}\widehat{\bm\Gamma}^\xi(\underline{\bm\Sigma}^\xi)^{-1}\underline{\bm\Lambda}\underline{\bm P}^{-1}\underline{\bm\lambda}_i
\otimes
\underline{\bm\lambda}_i^\prime \underline{\bm P}^{-1}
\right]\nonumber\\
&-\frac T{\underline{\sigma}_i^2}
\left[
\underline{\bm P}^{-1}
\right] \nonumber\\
&+\frac T{\underline{\sigma}_i^4}
\left[
[\widehat{\bm\Gamma}^x]_{ii}
\underline{\bm P}^{-1}
\right]\nonumber\\
&-\frac T{\underline{\sigma}_i^4}
\left[
\left\{\underline{\bm\lambda}_i^\prime\underline{\bm P}^{-1}\underline{\bm\Lambda}^\prime(\underline{\bm\Sigma}^\xi)^{-1}[\widehat{\bm\Gamma}^x]_{\cdot i}
\right\}
\underline{\bm P}^{-1}
\right]\nonumber\\
&-\frac T{\underline{\sigma}_i^4}
\left[
\underline{\bm P}^{-1}\underline{\bm\Lambda}^\prime(\underline{\bm\Sigma}^\xi)^{-1}[\widehat{\bm\Gamma}^x]_{\cdot i}
\otimes
\underline{\bm\lambda}_i^\prime\underline{\bm P}^{-1}
\right]\nonumber\\
&-\frac T{\underline{\sigma}_i^2}
\left[
\underline{\bm P}^{-1}\underline{\bm\Lambda}^\prime(\underline{\bm\Sigma}^\xi)^{-1}\widehat{\bm\Gamma}^x(\underline{\bm\Sigma}^\xi)^{-1}\underline{\bm\Lambda}\underline{\bm P}^{-1}
\right],
\end{align}
where $[\widehat{\bm\Gamma}^x]_{i\cdot}$ and $[\widehat{\bm\Gamma}^x]_{\cdot i}$ are the $i$th row and column of $\widehat{\bm\Gamma}^x$, respectively, and $[\widehat{\bm\Gamma}^x]_{ii}$ is the $i$th term on its diagonal.
Since,
\begin{align}
&\left\Vert \underline{\bm\lambda}_i\right\Vert \le M_\Lambda,\quad \left\Vert \underline{\bm\Lambda}\right\Vert = O(\sqrt n),\nonumber\\
&\left\Vert (\underline{\bm\Sigma}^\xi)^{-1}\right\Vert \le \frac 1{C_\xi},\quad \left\vert \frac 1{\sigma_i^2}\right\vert\le \frac 1{C_\xi},\nonumber\\
&\left\Vert \underline{\bm P}^{-1}\right\Vert =\left\Vert \left\{\mathbf I_r+\underline{\bm \Lambda}^\prime (\underline{\bm\Sigma}^\xi)^{-1}\underline{\bm \Lambda} \right\}^{-1}\right\Vert = O\left(\frac 1n\right),\nonumber\\
&\left\Vert \widehat{\bm\Gamma}^x\right\Vert \le \left\Vert \widehat{\bm\Gamma}^x-\bm\Gamma^x\right\Vert+\left\Vert {\bm\Gamma}^x\right\Vert = O_{\mathrm P}\left(\frac n{\sqrt T}\right) +O(n),\nonumber\\
& \left\Vert [\widehat{\bm\Gamma}^x]_{i\cdot }\right\Vert = \left\Vert [\widehat{\bm\Gamma}^x]_{i\cdot }-[{\bm\Gamma}^x]_{i\cdot }\right\Vert+\left\Vert [{\bm\Gamma}^x]_{i\cdot }\right\Vert= O_{\mathrm P}\left(\frac{\sqrt n}{\sqrt T}\right) +O(\sqrt n),
\end{align}
because of Assumption (ref)(a), Lemma (ref)(i), Assumption (ref)(a), Lemma (ref)(i), and Lemma (ref)(vi), respectively
From (ref) and (ref) it is clear that the leading term in $\bm h_{ii}(\bm{\mathcal X};\underline{\bm\varphi})$ is the last one, which is $O_{\mathrm P}(T)$ while all others are $o_{\mathrm P}(T)$, specifically,
\begin{align}
\frac 1T \bm h_{ii}(\bm{\mathcal X};\underline{\bm\varphi})&= -\frac 1{\underline{\sigma}_i^2}
\left[
\underline{\bm P}^{-1}\underline{\bm\Lambda}^\prime(\underline{\bm\Sigma}^\xi)^{-1}\widehat{\bm\Gamma}^x(\underline{\bm\Sigma}^\xi)^{-1}\underline{\bm\Lambda}\underline{\bm P}^{-1}
\right]+O_{\mathrm P}\left(\frac 1n\right)+O\left(\frac 1{n^2}\right)\nonumber\\
&=\frac 1T \bm h^*_{ii}(\bm{\mathcal X};\underline{\bm\varphi})+O_{\mathrm P}\left(\frac 1n\right)+O\left(\frac 1{n^2}\right), \;\text{say.}
\end{align}
From (ref)
\begin{align}
\text{vec}\left(\mathrm d \bm h_{ii}^{*\prime} (\bm{\mathcal X};\underline{\bm\varphi})\right) =&\,
-\frac T{\underline{\sigma}_i^2}\left\{
\underline{\bm P}^{-1}\underline{\bm\Lambda}^\prime(\underline{\bm\Sigma}^\xi)^{-1}\widehat{\bm\Gamma}^x(\underline{\bm\Sigma}^\xi)^{-1}\underline{\bm\Lambda}
\otimes
\mathbf I_r
\right\}\text{vec}\left(\mathrm d\underline{\bm P}^{-1}\right)\nonumber\\
&-\frac T{\underline{\sigma}_i^2}\left\{
\underline{\bm P}^{-1}\underline{\bm\Lambda}^\prime(\underline{\bm\Sigma}^\xi)^{-1}\widehat{\bm\Gamma}^x(\underline{\bm\Sigma}^\xi)^{-1}
\otimes
\underline{\bm P}^{-1}
\right\} \text{vec}\left(\mathrm d\underline{\bm\Lambda}^\prime\right)\nonumber\\
&-\frac T{\underline{\sigma}_i^2}\left\{
\underline{\bm P}^{-1}
\otimes
\underline{\bm P}^{-1}\underline{\bm\Lambda}^\prime(\underline{\bm\Sigma}^\xi)^{-1}\widehat{\bm\Gamma}^x(\underline{\bm\Sigma}^\xi)^{-1}
\right\}\bm C_{n,r}\text{vec}\left(\mathrm d\underline{\bm\Lambda}^\prime\right)\nonumber\\
&-\frac T{\underline{\sigma}_i^2}\left\{\mathbf I_r
\otimes
\underline{\bm P}^{-1}\underline{\bm\Lambda}^\prime(\underline{\bm\Sigma}^\xi)^{-1}\widehat{\bm\Gamma}^x(\underline{\bm\Sigma}^\xi)^{-1}\underline{\bm\Lambda}
\right\}\text{vec}\left(\mathrm d\underline{\bm P}^{-1}\right).
\end{align}
And, since
\begin{equation}
\text{vec}\left(\mathrm d \bm h_{ii}^{*\prime} (\bm{\mathcal X};\underline{\bm\varphi})\right) =\left(
\frac{\partial \text{vec}(\bm h_{ii}^* (\bm{\mathcal X};\underline{\bm\varphi}))^\prime}
{\partial \text{vec}(\underline{\bm\Lambda})^\prime}
\right)^\prime \text{vec}\left(\mathrm d\underline{\bm\Lambda}^\prime\right),\nonumber
\end{equation}
from (ref) and (ref),
\begin{align}
\left(
\frac{\partial \text{vec}(\bm h_{ii}^* (\bm{\mathcal X};\underline{\bm\varphi}))^\prime}
{\partial \text{vec}(\underline{\bm\Lambda})^\prime}
\right)^\prime=&\,
\frac T{\underline{\sigma}_i^2}\left[
\underline{\bm P}^{-1}\underline{\bm\Lambda}^\prime(\underline{\bm\Sigma}^\xi)^{-1}\widehat{\bm\Gamma}^x(\underline{\bm\Sigma}^\xi)^{-1}\underline{\bm\Lambda}\,\underline{\bm P}^{-1}\underline{\bm\Lambda}^\prime(\underline{\bm\Sigma}^\xi)^{-1}
\otimes
\underline{\bm P}^{-1}
\right]\nonumber\\
&+\frac T{\underline{\sigma}_i^2}\left[
\underline{\bm P}^{-1}\underline{\bm\Lambda}^\prime(\underline{\bm\Sigma}^\xi)^{-1}\widehat{\bm\Gamma}^x(\underline{\bm\Sigma}^\xi)^{-1}\underline{\bm\Lambda}\,\underline{\bm P}^{-1}
\otimes \underline{\bm P}^{-1}\underline{\bm\Lambda}^\prime(\underline{\bm\Sigma}^\xi)^{-1}
\right]\bm C_{n,r}\nonumber\\
&-\frac T{\underline{\sigma}_i^2}\left[
\underline{\bm P}^{-1}\underline{\bm\Lambda}^\prime(\underline{\bm\Sigma}^\xi)^{-1}\widehat{\bm\Gamma}^x(\underline{\bm\Sigma}^\xi)^{-1}
\otimes
\underline{\bm P}^{-1}
\right]\nonumber\\
&-\frac T{\underline{\sigma}_i^2}\left[
\underline{\bm P}^{-1}
\otimes
\underline{\bm P}^{-1}\underline{\bm\Lambda}^\prime(\underline{\bm\Sigma}^\xi)^{-1}\widehat{\bm\Gamma}^x(\underline{\bm\Sigma}^\xi)^{-1}
\right]\bm C_{n,r}\nonumber\\
&+\frac T{\underline{\sigma}_i^2}\left[ \underline{\bm P}^{-1}\underline{\bm\Lambda}^\prime(\underline{\bm\Sigma}^\xi)^{-1}
\otimes
\underline{\bm P}^{-1}\underline{\bm\Lambda}^\prime(\underline{\bm\Sigma}^\xi)^{-1}\widehat{\bm\Gamma}^x(\underline{\bm\Sigma}^\xi)^{-1}\underline{\bm\Lambda}\,\underline{\bm P}^{-1}
\right]\nonumber\\
&+\frac T{\underline{\sigma}_i^2}\left[
\underline{\bm P}^{-1}
\otimes
\underline{\bm P}^{-1}\underline{\bm\Lambda}^\prime(\underline{\bm\Sigma}^\xi)^{-1}\widehat{\bm\Gamma}^x(\underline{\bm\Sigma}^\xi)^{-1}\underline{\bm\Lambda}\,\underline{\bm P}^{-1}\underline{\bm\Lambda}^\prime(\underline{\bm\Sigma}^\xi)^{-1}
\right]\bm C_{n,r},
\end{align}
which is an $r^2\times nr$ matrix. From (ref) we have
\begin{align}
\frac{\partial \text{vec}(\bm h^*_{ii}(\bm{\mathcal X};\underline{\bm\varphi}))^\prime}{\partial \underline{\bm\lambda}_i^\prime}=&\,
\frac {2T}{\underline{\sigma}_i^4}\left[
\underline{\bm P}^{-1}\underline{\bm\Lambda}^\prime(\underline{\bm\Sigma}^\xi)^{-1}\widehat{\bm\Gamma}^x(\underline{\bm\Sigma}^\xi)^{-1}\underline{\bm\Lambda}\,\underline{\bm P}^{-1}\underline{\bm\lambda}_i
\otimes
\underline{\bm P}^{-1}
\right]\nonumber\\
&+\frac {2T}{\underline{\sigma}_i^4}\left[
\underline{\bm P}^{-1}\underline{\bm\lambda}_i
\otimes
\underline{\bm P}^{-1}\underline{\bm\Lambda}^\prime(\underline{\bm\Sigma}^\xi)^{-1}\widehat{\bm\Gamma}^x(\underline{\bm\Sigma}^\xi)^{-1}\underline{\bm\Lambda}\,\underline{\bm P}^{-1}
\right]\nonumber\\
&-\frac {2T}{\underline{\sigma}_i^4}\left[
\underline{\bm P}^{-1}\underline{\bm\Lambda}^\prime(\underline{\bm\Sigma}^\xi)^{-1}[\widehat{\bm\Gamma}^x]_{\cdot i}
\otimes
\underline{\bm P}^{-1}
\right].
\end{align}
By using (ref) into (ref) and because of (ref), we have
\begin{equation}
\left\Vert \text{vec}\left(\bm t_{iii}(\bm{\mathcal X};\bm\varphi)\right)\right\Vert = \left\Vert \bm t_{iii}(\bm{\mathcal X};\bm\varphi)\right\Vert = O_{\mathrm P}\left(\frac T{ n}\right).
\end{equation}
Moreover, from (ref) it immediately follows that $\text{vec}\left(\bm T(\bm{\mathcal X}|\bm{\mathcal F};{\bm\varphi})\right) = \mathbf 0_{n^3r^3}$.
and, thus,
$\text{vec}\left(\bm t_{iii}(\bm{\mathcal X}|\bm{\mathcal F};{\bm\varphi})\right) = \mathbf 0_{r^3} $, or, equivalently,
\begin{equation}
\bm t_{iii}(\bm{\mathcal X}|\bm{\mathcal F};{\bm\varphi}) = \mathbf 0_{r\times r\times r},
\end{equation}
i.e., a 3rd order tensor of zeros.
Finally,
\begin{align}
\left\Vert
\bm h_{ii}(\bm{\mathcal X};\widehat{\bm{\varphi}}^{\text{\tiny QML,E}})-\bm h_{ii}(\bm{\mathcal X}|\bm{\mathcal F};\widehat{\bm{\varphi}}^{\text{\tiny QML,E}})
\right\Vert
\le&\, \left\Vert
\bm h_{ii}(\bm{\mathcal X};{\bm{\varphi}})-\bm h_{ii}(\bm{\mathcal X}|\bm{\mathcal F};{\bm{\varphi}})
\right\Vert\nonumber\\
&+\left\Vert
\bm t_{iii}(\bm{\mathcal X};{\bm{\varphi}})-\bm t_{iii}(\bm{\mathcal X}|\bm{\mathcal F};{\bm{\varphi}})
\right\Vert\, \left\Vert \widehat{\bm\lambda}_i^{\text{\tiny QML,E}}-{\bm\lambda}_i \right\Vert+O_{\mathrm P}\left(\left\Vert \widehat{\bm\lambda}_i^{\text{\tiny QML,E}}-{\bm\lambda}_i \right\Vert^2\right)\nonumber\\
=&\,O_{\mathrm P}\left(\max\left(\frac {T}n,\frac {\sqrt T}{\sqrt{n}}\right)\right)+O_{\mathrm P}\left(\frac {T}{n}\right)O_{\mathrm P}\left(\max\left(\frac {1}n,\frac {1}{\sqrt{T}}\right)\right),\nonumber
\end{align}
which follows from part (b), (ref), (ref), and (ref) in the proof of Corollary (ref), which, in turn, holds even if we use the log-likelihood (ref) in place of thel log-likelihood (ref). This completes the proof. $\Box$
\setcounter{equation}{0}
\section{Auxiliary propositions}
\begin{prop}
Under Assumptions (ref) and (ref),
\begin{compactenum}[(a)]
• $\frac{\bm\Lambda^\prime\bm\Lambda}{n} = \frac{\mathbf M^\chi}{n}$, for all $n\in\mathbb N$;
• $\bm\Lambda=\mathbf V^\chi(\mathbf M^\chi)^{1/2}$, for all $n\in\mathbb N$.
\end{compactenum}
\end{prop}
\textbf{Proof of Proposition (ref).} For part (a), first notice that Assumptions (ref)(c) and (ref)(b), imply $\bm\Gamma^F=\mathbf I_r$ which, in turn, implies $\bm\Gamma^\chi=\bm\Lambda\bm\Lambda^\prime$. Therefore, since the non-zero eigenvalues of $\frac{\bm\Gamma^\chi}n$ are the same as the $r$ eigenvalues of $\frac{\bm\Lambda^\prime\bm\Lambda}{n}$, which is diagonal by Assumption (ref)(a). Then, we must have, for all $n\in\mathbb N$, $\frac{\bm\Lambda^\prime\bm\Lambda}{n} = \frac{\mathbf M^\chi}{n}$. Equivalently, for a given $T$, from Assumption (ref)(b), we have $\widehat{\bm\Gamma}^\chi=\frac 1T \sum_{t=1}^T\bm\chi_t\bm\chi_t^\prime=\bm\Lambda\bm\Lambda^\prime$ which implies that $\frac{\bm\Lambda^\prime\bm\Lambda}{n} = \frac{\widehat{\mathbf M}^\chi}{n}$. However, since $\widehat{\bm\Gamma}^\chi = {\bm\Gamma}^\chi$ it must be that $\frac{\widehat{\mathbf M}^\chi}{n}=\frac{{\mathbf M}^\chi}{n}$. This proves part (a).
For part (b), since $\bm\Gamma^\chi=\mathbf V^\chi\mathbf M^\chi\mathbf V^{\chi\prime}$, it must be that
\begin{equation}
\bm\Lambda\mathbf K_*= \mathbf V^\chi(\mathbf M^\chi)^{1/2},
\end{equation}
for some $r\times r$ invertible $\mathbf K_*$. Now, since $\text{rk}(\frac{\bm\Lambda} {\sqrt n})=r$ for all $n>N$ (see the proof of Proposition (ref)):
\begin{equation}
\mathbf K_*= (\bm\Lambda^\prime\bm\Lambda)^{-1}\bm\Lambda^\prime\mathbf V^\chi(\mathbf M^\chi)^{1/2}=(\mathbf M^\chi)^{-1}\bm\Lambda^\prime\mathbf V^\chi(\mathbf M^\chi)^{1/2},
\end{equation}
which is also obtained by linear projection, and
\begin{equation}
\mathbf K^{-1}_*= (\mathbf M^\chi)^{-1/2}{\mathbf V^{\chi\prime}\bm\Lambda}.
\end{equation}
Notice that since $\mathbf K_*$ is a special case of the matrix $\mathbf K$ defined in (ref) in the proof of Proposition (ref), then $\mathbf K_*$
is finite and positive definite because of Lemma (ref)(i) and (ref)(ii), respectively, hence, $\mathbf K^{-1}_*$ in (ref) is well defined.
Moreover, from (ref), because of Assumption (ref)(b) and part (a):
\begin{align}
\mathbf K_*\mathbf K_*^\prime&=(\mathbf M^\chi)^{-1}\bm\Lambda^\prime\mathbf V^\chi\mathbf M^\chi\mathbf V^{\chi\prime}\bm\Lambda(\mathbf M^\chi)^{-1}\nonumber\\
&=(\mathbf M^\chi)^{-1}\bm\Lambda^\prime\bm\Gamma^\chi\bm\Lambda(\mathbf M^\chi)^{-1}\nonumber\\
&=(\mathbf M^\chi)^{-1}\bm\Lambda^\prime\bm\Lambda\bm\Lambda^\prime\bm\Lambda(\mathbf M^\chi)^{-1}\nonumber\\
&=\mathbf I_r.
\end{align}
So because of (ref), we have that $\mathbf K_*$ is an orthogonal matrix, i.e., $\mathbf K_*=\mathbf K_*^{-1}$.
Finally, by (ref) we also have
\begin{equation}
\mathbf V^\chi = \bm\Lambda\mathbf K_*(\mathbf M^\chi)^{-1/2}
\end{equation}
and by substituting (ref) into (ref), because of (ref),
\[
\mathbf K_*^{-1} = (\mathbf M^\chi)^{-1}\mathbf K_*^\prime \bm\Lambda^\prime \bm\Lambda=(\mathbf M^\chi)^{-1}\mathbf K_*^{-1}\bm\Lambda^\prime \bm\Lambda,
\]
which is equivalent to:
\begin{equation}
\mathbf K_*^{-1}\bm\Lambda^\prime \bm\Lambda \mathbf K_*= \mathbf M^\chi,
\end{equation}
and by part (a), we must have $\mathbf K_*=\mathbf I_r$. Alternatively, by multiplying on the right both sides of (ref) by their transposed:
\begin{equation}
\mathbf K_*^\prime\bm\Lambda^\prime \bm\Lambda \mathbf K_*= \mathbf M^\chi,
\end{equation}
since eigenvectors are normalized, and again by part (a), we must have $\mathbf K_*=\mathbf I_r$. So from (ref) or (ref) we prove part (b). This completes the proof. $\Box$
\begin{prop}
Under Assumptions (ref) through (ref), as $n,T\to\infty$
$\min(n,\sqrt T)\left\Vert\frac{\widehat{\bm\Lambda}-\bm\Lambda\bm{\mathcal H}}{\sqrt n}\right\Vert= O_{\mathrm P}(1)$,
where $\bm{\mathcal H}=(\bm\Lambda^\prime\bm\Lambda)^{-1}\bm\Lambda^\prime\mathbf V^\chi (\mathbf M^\chi)^{1/2}\mathbf J$
and $\mathbf J$ is an $r\times r$ diagonal matrix with entries $\pm 1$.
\end{prop}
\begin{proof}
Notice that $\text{rk}\left(\frac{\bm\Lambda}{\sqrt n}\right)=r$ for all $n$, since $\text{rk}(\bm\Gamma^F)=r$ by Assumption (ref)(b) and $\text{rk}\left(\frac{\bm\Gamma^\chi}n \right)=r$ by Lemma (ref)(iv).
Indeed, $\text{rk}\left(\frac{\bm\Gamma^\chi}n \right)\le \min (\text{rk}(\bm\Gamma^F),\text{rk}\left(\frac{\bm\Lambda}{\sqrt n}\right))$.
This holds for all $n>N$ and since eigenvalues are an increasing sequence in $n$. Therefore, $(\frac{\bm\Lambda^\prime\bm\Lambda}n)^{-1}$ is well defined for all $n$ and $\bm\Lambda$ admits a left inverse.
Second, since
\begin{equation}
\frac{\bm\Gamma^\chi}n = \mathbf V^\chi\frac{\mathbf M^\chi}n\mathbf V^{\chi\prime}= \frac{\bm\Lambda}{\sqrt n}\bm\Gamma^F\frac{\bm\Lambda^\prime}{\sqrt n}.
\end{equation}
the columns of $\frac{\mathbf V^\chi(\mathbf M^\chi)^{1/2}}{\sqrt n}$ and the columns of $\frac{\bm\Lambda(\bm\Gamma^F)^{1/2}}{\sqrt n}$ must span the same space.
So there exists an $r\times r$ invertible matrix $\mathbf K$ such that
\begin{equation}
{\bm\Lambda(\bm\Gamma^F)^{1/2}}\mathbf K= {\mathbf V^\chi(\mathbf M^\chi)^{1/2}}
\end{equation}
Therefore, from (ref)
\begin{equation}
\mathbf K=(\bm\Gamma^F)^{-1/2} (\bm\Lambda^\prime\bm\Lambda)^{-1}\bm\Lambda^\prime\mathbf V^\chi(\mathbf M^\chi)^{1/2}
\end{equation}
which is also obtained by linear projection, and also
\begin{equation}
\mathbf K^{-1}= (\mathbf M^\chi)^{-1/2}{\mathbf V^{\chi\prime}\bm\Lambda}(\bm\Gamma^F)^{1/2}
\end{equation}
which are both finite and positive definite because of Lemma (ref).
Now, from (ref) and (ref)
\begin{align}
{\mathbf V}^\chi ( {\mathbf M}^\chi)^{1/2}= \bm\Lambda(\bm\Gamma^F )^{1/2}\mathbf K
= \bm\Lambda(\bm\Lambda^\prime\bm\Lambda)^{-1}\bm\Lambda^\prime\mathbf V^\chi(\mathbf M^\chi)^{1/2}
\end{align}
Let
\begin{equation}
\bm{\mathcal H}=(\bm\Lambda^\prime\bm\Lambda)^{-1}\bm\Lambda^\prime\mathbf V^\chi(\mathbf M^\chi)^{1/2}\mathbf J,
\end{equation}
which is finite and positive definite because of Lemma (ref).
Now,
because of Lemmas (ref)(iii), (ref)(iv), (ref)(i), using (ref), (ref), and (ref),
\begin{align}
\left\Vert\frac { \widehat{\bm\Lambda} - \bm\Lambda\bm{\mathcal H}}{\sqrt n}\right\Vert=&\,
\left\Vert\widehat{\mathbf V}^x\left(\frac{\widehat{\mathbf M}^x}n\right)^{1/2} - {\mathbf V}^\chi \left(\frac{ {\mathbf M}^\chi}n\right)^{1/2} \mathbf J\right\Vert=\ \left\Vert\widehat{\mathbf V}^x\left(\frac{\widehat{\mathbf M}^x}n\right)^{1/2} - {\mathbf V}^\chi\mathbf J \left(\frac{ {\mathbf M}^\chi}n\right)^{1/2} \right\Vert\nonumber\\
\le&\, \left\Vert \widehat{\mathbf V}^x-{\mathbf V}^\chi\mathbf J\right\Vert\,\left\Vert\frac{\mathbf M^\chi}{n} \right\Vert+
\left\Vert\frac 1 {\sqrt n}\left\{\left(\widehat{\mathbf M}^x\right)^{1/2}-\left(\mathbf M^\chi\right)^{1/2}\right\}\right\Vert \,\left\Vert \mathbf V^\chi\right\Vert\nonumber\\
&+ \left\Vert \widehat{\mathbf V}^x-{\mathbf V}^\chi\mathbf J\right\Vert\,
\left\Vert\frac 1 {\sqrt n}\left\{\left(\widehat{\mathbf M}^x\right)^{1/2}-\left(\mathbf M^\chi\right)^{1/2}\right\}\right\Vert\nonumber\\
&= O_{\mathrm P}\left(\max\left(\frac 1 n,\frac 1{\sqrt T}\right)\right)+ O_{\mathrm P}\left(\max\left(\frac 1 {n^2},\frac 1{T}\right)\right), \nonumber
\end{align}
since $\Vert \mathbf V^\chi\Vert=1$ because eigenvectors are normalized.
This completes the proof.
\end{proof}
\begin{prop}
Under Assumptions (ref) through (ref) the terms in (ref) are such that, as $n,T\to\infty$,
\begin{compactenum}[(a)]
• $\sqrt {nT}\left\Vert \text{\upshape (1.a)}\right\Vert = O_{\mathrm P}(1)$;
• $\sqrt T \left\Vert \text{\upshape (1.b)}\right\Vert = O_{\mathrm P}(1)$;
• $\min(n,\sqrt{nT})\left\Vert \text{\upshape (1.c)}\right\Vert = O_{\mathrm P}(1)$;
• $\min(\sqrt{nT},T)\left\Vert \text{\upshape (1.d)}\right\Vert = O_{\mathrm P}(1)$;
• $\min(\sqrt{nT},T)\left\Vert \text{\upshape (1.e)}\right\Vert = O_{\mathrm P}(1)$;
• $\min(n,\sqrt{nT},T)\left\Vert \text{\upshape (1.f)}\right\Vert = O_{\mathrm P}(1)$;
\end{compactenum}
uniformly in $i$.
\end{prop}
\begin{proof}
For part (a), for any $i=1,\ldots,n$, by Assumption (ref)(a),
\begin{equation}
\left\Vert\frac 1{nT}{\bm\lambda}_i^\prime\sum_{t=1}^T\sum_{j=1}^n\mathbf F_t\xi_{jt}{\bm\lambda}_j^\prime\right\Vert \le \left\Vert \bm\lambda_i\right\Vert\,\left\Vert\frac 1{nT}\sum_{t=1}^T\sum_{j=1}^n\mathbf F_t\xi_{jt}{\bm\lambda}_j^\prime \right\Vert\le M_\Lambda\left\Vert\frac 1{nT}\sum_{t=1}^T\sum_{j=1}^n\mathbf F_t\xi_{jt}{\bm\lambda}_j^\prime \right\Vert.
\end{equation}
Then, by Assumptions (ref)(a) and (ref)
\begin{align}
\mathbb{E}\left[\left\Vert\frac 1{nT}\sum_{t=1}^T\sum_{j=1}^n\mathbf F_t\xi_{jt}{\bm\lambda}_j^\prime \right\Vert^2\right]
&\le \mathbb{E}\left[\left\Vert\frac 1{nT}\sum_{t=1}^T\sum_{j=1}^n\mathbf F_t\xi_{jt}{\bm\lambda}_j^\prime \right\Vert_F^2\right]
=\frac 1{n^2T^2} \sum_{k=1}^r\sum_{h=1}^r \mathbb{E}\left[\left(\sum_{t=1}^T F_{kt}\sum_{j=1}^n \xi_{jt}[\bm\Lambda_{jh}]\right)^2\right] \nonumber\\
&\le \frac{r^2}{n^2T^2} \max_{h,k=1,\ldots, r}\sum_{t=1}^T\sum_{s=1}^T \mathbb{E}\left[F_{kt} \left(\sum_{j=1}^n \xi_{jt}[\bm\Lambda_{jh}]\right) F_{ks}\left(\sum_{\ell=1}^n \xi_{\ell s}[\bm\Lambda_{\ell h}]\right) \right]\nonumber\\
&= \frac{r^2}{n^2T^2} \max_{h,k=1,\ldots, r}\sum_{t=1}^T\sum_{s=1}^T \mathbb{E}[F_{kt} F_{ks}]\,\sum_{j=1}^n\sum_{\ell=1}^n \mathbb{E}\left[\xi_{jt} \xi_{\ell s} \right] [\bm\Lambda_{jh}][\bm\Lambda_{\ell h}]\nonumber\\
&\le \left\{\frac{r^2}{T}\max_{k=1,\ldots, r}\sum_{t=1}^T\sum_{s=1}^T \mathbb{E}\left[F_{kt}F_{ks}\right]\right\}
\,\left\{\max_{t,s=1\ldots,T} \frac{M_{\Lambda}^2}{n^2T}\sum_{j=1}^n\sum_{\ell=1}^n\left\vert \mathbb{E}[\xi_{jt}\xi_{\ell s}]\right\vert\right\}.
\end{align}
Now, by Cauchy-Schwarz inequality, for any $k=1,\ldots, r$,
\begin{equation}
\left\vert \frac{1}{T}\sum_{t=1}^T\sum_{s=1}^TF_{kt}F_{ks}\right\vert\le \left(\frac 1 T\sum_{t=1}^T F_{kt}^2\right)^{1/2}\left(\frac 1 T\sum_{s=1}^T F_{ks}^2\right)^{1/2},
\end{equation}
and, by Assumption (ref)(b), using (ref) and again Cauchy-Schwarz inequality,
\begin{align}
\max_{k=1,\ldots, r}\frac{1}{T}\sum_{t=1}^T\sum_{s=1}^T\mathbb{E}[F_{kt}F_{ks}]&
=\max_{k=1,\ldots, r}\mathbb{E}\left[\frac{1}{T}\sum_{t=1}^T\sum_{s=1}^TF_{kt}F_{ks}\right]\le\max_{k=1,\ldots, r}\mathbb{E}\left[\left\vert \frac{1}{T}\sum_{t=1}^T\sum_{s=1}^TF_{kt}F_{ks}\right\vert\right]\nonumber\\
&\le \max_{k=1,\ldots, r} \mathbb{E}\left[\left(\frac 1 T\sum_{t=1}^T F_{kt}^2\right)^{1/2}\left(\frac 1 T\sum_{s=1}^T F_{ks}^2\right)^{1/2}\right]\nonumber\\
&\le \max_{k=1,\ldots, r}
\left(\mathbb{E}\left[\left(\frac 1 T\sum_{t=1}^T F_{kt}^2\right)\right] \right)^{1/2}
\left(\mathbb{E}\left[\left(\frac 1 T\sum_{s=1}^T F_{ks}^2\right)\right] \right)^{1/2}\nonumber\\
&\le \max_{k=1,\ldots, r}
\left(\frac 1 T\sum_{t=1}^T \mathbb{E}[F_{kt}^2] \right)^{1/2}
\left(\frac 1 T\sum_{s=1}^T \mathbb{E}[F_{ks}^2] \right)^{1/2}\nonumber\\
&=\max_{k=1,\ldots, r} \frac{1}{T}\sum_{t=1}^T \mathbb{E}[F_{kt}^2]\le \max_{t=1,\ldots,T}\max_{k=1,\ldots, r}\mathbb{E}[F_{kt}^2]\nonumber\\
&= \max_{k=1,\ldots, r} \bm\eta_k^\prime \bm\Gamma_F\bm\eta_k \le \Vert\bm\Gamma^F\Vert\le M_F,
\end{align}
since $M_F$ is independent of $t$ and where $\bm\eta_k$ is an $r$-dimensional vector with one in the $k$th entry and zero elsewhere. And, because of Lemma (ref)(ii)
\begin{equation}
\max_{t,s=1,\ldots,T}\frac 1{n^2T}\sum_{j=1}^n\sum_{\ell=1}^n\left\vert \mathbb{E}[\xi_{jt}\xi_{\ell s}]\right\vert\le \frac{M_\xi}{nT},
\end{equation}
since $M_\xi$ is independent of $n$ and $t$.
By substituting (ref) and (ref) into (ref),
\begin{equation}
\mathbb{E}\left[\left\Vert\sum_{t=1}^T\sum_{j=1}^n\mathbf F_t\xi_{jt}{\bm\lambda}_j^\prime \right\Vert^2\right] \le \frac{r^2M_FM_\Lambda^2M_\xi}{nT}.
\end{equation}
By substituting (ref) into (ref), we prove part (a).
For part (b), for any $i=1,\ldots,n$, because of Lemma (ref)(i),
\begin{equation}
\left\Vert\frac 1{nT} \sum_{t=1}^T \xi_{it}\mathbf F_t^\prime\sum_{j=1}^n\bm\lambda_j\bm\lambda_j^\prime\right\Vert=
\left\Vert\frac 1{nT} \sum_{t=1}^T \xi_{it}\mathbf F_t^\prime(\bm\Lambda^\prime\bm\Lambda)\right\Vert\le
\left\Vert\frac 1{T} \sum_{t=1}^T \xi_{it}\mathbf F_t\right\Vert\,\left\Vert
\frac{\bm\Lambda}{\sqrt n}
\right\Vert^2\le \left\Vert\frac 1{T} \sum_{t=1}^T \xi_{it}\mathbf F_t\right\Vert M_\Lambda^2.
\end{equation}
Then, by Assumptions (ref) and (ref)(b) and using (ref)
\begin{align}
\mathbb{E}\left[\left\Vert\frac 1{T} \sum_{t=1}^T \xi_{it}\mathbf F_t\right\Vert^2\right]&=\frac 1{T^2} \sum_{j=1}^r \mathbb{E}\left[\left(\sum_{t=1}^T \xi_{it}F_{jt}\right)^2\right]\nonumber\\
&\le \frac{r}{T^2}\max_{j=1,\ldots, r}\sum_{t=1}^T\sum_{s=1}^T \mathbb{E}[\xi_{it}F_{jt}\xi_{is}F_{js}]=
\frac{r}{T^2}\max_{j=1,\ldots, r}\sum_{t=1}^T\sum_{s=1}^T \mathbb{E}[F_{jt}F_{js}]\,\mathbb{E}[\xi_{it}\xi_{is}]\nonumber\\
&\le \left\{\frac{r}{T}\max_{j=1,\ldots, r}\sum_{t=1}^T\sum_{s=1}^T \mathbb{E}[F_{jt}F_{js}]\right\}
\left\{\frac 1T \max_{t,s=1,\ldots, n} \left\vert\mathbb{E}[\xi_{it}\xi_{is}]\right\vert\right\}\nonumber\\
&\le rM_F\frac{ M_{ii}}T \max_{t,s=1,\ldots, n}\rho^{|t-s|} \le \frac{rM_F M_{\xi}}T,
\end{align}
since $M_\xi$ is independent of $i$. Or, equivalently, by Lemma (ref)(iii) and Cauchy-Schwarz inequality
\begin{align}
\mathbb{E}\left[\left\Vert\frac 1{T} \sum_{t=1}^T \xi_{it}\mathbf F_t\right\Vert^2\right]
&\le \frac{r}{T^2}\max_{j=1,\ldots, r}\sum_{t=1}^T\sum_{s=1}^T \mathbb{E}[F_{jt}F_{js}]\,\mathbb{E}[\xi_{it}\xi_{is}]\nonumber\\
&\le \left\{\frac{r}{T}\max_{j=1,\ldots, r}\max_{t,s=1,\ldots, n} \vert\mathbb{E}[F_{jt}F_{js}]\vert\right\}
\left\{\frac 1T \sum_{t=1}^T\sum_{s=1}^T\left\vert\mathbb{E}[\xi_{it}\xi_{is}]\right\vert\right\}\nonumber\\
&\le \frac r T \max_{j=1,\ldots, r}\max_{t,s=1,\ldots, n} \mathbb{E}[F_{jt}^2] \frac{M_\xi(1+\rho)}{1-\rho}
\le \frac{r M_F M_{3\xi}}{T},
\end{align}
since $M_F$ is independent of $t$ and $M_{3\xi}$ is independent of $i$. Notice that $M_\xi\le M_{3\xi}$. Notice that (ref) is a special case of Lemma (ref)(i). By substituting (ref), or (ref), into (ref), we prove part (b).
For part (c), for any $i=1,\ldots,n$, because of Assumption (ref)(a),
\begin{align}
\left\Vert\frac 1{nT} \sum_{t=1}^T\sum_{j=1}^n\xi_{it}\xi_{jt} \bm\lambda_j^\prime\right\Vert &=\left\{\sum_{k=1}^r \left(\frac 1{nT} \sum_{t=1}^T\sum_{j=1}^n\xi_{it}\xi_{jt} \lambda_{jk}\right)^2\right\}^{1/2}\le \sqrt rM_\Lambda\left\vert\frac 1{nT} \sum_{t=1}^T\sum_{j=1}^n\xi_{it}\xi_{jt} \right\vert\nonumber\\
&\le\sqrt r M_\Lambda\left\{ \left\vert\frac 1{nT} \sum_{t=1}^T\sum_{j=1}^n\left\{\xi_{it}\xi_{jt}-\mathbb{E}[\xi_{it}\xi_{jt}]\right\} \right\vert+\left\vert\frac 1{nT} \sum_{t=1}^T\sum_{j=1}^n\mathbb{E}[\xi_{it}\xi_{jt}] \right\vert\right\}.
\end{align}
Then, by Assumption (ref)(b),
\begin{align}
\left\vert\frac 1{nT} \sum_{t=1}^T\sum_{j=1}^n\mathbb{E}[\xi_{it}\xi_{jt}] \right\vert\le
\frac 1{nT}\sum_{t=1}^T\sum_{j=1}^n\left\vert\mathbb{E}[\xi_{it}\xi_{jt}]\right\vert\le
\max_{t=1,\ldots,T}\frac 1n\sum_{j=1}^n\left\vert\mathbb{E}[\xi_{it}\xi_{jt}]\right\vert\le\frac 1n \sum_{j=1}^n M_{ij} \le \frac{M_\xi}n,
\end{align}
since $M_\xi$ is independent of $i$ and $t$. Moreover, by Assumption (ref)(c),
\begin{equation}
\mathbb{E}\left[\left\vert\frac 1{nT} \sum_{t=1}^T\sum_{j=1}^n\left\{\xi_{it}\xi_{jt}-\mathbb{E}[\xi_{it}\xi_{jt}]\right\} \right\vert^2\right]\le \frac{K_\xi}{nT}.
\end{equation}
By substituting (ref) and (ref) into (ref), we prove part (c).
For part (d), for any $i=1,\ldots,n$, because of Assumption (ref)(a)
\begin{align}
\left\Vert\frac 1{nT}{\bm\lambda}_i^\prime\sum_{t=1}^T\sum_{j=1}^n\mathbf F_t\xi_{jt}(\widehat{\bm\lambda}_j^\prime-{\bm\lambda}_j^\prime\bm{\mathcal H})\right\Vert &\le M_\Lambda\left\Vert\frac 1{nT}\sum_{t=1}^T\sum_{j=1}^n\mathbf F_t\xi_{jt}(\widehat{\bm\lambda}_j^\prime-{\bm\lambda}_j^\prime\bm{\mathcal H}) \right\Vert\nonumber\\
&= M_\Lambda\left\Vert\frac {\bm F^\prime\bm \Xi (\widehat{\bm\Lambda}-\bm\Lambda\bm{\mathcal H} )}{nT} \right\Vert\le M_\Lambda \left\Vert\frac {\bm F^\prime\bm \Xi}{\sqrt n T} \right\Vert\, \left\Vert\frac{\widehat{\bm\Lambda}-\bm\Lambda\bm{\mathcal H}}{\sqrt n}\right\Vert.
\end{align}
Then, by using Lemma (ref) and Proposition (ref) in (ref), we prove part (d).
For part (e), for any $i=1,\ldots,n$,
\begin{align}
\left\Vert\frac 1{nT} \sum_{t=1}^T \xi_{it}\mathbf F_t^\prime\sum_{j=1}^n\bm\lambda_j(\widehat{\bm\lambda}_j^\prime-{\bm\lambda}_j^\prime\bm{\mathcal H})\right\Vert&=
\left\Vert\frac 1{nT} \sum_{t=1}^T \xi_{it}\mathbf F_t^\prime\bm\Lambda^\prime(\widehat{\bm\Lambda}-\bm\Lambda\bm{\mathcal H})\right\Vert\nonumber\\
&\le \left\Vert\frac 1{T} \sum_{t=1}^T \xi_{it}\mathbf F_t\right\Vert\,\left\Vert
\frac{\bm\Lambda}{\sqrt n}
\right\Vert \, \left\Vert\frac{\widehat{\bm\Lambda}-\bm\Lambda\bm{\mathcal H}}{\sqrt n}\right\Vert .
\end{align}
By substituting part (ii), Lemma (ref)(i), and part (a) of Proposition (ref) into (ref), we prove part (e).
Finally, for part (f), for any $i=1,\ldots,n$, let $\bm\zeta_i=(\xi_{i1}\cdots \xi_{iT})^\prime$, then
\begin{align}
\left\Vert\frac 1{nT} \sum_{t=1}^T\sum_{j=1}^n\xi_{it}\xi_{jt} (\widehat{\bm\lambda}_j^\prime-{\bm\lambda}_j^\prime\bm{\mathcal H})\right\Vert &= \left\Vert\frac{\bm\zeta_i^\prime\bm\Xi\left(\widehat{\bm\Lambda}-\bm\Lambda\bm{\mathcal H}\right)}{nT} \right\Vert \le \left\Vert\frac{\bm\zeta_i^\prime\bm\Xi}{\sqrt n T}\right\Vert\, \left\Vert\frac{\widehat{\bm\Lambda}-\bm\Lambda\bm{\mathcal H}}{\sqrt n}\right\Vert.
\end{align}
Then, by the $C_r$-inequality with $r=2$,
\begin{align}
\left\Vert\frac{\bm\zeta_i^\prime\bm\Xi}{\sqrt n T}\right\Vert^2&=
\left\Vert\frac 1 {\sqrt n T} \sum_{t=1}^T \xi_{it} \bm\xi_t^\prime\right\Vert^2
=\frac 1{n} \sum_{j=1}^n \left(\frac 1T\sum_{t=1}^T \xi_{it}\xi_{jt}\right)^2\nonumber\\
& =\frac 1{n} \sum_{j=1}^n \left(\frac 1T\sum_{t=1}^T \left\{\xi_{it}\xi_{jt}-\mathbb{E}[\xi_{it}\xi_{jt}]\right\}+\frac 1T\sum_{t=1}^T \mathbb{E}[\xi_{it}\xi_{jt}]\right)^2\nonumber\\
&\le\frac 2{n} \sum_{j=1}^n\left\{ \left(\frac 1T\sum_{t=1}^T \left\{\xi_{it}\xi_{jt}-\mathbb{E}[\xi_{it}\xi_{jt}]\right\}\right)^2+\left(\frac 1T\sum_{t=1}^T \mathbb{E}[\xi_{it}\xi_{jt}]\right)^2\right\}
\end{align}
By taking the expectation of (ref), and because of Assumption (ref)(c) and Lemma (ref)(v),
\begin{align}
\mathbb{E}\left[\left\Vert\frac{\bm\zeta_i^\prime\bm\Xi}{\sqrt n T}\right\Vert^2\right]&=\frac 2{n} \sum_{j=1}^n \mathbb{E}\left[
\left(\frac 1T\sum_{t=1}^T \left\{\xi_{it}\xi_{jt}-\mathbb{E}[\xi_{it}\xi_{jt}]\right\}\right)^2
\right]+\frac 2{n} \sum_{j=1}^n \left(\frac 1T\sum_{t=1}^T \mathbb{E}[\xi_{it}\xi_{jt}]\right)^2\nonumber\\
&\le 2\max_{j=1,\ldots, n}\mathbb{E}\left[
\left(\frac 1T\sum_{t=1}^T \left\{\xi_{it}\xi_{jt}-\mathbb{E}[\xi_{it}\xi_{jt}]\right\}\right)^2
\right]+\frac{2}{n}\sum_{i=1}^n\left(\frac 1T\sum_{t=1}^T \mathbb{E}[\xi_{it}\xi_{jt}]\right)^2\nonumber\\
&\le \frac{2K_\xi}{T} + \frac{2}{nT^2}\sum_{i=1}^n\sum_{t=1}^T \mathbb{E}[\xi_{it}\xi_{jt}]\sum_{s=1}^T\mathbb{E}[\xi_{is}\xi_{js}]\nonumber\\
&\le \frac{2K_\xi}{T} + \max_{t=1,\ldots ,T}\frac{2}{n}\sum_{i=1}^n \left\vert\mathbb{E}[\xi_{it}\xi_{jt}]\right\vert \max_{s=1,\ldots,T}\max_{i,j=1,\ldots, n} \mathbb{E}[\xi_{is}\xi_{js}]\nonumber\\
&\le \frac{2K_\xi}{T} + \frac{2}n \sum_{i=1}^n M_{ij} \max_{i,j=1,\ldots, n} \bm\varepsilon_i^\prime\bm\Gamma^\xi\bm\varepsilon_j\le \frac{2K_\xi}{T}+ \frac{2M_\xi}{n}\Vert \bm\Gamma^\xi\Vert\le \frac{2K_\xi}{T}+ \frac{2M_\xi M_{2\xi}}{n},
\end{align}
since $K_\xi$ is independent of $j$ and $M_\xi$ is independent of $i$, $j$, $t$, and $s$ and where $\bm\varepsilon_i$ is an $n$-dimensional vector with one in the $i$th entry and zero elsewhere.
By substituting (ref) and part (a) of Proposition (ref) into (ref), we prove part (f). This completes the proof.
\end{proof}
\begin{prop}
Under Assumptions (ref) through (ref),
\begin{compactenum}[(a)]
• $
\left\Vert \bm{\mathcal H}^\prime\left(\frac{\bm\Lambda^\prime\bm\Lambda}{n}\right)-\bm Q_0\right\Vert = o(1)$, as $n\to\infty$;
• $\left\Vert \bm{\mathcal H}^\prime\left(\frac{\bm\Lambda^\prime\bm\Lambda}{n}\right)-\left(\frac{\widehat{\bm\Lambda}^\prime\bm\Lambda}{n}\right)\right\Vert = o_{\mathrm P}(1)$, as $n,T\to\infty$.
\end{compactenum}
where
$\bm Q_0=\bm V_0\bm{\mathcal J}_0\bm\Upsilon_0^\prime (\bm\Gamma^F)^{-1/2}$, and where
$\bm{\mathcal J}_0$ is an $r\times r$ diagonal matrix with entries $\pm 1$, $\bm\Upsilon_0$ is the $r\times r$ matrix having as columns the normalized eigenvectors of $(\bm\Gamma^F)^{1/2}\bm\Sigma_\Lambda(\bm\Gamma^F)^{1/2}$, and $\bm V_0$ is the $r\times r$ matrix of corresponding eigenvalues sorted in descending order.
\end{prop}
\begin{proof} Start with part (a). From (ref) and (ref)
\begin{equation}
\bm{\mathcal H}^\prime\left(\frac{\bm\Lambda^\prime\bm\Lambda}{n}\right) = \mathbf J(\mathbf M^\chi)^{1/2} \mathbf V^{\chi\prime}\bm\Lambda(\bm\Lambda^\prime\bm\Lambda)^{-1} \left(\frac{\bm\Lambda^\prime\bm\Lambda}{n}\right) =\mathbf J\left(\frac{\mathbf M^\chi}n\right)^{1/2} \frac{\mathbf V^{\chi\prime}\bm\Lambda}{\sqrt n}=\mathbf J\left(\frac{\mathbf M^\chi}n\right) \mathbf K^{-1}(\bm\Gamma^F)^{-1/2}.
\end{equation}
Then, from (ref)
\begin{equation}
\mathbf V^{\chi} = \bm\Lambda (\bm\Gamma^F)^{1/2}\mathbf K ({\mathbf M^\chi})^{-1/2}
\end{equation}
thus, from (ref) and (ref)
\begin{align}
\mathbf J\left(\frac{\mathbf M^\chi}n\right) \mathbf K^{-1} &=\mathbf J \left(\frac{\mathbf M^\chi}n\right) (\mathbf M^\chi)^{-1/2}{\mathbf V^{\chi\prime}\bm\Lambda}(\bm\Gamma^F)^{1/2}\nonumber\\
&=
\mathbf J\left(\frac{\mathbf M^\chi}n\right) ({\mathbf M^\chi})^{-1}\mathbf K^\prime(\bm\Gamma^F)^{1/2}\bm\Lambda^\prime \bm\Lambda(\bm\Gamma^F)^{1/2}\nonumber\\
&= \mathbf J\mathbf K^\prime(\bm\Gamma^F)^{1/2}\frac{\bm\Lambda^\prime \bm\Lambda}n(\bm\Gamma^F)^{1/2}.
\end{align}
And, from (ref)
\begin{align}
\mathbf J\mathbf K\mathbf K^\prime\mathbf J &= \mathbf J (\bm\Gamma^F)^{-1/2} (\bm\Lambda^\prime\bm\Lambda)^{-1}\bm\Lambda^\prime\mathbf V^\chi(\mathbf M^\chi)^{1/2}
(\mathbf M^\chi)^{1/2}\mathbf V^{\chi\prime}\bm\Lambda(\bm\Lambda^\prime\bm\Lambda)^{-1}(\bm\Gamma^F)^{-1/2}
\mathbf J\nonumber\\
&=\mathbf J (\bm\Gamma^F)^{-1/2} (\bm\Lambda^\prime\bm\Lambda)^{-1}\bm\Lambda^\prime\bm\Gamma^\chi\bm\Lambda(\bm\Lambda^\prime\bm\Lambda)^{-1}(\bm\Gamma^F)^{-1/2}
\mathbf J\nonumber\\
&=\mathbf J (\bm\Gamma^F)^{-1/2} (\bm\Lambda^\prime\bm\Lambda)^{-1}\bm\Lambda^\prime\bm\Lambda\bm\Gamma^F\bm\Lambda^\prime\bm\Lambda(\bm\Lambda^\prime\bm\Lambda)^{-1}(\bm\Gamma^F)^{-1/2}
\mathbf J\nonumber\\
&= \mathbf J (\bm\Gamma^F)^{-1/2} \bm\Gamma^F(\bm\Gamma^F)^{-1/2}
\mathbf J=\mathbf I_r.
\end{align}
Therefore, the columns of $\mathbf J\mathbf K$ are the normalized eigenvectors of $(\bm\Gamma^F)^{1/2}\frac{\bm\Lambda^\prime \bm\Lambda}n(\bm\Gamma^F)^{1/2}$ with eigenvalues $\frac{\mathbf M^\chi}n$ (notice that $\mathbf J\left(\frac{\mathbf M^\chi}n\right)=\left(\frac{\mathbf M^\chi}n\right)\mathbf J$). Moreover, by Assumption (ref)(a)
\begin{equation}
\lim_{n\to\infty} \left\Vert(\bm\Gamma^F)^{1/2}\frac{\bm\Lambda^\prime \bm\Lambda}n(\bm\Gamma^F)^{1/2}-(\bm\Gamma^F)^{1/2}\bm\Sigma_\Lambda(\bm\Gamma^F)^{1/2}\right\Vert=0.
\end{equation}
Letting, $\bm V_0$ be the matrix of eigenvalues of $(\bm\Gamma^F)^{1/2}\bm\Sigma_\Lambda(\bm\Gamma^F)^{1/2}$ sorted in descending order, from (ref) we also have (this is proved also in Lemma (ref)(i))
\begin{equation}
\lim_{n\to\infty} \left\Vert\frac{\mathbf M^\chi}n - \bm V_0\right\Vert=0.
\end{equation}
Let $\bm\Upsilon_0$ be the normalized eigenvectors of $(\bm\Gamma^F)^{1/2}\bm\Sigma_\Lambda(\bm\Gamma^F)^{1/2}$
Hence, by continuity of eigenvectors, and since the eigenvalues $\frac{\mathbf M^\chi}n$ are distinct because of Assumption (ref), from yu15, by Lemma (ref)(iii) and using (ref) it follows that
\begin{equation}
\lim_{n\to\infty} \left\Vert\mathbf J\mathbf K-\bm\Upsilon_0\bm{\mathcal J}_0\right\Vert \le \lim_{n\to\infty}\frac{2^{3/2}\sqrt r \left\Vert(\bm\Gamma^F)^{1/2}\frac{\bm\Lambda^\prime \bm\Lambda}n(\bm\Gamma^F)^{1/2}-(\bm\Gamma^F)^{1/2}\bm\Sigma_\Lambda(\bm\Gamma^F)^{1/2}\right\Vert}{\mu_r(\bm V_0)}=0,
\end{equation}
where $\bm{\mathcal J}_0$ is an $r\times r$ diagonal matrix with entries $\pm 1$, which is in general different from $\mathbf J$. Finally, since $\bm\Upsilon_0$ is an orthogonal matrix, we have $\bm\Upsilon_0^{-1}=\bm\Upsilon_0^\prime$, so from (ref)
\begin{equation}
\lim_{n\to\infty} \left\Vert\mathbf K^{-1}\mathbf J-\bm{\mathcal J}_0\bm\Upsilon_0^\prime\right\Vert =0.
\end{equation}
By using (ref) and (ref) into (ref) and since $\mathbf K^{-1}\mathbf J=\mathbf J\mathbf K^{-1}$, we have
\begin{equation}\nonumber
\lim_{n\to\infty}\left\Vert\bm{\mathcal H}^\prime\left(\frac{\bm\Lambda^\prime\bm\Lambda}{n}\right) -\bm V_0\bm{\mathcal J}_0\bm\Upsilon_0^\prime(\bm\Gamma^F)^{-1/2}\right\Vert=0.
\end{equation}
By defining $\bm Q_0=\bm V_0\bm{\mathcal J}_0\bm\Upsilon_0^\prime(\bm\Gamma^F)^{-1/2}$, we prove part (a).
Part (b) follows from Proposition (ref) and Lemma (ref)(i), and since
\begin{equation}
\left\Vert \frac{\widehat{\bm\Lambda}^\prime\bm\Lambda}{n}-\bm {\mathcal H}^\prime\frac{\bm\Lambda^\prime\bm\Lambda}{n}\right\Vert \le
\left\Vert \frac{\widehat{\bm\Lambda}^\prime-\bm {\mathcal H}^\prime\bm\Lambda^\prime}{\sqrt n}\right\Vert\,\left\Vert
\frac{\bm\Lambda}{\sqrt n}\right\Vert = o_{\mathrm P}(1).
\end{equation}
This completes the proof.
\end{proof}
\begin{prop}
Under Assumptions (ref) through (ref), as $n,T\to\infty$,
$\min(n,\sqrt{nT},T)\left\Vert \frac{(\bm\Lambda\widehat{\mathbf H}-\widehat{\bm\Lambda})^\prime\widehat{\bm\Lambda}}{n}\right\Vert=O_{\mathrm {P}}(1)$, where $\widehat{\mathbf H}$ is defined in (ref).
\end{prop}
\begin{proof}We have,
\begin{align}
\left\Vert \frac{(\bm\Lambda\widehat{\mathbf H}-\widehat{\bm\Lambda})^\prime\widehat{\bm\Lambda}}{n}\right\Vert&\le
\left\{\left\Vert
\frac{(\widehat{\bm\Lambda}-\bm\Lambda\widehat{\mathbf H})^\prime\bm\Lambda\widehat{\mathbf H}}{n}\right\Vert+\left\Vert \frac{(\widehat{\bm\Lambda}-\bm\Lambda\widehat{\mathbf H})^\prime (\widehat{\bm\Lambda}-\bm\Lambda\widehat{\mathbf H})}{n}\right\Vert\right\}\nonumber\\
&=\left\{\bm I +\bm{II}\right\},\;\text{say.}
\end{align}
First, consider $\bm I$ in (ref). From (ref)
\begin{align}
\bm I=&\,\frac {(\widehat{\bm\Lambda}-\bm\Lambda\widehat{\mathbf H})^\prime \bm\Lambda \widehat{\mathbf H}}n\nonumber\\
=&\, \left(\frac{\widehat{\mathbf M}^x}{n}\right)^{-1}\bm{\mathcal H}^\prime\frac{\bm\Lambda^\prime\bm\Lambda}{n}
\frac{\bm F^\prime\bm\Xi\bm\Lambda \widehat{\mathbf H}}{ nT}
+\left(\frac{\widehat{\mathbf M}^x}{n}\right)^{-1}\bm{\mathcal H}^\prime\frac{\bm\Lambda^\prime\bm\Xi^\prime\bm F}{nT}\frac{\bm\Lambda^\prime\bm\Lambda \widehat{\mathbf H}}{n} + \left(\frac{\widehat{\mathbf M}^x}{n}\right)^{-1}\bm{\mathcal H}^\prime \frac{\bm\Lambda^\prime\bm\Xi^\prime\bm\Xi\bm\Lambda \widehat{\mathbf H}}{n^2T}\nonumber\\
&+\left(\frac{\widehat{\mathbf M}^x}{n}\right)^{-1}\frac{(\widehat{\bm\Lambda}-\bm\Lambda\bm{\mathcal H})^\prime}{\sqrt n}\frac{\bm \Xi^\prime\bm F\bm\Lambda^\prime}{nT}\frac{\bm\Lambda \widehat{\mathbf H}}{\sqrt n}
+\left(\frac{\widehat{\mathbf M}^x}{n}\right)^{-1}\frac{(\widehat{\bm\Lambda}-\bm\Lambda\bm{\mathcal H})^\prime}{\sqrt n}\frac{\bm\Lambda\bm F^\prime\bm\Xi}{nT}\frac{\bm\Lambda \widehat{\mathbf H}}{\sqrt n}\nonumber\\
&+\left(\frac{\widehat{\mathbf M}^x}{n}\right)^{-1}\frac{(\widehat{\bm\Lambda}-\bm\Lambda\bm{\mathcal H})^\prime}{\sqrt n}\frac{\bm\Xi^\prime\bm \Xi}{nT}\frac{\bm\Lambda \widehat{\mathbf H}}{\sqrt n}
= \bm I_a+\bm I_b+\bm I_c+\bm I_d+\bm I_e+\bm I_f, \;\text{say.}
\end{align}
Then, because of (ref) and (ref) in the proof of Proposition (ref)(a),
\begin{align}
\mathbb{E}\left[\left\Vert \frac{\bm F^\prime\bm\Xi\bm\Lambda}{ nT}\right\Vert^2\right] =\frac 1{n^2T^2}\sum_{k=1}^r\sum_{h=1}^r \mathbb{E}\left[\left(\sum_{t=1}^T F_{kt}\sum_{j=1}^n \xi_{jt}\lambda_{jh}\right)^2\right] \le \frac{r^2M_FM_\Lambda^2M_\xi}{nT}.
\end{align}
Therefore, by Assumption (ref)(a), Lemma (ref)(iv), (ref)(i), (ref)(i),
and using (ref) , we get
\begin{align}
\Vert \bm I_a\Vert \le \left\Vert \left(\frac{\widehat{\mathbf M}^x}{n}\right)^{-1}\right\Vert\,\left\Vert \bm{\mathcal H}\right\Vert\,\left\Vert\frac{\bm\Lambda^\prime\bm\Lambda}{n}\right\Vert\,
\left\Vert\frac{\bm F^\prime\bm\Xi\bm\Lambda }{ nT}\right\Vert\,\left\Vert \widehat{\mathbf H}\right\Vert=O_{\mathrm P}\left(\frac 1{\sqrt {nT}}\right),\\
\Vert \bm I_b\Vert\le \left\Vert \left(\frac{\widehat{\mathbf M}^x}{n}\right)^{-1}\right\Vert\,\left\Vert \bm{\mathcal H}\right\Vert\,\left\Vert \frac{\bm\Lambda^\prime\bm\Xi^\prime\bm F}{nT}\right\Vert \,\left\Vert \frac{\bm\Lambda^\prime\bm\Lambda}n \right\Vert\,\left\Vert \widehat{\mathbf H}\right\Vert
=O_{\mathrm P}\left(\frac 1{\sqrt {nT}}\right).
\end{align}
Moreover, because of Lemma (ref)(i), (ref)(iii), (ref)(iv), (ref)(i), and (ref)(i),
\begin{align}
\Vert \bm I_c\Vert \le \left\Vert\left(\frac{\widehat{\mathbf M}^x}{n}\right)^{-1}\right\Vert\,\left\Vert\bm{\mathcal H}\right\Vert\,
\left\Vert\frac{\bm\Lambda}{\sqrt n}\right\Vert\,
\left\Vert\frac{\bm\Lambda^\prime\bm\Xi^\prime\bm\Xi }{n^{3/2}T}\right\Vert\,
\left\Vert\widehat{\mathbf H}\right\Vert=O_{\mathrm P}\left(\max\left(\frac 1n,\frac 1{\sqrt {nT}}\right)\right).
\end{align}
Similarly, because of Proposition (ref)(a), Lemma (ref)(i), (ref)(i), (ref)(iv), (ref)(iv), and (ref)(i),
\begin{align}
\Vert \bm I_d\Vert &\le \left\Vert\left(\frac{\widehat{\mathbf M}^x}{n}\right)^{-1}\right\Vert\,\left\Vert \frac{\widehat{\bm\Lambda}-\bm\Lambda\bm{\mathcal H}}{\sqrt n}\right\Vert \,\left\Vert \frac{\bm \Xi^\prime\bm F\bm}{\sqrt nT}\right\Vert\, \left\Vert \frac{\bm\Lambda^\prime \bm\Lambda }{n}
\right\Vert\, \left\Vert \widehat{\mathbf H}\right\Vert= O_{\mathrm P}\left(\max\left(\frac{1}{n\sqrt {T}},\frac 1{T}\right)\right), \\
\Vert \bm I_e\Vert &\le \left\Vert\left(\frac{\widehat{\mathbf M}^x}{n}\right)^{-1}\right\Vert\,\left\Vert \frac{\widehat{\bm\Lambda}-\bm\Lambda\bm{\mathcal H}}{\sqrt n}\right\Vert \,\left\Vert \frac{\bm \Xi^\prime\bm F\bm}{\sqrt nT}\right\Vert\, \left\Vert \frac{\bm\Lambda }{\sqrt n}
\right\Vert^2\, \left\Vert \widehat{\mathbf H}\right\Vert= O_{\mathrm P}\left(\max\left(\frac{1}{n\sqrt {T}},\frac 1{T}\right)\right),
\end{align}
and, because of Proposition (ref)(a), Lemma (ref)(i), (ref)(iii), (ref)(iv), and (ref)(i),
\begin{align}
\Vert \bm I_f\Vert \le \left\Vert\left(\frac{\widehat{\mathbf M}^x}{n}\right)^{-1}\right\Vert\,\left\Vert \frac{\widehat{\bm\Lambda}-\bm\Lambda\bm{\mathcal H}}{\sqrt n}\right\Vert \,\left\Vert \frac{\bm \Xi^\prime \bm\Xi}{nT}\right\Vert\, \left\Vert \frac{\bm\Lambda }{\sqrt n}
\right\Vert\, \left\Vert \widehat{\mathbf H}\right\Vert
= O_{\mathrm P}\left(\max\left(\frac 1{n^2},\frac{1}{n\sqrt {T}},\frac 1{T}\right)\right).
\end{align}
By using (ref), (ref), (ref), (ref), (ref), and (ref) into (ref)
\begin{equation}
\Vert\bm{I}\Vert = O_{\mathrm P}\left(\max\left(\frac 1 {n},\frac 1{\sqrt{nT}},\frac 1T\right)\right).
\end{equation}
Second, consider $\bm {II}$ in (ref). From Proposition (ref)(b),
\begin{equation}
\Vert\bm{II}\Vert \le \frac 1 {n}\left\Vert \widehat{\bm\Lambda}-{\bm\Lambda}\widehat{\mathbf H}\right\Vert^2
= O_{\mathrm P}\left(\max\left(\frac{1}{n^2},\frac 1{T}\right)\right).
\end{equation}
And, by using (ref) and (ref) in (ref), because of Lemma (ref)(ii), and Lemma (ref)(i), we complete the proof.
\end{proof}
\begin{prop}
Under Assumptions (ref) through (ref), and Assumption (ref), if $\sqrt T/n\to0$ and $\sqrt n/T\to 0$, as $n,T\to\infty$,
$\min(\sqrt{n},\sqrt T)\Vert \widehat{\mathbf H}-\bm J\Vert = o_{\mathrm P}(1)$;
where $\bm J$ is a diagonal $r\times r$ matrix with entries $\pm 1$.
\end{prop}
\begin{proof}
From Proposition (ref) and by using (ref)
\begin{equation}
\frac{\widehat{\bm\Lambda}^\prime\bm\Lambda\widehat{\mathbf H}}{n}=\frac{\widehat{\bm\Lambda}^\prime(\bm\Lambda\widehat{\mathbf H}-\widehat{\bm\Lambda}+\widehat{\bm\Lambda})}{n}=\frac{\widehat{\mathbf M}^x}{n}+O_{\mathrm P}\left(\max\left(\frac 1n,\frac 1{\sqrt {nT}}, \frac 1T\right)\right).
\end{equation}
Or, equivalently,
\begin{equation}
\left(\frac{\widehat{\mathbf M}^x}{n}\right)^{-1}\frac{\widehat{\bm\Lambda}^\prime\bm\Lambda}{n}=\widehat{\mathbf H}^{-1}+O_{\mathrm P}\left(\max\left(\frac 1n,\frac 1{\sqrt {nT}}, \frac 1T\right)\right).
\end{equation}
Moreover, by Assumption (ref) and using (ref) in (ref) we have
\begin{equation}
\widehat{\mathbf H}^\prime = \left(\frac{\widehat{\mathbf M}^x}{n}\right)^{-1}\frac{\widehat{\bm\Lambda}^\prime\bm\Lambda}{n}=\widehat{\mathbf H}^{-1}+O_{\mathrm P}\left(\max\left(\frac 1n,\frac 1{\sqrt {nT}}, \frac 1T\right)\right).
\end{equation}
Therefore, because of (ref), as $n,T\to\infty$, $\widehat{\mathbf H}$ is an $r\times r$ orthogonal matrix thus it has eigenvalues $\pm 1$.
Moreover, because of Proposition (ref)
\begin{equation}
\frac{\widehat{\bm\Lambda}^\prime\bm\Lambda}{n}= \frac{(\widehat{\bm\Lambda}-\bm\Lambda\widehat{\mathbf H}+\bm\Lambda\widehat{\mathbf H})^\prime\bm\Lambda}{n} =\frac{\widehat{\mathbf H}^\prime\bm\Lambda^\prime\bm\Lambda}{n}+ O_{\mathrm P}\left(\max\left(\frac 1n,\frac 1{\sqrt {nT}}, \frac 1T\right)\right).
\end{equation}
Thus, from (ref) and (ref) and by Lemma (ref)(iv),
\begin{equation}
\widehat{\mathbf H}^\prime = \left(\frac{\widehat{\mathbf M}^x}{n}\right)^{-1}\frac{\widehat{\bm\Lambda}^\prime\bm\Lambda}{n}=\left(\frac{\widehat{\mathbf M}^x}{n}\right)^{-1}\frac{\widehat{\mathbf H}^\prime\bm\Lambda^\prime\bm\Lambda}{n}+ O_{\mathrm P}\left(\max\left(\frac 1n,\frac 1{\sqrt {nT}}, \frac 1T\right)\right).
\end{equation}
And from (ref) it follows that
\begin{equation}
\left(\frac{\widehat{\mathbf M}^x}{n}\right)\widehat{\mathbf H}^\prime = \widehat{\mathbf H}^\prime\frac{\bm\Lambda^\prime\bm\Lambda}{n}+ O_{\mathrm P}\left(\max\left(\frac 1n,\frac 1{\sqrt {nT}}, \frac 1T\right)\right).
\end{equation}
So, because of (ref), as $n,T\to\infty$, the columns of $\widehat{\mathbf H}$ are the eigenvectors of $\frac{\bm\Lambda^\prime\bm\Lambda}{n}$ with eigenvalues $\frac{\widehat{\mathbf M}^x}{n}$. The eigenvectors are normalized since $\widehat{\mathbf H}$ is orthogonal, as $n,T\to\infty$. Moreover, under Assumption (ref), $\frac{\bm\Lambda^\prime\bm\Lambda}{n}$ is diagonal, so, as $n,T\to\infty$, $\widehat{\mathbf H}$ must be diagonal with eigenvalues $\pm 1$. By noticing that, if $\sqrt T/n\to0$ and $\sqrt n/T\to 0$, then $\max\left(\frac 1n,\frac 1{\sqrt {nT}}, \frac 1T\right)=o_{\mathrm P}\left(\max\left(\frac 1{\sqrt {n}}, \frac 1{\sqrt T}\right)\right)$, we complete the proof.
\end{proof}
\setcounter{equation}{0}
\section{Auxiliary lemmata}
\begin{lem}
Under Assumptions (ref) and (ref):
\begin{compactenum}
• for all $n\in\mathbb N$ and $T\in\mathbb N$, $\frac 1{nT}\sum_{i,j=1}^n\sum_{t,s=1}^T \vert\mathbb{E}_{}[\xi_{it}\xi_{js}]\vert \le M_{1\xi}$, for some finite positive real $M_{1\xi}$ independent of $n$ and $T$;
• for all $n\in\mathbb N$ and $t\in\mathbb Z$, $\frac 1{n}\sum_{i,j=1}^n \vert\mathbb{E}_{}[\xi_{it}\xi_{jt}]\vert \le M_{2\xi}$, for some finite positive real $M_{2\xi}$ independent of $n$ and $t$;
• for all $i\in\mathbb N$ and $T\in\mathbb N$, $\!\frac 1{T}\sum_{t,s=1}^T \vert\mathbb{E}_{}[\xi_{it}\xi_{is}]\vert \le M_{3\xi}$, $\!$for some finite positive real $M_{3\xi}$ independent of $i$ and $T$;
•
for all $j=1,\ldots,r$, $\underline C_j\!\le \lim\inf_{n\to\infty} \frac{ \mu_{j}^\chi}n\le\lim\sup_{n\to\infty} \frac{\mu_{j}^\chi}n\le\! \overline C_j$, $\!\!$
for some finite positive reals $\underline C_j$ and $\overline C_j$;
• for all $n\in\mathbb N$, $\mu_{1}^\xi \le M_{2\xi}$, where $M_\xi$ is defined in part (ii);
• for all $j=1,\ldots,r$, $\underline C_j\le \lim\inf_{n\to\infty} \frac{ \mu_{j}^x}n\le\lim\sup_{n\to\infty} \frac{\mu_{j}^x}n\le \overline C_j$, and for all $n\in\mathbb N$, $\mu_{r+1}^x \le M_\xi$, where $M_\xi$ is defined in Assumption (ref)(b).
\end{compactenum}
\end{lem}
\begin{proof} Using Assumptions (ref)(a) and (ref)(b), we have:
\begin{align}
\frac 1{nT}\sum_{i,j=1}^n\sum_{t,s=1}^T \vert\mathbb{E}_[\xi_{it}\xi_{js}]\vert &=\frac 1{n}\sum_{i,j=1}^n\sum_{k=-(T-1)}^{T-1} \left(1-\frac{\vert k\vert}{T}\right) \vert\mathbb{E}_[\xi_{it}\xi_{j,t-k}]\vert\nonumber\\
&\le \frac 1n\sum_{i=1}^n \sum_{k=-\infty}^{\infty} \rho^{|k|}\sigma_i^2 + \max_{j=1,\ldots, n}\sum_{j=1, j\ne i}^n\sum_{k=-\infty}^{\infty} \rho^{\vert k\vert} M_{ij}\nonumber\\
&\le \frac{C_\xi^\prime(1+\rho)}{1-\rho}+\frac{M_\xi(1+\rho)}{1-\rho}. \nonumber
\end{align}
Similarly,
\begin{align}
\frac 1{n}\sum_{i,j=1}^n \vert\mathbb{E}_[\xi_{it}\xi_{jt}]\vert &\le \frac 1n\sum_{i=1}^n \sigma_i^2 + \max_{j=1,\ldots,n}\sum_{j=1, j\ne i}^n M_{ij}\le C_\xi^\prime+ M_\xi,\nonumber
\end{align}
and
\begin{align}
\frac 1{T}\sum_{t,s=1}^T \vert\mathbb{E}_[\xi_{it}\xi_{is}]\vert &= \sum_{k=-(T-1)}^{T-1} \left(1-\frac{\vert k\vert}{T}\right) \vert\mathbb{E}_[\xi_{it}\xi_{i,t-k}]\vert\le \sum_{k=-\infty}^{\infty} \rho^{\vert k\vert} \sigma_i^2 \le \frac{C_\xi^\prime (1+\rho)}{1-\rho}.\nonumber
\end{align}
Defining, $M_{1\xi}=\frac{(C_\xi^\prime+M_\xi)(1+\rho)}{1-\rho}$ and $M_{2\xi}=C_\xi^\prime+M_\xi$, and $M_{3\xi}=\frac{C_\xi^\prime(1+\rho)}{1-\rho}$, we prove parts (i), (ii), and (iii).
For part (iv), by MK04, for all $j=1,\ldots, r$, we have
\begin{equation}
\frac{\mu_r(\bm\Lambda^\prime\bm\Lambda)}n\mu_j(\bm\Gamma^F) \le \frac{\mu_{j}^\chi}n \le \frac{\mu_j(\bm\Lambda^\prime\bm\Lambda)}n\ \mu_1(\bm\Gamma^F).
\end{equation}
The proof then follows from Assumption (ref)(a) which, by continuity of eigenvalues, implies that, for any $j=1,\ldots, r$, as $n\to\infty$
\[
\lim_{n\to\infty}\frac{\mu_j(\bm\Lambda^\prime\bm\Lambda)}n = \mu_j(\bm\Sigma_\Lambda).
\]
with
$$
0<m_\Lambda^2\le \mu_r(\bm\Sigma_\Lambda)\le\mu_1(\bm\Sigma_\Lambda)\le M_\Lambda^2<\infty ,$$
and by Assumption (ref)(b) and Assumption (ref)(c) which imply
$$
0<m_F\le \mu_r(\bm\Gamma^F)\le\mu_1(\bm\Gamma^F)\le M_F<\infty.
$$
For part (v), by Assumption (ref)(b):
\begin{align}
\Vert\bm\Gamma^\xi\Vert\le \max_{i=1,\ldots,n}\sum_{j=1}^n \vert \mathbb{E}[\xi_{it}\xi_{jt}]\vert \le
\max_{i=1,\ldots, n}\sigma_i^2+ \max_{i=1,\ldots,n}\sum_{j=1,j\ne i}^n M_{ij}\le C_\xi^\prime + M_{\xi}.\nonumber
\end{align}
Part (vi) follows from parts (iv) and (v) and Weyl's inequality. This completes the proof.
\end{proof}
\begin{lem}
Under Assumptions (ref) through (ref), for all $t=1,\ldots, T$ and all $n,T\in\mathbb N$
\begin{compactenum}[(i)]
• $\Vert\frac{\bm\Lambda}{\sqrt n}\Vert=O(1)$;
• $\Vert \mathbf F_t\Vert = O_{\mathrm {ms}}(1)$ and $\Vert\frac{\bm F}{\sqrt T}\Vert=O_{\mathrm {ms}}(1)$.
\end{compactenum}
\end{lem}
\begin{proof}
By Assumption (ref)(a), which holds for all $n\in\mathbb N$,
\begin{equation}
\sup_{n\in\mathbb N}\left\Vert\frac{\bm\Lambda}{\sqrt n}\right\Vert^2\le \sup_{n\in\mathbb N}\left\Vert\frac{\bm\Lambda}{\sqrt n}\right\Vert^2_F =\sup_{n\in\mathbb N}\frac 1n\sum_{j=1}^r\sum_{i=1}^n \lambda_{ij}^2 \le\sup_{n\in\mathbb N} \max_{i=1,\ldots, n} \Vert\bm\lambda_i\Vert^2\le M_\Lambda^2,\nonumber
\end{equation}
since $M_\Lambda$ is independent of $i$. This proves part (i).
By Assumption (ref)(b), which holds for all $T\in\mathbb N$ because of stationarity (see also part (i) of Lemma (ref)),
\begin{align}
\sup_{T\in\mathbb N}\max_{t=1,\ldots ,T}\mathbb{E}[\Vert\mathbf F_t\Vert^2]&=
\sup_{T\in\mathbb N} \max_{t=1,\ldots ,T}\sum_{j=1}^r \mathbb{E}[F_{jt}^2]\le
r\sup_{T\in\mathbb N}\max_{t=1,\ldots ,T}\max_{j=1,\ldots,r} \mathbb{E}[F_{jt}^2] \nonumber\\
&\le r\sup_{T\in\mathbb N}\max_{t=1,\ldots ,T}\max_{j=1,\ldots,r} \bm\eta_j^\prime\bm\Gamma^F\bm\eta_j\le r \Vert\bm\Gamma^F\Vert\le
r M_F,
\end{align}
since $M_F$ is independent of $t$ and where $\bm\eta_j$ is an $r$-dimensional vector with one in the $j$th entry and zero elsewhere. Therefore, from (ref):
\begin{align}
\sup_{T\in\mathbb N}\mathbb{E}\left[\left\Vert\frac{\bm F}{\sqrt T}\right\Vert^2\right]\le\sup_{T\in\mathbb N} \mathbb{E}\left[\left\Vert\frac{\bm F}{\sqrt T}\right\Vert^2_F\right] =\sup_{T\in\mathbb N}\frac 1T\sum_{j=1}^r\sum_{t=1}^T \mathbb{E}[[\bm F]_{tj}^2] \le\sup_{T\in\mathbb N} \max_{t=1,\ldots, T} \mathbb{E}[\Vert\mathbf F_t\Vert^2]
\le r M_F.\nonumber
\end{align}
This proves part (ii) and it completes the proof.
\end{proof}
\begin{lem}
Under Assumptions (ref) through (ref), for all $n,T\in\mathbb N$
$\sqrt T\left\Vert\frac{\bm F^\prime\bm \Xi}{\sqrt n T} \right\Vert = O_{\mathrm{ms}}(1)$.
\end{lem}
\begin{proof} By Assumption (ref), Lemma (ref)(iii) and Cauchy-Schwarz inequality
\begin{align}
\mathbb{E}\left[ \left\Vert\frac {\bm F^\prime\bm \Xi}{\sqrt n T} \right\Vert^2\right] &=\mathbb{E}\left[ \left\Vert\frac 1{\sqrt nT}\sum_{t=1}^T \mathbf F_t\bm\xi_t^\prime \right\Vert^2\right]\le \mathbb{E}\left[ \left\Vert\frac 1{\sqrt nT}\sum_{t=1}^T \mathbf F_t\bm\xi_t^\prime \right\Vert_F^2\right]\nonumber\\
&=\frac 1{nT^2}\sum_{j=1}^r\sum_{i=1}^n \mathbb{E}\left[\left(\sum_{t=1}^T F_{jt}\xi_{it}\right)^2\right]\nonumber\\
&\le\frac r{T^2}\max_{j=1\ldots, r} \max_{i=1\ldots, n} \sum_{t=1}^T\sum_{s=1}^T \mathbb{E}[\xi_{it}F_{jt}\xi_{is}F_{js}]\nonumber\\
&=\frac r{T^2}\max_{j=1\ldots, r} \max_{i=1\ldots, n} \sum_{t=1}^T\sum_{s=1}^T \mathbb{E}[F_{jt}F_{js}]\, \mathbb{E}[\xi_{it}\xi_{is}]\nonumber\\
&\le \left\{\frac{r}{T}\max_{j=1,\ldots, r}\max_{t,s=1,\ldots, n} \vert\mathbb{E}[F_{jt}F_{js}]\vert\right\}
\left\{\frac 1T \sum_{t=1}^T\sum_{s=1}^T\left\vert\mathbb{E}[\xi_{it}\xi_{is}]\right\vert\right\}\nonumber\\
&\le \frac r T \max_{j=1,\ldots, r}\max_{t,s=1,\ldots, n} \mathbb{E}[F_{jt}^2] \frac{M_\xi(1+\rho)}{1-\rho}\le\frac{rM_F M_{3\xi}}T,\nonumber
\end{align}
since $M_F$ is independent of $t$ and $s$ and $M_{3\xi}$ is independent of $i$. This completes the proof.
\end{proof}
\begin{lem}
Under Assumptions (ref) through (ref), for all $n,T\in\mathbb N$
\begin{compactenum}[(i)]
• $\min(n,\sqrt T)\left\Vert\frac{\bm \Xi^\prime\bm \Xi}{nT} \right\Vert = O_{\mathrm{ms}}(1)$;
• $\min(n,\sqrt {nT})\left\Vert\frac{\bm\Lambda^\prime\bm \Xi^\prime\bm \Xi}{n^{3/2}T} \right\Vert = O_{\mathrm{ms}}(1)$;
• $\left\Vert\frac{\bm \Lambda^\prime\bm \Lambda}{n} \right\Vert = O(1)$.
\end{compactenum}
\end{lem}
\begin{proof}
For part (i), first notice that, by Lemmas (ref)(v) and (ref)(ii)
\begin{equation}
\left\Vert\frac{\bm \Xi^\prime\bm \Xi}{nT} \right\Vert \le \left\Vert\frac{\bm \Xi^\prime\bm \Xi}{nT}-\frac{\bm\Gamma^\xi}{n} \right\Vert +\left\Vert\frac{\bm\Gamma^\xi}{n}\right\Vert
= \left\Vert\frac{\bm \Xi^\prime\bm \Xi}{nT}-\frac{\bm\Gamma^\xi}{n} \right\Vert +\frac{\mu_1^\xi}{n}=O_{\mathrm P}\left(\frac 1{\sqrt T}\right)+O\left(\frac 1{n}\right).
\end{equation}
Similarly, for part (ii) by Lemmas (ref)(i) and (ref)(v)
\begin{equation}
\left\Vert\frac{\bm\Lambda^\prime\bm \Xi^\prime\bm \Xi}{n^{3/2}T} \right\Vert \le \left\Vert\frac{\bm\Lambda^\prime\bm \Xi^\prime\bm \Xi}{n^{3/2}T}-\frac{\bm\Lambda^\prime\bm\Gamma^\xi}{n^{3/2}} \right\Vert +\left\Vert \frac{\bm\Lambda}{\sqrt n}\right\Vert\,\left\Vert\frac{\bm\Gamma^\xi}{n}\right\Vert=
\left\Vert\frac{\bm\Lambda^\prime\bm \Xi^\prime\bm \Xi}{n^{3/2}T}-\frac{\bm\Lambda^\prime\bm\Gamma^\xi}{n^{3/2}} \right\Vert +O\left(\frac 1{n}\right).
\end{equation}
Then, because of Assumption (ref)(c),
\begin{align}
\mathbb{E}\left[\left\Vert\frac{\bm\Lambda^\prime\bm \Xi^\prime\bm \Xi}{n^{3/2}T}-\frac{\bm\Lambda^\prime\bm\Gamma^\xi}{n^{3/2}} \right\Vert ^2\right]&\le
\mathbb{E}\left[\left\Vert\frac{\bm\Lambda^\prime\bm \Xi^\prime\bm \Xi}{n^{3/2}T}-\frac{\bm\Lambda^\prime\bm\Gamma^\xi}{n^{3/2}} \right\Vert ^2_F\right]\nonumber\\
&= \sum_{k=1}^r\sum_{j=1}^n\mathbb{E}\left[\left\vert\frac 1{n^{3/2}T}
\sum_{i=1}^n \sum_{t=1}^T \left\{\lambda_{ik}\xi_{it}\xi_{jt}-\lambda_{ik}\mathbb{E}[\xi_{it}\xi_{jt}]\right\}
\right\vert^2
\right]\nonumber\\
&\le \frac{rM_\Lambda n}{n^2T}
\max_{j=1,\ldots, n}
\mathbb{E}\left[\left\vert\frac 1{\sqrt{nT}}
\sum_{i=1}^n \sum_{t=1}^T \left\{\xi_{it}\xi_{jt}-\mathbb{E}[\xi_{it}\xi_{jt}]\right\}
\right\vert^2
\right]\le \frac{rM_\Lambda K_\xi}{nT},
\end{align}
since $K_\xi$ is independent of $j$. By using (ref) into (ref),
we prove part (ii).
Part (iii) follows from Lemma (ref)(i) since
\[
\left\Vert\frac{\bm \Lambda^\prime\bm \Lambda}{n} \right\Vert\le \left\Vert\frac{\bm \Lambda}{\sqrt n} \right\Vert^2\le M_\Lambda^2.
\]
This completes the proof.
\end{proof}
\begin{lem}
$\,$
\begin{compactenum}[(i)]
• Under Assumption (ref), for all $T\in\mathbb N$, $\sqrt T\left\Vert\frac{\bm F^\prime\bm F}{T} -\bm\Gamma^F\right\Vert = O_{\mathrm{ms}}(1)$;
• Under Assumption (ref), for all $n,T\in\mathbb N$, $\sqrt T \left\Vert\frac{\bm \Xi^\prime\bm \Xi}{nT}-\frac{\bm\Gamma^\xi}{n} \right\Vert= O_{\mathrm{ms}}(1)$.
\end{compactenum}
\end{lem}
\begin{proof}
Part (i) is direct consequence of Assumption (ref)(c-ii), since
\begin{align}
\mathbb{E}\left[\left\Vert\frac{\bm F^\prime\bm F}{T}-\bm\Gamma^F\right\Vert^2\right]&=
\mathbb{E}\left[\left\Vert \frac 1T\sum_{t=1}^T \mathbf F_t\mathbf F_t^\prime-\bm\Gamma^F\right\Vert^2\right]\le
\mathbb{E}\left[\left\Vert \frac 1T\sum_{t=1}^T \left\{\mathbf F_t\mathbf F_t^\prime-\bm\Gamma^F\right\}\right\Vert^2_F\right]\nonumber\\
&=\frac 1{T^2} \sum_{j=1}^r\sum_{k=1}^r\mathbb{E}\left[\left(\sum_{t=1}^T \left\{F_{jt}F_{kt}-\mathbb{E}[F_{jt}F_{kt}] \right\}\right)^2\right]
\le \frac {r^2C_F}{T},\nonumber
\end{align}
since $C_F$ is independent of $j$ and $k$.
For part (ii), by Assumption (ref)(c), letting $\gamma_{ij}^\xi$ be the $(i,j)$the entry of $\bm\Gamma^\xi$, we have
\begin{align}
\mathbb{E}\left[\left\Vert\frac{\bm \Xi^\prime\bm \Xi}{nT}-\frac{\bm\Gamma^\xi}{n} \right\Vert ^2\right]&=\mathbb{E}\left[\left\Vert \frac 1{nT} \sum_{t=1}^T\bm\xi_t\bm\xi_t^\prime -\frac{\bm\Gamma^\xi}n \right\Vert^2\right]\le \mathbb{E}\left[\left\Vert \frac 1{nT} \sum_{t=1}^T\left\{\bm\xi_t\bm\xi_t^\prime -\bm\Gamma^\xi\right\} \right\Vert^2_F\right]\nonumber\\
&=\frac 1{n^2T^2}\sum_{i,j=1}^n\mathbb{E}\left[\left(\sum_{t=1}^T \left\{\xi_{it}\xi_{jt}-\gamma_{ij}^\xi\right\}\right)^2\right]\nonumber\\
&\le \max_{i,j=1,\ldots, n}\frac 1{T^2} \mathbb{E}\left[\left(\sum_{t=1}^T \left\{\xi_{it}\xi_{jt}-\gamma_{ij}^\xi\right\}\right)^2\right]
\le \frac {K_\xi}{T},\nonumber
\end{align}
since $K_\xi$ is independent of $i$ and $j$. This completes the proof.
\end{proof}
\begin{lem}
Under Assumptions (ref) through (ref),
for all $n,T\in\mathbb N$
\begin{compactenum}
• $\frac {\sqrt T} n\Vert \widehat{\bm\Gamma}^x-\bm\Gamma^x\Vert = O_{\mathrm {ms}}(1)$;
• $\frac {\min(n,\sqrt T)} n\Vert \widehat{\bm\Gamma}^x-\bm\Gamma^\chi\Vert = O_{\mathrm {ms}}(1)$;
• $\frac {\min(n,\sqrt T)} n \Vert \widehat{\mathbf M}^x-\mathbf M^\chi\Vert=O_{\mathrm {ms}}(1)$;
• as $n\to\infty$, ${\min(n,\sqrt T)} \Vert\widehat{\mathbf V}^x-\mathbf V^\chi\mathbf J\Vert = O_{\mathrm {ms}}(1)$, where $\mathbf J$ is $r\times r$ diagonal with entries $\pm 1$.
\end{compactenum}
\end{lem}
\begin{proof}
For part (i), from Lemma (ref)(i), (ref)(ii) and (ref)(i):
\begin{align}
\mathbb{E}\left[\left\Vert\frac 1n\left( \widehat{\bm\Gamma}^x-\bm\Gamma^x\right)\right\Vert^2\right]
& =
\mathbb{E}\left[\left\Vert \frac 1n\left\{ \bm\Lambda\left( \frac 1T\sum_{t=1}^T \mathbf F_t\mathbf F_t^\prime-\bm\Gamma^F\right)\bm\Lambda^\prime
+ \frac 1T\sum_{t=1}^T\bm\xi_t\bm\xi_t^\prime-\bm\Gamma^\xi\right\}\right\Vert^2
\right]\nonumber\\
& \le
\mathbb{E}\left[\left\Vert \frac 1n\left\{ \bm\Lambda\left( \frac 1T\sum_{t=1}^T \mathbf F_t\mathbf F_t^\prime-\bm\Gamma^F\right)\bm\Lambda^\prime
\right\}\right\Vert^2\right]
+\mathbb{E}\left[\left\Vert \frac 1n\left( \frac 1T\sum_{t=1}^T\bm\xi_t\bm\xi_t^\prime-\bm\Gamma^\xi\right)\right\Vert^2
\right]\nonumber\\
&\le \left\Vert\frac{\bm\Lambda}{\sqrt n}\right\Vert^4 \, \mathbb{E}\left[\left\Vert \frac 1T\sum_{t=1}^T \mathbf F_t\mathbf F_t^\prime-\bm\Gamma^F\right\Vert^2\right]+\mathbb{E}\left[\left\Vert \frac 1n\left( \frac 1T\sum_{t=1}^T\bm\xi_t\bm\xi_t^\prime-\bm\Gamma^\xi\right)\right\Vert^2
\right]\nonumber\\
&\le \frac{M_\Lambda^4 r^2 C_F+K_\xi}{T}.\nonumber
\end{align}
This proves part (i).
Part (ii), follows from
\begin{align}
\mathbb{E}\left[\left\Vert\frac 1n\left( \widehat{\bm\Gamma}^x-\bm\Gamma^\chi\right)\right\Vert^2\right] &\le \mathbb{E}\left[\left\Vert\frac 1n\left( \widehat{\bm\Gamma}^x-\bm\Gamma^x\right)\right\Vert^2\right] +\left\Vert\frac {\bm\Gamma^\xi}n\right\Vert^2 \le \frac{M_\Lambda^4K_F+K_\xi}{T}+
\frac{\mu_1^\xi}{n^2}\le \frac{M_\Lambda^4K_F+K_\xi}{T}+\frac{M_{2\xi}}{n^2}\nonumber\\
&\le M_1\max \left(\frac 1{n^2},\frac 1T\right),\;\text{say,}\nonumber
\end{align}
because of part (i) and Lemma (ref)(v) and where $M_1$ is a finite positive real independent of $n$ and $T$.
For part (iii), for any $j=1,\ldots, r$, because of Weyl's inequality and part (ii), it holds that
\begin{equation}
\vert\widehat\mu_j^x-\mu_j^\chi \vert\le \mu_1(\widehat{\bm\Gamma}^x-\bm\Gamma^\chi)
=\left\Vert\widehat{\bm\Gamma}^x-\bm\Gamma^\chi\right\Vert.
\end{equation}
Hence, from (ref) and part (ii),
\begin{align}
\mathbb{E}\left[ \left\Vert\frac 1n\left( \widehat{\mathbf M}^x-\mathbf M^\chi\right)\right\Vert^2\right] &\le
\mathbb{E}\left[ \left\Vert\frac 1n\left( \widehat{\mathbf M}^x-\mathbf M^\chi\right)\right\Vert^2_F\right] =
\frac 1{n^2} \sum_{j=1}^r \mathbb{E}\left[\left(\widehat\mu_j^x-\mu_j^\chi \right)^2\right]\nonumber\\
&\le \frac{r}{n^2} \mathbb{E}\left[\left(\widehat\mu_1^x-\mu_1^\chi \right)^2\right]\le r\mathbb{E}\left[\left\Vert\frac 1n\left(\widehat{\bm\Gamma}^x-\bm\Gamma^\chi\right)\right\Vert^2\right]\le \frac{rM_\Lambda^4K_F+K_\xi}{T}+\frac{rM_\xi}{n^2}\nonumber\\
&\le rM_1\max\left(\frac 1{n^2},\frac 1T\right).
\end{align}
This proves part (iii).
For part (iv), because of yu15, which is a special case of Davis Kahn Theorem, there exists an $r\times r$ diagonal matrix $\mathbf J$ with entries $\pm 1$ such that
\begin{equation}
\Vert\widehat{\mathbf V}^x-\mathbf V^\chi\mathbf J\Vert\le \frac{2^{3/2}\sqrt r\left\Vert \widehat{\bm\Gamma}^x-\bm\Gamma^\chi\right\Vert }{\min(\vert \mu_{0}^\chi-\mu_{1}^\chi \vert,\vert \mu_{r}^\chi-\mu_{r+1}^\chi \vert)},
\end{equation}
where $\mu_0^\chi=\infty$. This holds provided the eigenvalues $\mu_j^\chi$ are distinct as required by Assumption (ref).
Therefore, from (ref), part (ii) and Lemma (ref)(iv), and since $\mu_{r+1}^\chi =0$, as $n\to\infty$
\begin{align}
\min(n^2,T) \mathbb{E}\left[\Vert\widehat{\mathbf V}^x-\mathbf V^\chi\mathbf J\Vert^2\right]&\le \frac
{\min(n^2,T) 2^3 \frac r{n^2} \mathbb{E}\left[\left\Vert\widehat{\bm\Gamma}^x-\bm\Gamma^\chi \right\Vert^2\right]}
{\frac 1{n^2}\left\{\min(\vert \mu_{0}^\chi-\mu_{1}^\chi \vert,\vert \mu_{r}^\chi-\mu_{r+1}^\chi \vert)\right\}^2} \le \frac{8rM_1}{\underline C_r^2}= M_2,\;\text{say,}\nonumber
\end{align}
where $M_2$ is a finite positive real independent of $n$ and $T$.
This proves part (iv). This completes the proof.
\end{proof}
\begin{lem}
Let $\bm\varepsilon_i$ be the $n$-dimensional vector with one in entry $i$ and zero elsewhere. Under Assumptions (ref) through (ref), for all $i=1,\ldots, n$ and $T\in\mathbb N$ and as $n\to\infty$
\begin{compactenum}[(i)]
• $\frac {\min(\sqrt n,\sqrt T)}{\sqrt n} \Vert\bm\varepsilon^\prime_i (\widehat{\bm\Gamma}^x-\bm\Gamma^\chi)\Vert = O_{\mathrm {ms}}(1)$;
• $\sqrt n\Vert\mathbf v_i^\chi\Vert = O(1)$;
• ${\min(\sqrt n,\sqrt T)} \sqrt n \Vert\widehat{\mathbf v}^{x\prime}_i-\mathbf v^{\chi\prime}\mathbf J\Vert = O_{\mathrm {ms}}(1)$.
\end{compactenum}
\end{lem}
\begin{proof}
First notice that
\begin{align}
\max_{i=1,\ldots, n}\mathbb{E}\left[\left\Vert \frac 1{\sqrt n}\bm\varepsilon^\prime_i\left( \sum_{t=1}^T\bm\xi_t\bm\xi_t^\prime -\bm\Gamma^\xi\right) \right\Vert^2\right]
&\le\max_{i=1,\ldots, n}\frac 1{n}\sum_{j=1}^n\mathbb{E}\left[\left(\frac 1T\sum_{t=1}^T \xi_{it}\xi_{jt}-[\bm\Gamma^\xi]_{ij}\right)^2\right]\nonumber\\
&\le \max_{i,j=1,\ldots, n} \mathbb{E}\left[\left(\frac 1T\sum_{t=1}^T \xi_{it}\xi_{jt}-[\bm\Gamma^\xi]_{ij}\right)^2\right]
\le \frac {K_\xi}{T},
\end{align}
since $K_\xi$ is independent of $i$ and $j$. Therefore, from Lemma (ref)(i) and Lemma (ref)(i), and using (ref),
\begin{align}
\max_{i=1,\ldots, n} \mathbb{E}\left[\left\Vert\frac 1{\sqrt n}\bm\varepsilon^\prime_i\left( \widehat{\bm\Gamma}^x-\bm\Gamma^x\right)\right\Vert^2\right]
=&\,
\max_{i=1,\ldots, n} \mathbb{E}\left[\left\Vert \frac 1{\sqrt n}\left\{ \bm\lambda_i^\prime \left( \frac 1T\sum_{t=1}^T \mathbf F_t\mathbf F_t^\prime-\bm\Gamma^F\right)\bm\Lambda^\prime
+\bm\varepsilon^\prime_i\left( \frac 1T\sum_{t=1}^T\bm\xi_t\bm\xi_t^\prime-\bm\Gamma^\xi\right)\right\}\right\Vert^2
\right]\nonumber\\
\le&\,\max_{i=1,\ldots, n}\Vert \bm\lambda_i\Vert^2\, \left\Vert\frac{\bm\Lambda}{\sqrt n}\right\Vert^2 \, \mathbb{E}\left[\left\Vert \frac 1T\sum_{t=1}^T \mathbf F_t\mathbf F_t^\prime-\bm\Gamma^F\right\Vert^2\right]\nonumber\\
&+\max_{i=1,\ldots, n}\mathbb{E}\left[\left\Vert \frac 1{\sqrt n}\bm\varepsilon^\prime_i\left( \frac 1T\sum_{t=1}^T\bm\xi_t\bm\xi_t^\prime-\bm\Gamma^\xi\right)\right\Vert^2
\right]\nonumber\\
\le&\, \frac{M_\Lambda^4 r^2 C_F+K_\xi}{T},
\end{align}
since $M_\Lambda$ and $C_F$ are independent of $i$. Then, following the same arguments as Lemma (ref)(ii), because of (ref) and Lemma (ref)(v):
\begin{align}
\max_{i=1,\ldots, n}\mathbb{E}\left[\left\Vert\frac 1{\sqrt n}\bm\varepsilon^\prime_i\left( \widehat{\bm\Gamma}^x-\bm\Gamma^\chi\right)\right\Vert^2\right]& \le\max_{i=1,\ldots, n} \mathbb{E}\left[\left\Vert\frac 1{\sqrt n}\bm\varepsilon^\prime_i\left( \widehat{\bm\Gamma}^x-\bm\Gamma^x\right)\right\Vert^2\right] +\max_{i=1,\ldots, n}\left\Vert\bm\varepsilon^\prime_i\frac {\bm\Gamma^\xi}{\sqrt n}\right\Vert^2\nonumber\\
&\le \frac{M_\Lambda^4r^2C_F+K_\xi}{T}+\max_{i=1,\ldots, n} \Vert \bm\varepsilon_i\Vert^2\, \left\Vert\frac{\bm\Gamma^\xi}{\sqrt n}\right\Vert^2\nonumber\\
&=\frac{M_\Lambda^4r^2C_F+K_\xi}{T}+ \frac{\mu_1^\xi}{n} \le \frac{M_\Lambda^4r^2C_F+K_\xi}{T}+ \frac{M_{2\xi}}{n}\le M_1\max\left(\frac 1T,\frac 1n\right),\; \text{say,}\nonumber
\end{align}
since $\Vert \bm\varepsilon_i\Vert=1$ and where $M_1$ is a finite positive real independent of $n$ and $T$ defined in Lemma (ref)(ii). This proves part (i).
For part (ii), notice that for all $i=1,\ldots, n$ we must have:
\begin{equation}
\mathbb{V}\mathrm{ar}(\chi_{it}) =\bm\lambda_i^\prime\bm\Gamma^F\bm\lambda_i\le \Vert\bm\lambda_i\Vert^2\,\left\Vert\bm\Gamma^F\right\Vert
\le M_\Lambda^2 M_F,
\end{equation}
which is finite for all $i$ and $t$. So, since by Lemma (ref)(iv)
\[
\liminf_{n\to\infty}\max_{i=1,\ldots, n}\mathbb{V}\mathrm{ar}(\chi_{it})=\liminf_{n\to\infty}\max_{i=1,\ldots, n}\sum_{j=1}^r \mu_j^\chi [\mathbf V^\chi]_{ij}^2\ge
\liminf_{n\to\infty}\mu_r^\chi \max_{i=1,\ldots, n}\sum_{j=1}^r [\mathbf V^\chi]_{ij}^2 \ge \underline C_r n \max_{i=1,\ldots, n}\Vert\mathbf v_i^\chi\Vert^2,
\]
then, because of (ref), we must have, as $n\to\infty$,
\[
\underline C_r n \max_{i=1,\ldots, n}\Vert\mathbf v_i^\chi\Vert^2\le M_\Lambda^2 M_F
\]
which implies that, as $n\to\infty$,
\[
n \max_{i=1,\ldots, n}\Vert\mathbf v_i^\chi\Vert^2 \le M_V,
\]
for some finite positive real $M_V$ independent of $n$. This proves part (ii).
Finally, using the same arguments in Lemma (ref)(iv), from (ref), part (i) and Lemma (ref)(iv), and since $\mu_0^\chi=\infty$ and $\mu_{r+1}^\chi =0$, as $n\to\infty$
\begin{align}
\max_{i=1,\ldots, n}
\min(n,T)\mathbb{E}\left[\Vert\sqrt n(\widehat{\mathbf v}_i^{x\prime}-\mathbf v_i^{\chi\prime}\mathbf J)\Vert^2\right]
&=\max_{i=1,\ldots, n}\min(n,T)\mathbb{E}\left[\Vert\sqrt n\bm\varepsilon_i^\prime(\widehat{\mathbf V}^x-\mathbf V^\chi\mathbf J)\Vert^2\right]\nonumber\\
&\le\max_{i=1,\ldots, n}\ \frac
{\min(n,T) 2^3 \frac r{n^2} n \mathbb{E}\left[\left\Vert\bm\varepsilon_i^\prime(\widehat{\bm\Gamma}^x-\bm\Gamma^\chi) \right\Vert^2\right]}
{\frac 1{n^2}\left\{\min(\vert \mu_{0}^\chi-\mu_{1}^\chi \vert,\vert \mu_{r}^\chi-\mu_{r+1}^\chi \vert)\right\}^2}\nonumber\\
&=\max_{i=1,\ldots, n}\ \frac
{\min(n,T) 2^3 \frac r{n} \mathbb{E}\left[\left\Vert\bm\varepsilon_i^\prime(\widehat{\bm\Gamma}^x-\bm\Gamma^\chi) \right\Vert^2\right]}
{\frac 1{n^2}\left\{\min(\vert \mu_{0}^\chi-\mu_{1}^\chi \vert,\vert \mu_{r}^\chi-\mu_{r+1}^\chi \vert)\right\}^2}\nonumber\\
&\le \frac{8r M_1}{\underline C_r^2}= M_2,\nonumber
\end{align}
where $M_2$ is a finite positive real independent of $n$ and $T$ defined in Lemma (ref)(iv).
This proves part (iii) and completes the proof.
\end{proof}
\begin{lem} Under Assumptions (ref) through (ref), for all $n,T\in\mathbb N$,
\begin{compactenum}[(i)]
• $\Vert\frac{{\mathbf M}^\chi}{n}\Vert =O(1)$;
• $\Vert(\frac{{\mathbf M}^\chi}{n})^{-1}\Vert=O(1)$;
• $\Vert\frac{\widehat{\mathbf M}^x}{n}\Vert=O_{\mathrm{P}}(1)$;
• $\Vert(\frac{\widehat{\mathbf M}^x}{n})^{-1}\Vert=O_{\mathrm{P}}(1)$.
\end{compactenum}
\end{lem}
\begin{proof}
Parts (i) and (ii) follow directly from Lemma (ref)(iv), indeed,
\[
\left\Vert\frac{{\mathbf M}^\chi}{n}\right\Vert =\frac{\mu_1^\chi}{n}\le \overline C_1,
\]
and
\[
\left\Vert\left(\frac{{\mathbf M}^\chi}{n}\right)^{-1}\right\Vert = \frac n{\mu_r^\chi}\le \frac 1{\underline C_r}.
\]
Both statements hold for all $n\in\mathbb N$ since the eigenvalues are an increasing sequence in $n$.
For part (iii), because of part (i) and Lemma (ref)(iii),
\begin{align}
\left\Vert\frac{\widehat{\mathbf M}^x}{n}\right\Vert\le
\left\Vert\frac{{\mathbf M}^\chi}{n}\right\Vert+
\left\Vert\frac{\widehat{\mathbf M}^x}{n}-\frac{{\mathbf M}^\chi}{n}\right\Vert\le \overline C_1+O_{\mathrm P}\left(\max\left(\frac1{\sqrt n},\frac 1{\sqrt T}\right)\right).\nonumber
\end{align}
For part (iv) just notice that, because of Lemma (ref)(iii) and part (ii), then $\frac{\widehat{\mathbf M}^x}{n}$ is positive definite with probability tending to one as $n,T\to\infty$.
This completes the proof.
\end{proof}
\begin{lem} Under Assumptions (ref) through (ref),
\begin{compactenum}[(i)]
• $\Vert\frac{\mathbf M^\chi}{n}-\bm V_0 \Vert=o(1)$, as $n\to\infty$ and $\Vert \bm V_0\Vert = O(1)$;
• $\Vert\frac{\widehat{\mathbf M}^x}{n}-\bm V_0 \Vert=o_{\mathrm {ms}}(1)$, as $n,T\to\infty$;
• $\Vert(\frac{\mathbf M^\chi}{n})^{-1}-\bm V_0^{-1} \Vert=o(1)$, as $n\to\infty$ and $\Vert \bm V_0^{-1}\Vert = O(1)$
• $\Vert(\frac{\widehat{\mathbf M}^x}{n})^{-1}-\bm V_0^{-1} \Vert=o_{\mathrm {ms}}(1)$, as $n,T\to\infty$,
\end{compactenum}
where $\bm V_0$ is $r\times r$ diagonal with entries the eigenvalues of $\bm\Sigma_\Lambda\bm\Gamma^F$ sorted in descending order.
\end{lem}
\begin{proof}
For part (i), first notice that the $r$ non-zero eigenvalues of $\frac{\bm\Gamma^\chi}{n}$ are also the $r$ eigenvalues of $(\bm\Gamma^F)^{1/2}\frac{\bm\Lambda^\prime\bm\Lambda}{n}(\bm\Gamma^F)^{1/2}$ which in turn are also the entries of $\bm V_0$. The proof follows from continuity of eigenvalues and since, because of Assumption (ref)(a), as $n\to\infty$,
\[
\left\Vert(\bm\Gamma^F)^{1/2}\frac{\bm\Lambda^\prime\bm\Lambda}{n}(\bm\Gamma^F)^{1/2}
-(\bm\Gamma^F)^{1/2}\bm\Sigma_\Lambda(\bm\Gamma^F)^{1/2}
\right\Vert=o(1).\nonumber
\]
Part (ii) is a consequence of part (i) and Lemma (ref)(iii).
For part (iii) notice that
\[
\left\Vert\left(\frac{\mathbf M^\chi}{n}\right)^{-1}-\bm V_0^{-1} \right\Vert \le\left \Vert\left(\frac{\mathbf M^\chi}{n}\right)^{-1}\right\Vert\,
\left\Vert\frac{\mathbf M^\chi}{n}-\bm V_0 \right\Vert
\, \left\Vert
\bm V_0^{-1}
\right\Vert,
\]
then the proof follows from part (i), Lemma (ref)(ii), and since $\bm V_0$ is positive definite since $\bm\Sigma_\Lambda$ and $\bm\Gamma^F$ are positive definite by Assumptions (ref)(a) and (ref)(b), respectively. This completes the proof.
\end{proof}
\begin{lem} Under Assumptions (ref) through (ref), as $n\to\infty$,
\begin{compactenum}[(i)]
• $\Vert{\mathbf K}\Vert=O(1)$;
• $\Vert{\mathbf K}^{-1}\Vert=O(1)$.
\end{compactenum}
\end{lem}
\begin{proof}
For part (a), from (ref) in the proof of Proposition (ref)(a) it follows that
\begin{equation}
\left\Vert\mathbf K-\mathbf J\bm\Upsilon_0\right\Vert =o(1),
\end{equation}
where $\bm\Upsilon_0$ is the $r\times r$ matrix of normalized eigenvectors of $(\bm\Gamma^F)^{1/2}\bm\Sigma_\Lambda(\bm\Gamma^F)^{1/2}$ and $\mathbf J$ is a diagonal matrix with entries $\pm 1$. Part (i) follows from the fact that $\Vert \mathbf J\bm\Upsilon_0\Vert = O(1)$, since $(\bm\Gamma^F)^{1/2}\bm\Sigma_\Lambda(\bm\Gamma^F)^{1/2}$ is finite. Likewise part (ii) follows from the fact that $\mathbf J$ is obviously positive definite and $\bm\Upsilon_0$ is also positive definite because the eigenvalues of $(\bm\Gamma^F)^{1/2}\bm\Sigma_\Lambda(\bm\Gamma^F)^{1/2}$ are distinct by Assumption (ref). This completes the proof.
\end{proof}
\begin{lem} Under Assumptions (ref) through (ref), as $n\to\infty$,
$\Vert{\bm{\mathcal H}}\Vert=O(1)$.
\end{lem}
\begin{proof}
From (ref) and (ref) in the proof of Proposition (ref) $\bm{\mathcal H}=(\bm\Gamma^F )^{1/2}\mathbf K\mathbf J$. Then, the proof follows immediately from Assumption (ref)(b), Lemma (ref)(i), and since $\mathbf J$ is obviously finite and positive definite.
\end{proof}
\begin{lem} Under Assumptions (ref) through (ref), as $n,T\to\infty$,
$\Vert\widehat{\mathbf H}\Vert=O_{\mathrm {P}}(1)$.
\end{lem}
\begin{proof}
From (ref), by Proposition (ref)(a), Lemma (ref)(v), (ref)(iv) (ref)(i), we have
\begin{align}
\left\Vert\widehat{\mathbf H}\right\Vert &\le \left\Vert\frac{\bm F^\prime\bm F}{T}\right\Vert\,
\left\Vert\frac{\bm\Lambda^\prime\widehat{\bm\Lambda}}{n}\right\Vert\,
\left\Vert\left(\frac{\widehat{\mathbf M}^x}{n}\right)^{-1}\right\Vert\le \left\Vert\frac{\bm F}{\sqrt T}\right\Vert^2\, \left\Vert\frac{\bm\Lambda}{\sqrt n}\right\Vert\, \left\Vert\frac{\widehat{\bm\Lambda}}{\sqrt n}\right\Vert\,\left\Vert\left(\frac{\widehat{\mathbf M}^x}{n}\right)^{-1}\right\Vert\nonumber\\
&\le\left\Vert\frac{\bm F}{\sqrt T}\right\Vert^2\, \left\Vert\frac{\bm\Lambda}{\sqrt n}\right\Vert^2\, \left\Vert\bm{\mathcal H}\right\Vert\, \left\Vert\left(\frac{\widehat{\mathbf M}^x}{n}\right)^{-1}\right\Vert+ \left\Vert\frac{\bm F}{\sqrt T}\right\Vert^2\, \left\Vert\frac{\bm\Lambda}{\sqrt n}\right\Vert\, \left\Vert\frac{\widehat{\bm\Lambda}-\bm\Lambda\bm{\mathcal H}}{\sqrt n}\right\Vert\,\left\Vert\left(\frac{\widehat{\mathbf M}^x}{n}\right)^{-1}\right\Vert\nonumber\\
&=
O_{\mathrm P}(1)+ O_{\mathrm P}\left(\max\left(\frac 1{\sqrt n},\frac 1{\sqrt T}\right)\right),\nonumber
\end{align}
which completes the proof.
\end{proof}
\begin{lem}
Under Assumptions (ref) through (ref) and letting $\bm\Sigma^\xi$ be the $n\times n$ diagonal matrix with diagonal entries the diagonal entries of $\bm\Gamma^\xi$, as $n\to\infty$,
$$
n^{-1}\left\Vert\{\mathbf I_r+{\bm\Lambda}^\prime({\bm\Sigma}^\xi)^{-1} {\bm\Lambda}\}^{-1}\left\{{\bm\Lambda}^\prime({\bm\Sigma}^\xi)^{-1}{\bm\Lambda}\right\}-\mathbf I_r\right\Vert= O\left(1\right).
$$
\end{lem}
\begin{proof}
First notice that for any two invertible matrices $\bm K$ and $\bm H$ we have
\begin{align}
(\bm H+ \bm K)^{-1}&= (\bm H + \bm K)^{-1}- \bm K^{-1} + \bm K^{-1}= (\bm H+ \bm K)^{-1}(\bm K - (\bm H + \bm K))\bm K^{-1}+ \bm K^{-1}\nonumber\\
&= (\bm H + \bm K)^{-1}(-\bm H)\bm K^{-1} + \bm K^{-1}= \bm K^{-1}- (\bm H + \bm K)^{-1}\bm H\bm K^{-1}.
\end{align}
Then, setting $\bm K=\bm \Lambda^\prime(\bm\Sigma^\xi)^{-1}\bm\Lambda $ and $\bm H=\mathbf I_r$ from (ref) we have
\begin{align}
\{\mathbf I_r+{\bm\Lambda}^\prime({\bm\Sigma}^\xi)^{-1} {\bm\Lambda}\}^{-1}=\{\bm \Lambda^\prime(\bm\Sigma^\xi)^{-1}\bm\Lambda \}^{-1}-\{\mathbf I_r+{\bm\Lambda}^\prime({\bm\Sigma}^\xi)^{-1} {\bm\Lambda}\}^{-1}\{\bm \Lambda^\prime(\bm\Sigma^\xi)^{-1}\bm\Lambda \}^{-1}
.
\end{align}
which implies
\begin{equation}
\{\mathbf I_r+{\bm\Lambda}^\prime({\bm\Sigma}^\xi)^{-1} {\bm\Lambda}\}^{-1}\{\bm \Lambda^\prime(\bm\Sigma^\xi)^{-1}\bm\Lambda \} = \mathbf I_r-\{\mathbf I_r+{\bm\Lambda}^\prime({\bm\Sigma}^\xi)^{-1} {\bm\Lambda}\}^{-1}.
\end{equation}
Then, by Weyl's inequality:
\begin{equation}
\nu_r(\mathbf I_r+{\bm\Lambda}^\prime({\bm\Sigma}^\xi)^{-1} {\bm\Lambda}) \ge 1+\nu_r({\bm\Lambda}^\prime({\bm\Sigma}^\xi)^{-1} {\bm\Lambda})\ge \nu_r({\bm\Lambda}^\prime({\bm\Sigma}^\xi)^{-1} {\bm\Lambda}).
\end{equation}
From, (ref), we have
\begin{align}
\Vert \{\mathbf I_r+{\bm\Lambda}^\prime({\bm\Sigma}^\xi)^{-1} {\bm\Lambda}\}^{-1}\Vert&=\frac 1{\nu_r(\mathbf I_r+{\bm\Lambda}^\prime({\bm\Sigma}^\xi)^{-1} {\bm\Lambda})}\le \frac 1{\nu_r({\bm\Lambda}^\prime({\bm\Sigma}^\xi)^{-1} {\bm\Lambda})}.
\end{align}
Now, the $r$ eigenvalues of ${\bm\Lambda}^\prime({\bm\Sigma}^\xi)^{-1} {\bm\Lambda}$ are also the $r$ non-zero eigenvalues of $\bm \Lambda\bm \Lambda^\prime \bm ({\bm\Sigma}^\xi)^{-1}$, and the $r$ non-zero eigenvalues of $\bm \Lambda\bm \Lambda^\prime$ are the $r$ eigenvalues of $\bm \Lambda^\prime\bm \Lambda$.
Therefore, because of MK04 and Proposition (ref)(a), and Lemma (ref)(iv):
\begin{align}
\nu_r({\bm\Lambda}^\prime({\bm\Sigma}^\xi)^{-1} {\bm\Lambda})&\ge \nu_r (\bm \Lambda\bm \Lambda^\prime)\nu_r(({\bm\Sigma}^\xi)^{-1})=\frac {\nu_r (\bm \Lambda^\prime\bm \Lambda)}{\nu_1({\bm\Sigma}^\xi)} = \frac {\mu_r^\chi}{\max_{i=1,\ldots, n}\sigma_i^2}\ge \frac{n\underline C_r}{C_\xi}.
\end{align}
By substituting (ref) into (ref) we have
\begin{equation}
\Vert \{\mathbf I_r+{\bm\Lambda}^\prime({\bm\Sigma}^\xi)^{-1} {\bm\Lambda}\}^{-1}\Vert\le \frac{C_\xi}{n\underline C_r}.
\end{equation}
Therefore, from (ref) and (ref),
\[
\left\Vert \{\mathbf I_r+{\bm\Lambda}^\prime({\bm\Sigma}^\xi)^{-1} {\bm\Lambda}\}^{-1}\{\bm \Lambda^\prime(\bm\Sigma^\xi)^{-1}\bm\Lambda \} - \mathbf I_r\right\Vert =\left\Vert \{\mathbf I_r+{\bm\Lambda}^\prime({\bm\Sigma}^\xi)^{-1} {\bm\Lambda}\}^{-1}\right\Vert\le \frac{C_\xi}{n\underline C_r},
\]
which completes the proof.
\end{proof}
$\,$