EconBase
← Back to paper

Approximate Factor Models with Weaker Loadings

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

82,748 characters · 15 sections · 43 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Approximate Factor Models with Weaker Loadings

abstractPervasive cross-section dependence is increasingly recognized as a characteristic of economic data and the approximate factor model provides a useful framework for analysis. Assuming a strong factor structure where $\bm \Lambda^{0'}\bm \Lambda^0/N^\alpha$ is positive definite in the limit when $\alpha=1$, early work established convergence of the principal component estimates of the factors and loadings up to a rotation matrix. This paper shows that the estimates are still consistent and asymptotically normal when $\alpha\in(0,1]$ albeit at slower rates and under additional assumptions on the sample size. The results hold whether $\alpha$ is constant or varies across factor loadings. The framework developed for heterogeneous loadings and the simplified proofs that can be also used in strong factor analysis are of independent interest.

JEL Classification: C30, C31

Keywords: principal components, low rank decomposition, weak factors, factor augmented regressions.

\thispagestyle{empty} \setcounter{page}{0} \baselineskip=18.0pt

Introduction

Starting with fhlr-restat and stock-watson-di-wp,stock-watson-jasa:02, a large body of research has been developed to estimate the latent common variations in large panels in which the $N$ units observed over $T$ periods are cross-sectionally correlated. A fundamental result shown in baing-ecta:02 is that the space spanned by the factors can be consistently estimated by the method of static principal components (PC) at rate $\min(\sqrt{N}, \sqrt{T})$. bai-ecta:03 then establishes $\sqrt{N}$ asymptotic normality of the estimated factors $\tilde {\bm F}$ up to a rotation matrix $\bm H$. The maintained assumption is that the factor structure is strong, meaning that if $\bm F^0$ and $\bm \Lambda^0$ are the latent factors and loadings, the matrices $\bm F^{0'} \bm F^0/T$ and $\bm \Lambda^{0'} \bm \Lambda^0/N$ are both positive definite in the limit. However, onatski-joe:12 shows that the PC estimates are inconsistent when $\bm \Lambda^{0'}\bm \Lambda^0$ (without dividing by $N$) has a positive definite limit. This has generated a good deal of interest in determining the number of less pervasive factors. Some assume large idiosyncratic variances, some assume that the entries of $\bm \Lambda^0$ are non-zero but small, while others assume a sparse $\bm \Lambda^0$ with many zero entries. See, for example, dgr-lasso, lettau-pelger:20, uematsu-yamagata:est, freyaldenhoven:22. Though the term `weak factors' is used in different ways, there is a presumption that the PC estimator has undesirable properties when the strong factor assumption fails. However, to our knowledge, there does not exist a clear statement of what those properties are.

In this paper, we consider the weaker condition that $\bm \Lambda^{0'}\bm \Lambda^0/N^\alpha$ has a positive definite limit with $\alpha\in (0,1]$. Since it is the strength of the loadings that is being weakened and positive definiteness of $\bm F^{0'}\bm F^0/T$ is maintained throughout, we use the terminology of {\em weaker loadings}. We obtain two results for average errors. The first result, which concerns the low rank component, is $\frac{1}{NT} \sum_{i=1}^N \sum_{t=1}^T \| \tilde C_{it}-C_{it}^0\|^2=O_p(\frac{1}{N})+O_p(\frac{1}{T})$. This result is somewhat surprising as it is the same as in the strong factor case. The second result pertains to the error rate in estimating the space spanned by the factors. This rate is of interest because it determines whether $\tilde {\bm F}$ can be treated as though it was known in factor augmented regressions. We obtain $\frac{1}{T}\sum_{t=1}^T \|\tilde F_t-\bm H^\prime F_t^0\|^2=O_p(\frac{1}{N^\alpha})+O_p((\frac{N^{1-\alpha}}{T})^2)$ which is asymptotically $o_p(1)$ when $\alpha>0$ and $ \frac{N^{1-\alpha}}{T}\rightarrow 0$. This result implies an error rate for the strong factor case of $\alpha=1$ of $\min(N,T^2)$, which is better than the rate of $\min(N,T)$ previously derived. This improvement is made possible by a different proof technique that also leads to significant simplifications, hence of independent interest. The simplifications come partly from using higher level assumptions, and partly from using approximations to the original rotation matrix $\bm H$ which also make it possible to conduct inference using a representation of the asymptotic variance that the user deems most convenient.

Our main result is that while the strong factor assumption of $\alpha=1$ yields the fastest convergence rates possible, and the estimates are inconsistent in the other extreme when $\alpha=0$, the principal component estimator for $\bm \Lambda$ and $\bm F$ continues to be consistent when $\alpha\in(0,1]$. In other words, except in the special case considered in onatski-joe:12, the PC estimates are consistent. We find that asymptotic normality of $\sqrt{N^\alpha}(\tilde F_t-\bm H^\prime F_t^0)$, $\sqrt{T}(\tilde\Lambda_i-\bm H^{-1} \Lambda_i^0)$, and $\min(\sqrt{N^\alpha},\sqrt{T})(\tilde C_{it}-C_{it}^0)$ do require $\alpha > 1/2$ along with some additional assumptions on $N$ and $T$, though $N$ is not required to grow at the same rate as $T$. However, $\alpha>0$ suffices for consistency of the individual loadings $\tilde\Lambda_i$, while $\alpha>1/3$ suffices for consistency of the individual factor estimates $\tilde F_t$. Thus consistent estimates can be obtained with weaker loadings than asymptotic normality.

It is natural to ask what happens when the loadings have varying strength. That is, instead of a constant $\alpha$, we now have $1 \ge \alpha_1\ge \alpha_2\ge \ldots\ge \alpha_r >0$. We show that in this case, what matters is the weakest loading, $\alpha_r$. Asymptotic normality now requires $\alpha_r>1/2$ but consistency of the individual estimates is possible without this requirement. Though the results are in agreement with the constant $\alpha$ case, setting up the framework is not so trivial as it requires using different normalization rates to study convergence of $\tilde {\bm F}$ to $\bm F^0\bm H$ while allowing the rotation matrix $\bm H$ to be consistent with the data generating process. The framework is more general than that of freyaldenhoven:22 or uematsu-yamagata:est which require specific assumptions on $\bm F^0 \bm H$ or $\bm H$, as discussed below.

The paper proceeds as follows. We start with the simpler case that $\alpha$ is the same for all factors and provide the complete distribution theory for $\tilde F_t$, $\tilde \Lambda_i$ and $\tilde C_{it}$. We then consider the general case when $\alpha$ varies. Section 2 sets up the econometric framework and presents three useful preliminary results. Section 3 studies consistent estimation of the factors, the loadings, and introduces four asymptotically equivalent rotation matrices. The distribution theory is given in Section 4. Implications of weaker loadings for factor augmented regressions are discussed. Section 5 studies the case of heterogeneous $\alpha$.

Throughout, matrices are written in bold-face to distinguish them from vectors. As a matter of notation, $\|\bm A\|^2=\sum_{i=1}^m\sum_{j=1}^n |A_{ij}|^2=\text{Tr}(\bm A\bm A')$ is the squared Frobenius norm of a $m\times n$ matrix $\bm A$, $\|\bm A\|^2_{sp}=\rho_{\max}(\bm A'\bm A)$ denotes the squared spectral norm of $\bm A$, where $\rho_{\max}(\bm B)$ denotes the largest eigenvalue of a positive semi-definite matrix $\bm B$. Note that $\|\bm A\|_{sp}\le \|\bm A\| \le \sqrt{q} \|\bm A\|_{sp}$, where $q=rank(\bm A)$. Thus when the rank $q$ is fixed, the two norms are equivalent in terms of asymptotic behavior.

The Econometric Setup

We use $i=1,\ldots N$ to index cross-section units and $t=1,\ldots T$ to index time series observations. Let $X_i=(X_{i1},\ldots X_{iT})^\prime$ be a $T\times 1$ vector of random variables and $\bm X=(X_1,X_2,\ldots, X_N)$ be a $T\times N$ matrix. The normalized data $\bm Z=\frac{\bm X}{\sqrt{NT}}$ admit singular value decomposition (svd) \[ \bm Z=\frac{\bm X}{\sqrt{NT}} =\bm U_{NT}\bm D_{NT}\bm V_{NT}^\prime=\bm U_{NT,k} \bm D_{NT,k} \bm V_{NT,k}^\prime+ \bm U_{NT,N-k} \bm D_{NT,N-k} \bm V_{NT,N-k}^\prime\] where $\bm U_{NT}'\bm U_{NT}=\bm I_T$ and $\bm V_{NT}'\bm V_{NT}=\bm I_N$. In the above, $\bm D_{NT,k}$ is a diagonal matrix of $k$ singular values $d_{NT,1},\ldots, d_{NT,k}$ arranged in descending order, $\bm U_{NT,k}, \bm V_{NT,k}$ are the corresponding left and right singular vectors respectively. By the eckart-young theorem, the best rank $k$ approximation of $\bm Z$ is $\bm U_{NT,k}\bm D_{NT,k}\bm V_{NT,k}^\prime$. This is obtained without imposing probabilistic assumptions on the data.

We represent the data using a static factor model with $r$ factors. In matrix form,

eqnarray[eqnarray omitted — 84 chars of source]

To simplify notation, the subscripts indicating that $\bm F$ is $T\times r$ and $\bm \Lambda$ is $N\times r$ will be suppressed when the context is clear. The common component $\bm C=\bm F\bm \Lambda^\prime$ has reduced rank $r$ because $\bm F$ and $\bm \Lambda$ both have rank $r$. The $N\times N$ covariance matrix of $\bm X$ takes the form \[ \bm\Sigma_X=\bm \Lambda\bm\Sigma_F \bm\Lambda^\prime + \bm\Sigma_e=\bm \Sigma_C+\bm \Sigma_e .\]

A strict factor model obtains when the errors $e_{it}$ are cross-sectionally and serially uncorrelated so that $\bm\Sigma_e$ is a diagonal matrix. The classical factor model studied in anderson-rubin uses the stronger assumption that $e_{it}$ is iid and normally distributed. For economic analysis, this error structure is overly restrictive. We work with the approximate factor model formulated in chamberlain-rothschild which allows the idiosyncratic errors to be weakly correlated in both the cross-section and time series dimensions. In such a case, $\bm \Sigma_e$ need not be a diagonal matrix.

Let $\bm F^0$ and $\bm \Lambda^0$ be the true values of $\bm F$ and $\bm \Lambda$. The model for unit $i$ at time $t$ as \[ x_{it}=\Lambda_i^{0\prime} F^0_t+e_{it}.\] Letting $e_i^\prime=(e_{i1},e_{i2},...,e_{iT})$ and $e_t^\prime=(e_{1t},e_{2t},...,e_{Nt})$, the model for unit $i$ is

eqnarray*[eqnarray* omitted — 55 chars of source]

Estimation of $\bm F^0$ and $\bm \Lambda^0$ in an approximate factor model with $r$ factors proceeds by minimizing the sum of squared residuals:

eqnarray*[eqnarray* omitted — 248 chars of source]

As $\bm F$ and $\bm \Lambda$ are not separately identified, we impose the normalization restrictions

equation[equation omitted — 131 chars of source]

The solution is the (static) PC estimator defined as:

equation[equation omitted — 112 chars of source]

PC estimation of large dimensional approximate factor models must overcome two challenges not present in the classical factor analysis of anderson-rubin. The first pertains to the fact that the errors are now allowed to be cross-sectionally correlated. The second issue arises because the $T\times T$ covariance matrix of $\bm X$ and the $N\times N$ covariance of $\bm X'$ are of infinite dimensions when $N$ and $T$ are large. To study the properties of the PC estimates, we use $\bm X=\bm F^0\bm \Lambda^{0\prime}+\bm e$ to obtain:

eqnarray[eqnarray omitted — 280 chars of source]

But $\frac{1}{NT} \bm X\bm X^\prime=\bm U_{NT}\bm D_{NT}^2\bm U_{NT}^\prime$ and thus $ \frac{1}{NT} \bm X\bm X' \tilde {\bm F} =\tilde {\bm F} \bm D_{NT,r}^2$. It follows that

eqnarray[eqnarray omitted — 333 chars of source]

Rearranging terms yields

eqnarray[eqnarray omitted — 315 chars of source]

where

equation[equation omitted — 170 chars of source]

$ \gamma_{st}=E(\frac{1}{N}e_s^\prime e_t)=E(\frac{1}{N} \sum_{i=1}^N e_{is} e_{it})$, $\zeta_{st}=\frac{1}{N}e_s^\prime e_t-\gamma_{st}$, $\eta_{st}= \bm F_s^{0^T}\bm \Lambda^{0^T} e_t/N$, and $\xi_{st}= \bm F_t^{0^T} \bm \Lambda^{0^T} e_s/N$. stock-watson-jasa:02,baing-ecta:02,bai-ecta:03 established properties of the PC estimator by analyzing the four terms in ((ref)) under certain assumptions, and this is by and large the approach that the literature has taken. We work directly with the matrix norms of the terms in ((ref)). This makes it possible to obtain simpler proofs under more general assumptions.\footnote{An earlier version of the paper circulated as {\em Simpler proofs for approximate factor models of large dimensions} considers $\alpha=1$ only.}

Weaker Loadings: Homogeneous Case

The defining characteristic of an approximate factor model is that the first $r$ largest population eigenvalues of $\bm \Sigma_C$ diverge with $N$ while all remaining eigenvalues of $\bm \Sigma_C$ are zero, and all eigenvalues of $\bm \Sigma_e$ are bounded. Previous works model the `diverge with $N$' feature by assuming that $\bm F^{0^\prime} \bm F^0/T>0$ and $\bm \Lambda^{0^\prime} \bm \Lambda^0/N$ are positive definite in the limit. These two conditions have come to be known as the strong factor structure. onatski-joe:12 considers the other extreme that requires $\bm \Lambda^{0'}\bm \Lambda^0$ to have a positive limit and shows that the factor estimates are inconsistent. There are many ways to accommodate weaker factor structures. For example, dgr-lasso let the eigenvalues of $\bm \Sigma_e$ to be large relative to those of $\bm \Sigma_C$. As discussed in onatski-joe:12, such a setup can be rewritten in terms of weaker loadings defined as $\bm \Lambda^{0'}\bm \Lambda^0/N^{\alpha}$ with $1\ge \alpha>0$, and is the approach that we will follow.

\paragraph{Assumption A1:} Let $M<\infty$ not depending on $N$ and $T$ and define \[ \delta_{NT}=\min(\sqrt{N},\sqrt{T}).\]

itemize• Mean independence: $ E(e_{it}|\Lambda_i^0 ,F_t^0)=0$. • Weak (cross-sectional and serial) correlation in the errors. \begin{itemize} • $E \Big[\frac 1 {\sqrt{N}} \sum_{i=1}^N [e_{it}e_{is}-E(e_{it}e_{is})]\Big]^2 \le M $. • For all $i$, $\frac 1 T \sum_{t=1}^T\sum_{s=1}^T | E (e_{it}e_{is})| \le M $. For all $t$, $\frac 1 N \sum_{i=1}^N\sum_{j=1}^N | E (e_{it}e_{jt})|\le M$. • For all $t$, $\frac{1}{N\sqrt{T}} \|e_t^\prime \bm e'\|=O_p(\delta_{NT}^{-1})$ and for all $i$, $\frac{1}{T\sqrt{N}} \|e_i^\prime \bm e\|=O_p(\delta_{NT}^{-1})$. • $\|\bm e\|_{sp}^2=\rho_{\max}(\bm e'\bm e) =O_p(\max\{N,T\})$. \end{itemize}

\paragraph{Assumption A2:} (i) $E\|F_t^0||^4 \le M,$ $\mathrm{plim}_{T\rightarrow\infty} \frac{\bm F^{0'} \bm F^0}{T}=\bm \Sigma_F>0$;\\ (ii) $\|\Lambda_i^0\|\le M$, $\lim_{N\rightarrow\infty}\frac{\bm \Lambda^{0'}\bm \Lambda^0}{N^\alpha}=\bm \Sigma_\Lambda>0$, for some $\alpha>0$ with $\alpha\in (0,1]$; \\ (iii) the eigenvalues of $\bm\Sigma_\Lambda \bm \Sigma_F$ are distinct.

\paragraph{Assumption A3:} For each $t$, (i) $E\|N^{- {\alpha} /2}\sum_i \Lambda^0_i e_{it}\|^2\le M$, (ii) $\frac{1}{NT} e_t^\prime\bm e^\prime\bm F^0=O_p(\delta_{NT}^{-2})$; for each $i$, (iii) $E\|T^{-1/2} \sum_t F^0_t e_{it}||^2\le M$, (iv) $\frac{1}{N^\alpha T}e_i'\bm e\bm \Lambda^0=O_p(\frac 1 {N^\alpha}) +O_p(\frac 1 {\sqrt{T N^\alpha}})$; (v) $\bm \Lambda^{0'} \bm e ' \bm F^0 = \sum_{i=1}^N\sum_{t=1}^T \Lambda_i^0 F_t^{0'} e_{it} =O_p( \sqrt{N^\alpha T})$.

\paragraph{Assumption A4:} As $ N,T\rightarrow \infty$, $(\frac N {N^\alpha}) \frac 1 T \rightarrow 0$ for the same $\alpha$ in Assumption A2.

Assumption A1.i uses mean independence in place of moment conditions on $e_{it}$ as in previous work. Assumption A1.ii assumes weak time and cross-section dependence. Assumption A1(d) is a bound on the maximum eigenvalues of $\bm e^\prime\bm e$. For iid data with uniformly bounded fourth moments, the rate is implied by random matrix theory. moon-weidner:17 extend the case to data that are weakly correlated across $i$ and $t$. We use it to obtain simpler proofs. Assumption A2 implies $\|\bm F^0\|^2/T=O_p(1)$ and $\|\bm \Lambda^0\|^2/N^ {\alpha} =O_p(1)$. A2.(ii) allows the $r$ eigenvalues of $\bm \Lambda^{0'}\bm \Lambda^{0'}$ to diverge at a rate of $N^ {\alpha} $ with $1\ge \alpha>0$.

Assumption A2 entertains weaker loadings by allowing $\bm \Lambda^{0'}\bm \Lambda^0/N^\alpha$ to have a positive definite limit with $1\ge \alpha>0$ which nests the strong factor model as a special case. When $\alpha$ is constant, the strength of the loadings is homogeneous. Note that the strength of the factor loadings affects the normalization of $\bm \Lambda^{0^\prime}\bm \Lambda^0$ but not $\bm F^{0^\prime}\bm F^0$.

Assumptions A3.(ii) and (iv) reflect the more general setup that $1\ge \alpha > 0$. Parts of Assumption A3 also appear in baing-ecta:02. When the errors $e_{it}$ are independent, Assumptions A1 and A2 are enough to validate A3. The assumption should hold under weak cross-sectional and serial correlations. Assumption A3 implies

eqnarray[eqnarray omitted — 485 chars of source]

Allowing for weaker factors comes at a cost. As stated in Assumption A4, which is new, a small $\alpha$ must be compensated by a larger $T$. We assume that all variables in $\bm X$ are relevant in the sense of having a non-negligible common component. chao-swanson:0322,chao-swanson:0422 study selection of relevant variables in factor augmented regressions. To accommodate irrelevant variables, they assume $\frac{N}{N_1}\frac{1}{T}\rightarrow c$, $c>0$ and possibly $\infty$ where $N_1$ is the number of relevant variables. Assumption A4 rules this out.

The above assumptions are written for analyzing weak loadings as defined in Assumption A2. The framework can be adapted to study weak factors modeled as $\frac{\bm F ' \bm F}{T^\beta}$ being positive definite in the limit for any $1 \ge \beta >0$. Whether we have weak loadings or weak factors, the key feature is that some eigenvalues of $\Sigma_C$ will diverge at a rate slower than $N$. Though the analysis proceeds as though the panel $X$ consists of data with i indexing units and t indexing time, the framework is also valid when $i$ and $t$ take on other interpretation provided that the data satisfy the assumptions above.

Useful Identities and Matrices

In the strong factor case, $\bm D_{NT,r}^2$ is a diagonal matrix of the $r$ largest eigenvalues of $\bm Z'\bm Z=\frac{1}{NT}\bm X\bm X^\prime$ . The singular values of $\bm Z$ are those of $\bm X$ divided by $\sqrt{NT}$. In practice, each column of $\bm Z$ is transformed to have unit variance so $d^2_{NT,j}$ is the fraction of variation in $\bm Z$ explained by factor $j$. The following lemma shows that to accommodate weaker factors, $\bm D^2_{NT,r}$ must be scaled up by $N^{1-\alpha}$ to have a limit matrix $\bm D^2_r$ that is full rank.

lemmaLet $\bm D_r^2$ be a diagonal matrix consisting of the the ordered eigenvalues of $\bm \Sigma_\Lambda \bm \Sigma_F$. Under Assumption A, we have \[ \Big(\frac N {N^{\alpha}} \Big) \bm D_{NT,r}^2 \smash{\mathop{\longrightarrow}\limits^p} \bm D_r^2 >0, \quad 1\ge \alpha >0. \]

\paragraph {Proof:} Proper normalization of $\bm D_{NT,r}^2$ is key to accommodating $1\ge \alpha>0$. The diagonal matrix $\frac{N}{N^\alpha}\bm D_{NT,r}^2$ consists of the $r$ largest eigenvalues of $\frac{1}{N^\alpha T} \bm X\bm X^\prime$, and

eqnarray*[eqnarray* omitted — 300 chars of source]

Since the largest eigenvalue of $\bm e\bm e'$ is of order $\max\{N,T \}$, the largest eigenvalue of the last matrix is bounded by $O_p(\frac N {N^\alpha} \frac 1 T )+ O_p(\frac 1 {N^{\alpha}}) \smash{\mathop{\longrightarrow}\limits^p} 0$. Furthermore, $\|\bm e\|_{sp}=O_p(\sqrt{\max\{N,T\}})$, $\|\bm F^0\|_{sp}=O_p(T^{1/2})$ and $\|\bm \Lambda^0\|_{sp}=O_p(N^{\alpha/2})$. A bound in spectral norm for the second matrix on the right hand side is \[\frac{\|\bm e\|_{sp}\|\bm F^0||_{sp}\|\bm \Lambda^0\|_{sp}}{T N^{\alpha}} \le O_p\bigg(\sqrt{\frac N {N^\alpha} \frac 1 T} \bigg)+ O_p(\frac 1 {N^{\alpha/2}}) \smash{\mathop{\longrightarrow}\limits^p} 0.\] The third matrix is the transpose of the second. Thus, the largest eigenvalues of the last three matrices converge to zero. By the matrix perturbation theorem, the $r$ largest eigenvalues of $\frac{1}{N^\alpha T} \bm X\bm X^\prime$ are determined by the first matrix on the right hand side. The eigenvalues of this matrix are the same as those of \[ \bigg( \frac{\bm\Lambda^{0\prime}\bm \Lambda^0}{N^\alpha}\bigg)\bigg(\frac{\bm F^{0'}\bm F^0 }{T}\bigg). \] This matrix converges to $\bm \Sigma_\Lambda \bm \Sigma_F$ whose eigenvalues are $\bm D_r^2$, proving the lemma. $\Box$

Next, we turn to two matrices that will play important roles subsequently. The first is the matrix $\tilde {\bm F^\prime }\bm F^0/T$. To obtain its limit, we multiply $(N/N^\alpha) \bm \tilde {\bm F}'$ on each side of ((ref)) and use the fact that $\bm \tilde {\bm F}'\bm \tilde {\bm F}=T$ to obtain

eqnarray[eqnarray omitted — 498 chars of source]

The right hand side converges to a positive definite matrix (thus invertible) by Lemma (ref). The last three matrices on the left hand side converges in probability to zero. In particular, \[ \|\frac{\bm \tilde {\bm F}' \bm F^0\bm \bm \Lambda^{0'} \bm e^{\prime} \tilde {\bm F}}{N^\alpha T^2}\|\le \|\frac{\bm \tilde {\bm F}' \bm F^0} T\| \|\bm \bm \Lambda^{0'} \bm e^{\prime}\| \|\tilde {\bm F}\| \frac 1 {N^\alpha T} =O_p(\frac 1 {N^{\alpha/2}}) =o_p(1) \] and

equation[equation omitted — 252 chars of source]

The limit on the left hand side is thus determined by the first matrix, ie.

eqnarray[eqnarray omitted — 234 chars of source]

The limit of $\tilde {\bm F}'\bm F^{0'}/T $ can be obtained from this representation.

lemmaUnder Assumption A, \begin{itemize} • $\tilde {\bm F}'\bm F^{0'}/T \smash{\mathop{\longrightarrow}\limits^p} \bm Q:=\bm D_{r} \bm \Upsilon' \bm \Sigma_{\Lambda}^{-1/2}$, where $\bm \Upsilon$ consists of the eigenvectors of the matrix $\bm \Sigma_F^{1/2}\bm \Sigma_\Lambda \bm \Sigma_F^{1/2}$ with $\bm \Upsilon ' \bm \Upsilon =I_r$. • For $\bm H_{NT,0}=\bigg(\frac{\bm \Lambda^{0^\prime}\bm \Lambda^0}{N}\bigg)\bigg(\frac{\bm F^{0^\prime}\tilde {\bm F}}{T}\bigg) \bm D_{NT,r}^{-2},$ we have $\bm H_{NT,0}\smash{\mathop{\longrightarrow}\limits^p} \bm Q^{-1}$. \end{itemize}

Part (i) is obtained by taking limit on each side of ((ref)) to yield $ \bm Q \bm \Sigma_\Lambda \bm Q' = \bm D_r^2 $. Since $\bm D_r^2$ is a positive definite matrix, it follows that $\bm Q $ is invertible. Matrix $\bm Q$ can be expressed as $\bm Q =\bm D_{r} \bm \Upsilon' \bm \Sigma_{\Lambda}^{-1/2}$ where $\bm \Upsilon$ consists of the orthonormal eigenvectors of the matrix $\bm \Sigma_F^{1/2}\bm \Sigma_\Lambda \bm \Sigma_F^{1/2}$ such that $\bm \Upsilon' \bm \Upsilon=I_r$ (Bai, 2003). Note that $\bm Q$ is unique up to a column sign change, just like $\tilde {\bm F}$ is determined up to a column sign change.

The rotation matrix $\bm H_{NT,0}$, first derived in stock-watson-di-wp, has been used to evaluate the precision of $\tilde {\bm F}$. bai-ecta:03 shows that $\bm H_{NT,0}\smash{\mathop{\longrightarrow}\limits^p} \bm Q^{-1}$ when $\alpha=1$. To accommodate weaker loadings, we consider \[\bm H_{NT,0}=\bigg(\frac{\bm \Lambda^{0^\prime}\bm \Lambda^0}{N^{\alpha} }\bigg)\bigg(\frac{\bm F^{0^\prime}\tilde {\bm F}}{T}\bigg) \Big(\frac N {N^\alpha} \bm D_{NT,r}^2\Big)^{-1}.\] By assumption, the first matrix on the right hand side is invertible while the last two matrices are invertible by the previous lemmas. Hence $\bm H_{NT,0} \smash{\mathop{\longrightarrow}\limits^p} \bm \Sigma_\Lambda \bm Q' \bm D_r^{-2}\equiv \bm Q^{-1}$. The matrix $\bm Q$ and its relation to $\bm H_{NT,0}$ are fundamental to the asymptotic theory in the strong factor case. Lemma (ref) shows that the relations are unaffected when weaker loadings are allowed.

Average Errors in Estimating the Factor Space

This section has three parts. Subsection 1 presents results for consistent estimation of the space spanned by the factors. Subsection 2 introduces four new rotation matrices. Subsection 3 uses these new matrices to show consistent estimation of the spanned by the loadings.

The Factors

To establish consistent estimation of $\tilde {\bm F}$ for $\bm F^0$ up to rotation by $\bm H_{NT,0}$, we multiply $\bm D_{NT,r}^{-2}$ to both sides of ((ref)) and use the definition of $\bm H_{NT,0}$ to obtain

eqnarray[eqnarray omitted — 556 chars of source]

This implies

eqnarray*[eqnarray* omitted — 530 chars of source]

But $ \frac 1 {\sqrt{T} N^\alpha } \| \bm \Lambda^{0'}\bm e'\| =O_p(\frac 1 {\sqrt{N^\alpha}}) $ by ((ref)) and $ \frac {\|\bm e \bm e^\prime \bm \tilde {\bm F} \|} {N^\alpha T^{3/2} }\le \frac{\rho_{\max}(\bm e \bm e^\prime) \|\bm \tilde {\bm F}\|}{N^\alpha T^{3/2} }= O_p(\frac 1 {N^{\alpha}}) + \frac N {N^{\alpha}} \frac 1 T O_p(1)$. Thus \[ \frac 1 {\sqrt{T}} \|\tilde {\bm F} -\bm F^0\bm H_{NT,0}\|=O_p(\frac 1 {\sqrt{N^\alpha}})+\frac 1 { T} \frac N {N^\alpha} O_p(1). \] Squaring it gives the following proposition.

propositionUnder Assumption A, the following holds: \begin{eqnarray*} \frac{1}{T} \|\tilde {\bm F} -\bm F^0\bm H_{NT,0}\|^2=\frac{1}{T} \sum_{t=1}^T\|\tilde F_t-\bm H_{NT,0}^{\prime}F_t^0\|^2 =O_p\Big(\frac 1 {N^{\alpha}}\Big) + \frac 1 {T^2} \bigg(\frac N {N^{\alpha}}\bigg)^2 O_p(1). \end{eqnarray*}

The result is stated in terms of squared Frobenius norm. The average error in estimating $\bm F$ vanishes at rate $O_p(\frac{1}{N^\alpha})+\frac{N^{2(1-\alpha)}}{T^2}O_p(1)$. For $\alpha=1$, Theorem 1 of baing-ecta:02 gives a convergence rate for the same quantity of $O_p(\frac{1}{N})+\frac{1}{T}O_p(1)$. The proposition here uses a different proof to obtain a faster convergence rate of $O_p(\frac 1 N) +\frac 1 {T^2} O_p(1) $ for the strong factor case of $\alpha=1$. Implications of the proposition will be discussed subsequently.

Equivalent Rotation Matrices

The rotation matrix $\bm H_{NT,0}$ is a product of three $r\times r$ matrices and it is not easy to interpret. However, we can rewrite ((ref)) as

equation[equation omitted — 447 chars of source]

As $\bm\tilde {\bm F}^\prime \bm \bm F^0/T=O_p(1)$, the product of $\frac {\tilde {\bm F}^\prime \bm F^0 } T$ and $\bm H_{NT,0}$ is an identity matrix up to an negligible term if it can be shown that the three terms inside the bracket are small. The next Lemma formalizes this result and shows that it also holds for four other rotation.

lemmaUnder Assumption A, \begin{itemize} • $ \bm H_{NT,0} =\Big(\frac {\tilde {\bm F}^\prime \bm F^0 } T \Big)^{-1}+ O_p\bigg(\frac 1 {\sqrt{N^\alpha T}}\bigg)+ \bigg[ O_p\bigg(\frac 1 {N^\alpha}\bigg) +\bigg(\frac N {N^{\alpha}}\bigg) \frac 1 T O_p(1)\bigg].$ • For $\ell=1,2,3,4$, $ \bm H_{NT,\ell}=\bm H_{NT,0} + O_p(\frac 1 {\sqrt{N^\alpha T}})+ O_p(\frac 1 {N^\alpha}) +(\frac N {N^{\alpha}}) \frac 1 T O_p(1)$, where \newline $\bm H_{NT,1}=(\bm \Lambda^{0^\prime}\bm \Lambda^0)(\tilde {\bm \Lambda}^\prime \bm \Lambda^0)^{-1}$, \newline $ \bm H_{NT,2} = (\bm F^{0^\prime} \bm F^0)^{-1}(\bm F^{0^\prime} \tilde {\bm F})$, \newline $ \bm H_{NT,3} = (\tilde {\bm F}^\prime \bm F^0)^{-1}(\tilde {\bm F}^\prime \tilde {\bm F})=(\tilde {\bm F}^\prime \bm F^0/T)^{-1}$, and \newline $\bm H_{NT,4} =(\bm \Lambda^{0^\prime} \tilde {\bm \Lambda}) (\tilde {\bm \Lambda}^\prime\tilde {\bm \Lambda})^{-1} = (\bm \Lambda^{0^\prime}\tilde {\bm \Lambda}/N) \bm D_{NT,r}^{-2}$. • $\bm H_{NT,\ell}\smash{\mathop{\longrightarrow}\limits^p} \bm Q^{-1}$. \end{itemize}

Part (i), shown in the Appendix, establishes the error in approximating $\bm H_{NT,0}$ by $(\frac{\tilde {\bm F}^\prime \bm F^0}{T})^{-1}$ while part (ii) considers four additional approximations that provide an intuitive interpretation of $\bm H_{NT} F^0_t$. For example, $\bm H_{NT,2}$ is the coefficient matrix from projecting $\tilde{\bm F}$ on the space spanned by $\bm F^0$ and $\bm H^\prime_{NT,2}F^0_t$ is asymptotically the fit from the projection. These alternative rotation matrices were used in baing-joe:19 for $\alpha=1$. The above Lemma shows that they can still be used in place of $\bm H_{NT,0}$ when $\alpha<1$, but the adequacy of approximation will depend on $\alpha$. Lemma (ref) allows for simpler proofs and helps to interpret the error in estimating $\tilde {\bm F}_t$ and $\tilde {\bm \Lambda}_i$. But for consistency proofs, the result $\bm H_{NT,\ell}\bm H_{NT,0}^{-1}=I_r+o_p(1)$ often suffices, and it is implied by Lemma (ref).

The Loadings and the Common Component

The PC estimator satisfies $ \frac 1 N \tilde {\bm \Lambda}^\prime \tilde {\bm \Lambda} = \bm D_{NT,r}^2 $ and we already have $ \frac 1 {N^\alpha} \tilde {\bm \Lambda}^\prime \tilde {\bm \Lambda} = \frac N {N^\alpha}\bm D_{NT,r}^2 \smash{\mathop{\longrightarrow}\limits^p} \bm D_r^2. $ We can now provide a simple consistency proof for $\tilde {\bm \Lambda}$. Multiply $\frac 1 T \tilde {\bm F}'$ to both sides of $\bm X =\bm F^{0^\prime}\bm \Lambda^{0'} +\bm e$ to obtain $ \frac 1 T \tilde {\bm F}^\prime\bm X =(\tilde {\bm F}^\prime \bm F^0/T) \bm \Lambda^{0'} + \tilde {\bm F}^\prime\bm e /T$. We have

eqnarray[eqnarray omitted — 297 chars of source]

and thus \[ \frac 1 {\sqrt{N}} \|\tilde {\bm \Lambda}^\prime - \bm H_{NT,3}^{-1} \bm \Lambda^{0'}\|\le \|H_{NT,0}\| \frac{\| \bm F^{0'} \bm e\|}{T\sqrt{N}} + \frac{ \|(\tilde {\bm F}- \bm F^0 \bm H_{NT,0})' \bm e\| } {T\sqrt{N} } \] The first term $\|\bm \bm F^0 \bm e \|/(T \sqrt{N})=O_p(1/\sqrt{T})$ by equation ((ref)). The second term is $O_p(\frac{1}{\sqrt{N^{1+\alpha}}})$, shown in the Appendix. Combining results and ignoring terms dominated by $O_p(T^{-1/2})$, we have \[ \frac 1 {\sqrt{N}} \|\tilde {\bm \Lambda}^\prime - \bm H_{NT,3}^{-1} \bm \Lambda^{0'}\| = O_p\bigg(\frac 1 {\sqrt{T}}\bigg)+ O_p\bigg(\frac 1 {\sqrt{N^{1+\alpha}}}\bigg). \]

Squaring gives the next proposition:

propositionUnder Assumption A, the following holds \begin{eqnarray*} \frac 1 N \|\tilde {\bm \Lambda}-\bm \Lambda^0 (\bm H_{NT,0}')^{-1}\|^2=\frac{1}{N} \sum_{i=1}^N \|\tilde\Lambda_i- \bm H_{NT,0}^{-1}\Lambda_i^0\|^2= O_p\bigg(\frac 1 T\bigg) + O_p\bigg(\frac 1 {N^{1+\alpha}}\bigg). \end{eqnarray*}

Note that replacing $\bm H_{NT,3}$ by $ \bm H_{NT,0}$ does not affect the rate analysis. In fact, we can use other rotation matrices to gain intuition. For example, $\bm H_{NT,1}^{-1}$ is obtained by regressing $\bm \tilde {\bm \Lambda}$ on $\bm \Lambda_0$. Hence $\bm \Lambda_0 (\bm H'_{NT,1})^{-1}$ is asymptotically the fit from projecting $\tilde{ \bm \Lambda}$ on the space spanned by $\bm \Lambda_0$, and $\tilde {\bm \Lambda}-\bm \Lambda^0 (\bm H_{NT,0}^\prime)^{-1}$ is the error from that projection.

While $\tilde {\bm F}$ and $\tilde {\bm \Lambda}$ only estimate $\bm F^0$ and $\bm \Lambda^0$ up to a rotation matrix, $\tilde{\bm C}$ does not depend on rotations and is directly comparable to $\bm C^0$.

propositionUnder Assumption A, \begin{eqnarray*} \frac{1}{NT} \|\tilde{\bm C}-\bm C^0\|^2=\frac{1}{NT}\sum_{i=1}^N\sum_{t=1}^T \|\tilde C_{it}-C^0_{it}\|^2=O_p(\delta_{NT}^{-2}). \end{eqnarray*}

\paragraph{Proof:} From $\tilde {\bm C}-\bm C_0=\tilde {\bm F} \bm \tilde {\bm \Lambda}^\prime-\bm F^0 \bm \Lambda^{0\prime}= (\tilde {\bm F}-\bm F^0\bm H)\bm \tilde {\bm \Lambda}'+ \bm F^0\bm H\bm \tilde {\bm \Lambda}'-\bm F^0 \bm \Lambda^{0\prime}$, we have

eqnarray*[eqnarray* omitted — 610 chars of source]

where the second inequality follows from Propositions 1 and 2. The term $\frac 1 {T^2} \frac N {N^\alpha} O_p(1)$ is dominated by $O_p(1/T)$ since $\frac N {N^\alpha} \frac 1 T \rightarrow 0$ by Assumption A.4. Note this is the same rate as the strong factor case. $\Box$

Distribution Theory

As we do not observe $\bm F^0$ or $\bm \Lambda^0$, we need an inferential theory for $\tilde F_t$, $\tilde\Lambda_i$, and $\tilde C_{it}=\tilde F_t \tilde\Lambda_i^\prime$. Theorems 1 and 2 of bai-ecta:03 establish that

subequations\begin{eqnarray} \sqrt{N}(\tilde F_t-\bm H_{NT,0}'F^0_t)&\smash{\mathop{\longrightarrow}\limits^d}& \mathcal N(0, \bm D_{r}^{-2}\bm Q\bm \Gamma_t\bm Q'\bm D_{r}^{-2})\\ \sqrt{T}(\tilde \Lambda_i-\bm H_{NT,0}^{-1}\Lambda_i^0)&\smash{\mathop{\longrightarrow}\limits^d}& \mathcal N(0,\bm Q^{'-1} \bm \Phi_i \bm Q^{-1}). \end{eqnarray}

by positing appropriate central limit theorems (CLTs) appropriate for $\alpha=1$. Assumption B accommodates weaker loadings.

\paragraph{Assumption B.} The following holds for each $i$ and $t$ as $N, T\rightarrow \infty$:

eqnarray*[eqnarray* omitted — 271 chars of source]

where $\Gamma_t$ and $\Phi_i$ are $r \times r$ positive definite matrices.

The first CLT in Assumption B involves random variables over the cross section, while the second CLT involves random variables over different time periods. In view of the assumed weak dependence of $e_{it}$ over $i$ and $t$, the two limiting distributions are independent and the convergence holds jointly. The first CLT uses a normalization of $N^{\alpha/2}$ instead of the usual $N^{1/2}$ and is consistent with Assumption A2(ii). If each $e_{it}$ is iid with $Ee_{it}^2 =\sigma^2$, then the variance of the first term is $\sigma^2 \frac 1 {N^\alpha} \Lambda^{0\prime}\Lambda^0$, which converges to a positive definite matrix. To see that CLT holds in spite of weaker loadings, suppose that $\Lambda_i^0 =\frac 1 {N^\tau} \delta_i$ where $\tau \in [0,1/2)$, and $\delta_i$ are either bounded constants or are iid with $E(\delta_i\delta_i')>0$. Then for $\alpha=1-2\tau$, $N^{-\alpha/2} \sum_{i=1}^N \Lambda_i^0 e_{it} =N^{-1/2} \sum_{i=1}^N \delta_i e_{it}$, which is asymptotically normal by the standard arguments.

To obtain comparable asymptotic normality results when $1\ge \alpha> 0$ requires additional assumptions. To see why, first consider the limiting distribution $\tilde {\bm \Lambda}$. From ((ref)), \[ \sqrt{T}(\tilde \Lambda_i -\bm H_{NT,3}^{-1} \Lambda_i^0) = \bm H_{NT,3}' \frac{ 1}{\sqrt{ T}} \sum_{t=1}^T F_t^0 e_{it} + \frac{(\tilde {\bm F}-\bm F^0\bm H_{NT,3})^\prime e_i}{\sqrt{T}}.\] The first term is asymptotically normal by Assumption B and $\bm H_{NT,3}^\prime\smash{\mathop{\longrightarrow}\limits^p} \bm Q^{\prime -1}$ by Lemma (ref). For the second term, we show in Lemma {(ref) part (iii) that for each $i$, \[ \frac 1 T e_i' (\tilde {\bm F}- \bm F^0 \bm H_{NT,\ell}) =O_p(\frac 1 {N^\alpha} ) + O_p(\frac 1 {\sqrt{TN^{\alpha}}})+O_p(\frac N {N^{\alpha}T} ).\] For $\sqrt{T}$ consistency of $\tilde {\bm \Lambda}$, we need $\sqrt{T}$ times the above three terms to be negligible. We thus require $\frac {\sqrt{T}} {N^\alpha} \rightarrow 0$ and $ \frac 1{\sqrt{T}} \frac N {N^\alpha} \rightarrow 0.$ A necessary condition for both to hold is $\alpha>1/2$. These are listed as Assumption C(i)-(iii) below. Note that the preceding equation is $o_p(1)$ when $\alpha>0$. Thus if we merely consider consistency of $\tilde\Lambda_i$, we only need $\alpha>0$. It is only for $\sqrt{T}$ convergence and asymptotic normality that we need $\alpha>1/2$.

Similarly, for the distribution of $\tilde F_t$, we multiply $ \tilde {\bm \Lambda} (\tilde {\bm \Lambda}^\prime\tilde {\bm \Lambda})^{-1}$ to both sides of $\bm X= \bm F^0\bm \bm \Lambda^{0'} +\bm e$ to obtain $\bm X \tilde {\bm \Lambda} (\tilde {\bm \Lambda}^\prime\tilde {\bm \Lambda})^{-1} =\bm F^0 \bm \Lambda^{0'}\tilde {\bm \Lambda} (\tilde {\bm \Lambda}^\prime\tilde {\bm \Lambda})^{-1} + \bm e\tilde {\bm \Lambda} (\tilde {\bm \Lambda}^\prime\tilde {\bm \Lambda})^{-1} $. Using the definition of $\bm H_{NT,4}$,

eqnarray[eqnarray omitted — 417 chars of source]

Thus

eqnarray*[eqnarray* omitted — 407 chars of source]

The first term is asymptotically normal by Assumption B. Now Lemma (ref) part (iv) shows that for each $t$, the average correlation between $e_t$ and the error from estimating $\tilde {\bm \Lambda}$ is

equation[equation omitted — 310 chars of source]

In the strong factor case when $\alpha=1$, having the three terms vanish as $N,T\rightarrow \infty$ suffice for $\sqrt{N}$ consistency of $\tilde F_t$. With weaker loadings, we can only get $\sqrt{N^\alpha}$ consistency of $\tilde F_t$, and this requires $N/\sqrt{N^\alpha}$ times each of the four terms to be negligible. That is, \[ \frac 1 T \frac{N^{3/2}}{N^ {\alpha} } \rightarrow 0, \quad N^{\frac 1 2 -\alpha}\rightarrow 0, \quad \frac 1 {T^{3/2}} \frac {N^2}{N^{3\alpha/2}}\rightarrow 0, \quad \frac N {N^{3 \alpha/ 2}} \frac 1 {\sqrt{T}} \rightarrow 0.\] The third and the fourth conditions hold if the first two along with C(iii) are satisfied. Thus, in addition to $\alpha>1/2$, we also need C(iv), $\frac{N}{T} N^{1/2-\alpha}\rightarrow 0$ to restrict the relation between $N$ and $T$. But if instead of root-$N^{\alpha}$ asymptotic normality we merely consider consistency of $\tilde F_t$, we only need $\alpha>1/3$. This follows upon multiplying ((ref)) by $\frac N {N^{\alpha}}.$

We collect the required conditions for $\sqrt{T}$ asymptotic normality of $\tilde {\bm \Lambda}$ and $\sqrt{N^\alpha}$ asymptotic normality of $\tilde {\bm F}$ into the following: \paragraph{Assumption C:} (i) $ \alpha >\frac 1 2$, (ii) $ \frac {\sqrt{T}} {N^\alpha} \rightarrow 0$, (iii) $ \frac 1{\sqrt{T}} \frac N {N^\alpha} \rightarrow 0$, and (iv) $ \frac 1 T \frac{N^{3/2}}{N^ {\alpha} } \rightarrow 0$.

The above conditions reduce to $\sqrt{T}/N\rightarrow 0$ and $\sqrt{N}/T\rightarrow 0$ when $\alpha=1$. In general, they are distinctive restrictions, though some conditions may be redundant depending on the value of $\alpha$. Under Assumption C, the preceding analysis implies

subequations\begin{eqnarray} \sqrt{N^ {\alpha} } (\tilde F_t -\bm H_{NT,4}' F_t^0) &=&\bigg(\frac{\tilde {\bm \Lambda}'\tilde {\bm \Lambda}}{N^ {\alpha} }\bigg)^{-1} \bm H_{NT,4}^{-1} \frac 1{\sqrt{N^ {\alpha} }} \sum_{i=1}^N \Lambda_i^0 e_{it} + o_p(1)\\ \sqrt{T} (\tilde \Lambda_i -\bm H_{3,NT}^{-1} \Lambda_i^0) &=& \bm H_{NT,3}'\frac 1 {\sqrt{T}} \sum_{t=1}^T F_t^0 e_{it} + o_p(1). \end{eqnarray}
propositionUnder Assumptions A, B, and C and the normalization that $\bm F'\bm F/T=I_r$ and $\bm\Lambda'\bm\Lambda$ is diagonal, we have, as $N,T\rightarrow\infty$, \begin{eqnarray*} \sqrt{N^ {\alpha} }(\tilde F_t-\bm H_{NT,4}'F^0_t)&\smash{\mathop{\longrightarrow}\limits^d}& \mathcal N(0, \bm D_{r}^{-2}\bm Q\bm \Gamma_t\bm Q'\bm D_{r}^{-2}),\\ \sqrt{T}(\tilde \Lambda_i-\bm H_{NT,3}^{-1}\Lambda_i^0)&\smash{\mathop{\longrightarrow}\limits^d}& \mathcal N(0,\bm Q^{'-1} \bm \Phi_i \bm Q^{-1}). \end{eqnarray*}

The Proposition establishes $(\sqrt{N^\alpha},\sqrt{T})$ asymptotic normality of $(\tilde F_t,\tilde\Lambda_i)$. Though the conditions are more restrictive than the strong factor case, the results are more general than previously understood. Conditions weaker than Assumption C may be possible. For example, if $T/N\rightarrow c\in(0,\infty)$, then $\alpha>1/2$ is sufficient for asymptotic normality.

For hypothesis testing, there is no need to know $\alpha$. The limiting variance $\frac{1}{N^\alpha} \bm D_r^{-2} \bm Q \bm \Gamma_t \bm Q^\prime \bm D_r^{-2}$ can be consistently estimated by $\bm D_{NT,r}^{-2}(\sum_{i=1}^N \tilde \Lambda_i \tilde \Lambda_i' \tilde e_{it}^2) \bm D_{NT,r}^{-2}$ if we assume no cross-section correlation in $e_{it}$. If cross-section and time dependence are allowed, the CS-HAC variance estimator developed in baing-ecta:06 can be used. The square root of the $k$-th diagonal element of this matrix gives the standard error of the $k$-th component of $\tilde F_t$. Similarly, $\frac{1}{T} \bm Q^{\prime^{-1}}\bm \Phi_i \bm Q^{-1}$ can be consistently estimated by $\frac 1 T \sum_{s=1}^T \tilde F_s \tilde F_s' \tilde e_{is}^2$ (assuming no serial correlation in $e_{it}$). These formulas are identical to the case of strong factors.

Although the limiting covariance matrices look different from those given in ((ref)) and ((ref)), they are mathematically identical because of the different ways to represent $\bm H_{NT}$ and thus $\bm Q$. Indeed, other representations of the asymptotic variances can be used. For example, since $ (\tilde {\bm \Lambda}^\prime\tilde {\bm \Lambda}/N^ {\alpha} )^{-1} \bm H_{NT,4}^{-1 } = (\bm \Lambda^{0^\prime} \tilde {\bm \Lambda} /N^ {\alpha} )^{-1} =\bm H_{1,NT}^{\prime} (\bm \Lambda^{0'}\bm \Lambda^0/N^ {\alpha} )^{-1}$ and $\bm H_{NT,\ell} \smash{\mathop{\longrightarrow}\limits^p} \bm Q^{-1}$ for all $H_{NT,\ell}$ considered in Lemma (ref), we also have

eqnarray[eqnarray omitted — 395 chars of source]

From $\bm H_{NT,3}'=\bm H_{NT,2}^{-1} (\bm F^{0^\prime} \bm F^0/T)^{-1}$ and using ((ref)), it also holds that

eqnarray[eqnarray omitted — 358 chars of source]

As the factor estimates are asymptotically normal regardless of the choice of the rotation matrix, one can use the most convenient representation for inference.

While there are many ways to represent the sampling error of $\tilde F_t$ and $\tilde \Lambda_i$, the properties of $\tilde C_{it}$ are invariant to the choice of $\bm H_{NT,\ell}$, so we can simply write $\bm H_{NT}$. By definition, $C^0_{it}=\Lambda_i^{0\prime} F^0_t $ and $\tilde C_{it}=\tilde\Lambda_i^\prime \tilde F_t^0 $. Adding and subtracting terms

eqnarray*[eqnarray* omitted — 345 chars of source]

Using ((ref)) and ((ref)), we have

eqnarray*[eqnarray* omitted — 391 chars of source]

This leads to the distribution theory for the estimated common components.

propositionUnder Assumptions A, B and C, we have, as $N,T\rightarrow\infty$, \begin{eqnarray*} \frac{\tilde C_{it}-C^0_{it}}{\sqrt{ \frac 1 {N^ {\alpha} } W^\Lambda_{NT,it} +\frac{1}{T} W^F_{NT,it} }}&\smash{\mathop{\longrightarrow}\limits^d}& \mathcal N(0,1) \end{eqnarray*} where $W^\Lambda_{it}=\Lambda_i^{0\prime} \bm \Sigma_\Lambda^{-1} \bm \Gamma_t \bm \Sigma_\Lambda^{-1} \Lambda_i^0$ and $W^F_{it}=F_t^{0\prime} \bm \Sigma_F^{-1}\bm \Phi_i \bm \Sigma_F^{-1} F_t^0$.

The proposition implies a convergence rate of $\min\{ \sqrt{N^{\alpha}}, \sqrt{T}\}$ for the estimated common components. To estimate the standard error of $\tilde C_{it}$, there is also no need to know $\alpha$. Assuming no cross-sectional correlation in $e_{it}$, the term $\frac 1 {N^ {\alpha} } W^\Lambda_{NT,it}$ can be consistently estimated by $\tilde \Lambda_i^{\prime} (\sum_{i=1}^N \tilde \Lambda_i \tilde \Lambda_i')^{-1} (\sum_{i=1}^N \tilde \Lambda_i \tilde \Lambda_i' \tilde e_{it}^2) (\sum_{i=1}^N \tilde \Lambda_i \tilde \Lambda_i')^{-1} \tilde \Lambda_i$ . Furthermore, assuming no serial correlation in $e_{it}$, $\frac{1}{T} W^F_{NT,it}$ can be consistently estimated by $\frac 1 T \tilde F_t' (\sum_{s=1}^T \tilde F_s \tilde F_s' \tilde e_{is}^2) \tilde F_t$. A variance estimator that allows $e_{it}$ to be correlated is discussed in baing-ecta:06.

Implications for Factor-Augmented Regressions

The results have implications for empirical work. Consider the infeasible regression

equation*[equation* omitted — 79 chars of source]

where $W_t$ is observed but $F^0_t$ is not, $\mathbb E(W_t\epsilon_{t+h})=0$ and $\mathbb E(F^0_t\epsilon_{t+h})=0$. The feasible regression upon replacing $F^0_t$ with $\tilde {\bm F}$ is

eqnarray[eqnarray omitted — 191 chars of source]

baing-ecta:06 shows under the strong factor assumption $\tilde {\bm F}$ can be used in a second step regression to obtain standard normal inference without the need for standard error adjustments if $\sqrt{T}/N\rightarrow 0$. To establish the comparable conditions when $1 \ge \alpha >0$, we need to analyze the correlation between the regressors $(\tilde F_t, W_t)$ and the errors $\epsilon_{t+h}$ as well as with the first step estimation error $(\bm H'_{NT}F_t^0-\tilde F_t)$.

lemmaLet $\hat z_t=(\hat F_t^{\prime}, W_t^\prime)^\prime$ be used in place of $z_t=(F_t^\prime,W_t^\prime)^\prime$ in the factor-augmented regression ((ref)), and denote $\delta^0=(\gamma^\prime \mathbf H_{NT}^{\prime -1} ,\beta^\prime)^\prime$. Suppose $\bm W^\prime \bm e \bm\Lambda^0=\sum_{i=1}^N \sum_{t=1}^T W_t \Lambda_i^{0\prime}e_{it}=O_p(\sqrt{N^\alpha T})$ and Assumptions A, B, and C hold. Then \[ \sqrt{T}(\hat\delta-\delta^0)\smash{\mathop{\longrightarrow}\limits^d} N(0,\mathbf J^{'^{-1}} \mathbf \Sigma_{zz}^{-1}\mathbf \Sigma_{zz,\epsilon}\mathbf \Sigma_{zz}^{-1}\mathbf J^{-1})\] where $\mathbf J$ is the limit of $\mathbf J_{NT}=\text{diag}(\mathbf H_{NT}', \mathbf I_{\text{dim}(W)})$.

To understand the result, we need to show that the errors in ((ref)) are asymptotically uncorrelated with the regressors. Notice that \[ \frac{1}{\sqrt{T}}\sum_{t=1}^T \tilde F_t \epsilon_{t+h} = \mathbf H_{NT}' \frac{1}{\sqrt{T}} \sum_{t=1}^T F^0_t \epsilon_{t+h} +\frac{1}{\sqrt{T}}\sum_{t=1}^T (\tilde F _t-\mathbf H_{NT}' F_t^0) \epsilon_{t+h}. \] We need to show that the second term is $o_p(1)$ and that $W_t$ is asymptotically uncorrelated with $(F_t^{0\prime}\mathbf H_{NT}- \tilde F_t^{\prime})$. More precisely we need to show that each of the following three terms is $o_p(1)$: \[ (i), \frac{1}{\sqrt{T}}\sum_{t=1}^T (\tilde F _t-\mathbf H_{NT}' F_t^0) \epsilon_{t+h}, \quad (ii), \frac{1}{\sqrt{T}}\sum_{t=1}^T \tilde F_t(F_t^{0\prime}\mathbf H_{NT}- \tilde F_t^{\prime}), \quad (iii), \frac{1}{\sqrt{T}}\sum_{t=1}^T W_t (F_t^{0\prime}\mathbf H_{NT}- \tilde F_t^{\prime}).\]

Note first that $ \frac 1 {\sqrt{T}} \epsilon' (\tilde {\bm F}- \bm F^0 \bm H_{NT,\ell}) =o_p(1)$ upon replacing $e_i=(e_{i1},..., e_{iT})'$ by $\epsilon =(\epsilon_{1+h},...\epsilon_{T+h})$ in Lemma (ref)(iii). Furthermore, $ \frac 1 {\sqrt{T}} \bm W' (\tilde {\bm F}- \bm F^0 \bm H_{NT,\ell}) =o_p(1)$ upon replacing $\bm F^0$ by $\bm W=(W_1,\ldots, W_T)^\prime$ in Lemma (ref)(i). Hence all three terms are $o_p(1)$ under the assumptions of the analysis. As a consequence $\tilde {\bm F}$ can be used in the augmented regression and yield standard normal inference as though it were $\bm F$, though the conditions $\alpha>1/2$ and $\sqrt{T}/N^\alpha\rightarrow 0$ are stronger than for the $\alpha=1$ case. However, if we only want consistent estimates instead of root-$T$ consistency and normality, Assumption A.4 is sufficient; there is no need for Assumption C or $\alpha>1/2$.

Heterogeneous $\alpha$

The foregoing analysis assumes that all loadings have the same strength as indicated by the constant $\alpha$. In this section, we allow $\alpha$ to vary across factors. Let $ 1\ge \alpha_1 \ge \alpha_2 \ge \cdots \ge \alpha_r>0$ so that the weakest loading has strength $\alpha_r>0$. Now define the $r\times r$ normalization matrix \[ \bm \bm B_N={\rm diag}(N^{\alpha_1/2},...,N^{\alpha_r/2}), \] noting that $\|\bm B_N\|\le N^{\alpha_1/2}$ and $\|\bm B_N^{-1}\|\le N^{-\alpha_r/2}$. Because of varying loadings strength, the bounds in Assumption A3 need to be replaced by the following:

\paragraph{Assumption A3':} For each $t$, (i) $E\|\bm B_N^{-1}\sum_i \Lambda^0_i e_{it}\|^2\le M$, (ii) $\frac{1}{NT} e_t^\prime\bm e^\prime\bm F^0=O_p(\delta_{NT}^{-2})$; for each $i$, (iii) $E\|T^{-1/2} \sum_t F^0_t e_{it}||^2\le M$, (iv) $\frac{1}{T}e_i'\bm e\bm \Lambda^0\bm B_N^{-1}=O_p(\frac 1 { N{^{\alpha_r/2}}}) +O_p(\sqrt{\frac {N^{\alpha_1}}{T N^{\alpha_r}} })$; (v) $\bm \Lambda^{0'} \bm e ' \bm F^0 = \sum_{i=1}^N\sum_{t=1}^T \Lambda_i^0 F_t^{0'} e_{it} =O_p( \sqrt{N^{\alpha_1} T})$.

Changes to Assumption A2 and A4 are also needed. We will maintain A2(i) and A2(iii), but A2.(ii) previously stated as $\frac{\bm \Lambda^{0'}\bm \Lambda^0}{N^\alpha}=\Sigma_{\Lambda}$ for $\alpha \in (0,1]$ is now replaced by \paragraph{Assumption A2'.(ii):} $ \bm B_N^{-1}\bm \bm \Lambda^{0'} \bm \bm \Lambda^0 \bm B_N^{-1} \smash{\mathop{\longrightarrow}\limits^p} \bm \Sigma_\Lambda>0$, where $\bm \Sigma_\Lambda$ is diagonal.

Instead of $\frac{N}{N^{\alpha}}\frac{1}T\rightarrow 0$, we need to replace Assumption A4 by \paragraph{Assumption A4':} As $N, T\rightarrow \infty$, $\frac{N}{N^{\alpha_r}}\frac{1}{T}\rightarrow 0$ as $N,T\rightarrow \infty$.

We will need new identities to accommodate varying $\alpha$. In Lemma (ref), we see that the first $r$ largest eigenvalues of $\frac{1}{N T} \bm X\bm X^\prime$ are determined by the matrix $\frac1 {NT} \bm F^0(\bm\Lambda^{0\prime}\bm \Lambda^0)\bm F^{0'}$. A matrix with equivalent eigenvalues that we now use is $\frac 1 N (\bm B_N^{-1}\bm \bm \Lambda^{0'} \bm \bm \Lambda^0 \bm B_N^{-1})(\bm B_N\bm F^{0'} \bm F^0/T) \bm B_N $. By assumption, $(\bm B_N^{-1} \bm \bm \Lambda^{0'} \bm \bm \Lambda^0 \bm B_N^{-1} )$ and $\bm F^{0'} \bm F^0/T$ both converge in probability to positive definite matrices. Furthermore, the first $r$ eigenvalues is the order of $\bm B_N^{2}/N$. Since $\bm D_r^2$ is a positive definite matrix, the comparable result to $\frac{N}{N^\alpha} \bm D_{NT,r}^2 \smash{\mathop{\longrightarrow}\limits^p} \bm D_r^2>0$ in Lemma (ref) is \[ N \bm B_N^{-2} \bm D_{NT,r}^2 \smash{\mathop{\longrightarrow}\limits^p} \bm D_r^2 >0. \]

Next, we need to find a result comparable to positive definiteness of $\bm Q$ in Lemma (ref), where $\bm Q$ is the limit of $\tilde {\bm F}'\bm F^0/T$. To proceed, we use $\frac{1}{NT} \bm X\bm X'\tilde {\bm F} =\tilde {\bm F} \bm D_{NT,r}^2$ and $\tilde {\bm F}'\tilde {\bm F}/T=\bm I_r$ to obtain

equation*[equation* omitted — 156 chars of source]

The leading term of the left hand side is \[( \bm B_N^{-1} \tilde {\bm F}' \bm F^0\bm B_N/T) (\bm B_N^{-1} \bm\Lambda^{0\prime}\bm \Lambda^0 \bm B_N^{-1}) (\bm B_N \bm F^{0'} \tilde {\bm F} \bm B_N^{-1} /T) \smash{\mathop{\longrightarrow}\limits^p} \bm D_r^2 \] Since $(\bm B_N^{-1} \bm\Lambda^{0\prime}\bm \Lambda^0 \bm B_N^{-1})\smash{\mathop{\longrightarrow}\limits^p} \bm \Sigma_\Lambda>0$, and $\bm D_r^2 >0$, the limit of $\bm B_N^{-1} \tilde {\bm F}' \bm F^0\bm B_N /T$ is invertible. We continue to denote its limit as $\bm Q$.

Average Error in Estimation of the Factor Space

Finding an appropriate definition of the rotation matrix to accommodate heterogeneous $\alpha$ is delicate because the moments have to be normalized in accordance with the strength of the loadings. We proceed by right multiplying $\bm B_N^{-1}$ to $ \frac 1 {TN} \bm X \bm X' \tilde {\bm F} = \tilde {\bm F} \bm D_{NT,r}^2 $ to get \[ \frac 1 {TN} \bm X \bm X' \tilde {\bm F}\bm B_N^{-1} = \tilde {\bm F} \bm D_{NT,r}^2 \bm B_N^{-1} \equiv \tilde {\bm F} \bm B_N (\bm B_N^{-2} \bm D_{NT,r}^2). \] Expanding $\bm X \bm X'$, and multiplying $N$ on each side gives

eqnarray*[eqnarray* omitted — 386 chars of source]

We now define

equation[equation omitted — 281 chars of source]

Note that $\| \bar \bm H_{NT}||=O_p(1)$. We obtain the following average errors in estimating the space spanned by the factors and the loadings.

propositionLet $\bar \bm H_{NT}$ be defined as in ((ref)) and $\bm H_{NT}= \bm B_N \bar \bm H_{NT} \bm B_N^{-1}$. Then under Assumptions A', \begin{itemize} • $ \frac 1 T \| (\tilde {\bm F} -\bm F^0 \bm H_{NT})\|^2 =O_p(\frac 1 {N^{\alpha_r}} ) + O_p( \frac {N^{2-2\alpha_r}} {T^2} )$; • $ \frac 1 N \|\tilde {\bm \Lambda}^\prime - \bm H_{NT,3}^{-1} \bm \Lambda^{0'}\|^2 = O_p(\frac {N^{\alpha_1-\alpha_r}} T )+\frac 1 {N^{(1+\alpha_r)}} {O_p(1)}.$ \end{itemize}

When $\alpha$ is homogeneous, $\bm B_N=\sqrt{N^\alpha} \bm I_r$ and $\bar \bm H_{NT}$ is $\bm H_{NT,0}= (\frac{\bm \Lambda^{0'}\bm \Lambda^0) }{N^\alpha}) $ $(\frac{\bm F^{0'}\tilde {\bm F} }{T})\frac{N^{\alpha}}{N}\bm D_{NT,r}^{-2}$ defined earlier since $ \frac{\bm \Lambda^{0'}\bm e' \tilde {\bm F} }{NT}\bm D_{NT,r}^{-2}$ is negligible. With heterogeneous $\alpha$, the second term is still negligible, but including it in $\bar {\bm H}_{NT}$ makes it possible to use arguments that lead to better convergence rates. In particular, the definition of $\bar \bm H_{NT}$ implies:

eqnarray[eqnarray omitted — 322 chars of source]

The Appendix shows that $\|a\|=O_p(\sqrt{T})$ and $\|b\|=\frac{\max(N,T)} {\sqrt{T} N^{\alpha_r/2}}$, which together imply

equation[equation omitted — 158 chars of source]

Since $ \bm H_{NT} \bm B_N= \bm B_N \bar \bm H_{NT}$, it follows that \[ \| \tilde {\bm F} -\bm F^0 \bm H_{NT}\| =\| (\tilde {\bm F} \bm B_N-\bm F^0 \bm B_N \bar \bm H_{NT})\bm B_N^{-1}\|\le \| (\tilde {\bm F} \bm B_N-\bm F^0 \bm B_N \bar \bm H_{NT})\| \|\bm B_N^{-1}\|. \] Combining ((ref)) with the fact that $\|\bm B_N^{-1}\|\le N^{-\alpha_r/2}$ yields \[ \frac 1 {\sqrt{T}}\| (\tilde {\bm F} -\bm F^0 \bm H_{NT})\| \le O_p(1)N^{-\alpha_r/2} +\frac { N^{1-\alpha_r}} T O_p(1). \] Part (i) of the Proposition follows. A generalization of Lemma (ref) concerning the factor augmented regression is that $\alpha_k>1/2$ and $\sqrt{T}/N^{\alpha_k}\rightarrow 0$ will be needed for standard normal inference if we use estimates of the largest $k$ factors in two step regressions as though $\bm F_1,\ldots, \bm F_k$ (where $k\le r$) were observable.

For part (ii), we have $\tilde {\bm \Lambda}^\prime = \bm H_{NT,3}^{-1} \bm \Lambda^{0'} + \tilde {\bm F}^\prime\bm e /T $, where $ \bm H_{NT,3}^{-1}=(\tilde {\bm F}^\prime \bm F^0/T)^{-1}$. Adding and subtracting terms

eqnarray*[eqnarray* omitted — 283 chars of source]

Taking norms, \[ \frac 1 {\sqrt{N}} \|\tilde {\bm \Lambda}^\prime - \bm H_{NT,3}^{-1} \bm \Lambda^{0'}\|\le \|\bm H_{NT}\| \frac{\| \bm F^{0'} \bm e\|}{T\sqrt{N}} + \frac{ \|\bm B_N^{-1}\| \|(\tilde {\bm F} \bm B_N- \bm F^0 \bm H_{NT} \bm B_N)' \bm e\| } {T\sqrt{N} }. \] Now $ \|\bm H_{NT}\| \frac{\| \bm F^{0'} \bm e\|}{T\sqrt{N}}=N^{(\alpha_1-\alpha_r)/2} O_p(1/\sqrt{T})$ since $\|\bm H_{NT}\|\le \|\bm B_N\|\|\bar \bm H_{NT}\| \bm B_N^{-1}\|\le N^{(\alpha_1-\alpha_r)/2}O_p(1)$. The Appendix shows that, under$\frac N {TN^{\alpha_r}} \rightarrow 0$ as in Assumption A4',

eqnarray[eqnarray omitted — 203 chars of source]

Thus $ \frac 1 {\sqrt{N}} \|\tilde {\bm \Lambda}^\prime - \bm H_{NT,3}^{-1} \bm \Lambda^{0'}\|\le N^{(\alpha_1-\alpha_r)/2}O_p(1/\sqrt{T})+\frac 1 {N^{(1+\alpha_r)/2}} {O_p(1)} $. Squaring gives the desired result.

The thrust of Proposition (ref) is that as far as consistent estimation of the factor space is concerned, the only loading strength that matters is that of the weakest, $\alpha_r$. Provided that $\alpha_r>0$, and $N^{1-\alpha_r}/T\rightarrow 0$, the average error in estimating the factor space will vanish, albeit at a slower rate than in the strong loadings case.

Distribution Theory

We will use $\bm H_{NT,3}$ to obtain distributional results but we first need to establish its relation to $\bm H_{NT}$ when $\alpha$ is heterogeneous.

lemmaThe following holds under Asumptions A', \begin{itemize} • $ \bm H_{NT,3}- \bm H_{NT} = N^{\frac 1 2 \alpha_1 -\alpha_r}O_p(1)+\frac { N^{1+\frac 1 2 (\alpha_1-3\alpha_r)} } {T} O_p(1)+ N^{\frac 1 2 (\alpha_1-3\alpha_r)} O_p(1)$. • $ \bm H_{NT} \bm H_{NT,3}^{-1}=I_r +N^{\frac 1 2 \alpha_1 -\alpha_r}O_p(1)+\frac { N^{1+\frac 1 2 (\alpha_1-3\alpha_r)} } {T} O_p(1)+ N^{\frac 1 2 (\alpha_1-3\alpha_r)} O_p(1) $ \end{itemize}

The lemma says $\bm H_{NT,3}- \bm H_{NT} = o_p(1)$ and $ \bm H_{NT} \bm H_{NT,3}^{-1}=I_r+o_p(1)$ if $\alpha_r>\alpha_1/2$ and $T$ is sufficiently large. We will use this result to prove consistency of $\tilde C_{it}$.

To derive the limiting distributions, we modify Assumptions B and C with the following:

\paragraph{Assumption B'.} The following holds for each $i$ and $t$ as $N, T\rightarrow \infty$:

eqnarray*[eqnarray* omitted — 254 chars of source]

\paragraph{Assumption C':} (i) $ \alpha_r >\frac 1 2$, (ii) $ \frac {\sqrt{T}} {N^{\alpha_r}} \rightarrow 0$, (iii) $ \frac 1{\sqrt{T}} \frac N {N^{\alpha_r}} \rightarrow 0$, and (iv) $ \frac 1 T \frac{N^{3/2}}{N^{\alpha_r}} \rightarrow 0$.

For the limiting distribution of $\tilde F_t$, from the $t$-th row of ((ref)), we have \[ \bm B_N( \tilde {\bm F}_t-\bm H_{NT}' \bm F^0_t) = (N \bm B_N^{-2} \bm D_{NT,r}^2)^{-1} (\bm B_N^{-1} \tilde {\bm F}'\bm F^0 \bm B_N/T) \bm B_N^{-1} \bm \Lambda^0 \bm e_t + (N \bm B_N^{-2} \bm D_{NT,r}^2)^{-1} \bm B_N^{-1} \tilde {\bm F}' \bm e \bm e_t/T. \] Now $(N \bm B_N^{-2} \bm D_{NT,r}^2)^{-1}\smash{\mathop{\longrightarrow}\limits^p} \bm D_r^{-2}$, $(\bm B_N^{-1} \tilde {\bm F}'\bm F^0 \bm B_N/T)\smash{\mathop{\longrightarrow}\limits^p} \bm Q$. Furthermore, by Assumption B', $ \bm B_N^{-1} \bm \Lambda^0 \bm e_t\smash{\mathop{\longrightarrow}\limits^d} N(0,\bm \Gamma_t)$. The first term is thus asymptotically normal. As shown in the Appendix, the second term is

equation[equation omitted — 148 chars of source]

which is $o_p(1)$ under Assumption C'. So we have \[ \bm B_N( \tilde {\bm F}_t-\bm H_{NT}' \bm F^0_t)\smash{\mathop{\longrightarrow}\limits^d} \mathcal N( 0, \bm D_{r}^{-2}\bm Q\bm \Gamma_t\bm Q'\bm D_{r}^{-2}). \]

For the distribution of $\tilde\Lambda_i$,

equation[equation omitted — 254 chars of source]

Multiplying by $\bm H_{NT}^{\prime-1} $ on each side, \[ \sqrt{T}\bm H_{NT}^{\prime-1} (\tilde \Lambda_i -\bm H_{NT,3}^{-1} \Lambda_i^0) = \frac{ 1}{\sqrt{ T}} \sum_{t=1}^T F_t^0 e_{it} + \frac{\bm H_{NT}^{\prime-1} \bm B_N^{-1}(\tilde {\bm F}\bm B_N-\bm F^0\bm H_{NT} \bm B_N)^\prime \bm e_i}{\sqrt{T}}.\] The first term is asymptotically normal by Assumption B'. Using Lemma (ref)(ii) in appendix, the second term is bounded by

equation*[equation* omitted — 262 chars of source]

which is $o_p(1)$ under Assumption C'. Summarizing results we have

propositionSuppose that Assumptions A', B', and C' hold. Then \begin{itemize} • $ \bm B_N (\tilde F_t-\bm H_{NT}'F^0_t)\smash{\mathop{\longrightarrow}\limits^d} \mathcal N(0, \bm D_{r}^{-2}\bm Q\bm \Gamma_t\bm Q'\bm D_{r}^{-2})$; • $\sqrt{T}\bm H_{NT}^{\prime-1} (\tilde \Lambda_i-\bm H_{NT,3}^{-1}\Lambda_i^0)\smash{\mathop{\longrightarrow}\limits^d} \mathcal N(0,\bm \Phi_i ).$ \end{itemize}

Part (i) of the Proposition says that $\tilde F_{1t}$ associated with the strongest loading will converge to the normal distribution at a faster rate of $\sqrt{\alpha_1}$ than the $\tilde F_{rt}$ which only converges at rate $\sqrt{\alpha_r}$.

Assumption C' is needed for asymptotic normality but is stronger than is necessary for individual consistency of $\tilde \Lambda_i$ and $\tilde F_t$. All that is needed for $\tilde \Lambda_i$ to be consistent is $\alpha_r>0$. To see this, we can divide ((ref)) by $T^{1/2}$. The first term becomes $\frac 1 {\sqrt{T}}\|\bm H_{NT}\|O_p(1) = \sqrt{\frac {N^{\alpha_1}}{T N^{\alpha_r}} } O_p(1)=o_p(1)$. Now $\|\bm B_N^{-1}\|=N^{-\alpha_r/2}$, and the second term is equal to $\bm B_N^{-1}$ multiplied by the term analyzed in Lemma (ref)(i) in the Appendix, and this term is $o_p(1)$ provided $\alpha_r>0$. Similarly, we only require $\alpha_r>1/3$ (together with $\frac {N^{3/2}} {TN^{3\alpha_r/2}}\rightarrow 0$) for $\tilde F_t$ to be consistent. To see this, we first multiply ((ref)) by $\bm B_N^{-1}$, which is $O(N^{-\alpha_r/2})$. Then \[ \bm B_N^{-2} \tilde {\bm F} '\bm e \bm e_t/T= \frac {N^{3/2}} {TN^{3\alpha_r/2}}O_p(1) + N^{\frac {1 -3\alpha_r} 2} O_p(1). \] Thus if $\alpha_r>1/3$ and $\frac {N^{3/2}} {TN^{3\alpha_r/2}}\rightarrow 0$, the above is $o_p(1)$, which implies $ \tilde {\bm F}_t-\bm H_{NT}' \bm F^0_t =o_p(1)$. Both results are similar to the homogenous case, but with $\alpha$ replaced by $\alpha_r$.

We next consider estimating the common component. Adding and subtracting terms,

eqnarray*[eqnarray* omitted — 309 chars of source]

The first two terms are both $o_p(1)$ by Proposition (ref). The last term is $o_p(1)$ by part (ii) of Lemma (ref).

We close this section with some remarks about our results in relation to those in the literature. The PC estimator imposes the normalization restrictions that $\bm \Lambda'\bm \Lambda$ is diagonal, and $\bm F'\bm F/T=I_r$. If the true data generating process coincides with these restrictions, that is, $\bm \Lambda^{0'}\bm \Lambda^0$ being diagonal and $\bm F^{0'}\bm F^0/T=I_r$, then all the rotation matrices introduced in this paper will be asymptotically an identity matrix as shown in baing-joe:13. Our analysis above has refrained from making the assumption that the true DGP coincides with the normalization restrictions so that we can compare $\tilde {\bm F}$ to $ \bm F^0\bm H_{NT}$, but not of $\tilde {\bm F}_k$ to $\bm F_k$. The challenge lies in being able to define $\bm H_{NT}$ in a way consistent with the data generating process so that $\frac{1}{T}\|\tilde {\bm F}-\bm F^0\bm H_{NT}\|^2$ remains the metric for assessing estimation error.

Our approach contrasts with a growing body of work on this problem. uematsu-yamagata:inf, uematsu-yamagata:est consider regularized estimation and inference of sparsity induced weak factors. They assume in our notation that $ E[\bm H \bm F^0_t \bm F^{0'}_t\bm H']=\bm I_r$ and $\bm H^{-1\prime} \bm \Lambda^{0'} \bm \Lambda^0 \bm H^{-1}$ is a diagonal matrix with $\bm D_{jj}^2 N^{\alpha_j}$ in the $j$-th diagonal, which implicitly make unspecified restrictions about $\bm H$ and $\bm F^0$. Even as it is, the average error bound for their estimator of $\bm F$ already restricts the relation between $\alpha_1$ and $\alpha_r$.\footnote{ They require that $\alpha_1+\max(1, \tau)/2<3 \alpha_r/2+\tau/2$ where $T=N^{\tau}$ for some $\tau>0$. As noted in their Remark 1, the upper bound of $\alpha_1-\alpha_r$ of 1/4 is attainable when $\alpha_1=1$ with $\tau=(3/4,1]$, the same as PC. Their lower bound of $\alpha_r$ is attained when $\alpha_1=\alpha_r$ and $\tau=2/3$,} Their assumptions applied to PC estimation yields a strict lower bound of $\alpha_r> 1/2$, stronger than the $\alpha_r>0$ result that we obtain. As the authors noted, without the implicit restrictions on $\bm H$ and $\bm F^0$ that are not innocuous, additional assumptions on $[\alpha_1,\alpha_r]$ would otherwise be required.

Most related to our result is freyaldenhoven:22 whose goal is to determine the number of (local) factors whose loadings are of varying strength. He assumes $\alpha_k>1/2$, $N/T\rightarrow c$, $\bm F^{0'}\bm F^0/T$ is truly an identity matrix and $\bm \Lambda^{0\prime}\bm \Lambda^0$ is truly diagonal with the implication that $\bm H$ is also an identity matrix. This simplifies the analysis as $\tilde {\bm F}_{kt}$ estimates $\bm F_{kt}$, not just a rotation of it. With this additional identifying assumption, he finds that $\alpha_k>1/2$ is required for consistency of $\tilde {\bm F}_{kt} $ for $\bm F_{kt}^0$. For consistency, we obtain the result of $\alpha_r>1/3$ without restricting $\bm H$ to be an identity matrix.

Simulation Experiments

To verify the asymptotic results, we conduct two simulation experiments with 5000 replications of data $X^0_{it}=\Lambda_i^{0'}F^0_t+\sigma_i e^0_{it}$ with $r=3$. Let $\bm D^2$ and $\bm B$ are diagonal matrices, $B_{jj}=N^{\alpha_j/2}$, $(j=1,2,3$). Two data generating processes are considered.

itemize• DGP1: $\bm F^0=\sqrt{T}\bm U\bm D$ and $\bm \Lambda^0= \bm V\bm B$ where $\bm U$ and $\bm V$ are random orthonormal $T\times r$ and $N\times r$ matrices respectively. Hence $\bm F^{0'} \bm F^0/T=\bm D^2$ and $\bm B^{-1} \bm \Lambda^{0'} \bm \Lambda^0 \bm B^{-1}$ is $\bm I_r$. • DGP2, $F_t^0\sim N(0,I_3)$, $ \Lambda^0_i\sim N(0,I_r) \bm D \bm B/\sqrt{N}$. Hence $\bm F^{0'}\bm F^0/T\approx I_r$ and \newline $\bm B^{-1} \bm \Lambda^{0'}\bm \Lambda^0\bm B^{-1} \approx \bm D^2$ is diagonal.

The common component $C_{it}^0=\Lambda_i^{0\prime} F_t^0$ has variance $\sigma^2_{Ci}$ calculated for each $i$ from the time series data. The importance of the common component is $\bar R^2_C$, defined as the ratio of the mean of the variance of the common component of each series to the mean of the variance of each series.

center[center omitted — 627 chars of source]

By design of $\bm B$, $F_1^0$ contributes more than $F_2^0$ and $F_3^0$ to the variations in the data. In the strong factor case, the parameterizations yield a common component that accounts for a bit over half of the total variations in the data in both DGPs and reduces to less than 10% when all loadings are equally weak. In the heterogeneous case, strong and weak loadings co-exist and the common component explains about one-third of the total variations in DGP1 and half in DGP2.

For each factor $j=1,\ldots, r$, we regress the $j$-th column of $\tilde F$ on $F^0$. The coefficients estimate the $j$-th column of $\bm H_{NT,3}$, so the $R^2$ from the regression is an assessment of fit. Likewise, the $j$-th column of $\tilde \Lambda$ is regressed on $\Lambda^0$ and the coefficients estimate the $j$-th row of $\bm H^{-1}_{NT,4}$. The residuals from these regressions are then the (non-normalized) estimation error. Tables (ref) and (ref) report both the $R^2$ for each $j$, as well as $M(\tilde{\bm F})$=trace(top)/trace(bottom), a multivariate measure of fit between $\bm F^0 $ and $\tilde {\bm F}$, where top= $\bm F^{0'} \tilde{\bm F} (\tilde {\bm F}^{'}\tilde{\bm F})^{-1} \tilde {\bm F'}\bm F^0$ and bottom = $\bm F^{0'}\bm F^0$. The statistic $M(\tilde {\bm \Lambda})$ is similarly defined. The last statistic is the average of correlation between $\tilde { C}_i$ and $ C^0_i$. The distributions of the estimation error for $\tilde F_t$ are shown in Figure (ref) for $t=100$. Figure (ref) shows the distributions of the estimation error for $\tilde \Lambda_i$ at $i=50$. In both cases, $T=500$ and $N=100$. Both distributions appear symmetric.

We consider eight configurations of $N$ and $T$. In the strong factor case, all statistics indicate that the factors are precisely estimated. Under Assumption A4, $N^{1-\alpha_r}/T$ must tend to zero. Hence for given $N$, a larger $T$ will give faster convergence of the estimates to a rotation of the true values. The results bear this out. In the homogeneous $\alpha$ case reported in the middle panel, the estimation errors are similar across factors. This is different for the weak-heterogeneous case in the bottom panel as estimates of $\bm F_1$ and $\bm \Lambda_1$ are more precise than for $\bm F_3$ and $\bm \Lambda_3$, as suggested by theory. The common component remains precisely estimated with errors closer to the stronger loadings case than the weak-homogeneous case because the dominant factor is strong.

So far, DGP1 assumes that the factors are orthogonal and DGP2 assumes that the factors are asymptotically orthogonal. Since our theory does not require the assumption that $\bm F^{0'}\bm F^0$ is a diagonal matrix, we also modify DGP1 so that $U$ and $V$ are no longer orthogonal. This is achieved by multiplying these orthogonal vectors into two $r\times r$ matrices with random elements below the diagonal. The results, shown in Table (ref), are similar to Table (ref).

Conclusion

A sizable literature has emerged since onatski-joe:12 shows that the PC estimates are inconsistent when the loadings are extremely weak in the sense that $\frac{\bm\Lambda'\bm \Lambda}{N^\alpha} \rightarrow \bm \Sigma_\Lambda$ with $\alpha=0$. This paper establishes conditions under which the PC estimates are consistent and asymptotically normal in more moderate cases when $1\ge \alpha > 0$. The main conclusion is unchanged in the heterogeneous $\alpha$ case, where the one $\alpha$ that determines consistency is that of the weakest loading. The takeaway is that asymptotic normality requires $\alpha_r>1/2$, stronger than is needed for the consistency results.

Allowing $\alpha$ to take on a range of values naturally raises questions about the different criteria available to determine the number of factors. If we want to estimate the number of factors with $\alpha>0$, the criteria in baing-ecta:02 remain useful. While the criteria of onatski-restat and ahn-horenstein:13 try to better separate the bounded from the diverging eigenvalues of $\bm X\bm X'$, baing-joe:19 seek to isolate the factors with a tolerated level of explanatory power by singular value thresholding. This criteria will return an estimate of the `minimum rank', a concept of long standing interest in classical factor analysis, see, e.g., tenberge-kiers. For documenting the number of factors with different strength, methods of bailey-kapetanios-pesaran:16 and uematsu-yamagata:est are available. The criteria in freyaldenhoven:22 determine the number factors with $\alpha_r>1/2$. But as seen above, the $\alpha$ required for consistent estimation of the factor space can be smaller than the one needed for asymptotic normality. Ultimately, the desired number of factors depends on the objective of the exercise and the assumptions that the researcher find defensible. It seems difficult to avoid taking a stand on what is meant by weak in practice.