EconBase
← Back to paper

High Dimensional Factor Analysis with Weak Factors

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

72,322 characters · 8 sections · 20 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

High Dimensional Factor Analysis with Weak Factors

\doparttoc \faketableofcontents

\def\spacingset#1{ {#1}} \spacingset{1}

\if11 \fi

\if01 \fi

abstractThis paper studies the principal components (PC) estimator for high dimensional approximate factor models with weak factors in that the factor loading ($\boldsymbol{\Lambda}^0$) scales sublinearly in the number $N$ of cross-section units, i.e., $\boldsymbol{\Lambda}^{0\top}\boldsymbol{\Lambda}^0 / N^\alpha$ is positive definite in the limit for some $\alpha \in (0,1)$. While the consistency and asymptotic normality of these estimates are by now well known when the factors are strong, i.e., $\alpha=1$, the statistical properties for weak factors remain less explored. Here, we show that the PC estimator maintains consistency and asymptotical normality for any $\alpha\in(0,1)$, provided suitable conditions regarding the dependence structure in the noise are met. This complements earlier result by onatski2012asymptotics that the PC estimator is inconsistent when $\alpha=0$, and the more recent work by bai2023approximate who established the asymptotic normality of the PC estimator when $\alpha \in (1/2,1)$. Our proof strategy integrates the traditional eigendecomposition-based approach for factor models with leave-one-out analysis similar in spirit to those used in matrix completion and other settings. This combination allows us to deal with factors weaker than the former and at the same time relax the incoherence and independence assumptions often associated with the later.

{\it Keywords: Approximate factor model, leave-one-out analysis, principal components, weak factors/loadings.}

\part

\spacingset{1.6}

Introduction

Approximate factor models are widely used in diverse fields such as economics, finance, biology, and psychology, to name a few. In these models observations of $N$ cross-section units over $T$ time points are represented as the sum of two unobserved components, a common component driven by systematic factors and an idiosyncratic noise component:

equation[equation omitted — 120 chars of source]

For many modern applications, of particular interest is the high dimensional setting when both $N$ and $T$ are large. In response, many estimation methods and inferential tools for the latent factors $\boldsymbol{F}^0:=(f_1^0,\ldots,f_T^0)^\top$, the loadings $\boldsymbol{\Lambda}^0 =(\lambda_1^0,\ldots,\lambda_N^0)^\top$ and the common component $\boldsymbol{M}^0 :=(\lambda_i^{0\top} f_t^0)_{1\le i\le N, 1\le t\le T}$ have been developed. See, e.g., stock1998diffusion,stock2002forecasting,forni2000generalized,bai2002determining,bai2003inferential,bai2008large,bai2019rank.

Arguably, the most natural and popular techniques are based on the principal components (PC) and their use can be traced back at least to connor1986performance,connor1988risk. Asymptotic properties of PC estimators for large dimensional factor model have also been well studied. See, e.g., bai2008large for a recent survey. A common and crucial premise underlying this rich literature is that the factor structure is strong in the sense that both $\boldsymbol{F}^{0\top}\boldsymbol{F}^0 /T$ and $\boldsymbol{\Lambda}^{0\top}\boldsymbol{\Lambda}^0 /N$ are positive definite in the limit. Although this is a reasonable assumption for some applications, it could be problematic for many others. In the past several years, there has been growing interest in the case when the explanatory power of the factors is weak relative to idiosyncratic noise. See, e.g., onatski2012asymptotics,onatski2018asymptotics,giglio2021test,uematsu2022estimation,armstrong2022robust,anatolyev2022factor,bai2023approximate.

To this end, consider a general weak factor structure where $\boldsymbol{F}^{0\top}\boldsymbol{F}^0 /T$ and $\boldsymbol{\Lambda}^{0\top}\boldsymbol{\Lambda}^0 / N^\alpha$ have positive definite limits for some $\alpha \in (0,1)$. The usual strong factor case corresponds to the choice of $\alpha=1$, under which both consistency and asymptotic normality of the PC estimator are now well known. See, e.g., bai2003inferential. On the other hand, onatski2012asymptotics showed that the PC estimator is inconsistent when $\alpha = 0$. More recently, bai2023approximate established the asymptotic normality of the PC estimator when $\alpha \in (1/2,1)$. However, the inferential theory of the PC estimator when $\alpha \in (0,1/2]$ remains unknown. The main objective of this work is to fill this gap and investigate the consistency and asymptotic normality of PC estimators when $\alpha \in (0,1/2]$.

To fix ideas, let us focus the discussion here on the case when $N\asymp T$ although our main development is more general. Our results indicate that, in particular, if the idiosyncratic terms $\epsilon_{it}$s are cross-sectionally and temporally independent, then the PC estimators of both the factor $f_t^0$ and the common component $m_{it}^0$ are asymptotically normal whenever $\alpha>0$. On the other hand, the asymptotic normality of the estimator for the loading $\lambda_i^0$ may depend on its $\ell_2$ norm, $\|\lambda_i^0\|$. Specifically, if $$\|\lambda_i^0\|\lesssim N^{(\alpha_i-1)/2}$$ for some $\alpha_i\le 1$, then its PC estimator is asymptotically normal if $\alpha > \alpha_i/2$. Note that if $\|\lambda_i^0\|$ is of the same order across all cross-section index $i$, in other words the loadings are incoherent, then $\alpha_i=\alpha$ and the asymptotic normality holds again whenever $\alpha>0$. Even if $\|\lambda_i^0\|$s are of different orders, as long as they are bounded, i.e., $\alpha_i=1$, we can still derive the inferential theory for all $i$ when $\alpha>1/2$.

It is worth noting that in deriving the asymptotic normality of $m_{it}^0=\lambda_i^{0\top}f_t^0$, bai2023approximate implicitly assume that $\|\lambda_i^0\|$ is bounded away from zero and infinity which amounts to setting $\alpha_i=1$. However, in light of the weak factor structure assumption, only a vanishing proportion of $\|\lambda_i^0\|$s, at most $N^\alpha$ out of $N$, can be bounded away from zero and hence their result can only be applied to small number of factor loadings if any. Our results, on the other hand, can be applied to more factor loadings.

We also investigate the impact of possible dependence among the noise $\epsilon_{it}$s. More specifically, if they are temporally independent but cross-sectionally dependent, we show that the PC estimator of $f_t^0$ is asymptotically normal for all $\alpha\in (0,1]$, whereas PC estimators of both $m_{it}^0$ and $\lambda_i^0$ are asymptotically normal if $\alpha>\max\{1/3,\alpha_i/2\}$. On the other hand, if $\epsilon_{it}$s are cross-sectionally independent but temporally dependent, then we show that PC estimators of both $f_t^0$ and $m_{it}^0$ are asymptotically normal when $\alpha>1/3$ whereas PC estimator of $\lambda_i^0$ is asymptotically normal when $\alpha>\alpha_i/2$. Moreover, if $\epsilon_{it}$s are cross-sectionally and temporally dependent, we show that PC estimators of both $m_{it}^0$ and $\lambda_i^0$ are asymptotically normal if $\alpha>\max\{1/3,\alpha_i/2\}$ while the PC estimator of $f_t^0$ is asymptotically normal when $\alpha > 1/3$.

These results offer an overall picture of the effect of the strength (or weakness) of the factor structure and the potential impact of the dependence structure of the noise terms. In general, to ensure the asymptotic normality of the PC estimates, weaker dependence among the noise is required for weaker factors.

table[table omitted — 1,476 chars of source]

Moreover, a similar pattern can be founded in the condition for consistency. When $\epsilon_{it}$s are temporally independent, all PC estimators are consistent as long as $\alpha > 0 $. On the other hand, if $\epsilon_{it}$s are cross-sectionally independent but temporally dependent, then we show that the PC estimator of $f_t^0$ is consistent when $\alpha > 1/4$ and that of $m_{it}^0$ is consistent if $\alpha > \max\{0,\alpha_i/4\}$, while that of $\lambda_i^0$ is consistent when $\alpha>0$. In addition, if $\epsilon_{it}$s are cross-sectionally and temporally dependent, we show that the PC estimator of $f_t^0$ is consistent when $\alpha > 1/4$ and that of $m_{it}^0$ is consistent when $\alpha > \max\{1/7,\alpha_i/4\}$ whereas that of $\lambda_i^0$ is consistent if $\alpha>0$.

table[table omitted — 1,117 chars of source]

Our proof strategy combines the traditional approach based on the eigendecomposition of the covariance matrix bai2002determining,bai2003inferential,bai2023approximate with the more recently developed leave-one-out analysis often used in the context of matrix completion abbe2020entrywise,ma2020implicit,chen2019inference,chen2020nonconvex,chen2020noisy. The leave-one-out analysis allows us to derive higher order approximations to the estimation error than the traditional approach which can be used to handle weaker factors. On the other hand, the insights from the traditional approach enables us to do away with the incoherence conditions of the common component and independence assumption of the noise that are often associated with the leave-one-out type of analysis. The technical insights into the advantage of either method may be of independent interests and beneficial to other related problems.

The remainder of this paper is organized as follows. Section (ref) introduces our model and discusses important features of our proof technique in comparison with the traditional approach in bai2003inferential,bai2023approximate. Section (ref) presents the asymptotic properties of the PC estimator for general weak factors when idiosyncratic noises are cross-sectionally and temporally independent. It shows the convergence rates of the estimator and the specific conditions for asymptotic normality. In addition, Section (ref) introduces the leave-neighbor-out technique which allows us to consider the case of dependent noises and studies the asymptotic properties of the PC estimator when the idiosyncratic noises are cross-sectionally or/and temporally dependent. Lastly, we conclude with a few remarks in Section (ref). All proofs are relegated to the Appendix.

In what follows, we use $\left\Vert\cdot\right\Vert_{\rm F}$ and $\left\Vert\cdot\right\Vert$ to denote the matrix Frobenius norm and spectral norm, respectively. For any vector $a$, $\left\Verta\right\Vert$ denotes its $\ell_2$ norm. $a \lesssim b $ and $b \gtrsim a$ mean $\left\verta\right\vert/\left\vertb\right\vert \leq C$ for some constant $C > 0$. $a \asymp b$ means $a \lesssim b $ and $a \gtrsim b$. In addition, $[K] = \{1 , \dots, K \}$ and $\mathcal{O}^{r \times r}$ is the set of $r\times r$ orthonormal matrices.

Factor Model and Method of PC

Denote by $\boldsymbol{X}=(x_{it})_{1\le i\le N, 1\le t\le T}$ and $\boldsymbol{E}=(\epsilon_{it})_{1\le i\le N, 1\le t\le T}$. Then the approximate factor model (ref) can be expressed in matrix form as: $$ \boldsymbol{X}=\boldsymbol{\Lambda}^0\boldsymbol{F}^{0\top}+\boldsymbol{E}. $$ We shall assume that \paragraph{Assumption A.} [Factors and Loadings]

itemize$T^{-1} \sum_{t=1}^T f_t^0 f_t^{0\top} \to_p \boldsymbol{\Sigma}_{\boldsymbol{F}}$ where $\boldsymbol{\Sigma}_{\boldsymbol{F}}$ is a $r \times r$ positive definite matrix; • $N^{-\alpha} \sum_{i=1}^N \lambda_i^0 \lambda_i^{0\top} \to \boldsymbol{\Sigma}_{\boldsymbol{\Lambda}}$ for some $\alpha \in (0,1]$ where $\boldsymbol{\Sigma}_{\boldsymbol{\Lambda}}$ is a $r \times r$ positive definite matrix; • The eigenvalues of $\boldsymbol{\Sigma}_{\boldsymbol{\Lambda}} \boldsymbol{\Sigma}_{\boldsymbol{F}}$ are distinct; • For all $t$, $\mathbb{E} \left\Vertf_t^0\right\Vert^2 \leq C$ for some constant $C>0$. In addition, for each $i$, there is a parameter $\alpha_i\le 1$ such that for some constant $C'>0$, $$ \left\Vert\lambda_i^0\right\Vert \leq C' N^{(\alpha_i - 1)/2}. $$

Here, $\alpha_i$ designates the order of $\lambda_i^0$. Under the setting $\sum_{i=1}^N \lambda_i^0\lambda_i^{0\top} \asymp N^\alpha$, some of the $\lambda_i^0$s, if not all, should decrease as $N$ increases when $\alpha < 1$. Because $\sum_{i=1}^N \left\Vert\lambda_i^0\right\Vert^2 \asymp N^\alpha$ by Assumption A(ii), if the orders of $\lambda_i^0$s are the same across units, we have $\alpha_i = \alpha$ for all $i$. On the other hand, if the orders of $\lambda_i^0$s are heterogeneous, $\lambda_i^0$s would spread around the average order parameter `$\alpha$' due to Assumption A(ii).

Under Assumption A, the common component $\boldsymbol{M}^0 = (m_{it}^0)_{1\le i\le N, 1\le t\le T}$ has reduced rank $r$ because the ranks of $\boldsymbol{\Lambda}^0$ and $\boldsymbol{F}^0$ are $r$. The method of PC proceeds to estimate $\boldsymbol{\Lambda}^0$ and $\boldsymbol{F}^0$ by minimizing the sum of the squared residuals:

equation[equation omitted — 209 chars of source]

subject to the normalization condition that $\boldsymbol{F}^\top\boldsymbol{F}/T=I_r$ and $\boldsymbol{\Lambda}^\top\boldsymbol{\Lambda}$ is diagonal where $\boldsymbol{F}=(f_1,\ldots,f_T)^\top$ and $\boldsymbol{\Lambda}=(\lambda_1,\ldots, \lambda_N)^\top$. The solution to (ref), denoted by $(\widehat{\boldsymbol{F}},\widehat{\boldsymbol{\Lambda}})$, can also be expressed in terms of the singular values and vectors of $\boldsymbol{X}$. More specifically, let $$ \boldsymbol{X} = \mathbf{U}\mathbf{D}\mathbf{V}^{\top} = \mathbf{U}_r \mathbf{D}_r \mathbf{V}_r^{\top} + \mathbf{U}_{N-r} \mathbf{D}_{N-r} \mathbf{V}_{N-r}^{\top} $$ be its singular value decomposition where $\mathbf{D}_r$ is a diagonal matrix of the top-$r$ singular values, $\mathbf{U}_r$, $\mathbf{V}_r$ are the corresponding left and right singular values, respectively. Then $\widehat{\boldsymbol{\Lambda}}=T^{-1/2}\mathbf{U}_r \mathbf{D}_r$ and $\widehat{\boldsymbol{F}}=\sqrt{T} \mathbf{V}_r$.

\paragraph{Eigenecomposition and Rotation.} Note that $\boldsymbol{X}^\top \boldsymbol{X} \widehat{\boldsymbol{F}} = \widehat{\boldsymbol{F}} \mathbf{D}_r^2$. We can derive from this identity that $$ \widehat{\boldsymbol{F}}- \boldsymbol{F}^{0} \mathbf{H}_{\rm BN,0} = \boldsymbol{F}^0 \boldsymbol{\Lambda}^{0\top} \boldsymbol{E} \widehat{\boldsymbol{F}} \mathbf{D}_r^{-2} + \boldsymbol{E}^{\top} \boldsymbol{\Lambda}^0 \boldsymbol{F}^{0\top} \widehat{\boldsymbol{F}}\mathbf{D}_r^{-2} + \boldsymbol{E}^\top \boldsymbol{E} \widehat{\boldsymbol{F}}\mathbf{D}_r^{-2}, $$ where $\mathbf{H}_{\rm BN,0}=\boldsymbol{\Lambda}^{0\top}\boldsymbol{\Lambda}^0\boldsymbol{F}^{0\top}\widehat{\boldsymbol{F}}\mathbf{D}_r^{-2}$; and multiplying both sides of $\boldsymbol{X}-\boldsymbol{\Lambda}^0\boldsymbol{F}^{0\top}=\boldsymbol{E}$ with $\widehat{\boldsymbol{F}}/T$ leads to $$ \widehat{\boldsymbol{\Lambda}} - \boldsymbol{\Lambda}^0 \mathbf{H}^{-1}_{\rm BN,1} = \boldsymbol{E}\boldsymbol{F}^0 \mathbf{H}_{\rm BN,0} /T + \boldsymbol{E}(\widehat{\boldsymbol{F}} - \boldsymbol{F}^0 \mathbf{H}_{\rm BN,0}) /T, $$ where $\mathbf{H}_{\rm BN,1} = (\boldsymbol{F}^{0\top} \widehat{\boldsymbol{F}}/T)^{-1}$. See, e.g., bai2003inferential,bai2023approximate. These decompositions are key to deriving the asymptotic properties of the PC estimates. For example, it follows immediately that, for any $1\le i\le N$, $$ T^{1/2}(\widehat{\lambda}_i-\mathbf{H}^{-\top}_{\rm BN,1}\lambda_i^{0})=T^{-1/2}\mathbf{H}^{\top}_{\rm BN,0}\boldsymbol{F}^{0\top}\mathbf{e}_i+T^{-1/2}(\widehat{\boldsymbol{F}}-\boldsymbol{F}^0\mathbf{H}_{\rm BN,0})^\top\mathbf{e}_i, $$ where $\mathbf{e}_i=(\epsilon_{i1},\ldots,\epsilon_{iT})^\top$. The first term on the right hand side is asymptotically normal by central limit theorem and it therefore suffices to show that the second term is of order $o_p(1)$ to claim the asymptotic normality of PC estimate $\widehat{\lambda}_i$.

To this end, we note that $$ \mathbf{e}_i^\top(\widehat{\boldsymbol{F}}-\boldsymbol{F}^0\mathbf{H}_{\rm BN,0})=\mathbf{e}_i^\top\boldsymbol{F}^0 \boldsymbol{\Lambda}^{0\top} \boldsymbol{E} \widehat{\boldsymbol{F}} \mathbf{D}_r^{-2} + \mathbf{e}_i^\top\boldsymbol{E}^{\top} \boldsymbol{\Lambda}^0 \boldsymbol{F}^{0\top} \widehat{\boldsymbol{F}}\mathbf{D}_r^{-2} + \mathbf{e}_i^\top\boldsymbol{E}^\top \boldsymbol{E} \widehat{\boldsymbol{F}}\mathbf{D}_r^{-2} . $$ The conventional approach proceeds to bound each term on the right hand side. More specifically, it can be shown that (see, e.g., bai2023approximate) $$ \left\Vert\mathbf{e}_i^\top \boldsymbol{E}^{\top} \boldsymbol{\Lambda}^0 \boldsymbol{F}^{0\top} \widehat{\boldsymbol{F}}\mathbf{D}_r^{-2} \right\Vert = O_p\left( \frac{T}{N^\alpha} + \frac{\sqrt{T}}{\sqrt{N^\alpha}} \right), $$ and $$ \left\Vert\mathbf{e}_i^\top \boldsymbol{E}^\top \boldsymbol{E} \widehat{\boldsymbol{F}}\mathbf{D}_r^{-2}\right\Vert = O_p\left( \frac{T}{N^\alpha} + \frac{N}{N^\alpha} \right). $$ This, however, means that when $N\asymp T$, $\alpha > 1/2$ is needed to prove asymptotic normality of $\widehat{\lambda}_i$. Interestingly, this requirement is not inherent to the PC estimates themselves but rather due to the limitation of this particular proof technique. We now describe two main ideas that enable us to handle weaker factor structures.

\paragraph{Alternative Matching Matrix.} Our first observation is that the matching matrix $\mathbf{H}_{\rm BN}$ can be unduly affected by the noise $\boldsymbol{E}$. To alleviate its impact, we shall seek an alternative matching matrix. To this end, consider a balanced version of singular vectors: $$\bm{Y}_r = \mathbf{U}_r \mathbf{D}_r^{1/2},\qquad {\rm and}\qquad \mathbf{Z}_r = \mathbf{V}_r \mathbf{D}_r^{1/2}.$$ It is clear that $ \widehat{\boldsymbol{\Lambda}} = T^{-1/2}\bm{Y}_r \mathbf{D}_r^{1/2}$ and $ \widehat{\boldsymbol{F}} = T^{1/2}\mathbf{Z}_r \mathbf{D}_r^{-1/2}$.

Similarly, let $\boldsymbol{M}^0=\mathbf{U}^0_r \mathbf{D}^0_r \mathbf{V}_r^{0\top}$ be its reduced singular value decomposition. Then $\bm{Y}_r$ and $\mathbf{Z}_r$ can be viewed as estimates of $\bm{Y}_r^0 = \mathbf{U}_r^0 (\mathbf{D}_r^{0})^{1/2}$ and $\mathbf{Z}_r^0 = \mathbf{V}_r^0 (\mathbf{D}_r^{0})^{1/2}$ respectively. More importantly, we can find a matrix $\mathbf{H}^0\in \mathbb{R}^{r\times r}$ that is independent of $\boldsymbol{E}$ such that $\boldsymbol{F}^0=T^{1/2}\mathbf{Z}_r^0\mathbf{H}^0$ and $\boldsymbol{\Lambda}^0=T^{-1/2}\bm{Y}_r^0(\mathbf{H}^{0})^{-\top}$. We then seek a refinement of $\mathbf{H}^0$ by rotating $(\bm{Y}_r,\mathbf{Z}_r)$ to match $(\bm{Y}_r^0, \mathbf{Z}_r^0)$: $$ \mathbf{O} = \operatorname*{arg\,min}_{\mathbf{R} \in \mathcal{O}^{r\times r}} \left\|

bmatrix[bmatrix omitted — 41 chars of source]

\mathbf{R} -

bmatrix[bmatrix omitted — 45 chars of source]

\right\|_{\rm F}^2. $$ Finally, we shall consider a matching matrix $\mathbf{H}=(\mathbf{D}_r^{1/2}\mathbf{O}\mathbf{H}^0)^{-1}$. This choice of matching matrix allows us to translate the estimation error of $\widehat{\boldsymbol{\Lambda}}$ and $\widehat{\boldsymbol{F}}$ into that of $\bm{Y}_r$ and $\mathbf{Z}_r$: $$ \widehat{\boldsymbol{\Lambda}}-\boldsymbol{\Lambda}^0\mathbf{H}^{-\top}=T^{-1/2}(\bm{Y}_r\mathbf{O}-\bm{Y}^0_r)\mathbf{O}^\top\mathbf{D}_r^{1/2}, $$ and $$ \widehat{\boldsymbol{F}}-\boldsymbol{F}^0\mathbf{H}=T^{1/2}(\mathbf{Z}_r\mathbf{O}-\mathbf{Z}^0_r)\mathbf{O}^\top\mathbf{D}_r^{-1/2}. $$

Note that $$ \bm{Y}_r=\boldsymbol{X}\mathbf{Z}_r(\mathbf{Z}_r^\top \mathbf{Z}_r)^{-1}= \boldsymbol{E} \mathbf{Z}_r(\mathbf{Z}_r^\top \mathbf{Z}_r)^{-1} + \bm{Y}^0_r\mathbf{Z}^{0\top}_r\mathbf{Z}_r (\mathbf{Z}_r^\top \mathbf{Z}_r)^{-1}. $$ We can write

gather[gather omitted — 505 chars of source]

where $\widetilde{\mathbf{Z}}_r = \mathbf{Z}_r \mathbf{O}$. Similar to before, the first term on the right hand side is asymptotic normal and it suffices to show that the remaining two terms are of smaller order. The last term can be bounded by virtue of Davis-Kahan type of bounds for $\widetilde{\mathbf{Z}}_r-\mathbf{Z}_r^0$. Bounding the second term turns out to be the key when considering weaker factors ($\alpha\le 1/2$).

\paragraph{Leave-one-out Analysis.} Recall that, with the new matching matrix, we have $$ \sqrt{T}(\widehat{\lambda}_i - \mathbf{H}^{-1} \lambda_i^0 ) = \mathbf{H}^{-1} (\boldsymbol{F}^{0\top} \boldsymbol{F}^{0}/T)^{-1} \boldsymbol{F}^{0\top} \mathbf{e}_i/\sqrt{T} + \mathbf{D}_r^{1/2} \mathbf{O} R_{1,i} + D_r^{1/2}\mathbf{O} R_{2,i}, $$ where $R_{1,i}$ and $R_{1,i}$ are the transpose of $i$-th row of $R_1$ and $R_2$, respectively. Since $\left\Vert\mathbf{D}_r^{1/2}\right\Vert = O_p(N^{\alpha/4} T^{1/4})$ and $\left\Vert\widetilde{\mathbf{Z}}_r - \mathbf{Z}_r^0\right\Vert = O_p ( N^{-\alpha/4} T^{-1/4}\max\{\sqrt{N},\sqrt{T}\})$, a naive bound for the second term is: $$ \left\Vert\mathbf{D}_r^{1/2} \mathbf{O} R_{1,i}\right\Vert \leq \left\Vert\mathbf{D}_r^{1/2}\right\Vert\left\Vert\mathbf{e}_i\right\Vert\left\Vert((\widetilde{\mathbf{Z}}_r^{\top}\widetilde{\mathbf{Z}}_r)^{-1}\widetilde{\mathbf{Z}}^{\top}_r - (\mathbf{Z}_r^{0 \top} \mathbf{Z}_r^0)^{-1}\mathbf{Z}_r^{0 \top} )\right\Vert = O_p\left( \frac{\max\{\sqrt{N},\sqrt{T}\}}{N^{\alpha/2}}\right). $$ This bound however is not tight.

Instead, we shall carry out a leave-one-out analysis to decouple the estimates and a particular noise term. Denote by $\boldsymbol{X}^{(-i)}$ a $N\times T$ matrix whose $i$-th row is $(m^{0}_{it})_{1\leq t \leq T}$ and other rows are $(x_{jt})_{1\leq t \leq T}$ for all $j \neq i$. That is, $\boldsymbol{X}^{(-i)}$ replaces the $i$-th row of $\boldsymbol{X}$ with $(m^{0}_{it})_{1\leq t \leq T}$ to remove the noises of the unit $i$. We shall apply the aforementioned operations to $\boldsymbol{X}^{(-i)}$ leading to corresponding balanced singular vectors $\bm{Y}_r^{(-i)}$ and $\mathbf{Z}_r^{(-i)}$, rotation matrix $\mathbf{O}^{(-i)}$, matching matrix $\mathbf{H}^{(-i)}$ and etc..

We can then write

equation[equation omitted — 140 chars of source]

where $$\Delta_{1}^{(-i)} = \widetilde{\mathbf{Z}}^{(-i)}_r(\widetilde{\mathbf{Z}}^{(-i)\top}_r\widetilde{\mathbf{Z}}^{(-i)}_r)^{-1} - \mathbf{Z}_r^0(\mathbf{Z}_r^{0\top} \mathbf{Z}_r^0)^{-1}$$ and $$\Delta_{2}^{(-i)} = \widetilde{\mathbf{Z}}_r(\widetilde{\mathbf{Z}}^\top_r\widetilde{\mathbf{Z}}_r)^{-1} - \widetilde{\mathbf{Z}}^{(-i)}_r(\widetilde{\mathbf{Z}}^{(-i)\top}_r\widetilde{\mathbf{Z}}^{(-i)}_r)^{-1},$$ $\Delta_{1,t}^{(-i)}$ and $\Delta_{2,t}^{(-i)}$ are the transpose of $t$-th row of $\Delta_{1}^{(-i)}$ and $\Delta_{2}^{(-i)}$, respectively. Note that $\Delta_{2,t}^{(-i)}$ is a higher order difference so that the second term on the right-hand side of (ref) is typically negligible. The first term can now be bounded by exploiting the potential independence between $\mathbf{e}_i$ and $\Delta_{1,t}^{(-i)}$. For simplicity, consider the case when $\epsilon_{it}$s are cross-sectionally independent, then $\mathbf{e}_i$ is independent of $\Delta_1^{(-i)}$. This implies that

gather[gather omitted — 288 chars of source]

Hence, the first term can be negligible even when $\alpha\le 1/2$.

Independent Noise

We now show how the ideas described in Section (ref) can be used to develop statistical properties for general weak factors. It is instructive to start with the case where the idiosyncratic noises $\epsilon_{it}$s are cross-sectionally and temporally independent.

\paragraph{Assumption B.} [Noise]

itemize$\mathbb{E}[\epsilon_{it} | \boldsymbol{M}^0] = 0$ and $\mathbb{E}[\epsilon_{it}^2 | \boldsymbol{M}^0] = \mathbb{E}[\epsilon_{it}^2] \leq C$ for some constant $C>0$. In addition, $(\epsilon_{it})_{i \leq N, t\leq T}$ is independent across $i$ and $t$. • With probability converging to 1, $\left\Vert\boldsymbol{E}\right\Vert \lesssim \max\{\sqrt{N},\sqrt{T} \}$.

We first consider the rate of convergences of the PC estimator defined in the previous section.

theorem[Convergence rate of PC estimator] Suppose that Assumptions A and B are satisfied. If $\max\{N, T \}=o(N^\alpha T)$, then \begin{eqnarray*} \left\Vert \widehat{f}_t - \mathbf{H}^{\top} f_t^0 \right\Vert &=& O_p \left( \frac{1}{\sqrt{N^\alpha}} + \frac{\max\{N, T \}}{N^\alpha T} \right),\\ \left\Vert\widehat{\lambda}_i - \mathbf{H}^{-1} \lambda_i^0 \right\Vert &=& O_p \left( \frac{1}{\sqrt{T}} + \sqrt{\frac{N^{\alpha_i}}{N}} \frac{\max\{N, T \}}{N^\alpha T} \right),\\ \left\Vert\widehat{m}_{it} - m^0_{it}\right\Vert &=& O_p\left( \frac{1}{\sqrt{T}} + \sqrt{\frac{N^{\alpha_i}}{N}} \frac{1}{\sqrt{N^\alpha}} + \sqrt{\frac{N^{\alpha_i}}{N}} \frac{\max\{N, T \}}{N^\alpha T} + \frac{\max\{\sqrt{N}, \sqrt{T} \}}{\sqrt{N^\alpha} T} \right). \end{eqnarray*}

The convergence rates of $\widehat{\lambda}_i$ and $\widehat{m}_{it}$ depend on the size of $\alpha_i$. As $\alpha_i$ decreases, the estimators converge to the corresponding parameters more quickly. In addition, since $N^{\alpha_i} \leq N$ for all $i$, the condition for the consistency of all three estimators for all $i$ and $t$ is $\alpha > 0$ and $\max\{N, T \}=o(N^\alpha T)$. Hence, if $N \asymp T$, the condition $\alpha > 0$ is sufficient for all estimators to be consistent. On the other hand, the traditional proof method requires $\alpha > 1/3$ for the consistency of $\widehat{f}_t$. Note that the condition $\alpha>0$ is also necessary in light of the results by onatski2012asymptotics.

Next, we present the asymptotic normality of the PC estimator. For this purpose, we need the following assumptions.

\paragraph{Assumption C.} [CLT for weak factors] As $N,T \rightarrow \infty$, $$ \frac{1}{\sqrt{N^\alpha}} \sum_{i=1}^N \lambda_i^0 \epsilon_{it} \to_d \mathcal{N}(0,\boldsymbol{\Phi}_{\boldsymbol{\Lambda},t}), \qquad{\rm and}\qquad \frac{1}{\sqrt{T}} \sum_{t=1}^T f_t^0 \epsilon_{it} \to_d \mathcal{N}(0,\boldsymbol{\Phi}_{\boldsymbol{F},i}),$$ where $\boldsymbol{\Phi}_{\boldsymbol{\Lambda},t}$ and $\boldsymbol{\Phi}_{\boldsymbol{F},i}$ are $r \times r$ positive definite matrices.

For the first CLT, we use a normalization of $N^{\alpha/2}$ instead of $N^{1/2}$ to be consistent with Assumption A(ii). The next assumption presents specific conditions for the size of $\alpha$, $\alpha_i$, $N$ and $T$.

\paragraph{Assumption D.} [Parameter size]

itemize• For the inference of $\lambda_i^0$, we assume that $$ \frac{\max\{N,T\}}{N^{\alpha}T} \rightarrow 0, \qquad {\rm and}\qquad \frac{\max\{N^2, T^2\}}{N^{(2\alpha - \alpha_i + 1)} T} \rightarrow 0. $$ • For the inference of $f_t^0$ and $m_{it}^0$, we assume that $$ \frac{\max\{N^2, T^2\}}{N^\alpha T^2} \rightarrow 0. $$

The following theorem provides the asymptotic normality of the PC estimator.

theorem[CLT for PC estimator] Suppose that Assumptions A, B and C are satisfied. \begin{itemize} • If Assumption D(i) holds, then $$ \sqrt{T} \left( \widehat{\lambda}_i - \mathbf{H}^{-1} \lambda_i^0 \right) \to_d \mathcal{N}\left(0, \mathcal{Q}^{-\top} \boldsymbol{\Phi}_{\boldsymbol{F},i} \mathcal{Q}^{-1} \right), $$ where $\mathcal{Q} = \mathcal{D} \mathcal{G}^{\top} \boldsymbol{\Sigma}_{\boldsymbol{\Lambda}}^{-1/2}$, $\mathcal{D}$ is the diagonal matrix with the square roots of the eigenvalues of $\boldsymbol{\Sigma}_{\boldsymbol{\Lambda}}^{1/2}\boldsymbol{\Sigma}_{\boldsymbol{F}}\boldsymbol{\Sigma}_{\boldsymbol{\Lambda}}^{1/2}$ and $\mathcal{G}$ is an eigenvector of $\boldsymbol{\Sigma}_{\boldsymbol{\Lambda}}^{1/2}\boldsymbol{\Sigma}_{\boldsymbol{F}}\boldsymbol{\Sigma}_{\boldsymbol{\Lambda}}^{1/2}$. • If Assumption D(ii) holds, then $$ \sqrt{N^\alpha} \left( \widehat{f}_t - \mathbf{H}^{\top} f_t^0 \right)\to_d \mathcal{N}\left(0, \mathcal{D}^{-2} \mathcal{Q} \boldsymbol{\Phi}_{\boldsymbol{\Lambda},t} \mathcal{Q}^{\top} \mathcal{D}^{-2} \right), $$ • If Assumption D(ii) holds and there are constants $c_1,c_2>0$ such that $$ \left\Vertf^0_t\right\Vert \geq c_1,\qquad {\rm and} \qquad \left\Vert\lambda^0_i\right\Vert \geq c_2 N^{(\alpha_i - 1)/2}, $$ with probability tending to one, then $$ \mathcal{V}_{it}^{-1/2} \left( \widehat{m}_{it} - m^0_{it} \right) \to_d\mathcal{N}\left(0,1 \right), $$ where $$\mathcal{V}_{it} = \frac{1}{N^{\alpha}} \lambda_i^{0\top} \boldsymbol{\Sigma}_{\boldsymbol{\Lambda}}^{-1} \boldsymbol{\Phi}_{\boldsymbol{\Lambda},t} \boldsymbol{\Sigma}_{\boldsymbol{\Lambda}}^{-1} \lambda_i^0 + \frac{1}{T} f_t^{0 \top} \boldsymbol{\Sigma}_{\boldsymbol{F}}^{-1} \boldsymbol{\Phi}_{\boldsymbol{F},i} \boldsymbol{\Sigma}_{\boldsymbol{F}}^{-1} f_t^0.$$ \end{itemize}

Note that to derive asymptotic normality of $m_{it}^0$, setting lower bounds for $\left\Vert\lambda_i^0\right\Vert$ and $\left\Vertf_t^0\right\Vert$ is necessary to avoid degenerate variances. In the weak factors setting, it is natural to allow the lower bound of $\left\Vert\lambda_i^0\right\Vert$ to decrease as $N \rightarrow \infty$. This is to be contrast with bai2023approximate who implicitly assume that $\left\Vert\lambda_i^0\right\Vert$ is bounded away from zero while explicitly positing $\boldsymbol{\Lambda}^{0\top} \boldsymbol{\Lambda}^0 \asymp N^\alpha$. In other words, the asymptotic normality they established can only be applied to a vanishing proportion of $\lambda_i^0$s. Indeed, the normalizing constant ($\mathcal{V}_{it}^{-1/2}$) derived under their assumptions, $O_p\left( \min\{ N^{\alpha/2} , T^{1/2} \} \right)$, is too small in general. On the other hand, we show here that the correct normalizing constant should be of the order $O_p\left( \min\{ N^{(\alpha+1-\alpha_i)/2} , T^{1/2} \}\right)$. For instance, when $\alpha_i = \alpha$, the order of $\mathcal{V}_{it}^{-1/2}$ becomes $O_p\left( \min\{ N^{1/2} , T^{1/2} \}\right)$.

When $N \asymp T$, for the asymptotic normality of the estimators for $f_t^0$ and $m_{it}^0$ to hold, the condition $\alpha > 0$ is sufficient. However, that of $\lambda_i^0$ requires an additional condition $2 \alpha > \alpha_i$. As noted above, if the orders of $\lambda_i^0$s are the same across units, it is satisfied since $\alpha_i = \alpha$ for all $i$.

On the other hand, if the orders of $\lambda_i^0$s are heterogeneous, $\alpha_i$s would spread around `$\alpha$'. In this case, the inferential theory for $\lambda_{i}^0$ is still valid as long as $\alpha_{i}$ is not too much larger than `$\alpha$'. For example, when $\alpha = 1/3$, we can derive the inferential theory for $\lambda_{i}^0$ whose order is $\alpha_{i} < 2/3$. When $\alpha = 1/4$, we can derive the inferential theory for $\lambda_{i}^0$ whose order is $\alpha_{i} < 1/2$. Hence, for the typical $\lambda_{i}^0$ whose order $\alpha_i$ is not too different from `$\alpha$', we can still derive the inferential theory. In addition, we can get the inference of $\lambda_i^0$s whose $\alpha_i$ are the same as or smaller than `$\alpha$' as long as $\alpha > 0$. Since $\alpha_i$s are spread around `$\alpha$', we can expect that a large portion of $\alpha_i$s would be the same as or smaller than `$\alpha$'.

Dependent Noise

Now we shall treat the case where the idiosyncratic noises are dependent temporally and/or cross-sectionally. Compared to the independent noise case, the main technical difficulty lies in the leave-one-out analysis: for example, the leave-one-out estimator which excludes noises of the time period $t$ in construction, is no longer independent of noises of the time period $t$ if the noises are temporally dependent. To this end, we shall consider a more general approach that leaves all neighbors of $t$ out.

\paragraph{Leave-neighbor-out Analysis.} To address this issue, we consider the leave-neighbor-out estimator, which is constructed from hypothetical outcomes that exclude noises of the ‘neighbor’ of time period $t$ from true outcomes. Let $\mathcal{N}_\delta(t) = (t-\delta, \dots, t, \dots, t + \delta )$ be the $\delta$-neighbor of the time period $t$. Denote by $\boldsymbol{X}^{(-\mathcal{N}_\delta(t))}$ a $N\times T$ matrix whose $s$-th columns with $s \in \mathcal{N}_\delta(t)$ are $(m^{0}_{is})_{1\leq i \leq N}$ and other columns are $(x_{is'})_{1 \leq i \leq N}$ for all $s' \notin \mathcal{N}_\delta(t)$. That is, $\boldsymbol{X}^{(-\mathcal{N}_\delta(t))}$ replaces the columns corresponding to $\mathcal{N}_\delta(t)$ with $(m^{0}_{is})_{1\leq i \leq N}$ to remove the noises of the neighbor of time period $t$. Then, we apply the aforementioned operations to $\boldsymbol{X}^{(-\mathcal{N}_\delta(t))}$ leading to corresponding balanced singular vectors $\bm{Y}_r^{(-\mathcal{N}_\delta(t))}$, $\mathbf{Z}_r^{(-\mathcal{N}_\delta(t))}$, rotation matrix $\mathbf{O}^{(-\mathcal{N}_\delta(t))}$, and matching matrix $\mathbf{H}^{(-\mathcal{N}_\delta(t))}$.

Similar to the above, the key task is to bound the following term:

align[align omitted — 594 chars of source]

where $\mathbf{e}_{s} = (\epsilon_{1s},\dots,\epsilon_{Ns})^\top$ and $$ \Delta_{1}^{(-\mathcal{N}_\delta(t))} = \widetilde{\bm{Y}}^{(-\mathcal{N}_\delta(t))}_r(\widetilde{\bm{Y}}^{(-\mathcal{N}_\delta(t))\top}_r\widetilde{\bm{Y}}^{(-\mathcal{N}_\delta(t))}_r)^{-1} - \bm{Y}_r^0(\bm{Y}_r^{0\top} \bm{Y}_r^0)^{-1}. $$ The first term can be bounded as in (ref) by conditioning on $\{(\mathbf{e}_{s})_{s \notin \mathcal{N}_\delta(t)}, \boldsymbol{M}^0\}$. On the other hand, the second term is an additional term in the leave-neighbor-out analysis. In the weak temporal dependence case where we can get a tight bound of $\sum_{i=1}^N \left(\mathbb{E}[\epsilon_{it} | (\mathbf{e}_{s})_{s \notin \mathcal{N}_\delta(t)} ] - \mathbb{E}[\epsilon_{it}]\right)^2$ when $\delta$ grows slowly, the second term can also be bounded tightly. In this manner, we utilize the leave-neighbor-out estimator to allow for temporally dependent noises. Moreover, symmetrically, we can also allow for the cross-sectionally dependent noises.

Temporal Dependence

We first consider the case when $\epsilon_{it}$s are temporally dependent but cross-sectionally independent. To this end, we shall assume the following.

\paragraph{Assumption B'.} [Noise]

itemize$\mathbb{E}[\epsilon_{it}] = 0$, $\mathbb{E}[\epsilon_{it}^6] \leq C$ for a constant $C>0$. $\boldsymbol{E}$ is independent of $\boldsymbol{M}^0$ and $(\mathbf{e}_{i})_{i \in [N]}$ is independent across $i$; • With probability converging to 1, $\|\boldsymbol{E}\| \lesssim \max\{\sqrt{N},\sqrt{T} \}$; • There is a constant $C>0$ such that for all $t \in [T]$, $\sum_{s=1}^T |\text{Cov}(\epsilon_{it},\epsilon_{is})| \leq C$, and for all $i \in [N]$, $t,k \in [T]$, $\sum_{s=1}^T |\text{Cov}(\epsilon_{it}\epsilon_{ik},\epsilon_{is}\epsilon_{ik})| \leq C$; • There is a sequence $\delta \to \infty$ such that $\delta \asymp (\log N)^\nu$ for some constant $\nu > 0$ and \begin{gather*} \sum_{j=1}^N \left( \mathbb{E}\left[\epsilon_{jt}\left| (\epsilon_{js})_{ s \in \mathcal{N}_\delta(t)^c}\right. \right] - \mathbb{E}[\epsilon_{jt}] \right)^2= O_p(N^{1/3}), \end{gather*} where $\mathcal{N}_\delta(t)^c = \{1,\dots, t-\delta-1\} \cup \{t+\delta +1 , \dots, T \}$.

For the convergence rate and inferential theory of $\widehat{\lambda}_i$, Assumptions B'(i) -- B'(iii) are sufficient. Assumptions B'(i) and B'(ii) are similar to the conditions for the noise in the independence case, and Assumption B'(iii) is a typical weak temporal dependence assumption. On the other hand, Assumption B'(iv) is an additional weak dependence condition for $\widehat{f}_t$ and $\widehat{m}_{it}$ to utilize the leave-neighbor-out analysis. It requires the dependence between $\epsilon_{it}$ and other $\epsilon_{is}$s on the outside of the neighbor of $t$ to be sufficiently weak. In particular, the following lemma shows that it holds for the usual autoregressive or moving average model.

lemma[Examples of Assumption B'(iv)] \begin{itemize} • For each $i \in [N]$, let $\epsilon_{it}$ be a $MA(q)$ process with $q \lesssim (\log N)^\nu$ such that $$ \epsilon_{it} = \sum_{k=0}^q {\phi}^{(i)}_{k} u_{i,t-k} , $$ where $(u_{is})_{s\in [T]}$ are serially independent white noises. Then, if $\delta = C \lceil (\ln N)^\nu \rceil$ for some large constant $C>0$, we have for all $i \in [N]$, $$ \mathbb{E}\left[\epsilon_{it}\left| (\epsilon_{is})_{ s \in \mathcal{N}_\delta(t)^c}\right. \right] - \mathbb{E}[\epsilon_{it}] = 0. $$ • For each $i$, $\epsilon_{it}$ is a stationary AR(p) process such that $$ \epsilon_{it} = \phi^{(i)}_1 \epsilon_{i,t-1} + \cdots + \phi^{(i)}_p \epsilon_{i,t-p} + u_{it}, \ \ \text{ where } u_{it} \sim i.i.d. \ \ \mathcal{N} (0,\sigma_{u,i}^2), $$ and there is a constant $0< \vartheta <1$ such that $\max_{1\leq i \leq N,1\leq k\leq p}\left\vert\psi^{(i)}_k\right\vert< \vartheta $, where $(\psi^{(i)}_1,\dots,\psi^{(i)}_p)$ are the roots of the characteristic polynomial $$ \psi^{p} - \phi_1^{(i)} \psi^{p-1} - \cdots - \phi^{(i)}_{p-1} \psi - \phi^{(i)}_{p} = 0. $$ If $\delta = C \lceil \ln N \rceil$ for some constant $C >0$, we have $$ \max_i \mathbb{E}\left[ \left( \mathbb{E}\left[\epsilon_{it}\left| (\epsilon_{is})_{ s \in \mathcal{N}_\delta(t)^c}\right. \right] - \mathbb{E}[\epsilon_{it}]\right)^2 \right] \lesssim N^{-1}. $$ \end{itemize}

We are now in position to state the statistical properties of the PC estimates. The following theorem provides the convergence rate of the PC estimator.

theorem[Convergence rate of PC estimator] Suppose that $\max\{N, T \}(\log N)^{2\nu}=o(N^\alpha T)$. (i) If Assumptions A and B'(i) -- B'(iii) are satisfied, then $$ \left\Vert\widehat{\lambda}_i - \mathbf{H}^{-1} \lambda_i^0 \right\Vert = O_p \left( \frac{1}{\sqrt{T}} + \sqrt{\frac{N^{\alpha_i}}{N}} \frac{\max\{N, T \}}{N^\alpha T} \right). $$ (ii) If Assumptions A and B' are satisfied, then \begin{eqnarray*} \left\Vert \widehat{f}_t - \mathbf{H}^\top f_t^0 \right\Vert &=& O_p \left( \frac{1}{\sqrt{N^\alpha}} + \frac{\max\{N, T \}N^{1/6}}{N^\alpha T} + \frac{\sqrt{N} \max\{N^{3/2}, T^{3/2} \}(\log N)^{\nu/2}}{N^{2\alpha} T^{3/2}} \right),\\ \left\Vert\widehat{m}_{it} - m^0_{it}\right\Vert &=& O_p\left( \frac{1}{\sqrt{T}} + \sqrt{\frac{N^{\alpha_i}}{N}} \frac{1}{\sqrt{N^\alpha}} + \sqrt{\frac{N^{\alpha_i}}{N}} \frac{\max\{N, T \}N^{1/6}}{N^\alpha T} \right.\\ & & \qquad + \frac{\sqrt{N^{\alpha_i}} \max\{N^{3/2}, T^{3/2} \}(\log N)^{\nu/2}}{N^{2\alpha} T^{3/2}} + \frac{\max\{N, T \}N^{1/6}}{N^\alpha T^{3/2}} \\ & & \left. \qquad + \frac{\sqrt{N} \max\{N^{3/2}, T^{3/2} \}(\log N)^{\nu/2}}{N^{2\alpha} T^{2}} \right) . \end{eqnarray*}

The convergence rate of $\widehat{\lambda}_i$ is the same as that in the independence case. On the other hand, the convergence rates of $\widehat{f}_t$ and $\widehat{m}_{it}$ are quite different from those in the independence case. To fix the idea, if we assume $N \asymp T$, the above results can be reduced to

eqnarray*[eqnarray* omitted — 646 chars of source]

Because $\alpha_i \leq 1$ for all $i$, the condition for the consistency of $\widehat{\lambda}_i$ is $\alpha > 0$ as in the independence case. On the other hand, that of $\widehat{f}_t $ becomes $\alpha > 1/4$ if we ignore the logarithmic terms. Moreover, that of $\widehat{m}_{it}$ becomes $\alpha > \max\{0,\alpha_i/4 \}$ in this case. Therefore, the conditions for the consistency of $\widehat{f}_t $ and $\widehat{m}_{it}$ become stronger compared to those in the independence case.

Next, we present the inferential theory for the PC estimator. For this purpose, we need the following assumption.

\paragraph{Assumption D'.} [Parameter size]

itemize• For the inference of $\lambda_i^0$, we assume that $$ \frac{\max\{N,T\}}{N^{\alpha}T} \rightarrow 0, \qquad {\rm and}\qquad \frac{\max\{N^2, T^2\}}{N^{(2\alpha - \alpha_i + 1)} T} \rightarrow 0. $$ If $N \asymp T$, it reduces to $\alpha > 0$ and $2\alpha - \alpha_i >0$. • For the inference of $f_t^0$, we assume that \begin{align*} &\frac{\max\{N^3,T^3\}(\log N)^{\nu}}{N^{(3\alpha-1)}T^{3}} \rightarrow 0, \quad \frac{\max\{ N, T \} (\log N)^{\nu}}{N^{(2\alpha - 2)} T^{3}}\rightarrow 0, \quad \frac{\max\{N^2, T^2\} (\log N)^{\nu}}{N^\alpha T^{2}} \rightarrow 0 . \end{align*} If $N \asymp T$, it reduces to $$ \frac{(\log N)^{\nu}}{N^{(3\alpha - 1)}} \rightarrow 0, \qquad {\rm and}\qquad \frac{(\log N)^{2\nu}}{ N^\alpha } \rightarrow 0. $$ • For the inference of $m_{it}^0$, we assume that \begin{align*} &\frac{ \max\{N^3,T^3 \} (\log N)^{\nu}}{N^{(3\alpha-1)} T^3 } \rightarrow 0, \qquad {\rm and}\qquad \frac{\max\{N^2 , T^2 \}(\log N)^{2\nu}}{ N^\alpha T^2 } \rightarrow 0. \end{align*} If $N \asymp T$, it reduces to \begin{align*} &\frac{ (\log N)^{\nu}}{N^{(3\alpha - 1)} } \rightarrow 0, \qquad {\rm and}\qquad \frac{(\log N)^{2\nu}}{ N^\alpha } \rightarrow 0. \end{align*}

The condition for the inference of $\lambda_i^0$ is the same as that of the independence case. We can derive the inferential theory of $\lambda_i^0$ in the case of $\alpha > 0$, as long as $\alpha_i$ is not too larger than `$\alpha$'. On the other hand, the condition for the inference of $f_t^0$ and $m_{it}^0$ is different from that of the independence case. If we consider the case of $N \asymp T$ and ignore logarithmic terms, we require $\alpha > 1/3$ for the inference of $f_t^0$ and $m_{it}^0$.

Here, the requirement $\alpha > 1/3$ shows that using the leave-neighbor-out analysis is not a free lunch. In deriving a tight bound of $$ \Delta_{2}^{(-\mathcal{N}_\delta(t))} = \widetilde{\bm{Y}}_r(\widetilde{\bm{Y}}^\top_r\widetilde{\bm{Y}}_r)^{-1} - \widetilde{\bm{Y}}^{(-\mathcal{N}_\delta(t))}_r(\widetilde{\bm{Y}}^{(-\mathcal{N}_\delta(t))\top}_r\widetilde{\bm{Y}}^{(-\mathcal{N}_\delta(t))}_r)^{-1}, $$ one key challenge is to obtain a tight bound of the following term: $$ \sum_{i=1}^N \epsilon_{is} Y^{(-\mathcal{N}_\delta(t))}_{i} \quad \text{for all $s \in \mathcal{N}_\delta(t) = (t - \delta, \dots, t , \dots, t + \delta )$}, $$ where $Y^{(-\mathcal{N}_\delta(t))}_{i}$ is the transpose of $i$-th row of $\bm{Y}^{(-\mathcal{N}_\delta(t))}_r$. Note that if $\mathcal{N}_\delta(t)= \{t\}$ and there is no temporal dependence in the noises, we can easily get a tight bound by the same token as in (ref). However, when noises are temporally dependent, this way does not work. Moreover, we cannot expliot the method in (ref) because, e.g., when $s = t \pm \delta$, $s$ is too close to $\mathcal{N}_\delta(t)^c$ to use the method in (ref). Specifically, we cannot get a tight bound of $\sum_{i=1}^N \left(\mathbb{E}[\epsilon_{is} | (\mathbf{e}_{s'} )_{s' \notin \mathcal{N}_\delta(t)} ] - \mathbb{E}[\epsilon_{is}]\right)^2$ if $s = t \pm \delta$.

Instead, we resort to the expansion of $Y^{(-\mathcal{N}(t))}_{i}$ similar to (ref) to obtain a tight bound of $\sum_{i=1}^N \epsilon_{is} Y^{(-\mathcal{N}(t))}_{i}$. However, as using the expansion of $\widehat{\boldsymbol{F}}$ to derive the asymptotic normality of $\widehat{\boldsymbol{\Lambda}}$ requires stronger factors ($\alpha > 1/2$) in the conventional approach, using expansion of $Y^{(-\mathcal{N}(t))}_{i}$ makes us to assume the stronger factors, $\alpha > 1/3$. Nonetheless, roughly speaking, we use the expansion in a more indirect way compared to the traditional method. This may be the reason why our condition, $\alpha > 1/3$, is still weaker than that of the conventional approach, $\alpha > 1/2$.

theorem[CLT for PC estimator] Suppose that Assumptions A and C are satisfied. \begin{itemize} • If Assumptions B'(i) -- B'(iii), and D'(i) hold, then $$ \sqrt{T} \left( \widehat{\lambda}_i - \mathbf{H}^{-1} \lambda_i^0 \right)\to_d \mathcal{N}\left(0, \mathcal{Q}^{-\top} \boldsymbol{\Phi}_{\boldsymbol{F},i} \mathcal{Q}^{-1} \right). $$ • If Assumptions B' and D'(ii) hold, then $$ \sqrt{N^\alpha} \left( \widehat{f}_t - \mathbf{H}^{\top} f_t^0 \right)\to_d \mathcal{N}\left(0, \mathcal{D}^{-2} \mathcal{Q} \boldsymbol{\Phi}_{\boldsymbol{\Lambda},t} \mathcal{Q}^{\top} \mathcal{D}^{-2} \right). $$ • If Assumptions B' and D'(iii) hold and there are constants $c_1,c_2>0$ such that $$ \left\Vertf^0_t\right\Vert \geq c_1,\qquad {\rm and} \qquad \left\Vert\lambda^0_i\right\Vert \geq c_2 N^{(\alpha_i - 1)/2}, $$ with probability tending to one, then $$ \mathcal{V}_{it}^{-1/2} \left( \widehat{m}_{it} - m^0_{it} \right) \to_d\mathcal{N}\left(0,1 \right). $$ \end{itemize}

Cross-Sectional Dependence Case

Next, we study the case where the idiosyncratic noises are cross-sectionally dependent. To this end, we first define the $\delta$-neighbor of the unit $i$. For each $i$, let $\mathcal{N}_\delta(i)$ be a subset of $\{1, \dots, N \}$ that contains the $\delta$ number of units in order of correlation with the noise of $i$. In other words, $\mathcal{N}_\delta(i)$ consists of $\delta$ units whose noise has a higher correlation with the noise of $i$. It would be a natural counterpart of $\mathcal{N}_\delta(t)$ in the temporal dependence case.

\paragraph{Assumption A'.} [Factors and Loadings]

itemize• Assumptions A(i) -- A(iii) are satisfied; • For all $t \in [T]$, $\mathbb{E}\left\Vertf_t^0\right\Vert^{2} \leq C$ for some constant $C>0$. Moreover, for each $i \in [N]$, there are parameters $\alpha_{1,i}, \alpha_{2,i} \leq 1$ so that for some constants $C_1,C_2> 0$, $$ \left\Vert\lambda_i^0\right\Vert \leq C_1 N^{(\alpha_{1,i} - 1)/2} ,\qquad {\rm and} \qquad \frac{1}{\left\vert\mathcal{N}_\delta(i)\right\vert} \sum_{j \in \mathcal{N}_\delta (i)} \left\Vert\lambda_j^0\right\Vert^2 \leq C_2 N^{(\alpha_{2,i} - 1)}. $$

Here, $\alpha_{2,i}$ designates the order of the average over the neighborhood of $\lambda_i^0$ while $\alpha_{1,i}$ designates the order of $\lambda_i^0$ itself. Note that, if the orders of $\lambda_i^0$s are the same across units, we have $\alpha_{1,i} = \alpha_{2,i} = \alpha$ for all $i$, because $\left\Vert\lambda_i^0\right\Vert^2 \asymp N^{\alpha - 1}$ for all $i$ by Assumption A(ii). Otherwise, $\alpha_{1,i}$s and $\alpha_{2,i}$s would spread around the average order parameter `$\alpha$'.

\paragraph{Assumption B”.} [Noise]

itemize$\mathbb{E}[\epsilon_{it}] = 0$, $\mathbb{E}[\epsilon_{it}^6] \leq C$ for a constant $C>0$. $\boldsymbol{E}$ is independent of $\boldsymbol{M}^0$ and $(\mathbf{e}_{t})_{t \in [T]}$ is independent across $t$; • With probability converging to 1, $\|\boldsymbol{E}\| \lesssim \max\{\sqrt{N},\sqrt{T} \}$; • There is a constant $C>0$ such that for all $i \in [N]$, $\sum_{j=1}^N |\text{Cov}(\epsilon_{it},\epsilon_{jt})| \leq C$, and for all $t \in [T]$, $i,l \in [N]$, $\sum_{j=1}^N |\text{Cov}(\epsilon_{it}\epsilon_{lt},\epsilon_{jt}\epsilon_{lt})| \leq C$; • For each $i \in [N]$, there is a sequence $\delta \to \infty$ such that $\delta \asymp (\log N)^\omega $ for some constant $\omega > 0$ and \begin{gather*} \sum_{s=1}^T \left( \mathbb{E}\left[\epsilon_{is}\left| (\epsilon_{js})_{ j \in \mathcal{N}_\delta(i)^c}\right. \right] - \mathbb{E}[\epsilon_{is}] \right)^2= O_p(T^{1/3}). \end{gather*}

Basically, this assumption is symmetric to Assumption B'. And similarly, for the convergence rate and inferential theory of $\widehat{f}_t$, Assumptions B”(i) -- B”(iii) are sufficient. On the other hand, Assumption B”(iv) is an additional condition for $\widehat{\lambda}_i$ and $\widehat{m}_{it}$. For example, if the noise of $i$ depends on the noises of at most $O((\log N)^{\omega^*})$ number of other units, it is satisfied with $\omega = \omega^*$ like the moving average process in the temporal dependence case.

The following theorem provides the convergence rate of the PC estimator.

theorem[Convergence rate of PC estimator] Suppose that $\max\{N, T \}(\log N)^{2\omega}=o(N^\alpha T)$. (i) If Assumptions A' and B”(i) -- B”(iii) are satisfied, then \begin{gather*} \left\Vert \widehat{f}_t - \mathbf{H}^\top f_t^0 \right\Vert = O_p \left( \frac{1}{\sqrt{N^\alpha}} + \frac{\max\{N, T \}}{N^\alpha T} \right). \end{gather*} (ii) If Assumptions A' and B” are satisfied, then \begin{eqnarray*} \left\Vert\widehat{\lambda}_i - \mathbf{H}^{-1} \lambda_i^0 \right\Vert &=& O_p \left( \frac{1}{\sqrt{T}} + \left[ \sqrt{\frac{N^{\alpha_{1,i}}}{N}} + \sqrt{\frac{N^{\alpha_{2,i}}}{N}} \right] \frac{\max\{N, T \}(\log N)^{\omega}}{N^\alpha T } \right.\\ & & \qquad \left. + \frac{\max\{\sqrt{N}, \sqrt{T} \}}{N^{\alpha/2} T^{5/6} } + \frac{ \max\{N^{3/2}, T^{3/2} \}(\log N)^{\omega/2}}{N^{3\alpha/2} T^{3/2}}\right),\\ \left\Vert\widehat{m}_{it} - m^0_{it}\right\Vert &=& O_p\left( \frac{1}{\sqrt{T}} + \sqrt{\frac{N^{\alpha_{1,i}}}{N}} \frac{1}{\sqrt{N^\alpha}} + \frac{ \max\{N^{3/2}, T^{3/2} \}(\log N)^{\omega/2}}{N^{3\alpha/2} T^{3/2}} + \frac{\max\{\sqrt{N}, \sqrt{T} \}}{N^{\alpha/2} T^{5/6} } \right.\\ & & \qquad \left. + \left[ \sqrt{\frac{N^{\alpha_{1,i}}}{N}} + \sqrt{\frac{N^{\alpha_{2,i}}}{N}} \right] \frac{\max\{N, T \}(\log N)^{\omega}}{N^\alpha T} \right) . \end{eqnarray*}

To fix the idea, if we assume that $N \asymp T$, the above results can be reduced to

eqnarray*[eqnarray* omitted — 706 chars of source]

The condition for the consistency of $\widehat{f}_t$ is $\alpha > 0$ as in the independence case. In addition, since $\alpha_{1,i}, \alpha_{2,i} \leq 1$ for all $i$, the condition for the consistency of $\widehat{\lambda}_i$ is $\alpha > 0$ if we ignore logarithmic terms. Similarly, the condition for the consistency of $\widehat{m}_{it}$ is $\alpha > 0$. Therefore, if we ignore logarithmic terms, the conditions for all estimators are reduced to $\alpha > 0$ as in the independence case.

Next, we present the inferential theory for the PC estimator. For the asymptotic normality of the estimator, we require the following additional assumption.

\paragraph{Assumption D”.} [Parameter size]

itemize• For the inference of $\lambda_i^0$, we assume that \begin{align*} \frac{\max\{N^3,T^3\}(\log N)^{\omega}}{N^{3\alpha}T^{2}} \rightarrow 0,\qquad \frac{\max\{N^2, T^2\} (\log N)^{2\omega}}{N^{(2\alpha - \max\{\alpha_{1,i},\alpha_{2,i}\} + 1)} T} \rightarrow 0,\qquad \frac{(\log N)^{2\omega}}{N^{\alpha}} \rightarrow 0. \end{align*} If $N \asymp T$, it reduces to $$ \frac{(\log N)^{\omega}}{N^{(3\alpha - 1)}} \rightarrow 0,\qquad \frac{(\log N)^{2\omega}}{N^{(2 \alpha - \max\{\alpha_{1,i},\alpha_{2,i}\})}} \rightarrow 0,\qquad \frac{(\log N)^{2\omega}}{N^{ \alpha}} \rightarrow 0. $$ • For the inference of $f_t^0$, we assume that $$ \frac{\max\{N^2, T^2\}}{N^\alpha T^{2}} \rightarrow 0. $$ If $N \asymp T$, it reduces to $\alpha > 0$. • For the inference of $m_{it}^0$, we assume that \begin{align*} \frac{\max\{N^3,T^3\}(\log N)^{\omega}}{N^{3\alpha}T^{2}} \rightarrow 0,\qquad \frac{\max\{N^2 , T^2 \}}{N^\alpha T^2 } \rightarrow 0,\qquad \frac{\max\{N, T\} (\log N)^{2\omega}}{N^{(2\alpha - \alpha_{2,i} + 1)} } \rightarrow 0 . \end{align*} If $N \asymp T$, it reduces to \begin{align*} & \frac{(\log N)^{\omega}}{N^{(3\alpha - 1)}} \rightarrow 0, \qquad {\rm and} \qquad \frac{(\log N)^{2\omega}}{N^{(2\alpha - \alpha_{2,i} )} } \rightarrow 0 . \end{align*}

The condition for the inference of $f_t^0$ is the same as that of the independence case and $\alpha > 0$ is enough. On the other hand, for the inference of $\lambda_i^0$, we require $\alpha > \frac{1}{3}$ and $2 \alpha - \max\{\alpha_{1,i},\alpha_{2,i}\} > 0$ if we consider the case of $N \asymp T$ and ignore logarithmic terms. Moreover, for the inference of $m_{it}^0$, we need $\alpha > \frac{1}{3}$ and $2 \alpha - \alpha_{2,i} > 0$. In the homogeneous order case where $\alpha_{1,i} = \alpha_{2,i} = \alpha$, the conditions for inference of $\lambda_i^0$ and $m_{it}^0$ can be reduced to $\alpha > \frac{1}{3}$. Generally speaking, if $\alpha_{1,i}$ and $\alpha_{2,i}$ are not too far from $\alpha$, $\alpha > \frac{1}{3}$ is enough for the inference of $\lambda_i^0$ and $m_{it}^0$.

theorem[CLT for PC estimator] Suppose that Assumptions A' and C are satisfied. \begin{itemize} • If Assumptions B” and D”(i) hold, then $$ \sqrt{T} \left( \widehat{\lambda}_i - \mathbf{H}^{-1} \lambda_i^0 \right)\to_d \mathcal{N}\left(0, (\mathcal{Q}^{\top})^{-1} \boldsymbol{\Phi}_{\boldsymbol{F},i} \mathcal{Q}^{-1} \right). $$ • If Assumptions B”(i) -- B”(iii) and D”(ii) hold, then $$ \sqrt{N^\alpha} \left( \widehat{f}_t - \mathbf{H}^{\top} f_t^0 \right)\to_d \mathcal{N}\left(0, \mathcal{D}^{-2} \mathcal{Q} \boldsymbol{\Phi}_{\boldsymbol{\Lambda},t} \mathcal{Q}^{\top} \mathcal{D}^{-2} \right). $$ • If Assumptions B” and D”(iii) hold and there are constants $c_1,c_2>0$ such that $$ \left\Vertf^0_t\right\Vert \geq c_1,\qquad {\rm and} \qquad \left\Vert\lambda^0_i\right\Vert \geq c_2 N^{(\alpha_{1,i} - 1)/2}, $$ with probability tending to one, then $$ \mathcal{V}_{it}^{-1/2} \left( \widehat{m}_{it} - m^0_{it} \right) \to_d\mathcal{N}\left(0,1 \right). $$ \end{itemize}

General Dependence Case

Lastly, we study the case where the idiosyncratic noises are cross-sectionally and temporally dependent. To this end, let $\mathcal{N}_{\delta_1}(i)$ and $\mathcal{N}_{\delta_2}(t)$ be the $\delta_1$-neighbor of the unit $i$ and $\delta_2$-neighbor of the time period $t$, respectively.

\paragraph{Assumption B”'.} [Noise]

itemize$\mathbb{E}[\epsilon_{it}] = 0$, $\mathbb{E}[\epsilon_{it}^6] \leq C$ for a constant $C>0$. $\boldsymbol{E}$ is independent of $\boldsymbol{M}^0$; • With probability converging to 1, $\|\boldsymbol{E}\| \lesssim \max\{\sqrt{N},\sqrt{T} \}$; • There is a constant $C>0$ such that for all $i,l \in [N]$ and $t,k \in [T]$, \begin{align*} &\sum_{j=1}^N \sum_{s=1}^T |Cov(\epsilon_{it},\epsilon_{js})| \leq C ,\ \ \sum_{j=1}^N \sum_{s=1}^T |Cov(\epsilon_{it}\epsilon_{lt},\epsilon_{js}\epsilon_{ls})| \leq C,\ \ \sum_{j=1}^N \sum_{s=1}^T |Cov(\epsilon_{it}\epsilon_{ik},\epsilon_{js}\epsilon_{jk})| \leq C; \end{align*} • For each $i \in [N]$ and $t \in [T]$, there are sequences $\delta_1,\delta_2 \to \infty$ such that $\delta_1 \asymp (\log N)^\omega $ for some constant $\omega > 0$ and $\delta_2 \asymp (\log N)^\nu $ for some constant $\nu > 0$, \begin{gather*} \sum_{s=1}^T \left( \mathbb{E}\left[\epsilon_{is}\left| (\mathbf{e}_{j})_{ j \in \mathcal{N}_{\delta_1}(i)^c}\right. \right] - \mathbb{E}[\epsilon_{is}] \right)^2= O_p(T^{1/3}),\\ \sum_{j=1}^N \left( \mathbb{E}\left[\epsilon_{jt}\left| (\mathbf{e}_{s})_{ s \in \mathcal{N}_{\delta_2}(t)^c}\right. \right] - \mathbb{E}[\epsilon_{jt}] \right)^2= O_p(N^{1/3}); \end{gather*} • For each $i \in [N]$, we have $$ \max_{1\leq t \leq T} \sum_{s=1}^T \left\vert \text{Cov}\left(\epsilon_{it}, \epsilon_{is} \left| (\mathbf{e}_{j})_{ j \in \mathcal{N}_{\delta_1}(j)^c} \right. \right)\right\vert =O_p \left( \max_{1\leq t \leq T} \text{Var} \left(\epsilon_{it} \left| (\mathbf{e}_{j})_{ j \in \mathcal{N}_{\delta_1}(t)^c} \right. \right) \right); $$ In addition, for each $t \in [T]$, we have $$ \max_{1\leq i \leq N} \sum_{j=1}^N \left\vert \text{Cov}\left(\epsilon_{it}, \epsilon_{jt}\left| (\mathbf{e}_{s})_{ s \in \mathcal{N}_{\delta_2}(t)^c} \right. \right)\right\vert =O_p \left( \max_{1\leq i \leq N}\text{Var} \left(\epsilon_{it} \left| (\mathbf{e}_{s})_{ s \in \mathcal{N}_{\delta_2}(t)^c} \right. \right) \right). $$

Assumptions B”'(i) - (iv) are the generalizations of Assumptions B' and B”. If we assume cross-sectional or temporal independence, these assumptions reduce to Assumptions B' or B”, respectively. Here, Assumption B”'(iii) can be comparable to Assumption A3 in bai2023approximate. To show Assumption A3 in bai2023approximate, one may need Assumption B”'(iii). On the other hand, Assumption B”'(v) is a new condition. It requires the conditional weak dependence conditioning on noises of the outside of the neighbor. In the special case where noises are dependent on a finite number of other noises, it reduces to the unconditional weak dependence and is satisfied by Assumption B”'(iii).

Then, the following theorem provides the convergence rate of the PC estimator.

theorem[Convergence rate of PC estimator] Suppose that $\max\{N, T \}(\log N)^{2\max\{\omega,\nu\}}=o(N^\alpha T)$. If Assumptions A' and B”' are satisfied, then \begin{align*} &\left\Vert\widehat{\lambda}_i - \mathbf{H}^{-1} \lambda_i^0 \right\Vert = O_p \left( \frac{1}{\sqrt{T}} + \left[ \sqrt{\frac{N^{\alpha_{1,i}}}{N}} + \sqrt{\frac{N^{\alpha_{2,i}}}{N}} \right] \frac{\max\{N, T \}(\log N)^{\omega}}{N^\alpha T } \right.\\ & \qquad\qquad\qquad\qquad\qquad \left. + \frac{\max\{\sqrt{N}, \sqrt{T} \}}{N^{\alpha/2} T^{5/6} } + \frac{ \max\{N^{3/2}, T^{3/2} \}(\log N)^{\omega/2}}{N^{3\alpha/2} T^{3/2}}\right),\\ &\left\Vert \widehat{f}_t - \mathbf{H}^\top f_t^0 \right\Vert = O_p \left( \frac{1}{\sqrt{N^\alpha}} + \frac{\max\{N, T \}N^{1/6}}{N^\alpha T} + \frac{\sqrt{N} \max\{N^{3/2}, T^{3/2} \}(\log N)^{\nu/2}}{N^{2\alpha} T^{3/2}} \right),\\ & \left\Vert\widehat{m}_{it} - m^0_{it}\right\Vert\\ & = O_p\left( \frac{1}{\sqrt{T}} + \sqrt{\frac{N^{\alpha_{1,i}}}{N}} \frac{1}{\sqrt{N^\alpha}} + \frac{ \max\{N^{2}, T^{2} \}(\log N)^{(\omega+\nu)/2}}{N^{3\alpha/2} T^{2}} + \sqrt{\frac{N^{\alpha_{1,i}}}{N}}\frac{\max\{N, T \}N^{1/6}}{N^\alpha T} \right.\\ & \qquad\quad \left. + \sqrt{\frac{N^{\alpha_{2,i}}}{N}} \frac{\max\{N, T \}(\log N)^{\omega}}{N^\alpha T} + \frac{\sqrt{N^{\alpha_{1,i}}} \max\{N^{3/2}, T^{3/2} \}(\log N)^{\nu/2}}{N^{2\alpha} T^{3/2}} + \frac{\max\{N^{2/3}, T^{2/3} \}}{N^{\alpha/2} T } \right.\\ & \qquad\quad + \sqrt{\frac{N^{\alpha_{2,i}}}{N}} \frac{\max\{N^2, T^2 \}N^{1/6}(\log N)^{\omega}}{N^{2\alpha} T^2 } + \frac{\max\{N^{3/2}, T^{3/2} \}N^{1/6}}{N^{3\alpha/2} T^{11/6}} \\ & \qquad\quad + \frac{\max\{N^{5/2}, T^{5/2} \}N^{1/6}(\log N)^{\omega/2}}{N^{5\alpha/2} T^{5/2}} + \frac{\sqrt{N^{\alpha_{2,i}}}\max\{N^{5/2}, T^{5/2} \}(\log N)^{\nu/2+\omega}}{N^{3\alpha} T^{5/2}} \\ & \left. \qquad\quad + \frac{\sqrt{N}\max\{N^{3}, T^{3} \}(\log N)^{(\omega+\nu)/2}}{N^{7\alpha/2} T^{3}} + \frac{\sqrt{N}\max\{N^{2}, T^{2} \}(\log N)^{\nu/2}}{N^{5\alpha/2} T^{7/3}} \right) . \end{align*}

To fix the idea, if we assume that $N \asymp T$, the above results can be reduced to

align*[align* omitted — 1,165 chars of source]

Because $\alpha_{1,i}, \alpha_{2,i} \leq 1$ for all $i$, the condition for the consistency of $\widehat{\lambda}_i$ is $\alpha > 0$ if we ignore logarithmic terms. In addition, that of $\widehat{f}_t $ is $\alpha > 1/4$ when we ignore the logarithmic terms. On the other hand, the condition for the consistency of $\widehat{m}_{it}$ is somewhat complicated. For the consistency, we require $\alpha > \max\{1/7, \alpha_{1,i}/4 , \alpha_{2,i}/6 \}$ if we ignore the logarithmic terms.

\paragraph{Assumption D”'.} [Parameter size]

itemize• For the inference of $\lambda_i^0$, we assume that \begin{align*} \frac{\max\{N^3,T^3\}(\log N)^{\omega}}{N^{3\alpha}T^{2}} \rightarrow 0,\qquad \frac{\max\{N^2, T^2\} (\log N)^{2\omega}}{N^{(2\alpha - \max\{\alpha_{1,i},\alpha_{2,i}\} + 1)} T} \rightarrow 0,\qquad \frac{(\log N)^{2\omega}}{N^{\alpha}} \rightarrow 0. \end{align*} If $N \asymp T$, it reduces to $$ \frac{(\log N)^{\omega}}{N^{(3\alpha - 1)}} \rightarrow 0,\qquad \frac{(\log N)^{2\omega}}{N^{(2 \alpha - \max\{\alpha_{1,i},\alpha_{2,i}\})}} \rightarrow 0,\qquad \frac{(\log N)^{2\omega}}{N^{ \alpha}} \rightarrow 0. $$ • For the inference of $f_t^0$, we assume that \begin{align*} &\frac{\max\{N^3,T^3\}(\log N)^{\nu}}{N^{(3\alpha-1)}T^{3}} \rightarrow 0, \quad \frac{\max\{ N, T \} (\log N)^{\nu}}{N^{(2\alpha - 2)} T^{3}}\rightarrow 0, \quad \frac{\max\{N^2, T^2\} (\log N)^{\nu}}{N^\alpha T^{2}} \rightarrow 0 . \end{align*} If $N \asymp T$, it reduces to $$ \frac{(\log N)^{\nu}}{N^{(3\alpha - 1)}} \rightarrow 0, \qquad {\rm and}\qquad \frac{(\log N)^{2\nu}}{ N^\alpha } \rightarrow 0. $$ • For the inference of $m_{it}^0$, we assume that \begin{align*} &\frac{\max\{N^4,T^4\}(\log N)^{\omega + \nu}}{N^{3\alpha}T^{3}} \rightarrow 0,\qquad \frac{\max\{N^{3/2},T^{3/2}\}}{N^{2\alpha}T}\rightarrow 0,\qquad \frac{\max\{N, T\} (\log N)^{2\omega}}{N^{(2\alpha - \alpha_{2,i} + 1)} } \rightarrow 0. \end{align*} If $N \asymp T$, it reduces to \begin{align*} & \frac{(\log N)^{\omega+\nu}}{N^{(3\alpha - 1 )} } \rightarrow 0,\qquad \frac{(\log N)^{2\omega}}{N^{(2\alpha - \alpha_{2,i} )} } \rightarrow 0 . \end{align*}

We then have the following asymptotic normality.

theorem[CLT for PC estimator] Suppose that Assumptions A', B”' and C are satisfied. \begin{itemize} • If Assumption D”'(i) hold, then $$ \sqrt{T} \left( \widehat{\lambda}_i - \mathbf{H}^{-1} \lambda_i^0 \right)\to_d \mathcal{N}\left(0, (\mathcal{Q}^{\top})^{-1} \boldsymbol{\Phi}_{\boldsymbol{F},i} \mathcal{Q}^{-1} \right). $$ • If Assumption D”'(ii) hold, then $$ \sqrt{N^\alpha} \left( \widehat{f}_t - \mathbf{H}^{\top} f_t^0 \right)\to_d \mathcal{N}\left(0, \mathcal{D}^{-2} \mathcal{Q} \boldsymbol{\Phi}_{\boldsymbol{\Lambda},t} \mathcal{Q}^{\top} \mathcal{D}^{-2} \right). $$ • If Assumption D”'(iii) hold and there are constants $c_1,c_2>0$ such that $$ \left\Vertf^0_t\right\Vert \geq c_1,\qquad {\rm and} \qquad \left\Vert\lambda^0_i\right\Vert \geq c_2 N^{(\alpha_{1,i} - 1)/2}, $$ with probability tending to one, then $$ \mathcal{V}_{it}^{-1/2} \left( \widehat{m}_{it} - m^0_{it} \right) \to_d\mathcal{N}\left(0,1 \right). $$ \end{itemize}

The condition for the inference of $f_t^0$ is the same as that of the temporal dependence case and we require $\alpha > 1/3$. Similarly, the condition for $\lambda_i^0$ is the same as that of the cross-sectional dependence case. For the inference of $\lambda_i^0$, we require $\alpha > \frac{1}{3}$ and $2 \alpha - \max\{\alpha_{1,i},\alpha_{2,i}\} > 0$ if we consider the case of $N \asymp T$ and ignore logarithmic terms. Moreover, for the inference of $m_{it}^0$, we need $ \alpha > \max\{1/3, \alpha_{2,i}/2 \}$. Hence, if $\alpha_{1,i}$ and $\alpha_{2,i}$ are not too far from $\alpha$, $\alpha > \frac{1}{3}$ is enough for the inference of $\lambda_i^0$ and $m_{it}^0$.

Lastly, it is noteworthy that our requirement of $\alpha$ for inference ($\alpha>1/3$) is weaker than that in bai2023approximate. However, this improvement is achieved at the cost of more restrictions in the dependence structure in the noise.

Concluding Remarks

This paper investigates the asymptotic properties of the PC estimator for high dimensional approximate factor model with weak factors. Assuming that $\boldsymbol{\Lambda}^{0\top}\boldsymbol{\Lambda}^0 / N^\alpha$ has a positive definite limit, we establish the consistency and asymptotic normality of the PC estimator for $\alpha \in (0,1)$, under some conditions about the dependence structure in the noise. In particular, we show the asymptotic normality of the estimator when $\alpha \in (0,1/2]$, which has not yet been clarified in the literature. Our proof method combines the conventional approach based on the eigendecomposition of the covariance matrix with the more recently developed leave-one-out analysis. However, unlike the existing literature using the leave-one-out technique, we allow for the dependence in the noises by exploiting the leave-neighbor-out estimator and do not require the incoherence condition, which is a common assumption in the literature. The technical understanding of this generalization may have independent value and be useful for other related issues.