EconBase
← Back to paper

How to Detect Network Dependence in Latent Factor Models? A Bias-Corrected CD Test

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

567,513 characters · 22 sections · 73 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.
This text was truncated for display. The citation measures were computed over the complete text.

How to Detect Network Dependence in Latent Factor Models? A Bias-Corrected CD Test

abstractIn a recent paper juodis2022incidental (JR) show that the application of the CD test proposed by pesaran2004general to residuals from panels with latent factors results in over-rejection. They propose a randomized test statistic to correct for over-rejection, and add a screening component to achieve power. This paper considers the same problem but from a different perspective, and shows that the standard CD test remains valid if the latent factors are weak. A bias-corrected version, CD$^{\text{*}}$, is proposed which is shown to be asymptotically standard normal under the null of error cross-sectional independence which has power against network type alternatives. This result is shown to hold for pure latent factor models as well as for panel regression models with latent factors. The case where the errors are serially correlated is also considered. Small sample properties of the CD$^{\text{*}}$ test are investigated by Monte Carlo experiments and are shown to have satisfactory small sample properties. In an empirical application, using the CD$^{\text{*}}$ test, it is shown that there remains spatial error dependence in a panel data model for real house price changes across 377 Metropolitan Statistical Areas in the U.S., even after the effects of latent factors are filtered out.

JEL Classifications:\ { C18, C23, C55}

Key Words:{ \ Latent factor models, strong and weak factors, error cross-sectional independence, spatial and network alternatives, size and power.}

\pagenumbering{gobble}

\pagenumbering{arabic}

Introduction

It is now quite standard to use latent multi-factor models to characterize and explain cross-sectional dependence in panels when the cross section dimension $(n)$ and the time series dimension $(T)$ are both large. However, due to uncertainty regarding the nature of error cross-sectional dependence, it is arguable whether error cross-sectional dependence is fully accounted for by latent factors. Some of the factors could be semi-strong, and the errors might have spatial or network features that are not necessarily captured by common factors alone. chudik2011weak provide an early discussion of the different sources of cross-sectional dependence. Diagnostic tests of error cross-sectional independence in panels are required to safeguard against estimation bias and unreliable inference. See, for example, bai2003inferential,bai2009panel, phillips2003dynamic,phillips2007bias, bai2006confidence, pesaran2006estimation, and pesaran2011large. Such tests are also helpful to researchers interested in network or spatial dependence once the influence of common factors are filtered out. See, for example, bailey2016two, shi2017spatial, aquaro2021estimation, and bai2021dynamic amongst others.

One standard test for error cross-sectional independence is the CD test proposed by pesaran2004general,pesaran2021general, which has been further developed. For example, hsiao2012diagnostic apply the CD test to panel data models with limited dependent variables, while pesaran2015testing uses it to test weak error cross-sectional dependence in large panels. In a recent paper juodis2022incidental (JR) show that the application of the CD test to the residuals from panels with latent factors is invalid and can result in over-rejection of the null of error cross-sectional independence.\footnote{In the empirical finance literature gagliardini2019diagnostic propose a diagnostic criterion to check if the errors from a (strong) factor model are weakly correlated or contain missing strong factor(s). These authors do not propose a test of cross-sectional error dependence but focus on the detection of potentially omitted strong factors in asset pricing models.} They propose a randomized CD test statistic as a solution. Their proposed test is constructed in two steps. First, they multiply the residuals from panel regressions with independent randomized weights to obtain their CD$_{W}$ statistic, which will have a zero mean by construction. In this way they avoid the over-rejection problem of the CD test, but by the very nature of the randomization process they recognize that the CD$_{W}$ test will lack power. To overcome the problem of lack of power, JR modify the CD$_{W}$ test statistic by adding to it a screening component proposed by fan2015power which is expected to tend to zero with probability approaching one under the null hypothesis, but to diverge at a reasonably fast rate under the alternative. This further modification of the CD$_{W}$ test is denoted by CD$_{W+}$. Accordingly, it is presumed that the CD$_{W+}$ test can overcome both over-rejection and the low power problems. However, JR do not provide a formal proof establishing conditions under which the screening component tends to zero under the null and diverges sufficiently fast under alternatives, including spatial or network dependence type alternatives. Also, our Monte Carlo simulations show that the CD$_{W+}$ test tends to over-reject when the errors are non-Gaussian and $n>>T$, and lacks power under spatial and network alternatives, which is likely to be particularly important in empirical applications.

In this paper we show that the standard CD test is in fact valid for testing error cross-sectional independence in panel data models with weak latent factors. However, when the latent factors are semi-strong or strong the use of the CD test will result in over-rejection and will no longer be valid, extending JR's results to panels with semi-strong latent factors. In short, whilst the CD$_{W+}$ test is a useful and welcome addition to testing for error cross-sectional independence, it would be interesting to develop a modified version of the test that simultaneously deals with the over-rejection problem and does not compromise power for a general class of alternatives. To that end, firstly we study testing for error cross-sectional independence in a pure latent factor model, and derive an explicit expression for the bias of the CD test statistic in terms of factor loadings and error variances. We then propose a bias-corrected version of the CD test statistic, denoted by $CD^{\ast}$, which is shown to have $\mathcal{N}(0,1)$ asymptotic distribution under the null hypothesis irrespective of whether the latent factors are weak or strong. When the latent factors are weak the correction tends to zero, $CD$ and $CD^{\ast}$ will be asymptotically equivalent. However, $CD-CD^{\ast}$ diverges if at least one of the underlying latent factors is (semi) strong. We show that under the null of cross-sectional independence, $CD^{\ast}$ converges to a standard normal distribution when $n$ and $T$ tend to infinity so long as $n/T\rightarrow\kappa$, where $0<\kappa<\infty$, and a test based on $CD^{\ast}$ will have the correct size asymptotically. In addition, it is shown that the CD$^{\text{*}}$ test has power against spatial and network type alternatives. In particular, we are the first to give a formal derivation of the power function for CD tests against general spatial and network alternatives, which can be applied equally to panel data models without latent factors and therefore supplement earlier research on CD tests\footnote{For instance, pesaran2004general, which is the unpublished version of pesaran2021general, also discusses the power of the CD test against spatial dependence in Section 8.2 of his paper, for a specific connection matrix with $n$ fixed as $T\rightarrow\infty$.}.

We then consider the application of the $CD^{\ast}$ to test error cross-sectional independence in the case of panel regression models with latent factors, discussed in pesaran2006estimation. It is shown that the asymptotic properties of the $CD^{\ast}$ in the case of pure latent factor models also carry over to panel regression models with latent factors. We also investigate the application of the $CD^{\ast}$ to panel data models with serially correlated errors, and consider the method proposed by baltagi2016testing as well as using an autoregressive distributed lag (ARDL) representation which transforms the model with serially correlated errors to one without error serial correlation.

The finite sample performance of the CD$^{\text{*}}$ test is investigated by Monte Carlo simulations in the case of pure latent factor models, panel regression models with latent factors with and without error serial correlation. It is found that the CD$^{\text{*}}$ test avoids the over-rejection problem under the null and has power against spatial and network alternatives, and has desirable small sample properties regardless of whether the errors are Gaussian or not, under different combinations of $n$ and $T$. We also find that both adjustments for dealing with error serial correlation considered in the paper give desirable small sample properties. Finally, as compared to JR's CD$_{W^{+}}$ test, the proposed bias-corrected CD test is better in controlling the size of the test and has much better power properties against spatial or network alternatives.

The use of the CD$^{\text{*}}$ test is illustrated by an empirical application in modeling real house price changes in the U.S. Because it is evident that real house price changes are driven by macroeconomic trends which can be modeled by latent factors, it is necessary to filter out these factors before testing for spillover effect. By applying the CD$^{\text{*}}$ test to real house price changes in the U.S. we are able to show significant existence of weak cross-sectional dependence in addition to latent factors.

The rest of the paper is organized as follows. Section (ref) sets out the latent factor model and its assumptions. Section (ref) introduces the estimation of latent factors and the CD test. The bias-corrected test, CD$^{\ast}$, is introduced in Section (ref) and its asymptotic distributions are derived under the null and the alternative hypotheses. The extension to more general panel regression models with observed covariates as well as latent factors are discussed in Section (ref). Adjustments to the CD$^{\text{*}}$ test for panels with serially correlated errors are discussed in Section (ref). Using Monte Carlo techniques, the small sample properties of CD, CD$^{\ast}$, and CD$_{W+}$ tests are discussed in Section (ref). An empirical illustration is provided in Section (ref). Proofs of the propositions and theorems are provided in an appendix. The auxiliary lemmas and the associated proofs are given in a supplement.

Notations: For the $n\times n$ matrix $\mathbf{A}=\left( a_{ij}\right) $, we denote its largest eigenvalue by $\mu_{max}\left( \mathbf{A}\right) $, its trace by $\mathrm{tr}\left( \mathbf{A}\right) =\sum_{i=1}^{n}a_{ii}$, its spectral norm by $\left\Vert \mathbf{A}\right\Vert =\mu_{\max}^{1/2}\left( \mathbf{A}^{\prime}\mathbf{A}\right) $, its maximum absolute column sum norm by $\left\Vert \mathbf{A}\right\Vert _{1}=\max_{1\leq j\leq n}\left( \sum_{i=1}^{n}\left\vert a_{ij}\right\vert \right) $, and its maximum absolute row sum norm by $\left\Vert \mathbf{A}\right\Vert _{\infty }=\max_{1\leq i\leq n}\left( \sum_{j=1}^{n}\left\vert a_{ij}\right\vert \right) $. We write $\mathbf{A}>\mathbf{0}$ when $\mathbf{A}$ is positive definite. For matrices $\mathbf{B}=\left( b_{ij}\right) $ and $\mathbf{C} =\left( c_{ij}\right) $, $\mathbf{B\odot C=C\odot B}$ denote Hadamard product with elements $b_{ij}c_{ij}$. $\rightarrow_{p}$ denotes convergence in probability, $\rightarrow_{d}$ convergence in distribution, and $\overset{a}{\thicksim}$ asymptotic equivalence in distribution. $O_{p}\left( \cdot\right) $ and $o_{p}\left( \cdot\right) $ denote the stochastic order relations. In particular, $o_{p}(1)$ indicates terms that tend to zero in probability as $(n,T)\rightarrow\infty$, such that $n/T\rightarrow\kappa$, where $0<\kappa<\infty$. $C$ and $c$ will be used to denote finite large and non-zero small positive numbers, respectively, that are bounded in $n$ and $T$. They can take different values at different instances. If $\left\{ f_{n}\right\} _{n=1}^{\infty}$ is any real sequence and $\left\{ g_{n}\right\} _{n=1}^{\infty}$ is a sequence of positive real numbers, then $f_{n}=O(g_{n})$, if there exists $C$ such that $\left\vert f_{n}\right\vert /g_{n}\leq C$ for all $n$. $f_{n}=o(g_{n})$ if $f_{n}/g_{n}\rightarrow0$ as $n\rightarrow\infty$. If $\left\{ f_{n}\right\} _{n=1}^{\infty}$ and $\left\{ g_{n}\right\} _{n=1}^{\infty}$ are both positive sequences of real numbers, then $f_{n}=\ominus\left( g_{n}\right) $ if there exists $n_{0} \geq1$ and positive finite constants $C_{0}$ and $C_{1}$, such that $\inf_{n\geq n_{0}}\left( f_{n}/g_{n}\right) \geq C_{0},$ and $\sup_{n\geq n_{0}}\left( f_{n}/g_{n}\right) \leq C_{1}$.

The latent factor model

To simplify the exposition and to highlight the main issue of concern, namely the presence of latent (unobserved) factors in the panel regression model, initially we focus on the approximate factor model, due to chamberlain1983arbitrage, and assume that for each unit $i=1,2,\ldots ,n$,

equation[equation omitted — 117 chars of source]

where

equation[equation omitted — 167 chars of source]

$\sup_{i}\sigma_{i}^{2}<C<\infty$ and $\inf_{i}\sigma_{i}^{2}>c>0$, $\left\{ w_{ij}:j=1,2,\ldots,n\right\} $ represent the strengths of connections of unit $i$ with the rest of units, $\mathbf{f}_{t}=\left( f_{1t},f_{2t} ,...,f_{m_{0}t}\right) ^{\prime}$ is an $m_{0}\times1$ vector of latent factors with $m_{0}$ fixed, and $\boldsymbol{\gamma}_{i}=\left( \gamma _{i1},\gamma_{i2},...,\gamma_{im_{0}}\right) ^{\prime}$ is the vector of associated factor loadings.

We make the following assumptions that are mostly standard in the analysis of latent factor models.

assumption(a) $\mathbf{f}_{t}$ is a covariance-stationary process with zero means and the covariance matrix, $E\left( \mathbf{f}_{t} \mathbf{f}_{t}^{\prime}\right) =\mathbf{\Sigma}_{ff}>\mathbf{0}$. (b) $T^{-1}\sum_{t=1}^{T}\left[ \Vert\mathbf{f}_{t}\Vert^{j}-E\left( \Vert\mathbf{f}_{t}\Vert^{j}\right) \right] {\rightarrow}_{p}0$, for $j=3,4$, as $T\rightarrow\infty$. (c) There exists $T_{0}$ such that for all $T>T_{0}$, $T^{-1}\sum_{t=1}^{T}\mathbf{f}_{t}\mathbf{f}_{t}^{\prime }=T^{-1}\mathbf{F}^{\prime}\mathbf{F=\Sigma}_{T,ff}>\mathbf{0}$, and $\mathbf{\Sigma}_{T,ff}\rightarrow_{p}E\left( T^{-1}\mathbf{F}^{\prime }\mathbf{F}\right) =\mathbf{\Sigma}_{ff}>\mathbf{0}$, where $\mathbf{F=(f} _{1},\mathbf{f}_{2},...,\mathbf{f}_{T})^{\prime}$. (d) There exist constants $r_{1}$, $C_{0}$ and $C_{1}>0$ such that \begin{equation} \sup_{j,t}\Pr\left( \left\vert f_{jt}\right\vert >a\right) \leq C_{0} \exp\left( -C_{1}a^{r_{1}}\right), \end{equation} all $a>0$.
assumption(a) $\varepsilon_{it}\sim IID\left( 0,1\right) $ for all $i$ and $t$, and there exist constants $r_{2}$, $C_{2}$ and $C_{3}>0$ such that \begin{equation} \sup_{i,t}\Pr\left( \left\vert \varepsilon_{it}\right\vert >a\right) \leq C_{2}\exp\left( -C_{3}a^{r_{2}}\right), \end{equation} for all $a>0$. (b) $\mu_{\max}\left( \mathbf{V}_{\varepsilon T}\right) =O_{p}(n/T)$, where $\mathbf{V}_{\varepsilon T}=T^{-1}\sum_{t=1} ^{T}\boldsymbol{\varepsilon}_{\circ t}\boldsymbol{\varepsilon}_{\circ t}^{\prime}$, and $\boldsymbol{\varepsilon}_{\circ t}=(\varepsilon _{1t},\varepsilon_{2t},...,\varepsilon_{nt})^{\prime}$. (c) $\varepsilon_{it}$ is distributed independently of $\mathbf{f}_{t^{\prime}}$, for all $i,t$ and $t^{\prime}$, and there exists $v_{0}>0$ such that for all $v=T-m_{0}>v_{0}$ \begin{equation} \inf_{i}\left( v^{-1}\boldsymbol{\varepsilon}_{i\circ}^{\prime}\mathbf{M} _{F}\boldsymbol{\varepsilon}_{i\circ}\right) >c>0, \end{equation} where $\boldsymbol{\varepsilon}_{i\circ}=(\varepsilon_{i1},\varepsilon _{i2},...,\varepsilon_{iT})^{\prime}$, and $\mathbf{M}_{F}=\mathbf{I} _{T}-\mathbf{F(F}^{\prime}\mathbf{F)}^{-1}\mathbf{F}^{\prime}$.
assumptionThe $m_{0}\times1$ vector of factor loadings $\boldsymbol{\gamma}_{i}$ is bounded such that $\sup_{i}\left\Vert \boldsymbol{\gamma}_{i}\right\Vert <C$, $n^{-1}\sum_{i=1}^{n} \boldsymbol{\gamma}_{i}\boldsymbol{\gamma}_{i}^{\prime}=\mathbf{\Sigma }_{n,\gamma\gamma}\rightarrow\mathbf{\Sigma}_{\gamma\gamma}>\mathbf{0}$, and \begin{equation} 1-\theta_{n}>0, \end{equation} for all $n>n_{0}$ and as $n\rightarrow\infty,$ where $\theta_{n}=1-n^{-1} \sum_{i=1}^{n}a_{i,n}^{2}$, with $ \text{ }a_{i,n}=1-\sigma_{i}\boldsymbol{\varphi}_{n}^{\prime} \boldsymbol{\gamma}_{i}, $ and $\boldsymbol{\varphi}_{n}=n^{-1}\sum_{i=1}^{n}\boldsymbol{\gamma} _{i}/\sigma_{i}.$
remarkThe above assumptions relate closely to those made in the literature on CD tests and high dimensional factor models. See, for example, the assumptions in pesaran2004general,pesaran2015testing,pesaran2021general, and assumptions in bai2003inferential.\ Part (a) of Assumption (ref) will be relaxed when we consider panel regression models with observed regressors. The sub-exponential type conditions ((ref)) and ((ref)) are needed for bounding the probabilities across all $i$, and are also adopted by fan2011high, fan2013large and chudik2018one.
remarkSince $\mathbf{M}_{F}$ is an idempotent matrix with rank $v$, then there exists the orthogonal transformation $\boldsymbol{\eta}_{i} =(\eta_{i1},\eta_{i2},...,\eta_{iv})^{\prime}=\mathbf{H} \boldsymbol{\varepsilon}_{i\circ}$ where $\mathbf{H}$ is a $v\times T$ matrix such that $v^{-1}\boldsymbol{\varepsilon}_{i\circ}^{\prime}\mathbf{M} _{F}\boldsymbol{\varepsilon}_{i\circ}=v^{-1}\boldsymbol{\eta}_{i}^{\prime }\boldsymbol{\eta}_{i}>0,$ and $E\left( \eta_{it}^{2}\right) =1$. See, for example, durbin1950testing. Therefore, there exists a finite $v_{0}$ such that for all $v>v_{0}$ condition ((ref)) is met, and as result we also have \begin{equation} E\left( \frac{\boldsymbol{\varepsilon}_{i\circ}^{\prime}\mathbf{M} _{F}\boldsymbol{\varepsilon}_{i\circ}}{v}\right) ^{-s}<\frac{1}{c^{s} }<C<\infty, \end{equation} for all $i$ and any fixed $s>0$.

The focus of this paper is on testing the null hypothesis of error cross-sectional independence:

equation[equation omitted — 47 chars of source]

where $\lambda_{T}$ is defined by equation ((ref)). For the analysis of power we consider local alternatives:

equation[equation omitted — 67 chars of source]

with $c_{\lambda}\neq0$. We also introduce the following assumption on $\mathbf{W}=\left( w_{ij}\right) $.

assumptionThe connection matrix $\mathbf{W}$ has bounded maximum absolute column and row sum norms: \begin{equation} \left\Vert \mathbf{W}\right\Vert _{1}=\sup_{j}\sum_{i=1}^{n}\left\vert w_{ij}\right\vert <C, and \left\Vert \mathbf{W}\right\Vert _{\infty }=\sup_{i}\sum_{j=1}^{n}\left\vert w_{ij}\right\vert <C, \end{equation} and $w_{ii}=0$ for all $i$.
remarkThe connection matrix does not need to be symmetric. To see this, we can consider a more generalized setup of idiosyncratic errors, \begin{equation} \frac{u_{it}}{\sigma_{i}}=\varepsilon_{it}+\lambda_{T}{\sum\limits_{j=1}^{n} }\frac{\rho_{i}\sigma_{j}\mathring{w}_{ij}}{\sigma_{i}}\varepsilon_{jt}, \end{equation} where $\left\vert \lambda_{T}\right\vert <C$ and $\left\vert \rho _{i}\right\vert <C$. It is clear that ((ref)) reduces to $u_{it}$ in ((ref)) by letting $w_{ij}=\rho_{i}\sigma_{i}^{-1}\sigma_{j}\mathring {w}_{ij}$. In this way, $\mathbf{W}$ need not be symmetric even if the connection matrix $\left( \mathring{w}_{ij}\right) $ is symmetric. Under local alternatives and Assumption (ref), the specification of ((ref)) allows for a wide range of spatial and network dependence characterized by the connection matrix. It is in accord with an early discussion in chudik2011weak that spatial dependence can be captured by a weak factor model, so long as the number of weak factors tends to infinity with $n$, which is ruled out in standard factor models where the number of latent factors is assumed to be fixed.
remarkIt is also easily seen that $\varepsilon_{it}\left( \lambda_{T}\right) $ defined by ((ref)) is sub-exponential for any $\left\vert \lambda _{T}\right\vert <C$. This follows since by Assumption (ref) $\left\{ \varepsilon_{it}\right\} $ are independently and identically distributed sub-exponential processes, and by Assumption (ref) $\sup_{i}\sum_{j=1}^{n}\left\vert w_{ij}\right\vert <C.$ For a proof see part (b) of Theorem 5.5 in goldie1998subexponential or Theorem 2.8.2 in vershynin2018high.

Estimation of latent factors and the CD test

Following the literature we use principal component (PC) analysis to estimate the latent factors and their loadings. Let $\mathbf{Y=(y}_{1},\mathbf{y} _{2},\ldots,\mathbf{y}_{n}\mathbf{)}$ be the $T\times n$ matrix of observations on $y_{it}$, where $\mathbf{y}_{i}=\left( y_{i1},y_{i2} ,\ldots,y_{iT}\right) ^{^{\prime}}$ and denote the first $m_{0}$ largest eigenvalues of $\mathbf{Y}^{^{\prime}}\mathbf{Y}$ by $(\hat{\rho}_{1} ,\hat{\rho}_{2},...,\hat{\rho}_{m_{0}})$, and its associated $n\times m_{0}$ matrix of orthonormal eigenvectors by $\mathbf{\hat{Q}}$. The PC estimators of factors $\mathbf{F}=\left( \mathbf{f}_{1},\mathbf{f}_{2},\ldots ,\mathbf{f}_{T}\right) ^{^{\prime}}$ and their loadings $\boldsymbol{\Gamma }=\left( \boldsymbol{\gamma}_{1},\boldsymbol{\gamma}_{2},\ldots ,\boldsymbol{\gamma}_{n}\right) ^{^{\prime}}$ are then given by

equation[equation omitted — 396 chars of source]

By construction $n^{-1}\boldsymbol{\hat{\Gamma}}^{\prime}\boldsymbol{\hat {\Gamma}}=\mathbf{I}_{m_{0}},$ and $T^{-1}\mathbf{\hat{F}}^{^{\prime} }\mathbf{\hat{F}=D}_{nT}$, where $\mathbf{D}_{nT}=(nT)^{-1}\mathrm{diag} (\hat{\rho}_{1},\hat{\rho}_{2},...,\hat{\rho}_{m_{0}})$. Under Assumptions (ref)-(ref) the asymptotic results derived by bai2003inferential for PCs continue to apply here, and $u_{it}$ can be consistently estimated by

equation[equation omitted — 111 chars of source]

The CD test is based on the standardized residuals,

equation[equation omitted — 98 chars of source]

where $\hat{\sigma}_{i,T}=\left( T^{-1}\sum_{t=1}^{T}\hat{u}_{it}^{2}\right) ^{1/2}=\left( T^{-1}\mathbf{y}_{i}^{\prime}\mathbf{M}_{\hat{F}}\mathbf{y} _{i}\right) ^{1/2}$, and $\mathbf{M}_{\hat{F}}=\mathbf{I}_{T}-\mathbf{\hat {F}(\hat{F}}^{\prime}\mathbf{\hat{F})}^{-1}\mathbf{\hat{F}}^{\prime}$. Only units with non-zero $\hat{\sigma}_{i,T}^{2}$ are included in the construction of the CD test, namely

equation[equation omitted — 70 chars of source]

The standard CD test statistic based on the residuals, ((ref)), is given by

equation[equation omitted — 122 chars of source]

where $\hat{\rho}_{ij,T}=T^{-1}\sum_{t=1}^{T}\tilde{\varepsilon}_{it,T} \tilde{\varepsilon}_{jt,T}$.

juodis2022incidental apply the CD test to a panel regression model with latent factors, assuming that all the factors are strong. They show in that case $CD=O_{p}\left( \sqrt{T}\right) $, and its use will lead to gross over-rejection of the null of error cross-sectional independence. To deal with the over-rejection problem, these authors propose a randomized CD test, CD$_{W+}$. However, as shown in Section (ref) of the supplement, the CD$_{W+}$ test is likely to over-reject and tends to lack power against spatial and network alternatives. See also Section (ref) for Monte Carlo evidence on the small sample performance of the CD$_{W+}$ test.

The bias-corrected CD test

The main reason for the failure of the standard CD test in the case of latent factor models lies in the fact that both the factors and their loadings are unobserved and need to be estimated, and the differences between $\boldsymbol{\hat{\gamma}}_{i}^{^{\prime}}\mathbf{\hat{f}}_{t}$ and $\boldsymbol{\gamma}_{i}^{^{\prime}}\mathbf{f}_{t}$ do not tend to zero at a sufficiently fast rate for the CD test to be valid. Since the errors from estimation of $\boldsymbol{\gamma}_{i}^{^{\prime}}\mathbf{f}_{t}$ are included in the residuals $\hat{u}_{it}$, the resultant CD statistic tends to over-state the degree of underlying error cross-sectional dependence. This problem also arises when latent factors are proxied by cross section averages, as is the case when panel data models are estimated using correlated common effect (CCE) estimators proposed by pesaran2006estimation, which we shall address below in Section (ref).

We propose a bias-corrected CD test statistic, which we denote by $CD^{\ast}$, that directly corrects the asymptotic bias of the CD test using the estimates of the factor loadings and error variances. To obtain the expression for the bias we first note under the null hypothesis of cross-sectional independence, $CD=z_{nT}+o_{p}(1)$ and

equation[equation omitted — 142 chars of source]

where

equation[equation omitted — 183 chars of source]

$\boldsymbol{\varphi}_{n}=n^{-1}\sum_{i=1}^{n}\boldsymbol{\gamma}_{i} /\sigma_{i}$,$\boldsymbol{\ }$which is established in the proof of Proposition (ref) in the Appendix. Since $a_{i,n}$ are given constants, then $E\left( \xi_{t,n}\right) =0$,

equation[equation omitted — 231 chars of source]

and

equation[equation omitted — 193 chars of source]

where $\kappa_{2}=E\left( \varepsilon_{it}^{4}\right) -3$. Clearly, when the errors are Gaussian then $E\left( \varepsilon_{it}^{4}\right) =3$, and the second term of $Var\left( \xi_{t,n}^{2}\right) $ defined by ((ref)) is exactly zero. But even for non-Gaussian errors the second term of $Var\left( \xi_{t,n}^{2}\right) $ is negligible when $n$ is sufficiently large. To see this note that under Assumptions (ref) and (ref) \[ \frac{1}{n^{2}}\sum_{i=1}^{n}a_{i,n}^{4}=\frac{1}{n^{2}}\sum_{i=1} ^{n}(1-\sigma_{i}\boldsymbol{\varphi}_{n}^{\prime}\boldsymbol{\gamma}_{i} )^{4}\leq\frac{C}{n}, \] where $C$ is a positive constant. Since $\varepsilon_{it}$ (and henceforth $\xi_{t,n}$) are assumed to be serially independent, then we can also compute the mean and the variance of $z_{nT}$ as

align*[align* omitted — 332 chars of source]

The above expressions for $E\left( z_{nT}\right) $ give the source of the asymptotic bias of $CD$ as $E\left( z_{nT}\right) $ rises with $\sqrt{T}$, unless \[ \lim_{n\rightarrow\infty}\omega_{n}^{2}=\lim_{n\rightarrow\infty}n^{-1} \sum_{i=1}^{n}\left( 1-\sigma_{i}\boldsymbol{\varphi}_{n}^{\prime }\boldsymbol{\gamma}_{i}\right) ^{2}=1. \] A bias-corrected version of $CD$ can be defined by

equation[equation omitted — 105 chars of source]

where

equation[equation omitted — 169 chars of source]

and $1-\theta_{n}>0$ by condition ((ref)). Also upon using ((ref))

equation[equation omitted — 328 chars of source]

The main difference between $CD$ and $CD^{\ast}\left( \theta_{n}\right) $ depends on the magnitude of $\sqrt{T}\theta_{n}$, which in turn depends on the strengths of the factor loadings. Following bailey2021measurement, we measure the strength of factor $j$ by $\alpha_{j},$ defined by the rate at which the sum of absolute values of factor loadings rises with $n$, namely

equation[equation omitted — 147 chars of source]

where $0\leq\alpha_{j}\leq1$. Using ((ref)) it is now easily established that $\theta_{n}$ $=\ominus\left( n^{\alpha-1}\right) $, where $\alpha=max_{j=1,2,...,m_{0}}(\alpha_{j})$, and $\theta_{n}$ does not tend to zero when there is at least one strong factor in the panel data model.\footnote{For a proof see Section (ref) of the supplement.}. Therefore, based on ((ref)), the relationship between $CD$ and $CD^{\ast}\left( \theta_{n}\right) $ is essentially controlled by the maximum factor strength $\alpha\ $as $\sqrt{T}\theta_{n}=O\left( T^{1/2}n^{\alpha-1}\right) $. Suppose now $T=\ominus\left( n^{d}\right) $ for some $d>0$, then $\sqrt{T}\theta_{n} =\ominus\left( n^{\alpha+d/2-1}\right) ,$ and the bias correction becomes negligible if $\alpha<1-d/2$. Under the required relative expansion rates of $n$ and $T$ entertained in this paper, we need to set $d=1$, and for this choice the bias correction term, $\sqrt{T}\theta_{n}$, becomes negligible if $\alpha<1/2$, and as a result $CD$ and $CD^{\ast}\left( \theta_{n}\right) $ will be asymptotically equivalent. In fact, the case of strong factors assumed in the PCA literature corresponds to $\alpha_{j}=1$ for $j=1,2,\ldots,m_{0}$, which is also fulfilled by Assumption (ref) and used in our mathematical derivations.

The theoretical results for $CD^{\ast}(\theta_{n})$ are summarized in the following proposition.

propositionSuppose that observations on $y_{it}$, for $i=1,2,\ldots,n,$ and $t=1,2,\ldots,T$ are generated from the pure latent factor model given by ((ref)) and ((ref)), where the number of factors, $m_{0}$, is known. Consider the statistic $CD^{\ast}\left( \theta_{n}\right) $ defined by ((ref)) and assume $(n,T)\rightarrow \infty$, such that $n/T\rightarrow\kappa,$ and $0<\kappa<\infty$. (a) Under the null hypothesis $H_{0}$, defined by ((ref)), and supposing that Assumptions (ref) to (ref) hold, then \begin{equation} CD^{\ast}\left( \theta_{n}\right) \rightarrow_{d}\mathcal{N}(0,1). \end{equation} (b) Under local alternatives $H_{1T}$, defined by ((ref)), and supposing that Assumptions (ref) to (ref) hold, then \begin{equation} CD^{\ast}\left( \theta_{n}\right) \rightarrow_{d}\mathcal{N}(\phi,1), \end{equation} where $\phi=\lim_{n\rightarrow\infty}\phi_{n}$ and \begin{equation} \phi_{n}=\frac{\sqrt{2}c_{\lambda}}{1-\theta_{n}}n^{-1}\mathbf{a}_{n}^{\prime }\mathbf{Wa}_{n}, \end{equation} $\mathbf{W}=\left( w_{ij}\right) $ is the connection matrix, $\mathbf{a} _{n}=\left( a_{1,n},a_{2,n},\ldots,a_{n,n}\right) ^{\prime}$, with $a_{i,n}$ and $\theta_{n}$ defined by ((ref)) and ((ref)), respectively.

For a proof see the Appendix.

The bias-corrected test statistic, $CD^{\ast}(\theta_{n}),$ depends on the unknown parameter, $\theta_{n}$, which can be estimated by

equation[equation omitted — 94 chars of source]

where

equation[equation omitted — 274 chars of source]

The following proposition establishes the probability order of the difference between $\hat{\theta}_{nT}$ and $\theta_{n}$.

propositionSuppose that observations on $y_{it}$, for $i=1,2,\ldots,n,$ and $t=1,2,\ldots,T$ are generated from the pure latent factor model given by ((ref)) and ((ref)), where the number of factors, $m_{0}$, is known, and $\lambda_{T}=c_{\lambda}T^{-1/2}$ with $\left\vert c_{\lambda}\right\vert <\infty$. Consider the term $\theta_{n}$ in the $CD^{\ast}\left( \theta_{n}\right) $ statistic given by ((ref)) and its estimator $\hat{\theta}_{nT}$ given by ((ref)). Let Assumptions (ref) to (ref) hold and $(n,T)\rightarrow\infty$, such that $n/T\rightarrow\kappa$, where $0<\kappa<\infty$. Then \begin{equation} \sqrt{T}\left( \hat{\theta}_{nT}-\theta_{n}\right) =o_{p}(1). \end{equation}

For a proof see the Appendix.

Consider now the following feasible version of $CD^{\ast}\left( \theta _{n}\right) $,

equation[equation omitted — 141 chars of source]

and note that in view of ((ref)) and ((ref)) we have

align*[align* omitted — 366 chars of source]

Also, \[ \frac{1-\theta_{n}}{1-\hat{\theta}_{nT}}=1+\frac{\sqrt{T}\left( \hat{\theta }_{nT}-\theta_{n}\right) }{\sqrt{T}\left( 1-\theta_{n}\right) -\sqrt {T}\left( \hat{\theta}_{nT}-\theta_{n}\right) }=1+o_{p}(1), \] and hence $CD^{\ast}\left( \hat{\theta}_{nT}\right) =CD^{\ast}(\theta _{n})+o_{p}(1)$. We refer to $CD^{\ast}\left( \hat{\theta}_{nT}\right) $ simply as $CD^{\ast}$ and the test based on it as the CD$^{\text{*}}$ test. The main result of the paper for pure latent factor models is summarized in the following theorem.

theoremSuppose that observations on $y_{it}$, for $i=1,2,\ldots ,n,$ and $t=1,2,\ldots,T$ are generated from the pure latent factor model given by ((ref)) and ((ref)), where the number of factors, $m_{0}$, is known. Consider the statistic $CD^{\ast}$ defined by ((ref)), and assume $(n,T)\rightarrow\infty$, such that $n/T\rightarrow\kappa,$ and $0<\kappa<\infty$. (a) Under the null hypothesis $H_{0}$, defined by ((ref)), and supposing that Assumptions (ref) to (ref) hold, then \[ CD^{\ast}\rightarrow_{d}\mathcal{N}(0,1). \] (b) Under local alternatives $H_{1T}$, defined by ((ref)), and supposing that Assumptions (ref) to (ref) hold, then \[ CD^{\ast}\rightarrow_{d}\mathcal{N}\left( \phi,1\right) , \] where $\phi=\lim_{n\rightarrow\infty}\phi_{n}$, and $\phi_{n}$ is defined by ((ref)).

For a proof see the Appendix.

This theorem establishes the conditions under which the proposed CD$^{\ast}$ test has the correct size asymptotically. It also shows that the CD$^{\ast}$ test has power against network alternatives if the limit of $\phi_{n}$ defined by ((ref)) is nonzero, namely so long as $\lim_{n\rightarrow\infty }n^{-1}\mathbf{a}_{n}^{\prime}\mathbf{Wa}_{n}\neq0$. This condition is likely to be satisfied if the connection matrix, $\mathbf{W}$, is not too sparse, although it must be sufficiently sparse so that Assumption (ref) is met. In the case where there are no latent factors, $\mathbf{a}_{n} =(1,1,...,1)^{\prime}$, it is sufficient that $n^{-1}\sum_{i=1}^{n} \sum_{j=1}^{n}w_{ij}\neq0$.

To our knowledge, this is the first paper to provide a formal derivation of the power function of CD tests against spatial and network alternatives, which applies equally to the CD test for panel data models without latent factors. Hence, our derivation of the power function can be used to supplement earlier research on CD tests.

As we shall see from the Monte Carlo results reported below, the CD$^{\ast}$ test performs well even if some of the latent factors happen to be weak with $\alpha_{j}\in(0,1/2]$ or semi-strong with $\alpha_{j}\in\left( 1/2,1\right) $. This is because when a factor is weak, it does not matter if its estimation by PCA is not consistent at the standard rate of $\delta_{nT}=\min (n^{1/2},T^{1/2})$, and its inclusion or exclusion from the analysis has no material impact on the $CD^{\ast}$ statistics for $n$ and $T$ sufficiently large. In view of this result, in the mathematical derivations it is sufficient to consider the case of strong factors, and let the weak factors to be absorbed in the error term.

However, it should be acknowledged that our derivations do not take account of the case when one or more of the factors are semi-strong. Such an extension is beyond the scope of the present paper, although recent studies by bai2023approximate and jiang2023revisiting show that PCA estimation is asymptotically valid for factor models so long as factor strengths are all above $1/2$. It is therefore reasonable to conjecture that the CD$^{\text{*}}$ test applied to PCA residuals will be asymptotically valid even if some of the factors are semi-strong, namely if $1/2<\alpha_{j}<1$.

In practice, the true number of factors, $m_{0}$, is unknown. In cases where the estimated number of factors, $\hat{m}$, is underestimated ($\hat{m}<m_{0} $), the CD$^{\text{*}}$ test has power against the missing strong factors. However, the rejection of the null hypothesis by the CD$^{\text{*}}$ test does not necessarily mean there are missing factors, since the rejection could be due to network error dependence. It is, therefore, important for the investigator to decide on the number of strong latent factors before the implementation of the proposed CD$^{\text{*}}$ test. To that end, we refer the reader to the information criterion approach advanced by bai2002determining and the eigenvalue ratio test of ahn2013eigenvalue, for example.

The CD$^{\text{*}}$ test for panel regression models with interactive effects

Consider now the factor model ((ref)) augmented with observed regressors

equation[equation omitted — 189 chars of source]

where $\mathbf{d}_{t}$ is a $k_{d}\times1$ vector of observed common factors, $\mathbf{x}_{it}$ is a $k_{x}\times1$ vector of unit-specific observed covariates, $\boldsymbol{\alpha}_{i}=(\alpha_{i1},\alpha_{i2},...,\alpha _{ik_{d}})^{\prime}$, and $\boldsymbol{\beta}_{i}=(\beta_{i1},\beta _{i2},...,\beta_{ik_{x}})^{\prime}$ are their associated unknown coefficients. To highlight the relevance of the CD test for this set up, model ((ref)) can be written alternatively as

equation[equation omitted — 146 chars of source]

where the errors, $v_{it}$, follow the factor structure

equation[equation omitted — 89 chars of source]

The CD test is applicable, without any modifications, to test the null hypothesis that the errors of the panel regression model, $v_{it}$, are cross-sectionally independent, so long as the regressors, $\mathbf{d}_{t}$ and $\mathbf{x}_{it}$, are strictly exogenous with respect to $v_{it}$. When the regressors are correlated with the errors, the least squares estimates of $v_{it}$ become inconsistent and the standard CD test will fail. One important example of endogeneity arises when both $y_{it}$ and $\mathbf{x}_{it}$ are driven by the same latent factor(s). pesaran2006estimation formalizes this form of endogeneity by assuming that

equation[equation omitted — 168 chars of source]

where $\mathbf{A}_{i}$ and $\boldsymbol{\Gamma}_{i}$\ are $k_{d}\times k_{x}$ and $m_{0}\times k_{x}$ factor loading matrices and $\boldsymbol{\varepsilon }_{xit}$ are distributed independently of $\mathbf{f}_{t}$. The system of equations ((ref)), ((ref)) and ((ref)) fully specify the dependence of $\mathbf{x}_{it}$ and $v_{it}$, and allows consistent estimation of $v_{it}$ which can then be used to test the hypothesis that $u_{it}$ are cross-sectionally independent in the pure latent factor model ((ref)). We now show that the CD$^{\text{*}}$ test applied to these residuals will be valid. To this end we make the following additional standard assumptions.

assumption(a) The $k_{d}\times1$\ vector $\mathbf{d}_{t}$ is a covariance stationary process, with absolute summable autocovariances and $\mathbf{d}_{t}$ is distributed independently of $\mathbf{f}_{t^{^{\prime}}},$ for all $t$ and $t^{^{\prime}}$, such that $T^{-1}\mathbf{D}^{^{\prime} }\mathbf{F}=O_{p}\left( T^{-1/2}\right) $, where $\mathbf{D}=\left( \mathbf{d}_{1},\mathbf{d}_{2},\ldots,\mathbf{d}_{T}\right) ^{^{\prime}}$ and $\mathbf{F}=\left( \mathbf{f}_{1},\mathbf{f}_{2},\ldots,\mathbf{f} _{T}\right) ^{^{\prime}}$ are matrices of observations on $\mathbf{d}_{t}$ and $\mathbf{f}_{t}$. (b) $\left( \mathbf{d}_{t},\mathbf{f}_{t}\right) $ is distributed independently of $u_{is}$ and $\boldsymbol{\varepsilon}_{xis}$ for all $i,t,s.$
assumptionThe unobserved factor loadings $\boldsymbol{\Gamma }_{i}$ are bounded, i.e. $\left\Vert \boldsymbol{\Gamma}_{i}\,\right\Vert _{2}<C$ for all $i.$
assumptionThe individual-specific errors $\varepsilon_{it}$ in ((ref)) and $\boldsymbol{\varepsilon}_{xi,t^{\prime}}$ are distributed independently for all $i,j,t$ and $t^{\prime},$ and $\boldsymbol{\varepsilon}_{xit}$ follows the linear stationary process $\boldsymbol{\varepsilon}_{xit}=\sum_{l=0}^{\infty }\mathbf{S}_{il}\boldsymbol{\eta}_{xi,t-l},$ where for each $i$, $\boldsymbol{\eta}_{xit}$ is a $k_{x}\times1$ vector of serially uncorrelated random variables with mean zero, the variance matrix $\mathbf{I}_{k_{x}}$, and finite fourth-order cumulants. For each $i$, the coefficient matrices $\mathbf{S}_{il}$ satisfy the condition \[ Var\left( \boldsymbol{\varepsilon}_{xit}\right) =\sum_{l=0}^{\infty }\mathbf{S}_{il}\mathbf{S}_{il}^{^{\prime}}=\boldsymbol{\Sigma}_{xi}, \] where $\boldsymbol{\Sigma}_{xi}$ is a positive definite matrix, such that $\sup_{i}\left\vert \left\vert \boldsymbol{\Sigma}_{xi}\right\vert \right\vert _{2}<C$.
assumptionLet $\tilde{\boldsymbol{\Gamma}}=E\left( \boldsymbol{\gamma}_{i},\boldsymbol{\Gamma}_{i}\right) .$ We assume that $Rank\left( \tilde{\boldsymbol{\Gamma}}\right) =m_{0}.$
assumptionConsider\ the cross-sectional averages of the individual-specific variables, $\mathbf{z}_{it}=\left( y_{it},\mathbf{x} _{it}^{^{\prime}}\right) ^{^{\prime}}$ defined by $\bar{\mathbf{z}} _{t}=n^{-1}\sum_{i=1}^{n}\mathbf{z}_{it}$, and let $\mathbf{\bar{M} }=\mathbf{I}_{T}-\mathbf{\bar{H}}\left( \mathbf{\bar{H}}^{^{\prime} }\mathbf{\bar{H}}\right) ^{-1}\mathbf{\bar{H}}^{^{\prime}}$, and $\mathbf{M}_{g}=\mathbf{I}_{T}-\mathbf{G}\left( \mathbf{G}^{^{\prime} }\mathbf{G}\right) ^{-1}\mathbf{G}^{^{\prime}}$, where $\mathbf{\bar{H} =}\left( \mathbf{D},\mathbf{\bar{Z}}\right) ,$ $\mathbf{G}=\left( \mathbf{D,F}\right) ,$ and $\mathbf{\bar{Z}=}\left( \mathbf{\bar{z}} _{1},\mathbf{\bar{z}}_{2},\ldots,\mathbf{\bar{z}}_{T}\right) ^{\prime}$ is the $T\times\left( k_{x}+1\right) $ matrix of observations on the cross-sectional averages. Let $\mathbf{X}_{i}=\mathbf{(x}_{i1},\mathbf{x} _{i2},...,\mathbf{x}_{iT})^{\prime}$, then the $k\times k$ matrices $\mathbf{\hat{\Psi}}_{i,T}=T^{-1}\mathbf{X}_{i}^{^{\prime}}\mathbf{\bar{M} X}_{i}$ and $\mathbf{\Psi}_{ig}=T^{-1}\mathbf{X}_{i}^{^{\prime}}\mathbf{M} _{g}\mathbf{X}_{i}$ are non-singular, and $\mathbf{\hat{\Psi}}_{i,T}^{-1}$ and $\mathbf{\Psi}_{ig}^{-1}$ have finite second-order moments for all $i.$ \begin{remark} The above assumptions are standard in the panel data models with multi-factor error structure. See, for example, pesaran2006estimation. But in our setup under Assumption (ref) we require the error term, $\varepsilon_{it}$, to be serially independent, since our focus is on testing $\varepsilon_{it}$ for cross-sectional independence, and this assumption is needed for asymptotic normality of the bias-corrected CD test. Later in Section (ref), we will consider models with serially correlated errors and show that the bias-corrected CD test remains valid. Nevertheless, we allow $\varepsilon_{xit}$, the errors in the $\mathbf{x}_{it}$ equations to be serially correlated. Assumption (ref) separates the observed and the latent factors, as in Assumption 11 of pesaran2011large. This assumption is required to obtain the probability order of estimated residuals needed for computation of $CD^{\ast}$ statistic. A necessary condition for the rank condition in Assumption (ref) to hold is $k_{x}\geq m_{0}-1$. \end{remark}

To estimate $v_{it}$ we first filter out the effects of observed covariates using the CCE estimators proposed in pesaran2006estimation, namely for each $i$ we estimate $\boldsymbol{\beta}_{i}$ by

equation[equation omitted — 210 chars of source]

and following pesaran2011large, estimate $\boldsymbol{\alpha}_{i}$ by

equation[equation omitted — 223 chars of source]

Then we have the following estimator of $v_{it}$

equation[equation omitted — 176 chars of source]

Using results in pesaran2011large (p. 189) it follows that under Assumptions (ref)-(ref)

equation[equation omitted — 170 chars of source]

Note when $\boldsymbol{\alpha}_{i}=\mathbf{0}$ and $\boldsymbol{\beta} _{i}=\mathbf{0}$, ((ref)) reduces to the pure latent factor model, ((ref)), where PCA can be applied to $v_{it}=y_{it}$ directly. In the case of panel regressions $\hat{v}_{it}$ can be used instead of $v_{it}$ to compute the bias-corrected CD statistic given by ((ref)). The errors involved will become asymptotically negligible in view of the fast rate of convergence of $\hat{v}_{it}$ to $v_{it}$, uniformly for each $i$ and $t$. Specifically, as in the case of the pure latent factor model, we first compute $m_{0}$ PCs of $\left\{ \hat{v}_{it};\text{ }i=1,\ldots,n;\text{ and }t=1,\ldots ,T\right\} $ and the associated factor loadings, $(\boldsymbol{\hat{\gamma} }_{i},\mathbf{\hat{f}}_{t}),$ subject to the normalization $n^{-1}\sum _{i=1}^{n}\boldsymbol{\hat{\gamma}}_{i}\boldsymbol{\hat{\gamma}}_{i} ^{^{\prime}}=\mathbf{I}_{m_{0}}$. The residuals

equation[equation omitted — 169 chars of source]

can then be used to compute the standard CD statistic, ((ref)), and its bias-corrected version, $CD^{\ast}$, using ((ref)).

remarkIt is important to bear in mind that $\hat{u}_{it}$ is not the same as the CCE residuals that result from running the panel regressions of $y_{it}$ on $(\mathbf{d}_{t},\mathbf{x}_{it},\mathbf{\bar{z}}_{t})$. As shown by juodis2022incidental, the standard CD test applied to the CCE residuals will result in over-rejection and is not recommended. In our approach, we filter out the latent factors from $\hat{v}_{it}$ and use the filtered residuals, $\hat{u}_{it}$, to compute the CD statistic and correct it, as in CD$^{\ast}$, to allow for errors associated with estimation of factors and their loadings.

The following theorem extends Theorem (ref) to panel regression models with observed regressors.

theoremSuppose that observations on $y_{it}$, for $i=1,2,\ldots ,n,$ and $t=1,2,\ldots,T$ are generated from the panel regression model defined by ((ref)),\ ((ref)) and ((ref)), where the number of latent factors in ((ref)), $m_{0}$, is known. Consider the statistic $CD^{\ast}$ given\ by ((ref)) using the filtered residuals defined by ((ref)). Suppose that $(n,T)\rightarrow\infty$, such that $n/T\rightarrow\kappa,$ and $0<\kappa<\infty$. (a) Under the null hypothesis $H_{0}$, defined by ((ref)), and supposing that Assumptions (ref) to (ref) and Assumptions (ref) to (ref) hold, then \[ CD^{\ast}\rightarrow_{d}\mathcal{N}(0,1). \] (b) Under local alternatives $H_{1T}$, defined by ((ref)), and supposing that Assumptions (ref) to (ref) hold, then \[ CD^{\ast}\rightarrow_{d}\mathcal{N}\left( \phi,1\right) , \] where $\phi=\lim_{n\rightarrow\infty}\phi_{n}$ and $\phi_{n}$ defined by ((ref)).

For a proof see the Appendix.

CD$^{\text{*}}$ tests for models with serially correlated errors

As shown by baltagi2016testing, when the errors $u_{it}$ in ((ref)) are serially correlated the variance of the standard CD test statistic is not unity (even asymptotically) and the test is no longer valid. The same also applies to the CD$^{\text{*}}$ test. To deal with this problem, we propose two solutions which involve different ways of adjusting the CD$^{\ast}$ test so that it will become applicable to panels with serially correlated errors. The first method closely follows the variance adjustment proposed by baltagi2016testing, in which $CD^{\ast}$ is scaled by $\varpi$ where

equation[equation omitted — 429 chars of source]

with $\boldsymbol{\tilde{\varepsilon}}_{i,T}=(\tilde{\varepsilon} _{i1,T},\tilde{\varepsilon}_{i2,T},\ldots,\tilde{\varepsilon}_{iT,T})^{\prime }$, $\tilde{\varepsilon}_{it,T}$ defined in ((ref)) and \[ \boldsymbol{\tilde{\varepsilon}}_{\left( ij\right) ,T}=\frac{1}{n-2} \sum_{1\leq\tau\neq i,j\leq n}\boldsymbol{\tilde{\varepsilon}}_{\tau,T}. \] The expression in ((ref)) is the equivalent to that provided in Theorem 3 of baltagi2016testing but the factor of $2$ in ((ref)) is missing in their paper. The same adjustment is also applied to the CD$_{W+}$ test to allow for serially correlated errors.

Alternatively, following pesaran2004general, we first transform the panel regression model to eliminate the error serial correlation and then apply the CD$^{\text{*}}$ test to the residuals of the transformed model. This is possible so long as the error serial correlation can be approximated by a finite order stationary autoregressive process. As a simple illustration consider the pure latent factor model $y_{it}=\gamma_{i}f_{t}+u_{it},$ in which factor $f_{t}$ and loading $\gamma_{i}$ are both latent, and the errors $u_{it}$ are generated as $AR(1)$ processes, $u_{it}=\rho_{i}u_{it-1} +\epsilon_{it},$ where $\rho_{i}$ is the autoregression coefficient and $\epsilon_{it}$ is serially independent, as well as being distributed independently of $f_{t^{\prime}}$ for all $i$ and $t,t^{\prime}=1,2,\ldots,T$. Testing the cross-sectional independence of $u_{it}$ is equivalent to testing the cross-sectional independence of $\epsilon_{it}$ in the following autoregressive distributed lag (ARDL) representation of $y_{it}$ \[ y_{it}=\rho_{i}y_{i,t-1}+\gamma_{i}f_{t}-\rho_{i}\gamma_{i}f_{t-1} +\epsilon_{it}, \] which can be written equivalently as a multi-factor AR panel regression model

equation[equation omitted — 145 chars of source]

where $\mathbf{\mathring{f}}_{t}=(f_{t},f_{t-1})^{\prime}$, and $\boldsymbol{\mathring{\gamma}}_{i}=\left( \gamma_{i},-\rho_{i}\gamma _{i}\right) ^{\prime}$. Since $y_{i,t-1}$ is weakly exogenous, the transformed model satisfies the setup of panel regression model ((ref)) with $\mathbf{\mathring{f}}_{t}$ viewed as a vector of latent variables with the associated factor loadings, $\boldsymbol{\mathring{\gamma}}_{i}$. It therefore follows that the CD$^{\text{*}}$ test can now be applied to test the cross-sectional independence of $\epsilon_{it}$ in ((ref)). We refer to this test as the ARDL adjusted CD$^{\text{*}}$ test.

The same approach can also be used for panels with observed covariates. In general, testing cross-sectional independence of $u_{it}$ in model ((ref)) is equivalent to testing the cross-sectional independence of $\epsilon_{it}$ in

equation[equation omitted — 256 chars of source]

where $\mathbf{h}_{t}$\ is an extended set of latent factors (that encompass $\mathbf{f}_{t}$), and $\mathbf{g}_{i}$ are the associated factor loadings. The number of lags $S$ is determined by the order of the AR specification assumed for $u_{it}$ in ((ref)).

The variance adjustment is simpler to implement but it requires theoretical justification in the context of panel data models with latent factors. The ARDL adjustment is theoretically justified so long as the underlying errors follow finite order AR processes. As we shall see both approaches work well in dealing with serially correlated errors, at least in the context of the limited MC designs that we are considering. Clearly, further theoretical and Monte Carlo investigations are needed for a better understanding of the relative merits of the two approaches.

Small sample properties of CD$^{\ast}$ and CD$_{W^{+}}$ tests

Data generating process

We consider the following data generating process

equation[equation omitted — 243 chars of source]

where $\varepsilon_{it}\left( \lambda\right) $ follows the first order spatial autoregressive process, $SAR\left( 1\right) $, such that

equation[equation omitted — 165 chars of source]

$\text{a}_{i}$ is a unit-specific effect$,$ $d_{t}$ is the observed common factor, $x_{it}$ is the observed regressor that varies across $i$ and $t$, $\mathbf{f}_{t}$ is the $m_{0}\times1$ vector of unobserved factors, $\boldsymbol{\gamma}_{i}$ is the vector of associated factor loadings. The scalar constants, $\sigma_{i}>0$, are generated as $\sigma_{i}^{2} =0.5+\frac{1}{2}\left( s_{i}^{2}-1\right) $, with $s_{i}^{2}\sim IID\chi ^{2}(2)$, which ensures that $E(\sigma_{i}^{2})=1$.

DGP under the null hypothesis

Under the null hypothesis, we set $\lambda=0$ and $c=1$, and consider both serially independent errors and serially correlated errors, which are generated by both Gaussian and non-Gaussian distributions:

itemize• Serially independent errors: Gaussian errors, $\varepsilon_{it}\sim IID\mathcal{N}(0,1)$; chi-squared distributed errors, $\varepsilon_{it}\sim IID\left( \frac{\chi^{2}(2)-2}{2}\right) $. • Serially correlated errors: $\varepsilon_{it}=\rho_{\varepsilon }\varepsilon_{it-1}+\sqrt{1-\rho_{\varepsilon}^{2}}$ $e_{\varepsilon it}$, for $i=1,2,\ldots,n$ and $t=1,2,\ldots,T$, where $\rho_{\varepsilon}=0.5$ and $e_{\varepsilon it}$ are generated as Gaussian errors, $e_{\varepsilon it}\sim IID\mathcal{N}\left( 0,1\right) $, or chi-squared distributed errors, $e_{\varepsilon it}\thicksim IID\left( \frac{\chi^{2}(2)-2}{2}\right) $.

The focus of the experiments is on testing the null hypothesis that $\varepsilon_{it}$ are cross-sectional independent, whilst allowing for the presence of $m_{0}$ unobserved factors, $\mathbf{f}_{t}=(f_{1t},f_{2t} ,...,f_{m_{0}t})^{\prime}$. We consider $m_{0}=1$ and $m_{0}=2$, and generate the factor loadings $\boldsymbol{\gamma}_{i}=(\gamma_{i1},\gamma_{i2} )^{\prime}$ as:

align*[align* omitted — 372 chars of source]

In the one-factor case ($m_{0}=1$), we only include $f_{1t}$ as the latent factor and denote its factor strength by $\alpha$. Three values of $\alpha$ are considered, namely $\alpha=1,2/3,1/2$, respectively representing strong, semi-strong and weak factors. Similarly, in the two-factor case ($m_{0}=2$), we include both $f_{1t}$ and $f_{2t}$ as the latent factors and consider the following combinations of factor strengths: $(\alpha_{1},\alpha_{2})=\left[ (1,1),(1,2/3),(2/3,1/2)\right] $. The intercepts a$_{i}$ are generated as $IID\mathcal{N}(1,2)$ and fixed thereafter. The observed common factor is generated as $d_{t}=\rho_{d}d_{t-1}+\sqrt{1-\rho_{d}^{2}}$ $v_{dt}$, with $\rho_{d}=0.8,$ and $v_{dt}\thicksim IID\mathcal{N}(0,1)$, thus ensuring that $E(d_{t})=0$ and $Var(d_{t})=1$. The observed unit-specific regressors, $x_{it}$, for $i=1,2,\ldots,n$ are generated to have non-zero correlations with the unobserved factors:

equation[equation omitted — 82 chars of source]

where $f_{jt}=r_{j}f_{j,t-1}+\sqrt{1-r_{j}^{2}}$ $v_{jt}$, with $r_{j}=0.9$ and $v_{jt}\sim IID\left( \frac{\chi^{2}(2)-2}{2}\right) $, for $j=1,2$. The factor loadings in ((ref)) are generated as $\gamma_{xi1}\sim IIDU\left( 0.25,0.75\right) $ and $\gamma_{xi2}\sim IIDU\left( 0.1,0.5\right) $. The error term of ((ref)) is generated as $e_{xit}=\rho_{i}e_{xi,t-1} +\sqrt{1-\rho_{i}^{2}}$ $v_{xit}\text{, }$where $\rho_{i}\sim IIDU(0,0.95)$ and $v_{xit}\thicksim IID\mathcal{N}(0,1)$.

We will examine the small sample properties of the CD and the bias-corrected CD tests for both the pure latent factor model and for the panel regression model which also includes observed covariates.

itemize• In the case of the pure latent factor model we set $\beta_{i1} =\beta_{i2}=0$. • In the case of the panel regression model with latent factors, we allow for heterogeneous slopes and generate the slopes of observed covariates, $d_{t}$ and $x_{it}$, as $\beta_{i1}\sim IID\mathcal{N}(\mu_{\beta1} ,\sigma_{\beta1}^{2}),$ and $\beta_{i2}\sim IID\mathcal{N}(\mu_{\beta2} ,\sigma_{\beta2}^{2})$ where $\mu_{\beta1}=\mu_{\beta2}=0.5$ and $\sigma_{\beta1}^{2}=\sigma_{\beta2}^{2}=0.25,$ respectively.

As our theoretical results show the null distributions of the CD and the bias-corrected CD tests do not depend on a$_{i}$, $\beta_{i1}$ and $\beta _{i2}$, it is therefore innocuous what values are chosen for these parameters. Moreover, the average fit of the panel is controlled in terms of the limiting value of the pooled R-squared defined by

equation[equation omitted — 207 chars of source]

Since the underlying processes, ((ref)) and ((ref)), are stationary and $E\left( \varepsilon_{it}^{2}\right) =1$, we have \[ \lim_{T\rightarrow\infty}PR_{nT}^{2}=PR_{n}^{2}=\frac{n^{-1}\sum_{i=1} ^{n}\sigma_{i}^{2}\left[ \beta_{i1}^{2}+\beta_{i2}^{2}Var\left( x_{it}\right) +m_{0}^{-1}\boldsymbol{\gamma}_{i}^{\prime}\boldsymbol{\gamma }_{i}+2Cov\left( x_{it},\boldsymbol{\gamma}_{i}^{^{\prime}}\mathbf{f} _{t}\right) \right] }{n^{-1}\sum_{i=1}^{n}Var\left( y_{it}\right) }, \] where $\boldsymbol{\gamma}_{i}=\left( \gamma_{i1},\gamma_{i2}\right) ^{\prime},$ $Var\left( x_{it}\right) =\boldsymbol{\gamma}_{xi}^{^{\prime} }\boldsymbol{\gamma}_{xi}+1,$ $Cov\left( x_{it},\boldsymbol{\gamma} _{i}^{^{\prime}}\mathbf{f}_{t}\right) =\boldsymbol{\gamma}_{xi}^{^{\prime} }\boldsymbol{\gamma}_{i},$ $\boldsymbol{\gamma}_{xi}=\left( \gamma _{xi1},\gamma_{xi2}\right) ^{\prime}$, and \[ Var\left( y_{it}\right) =\sigma_{i}^{2}\left[ \beta_{i1}^{2}+\beta_{i2} ^{2}Var\left( x_{it}\right) +m_{0}^{-1}\boldsymbol{\gamma}_{i}^{\prime }\boldsymbol{\gamma}_{i}+2m_{0}^{-1/2}Cov\left( x_{it},\boldsymbol{\gamma }_{i}^{^{\prime}}\mathbf{f}_{t}\right) +1\right] . \] Also since $\sigma_{i}^{2}$ and $\beta_{ij}$ are independently distributed and $E(\sigma_{i}^{2})=1$, it then readily follows that $\lim_{n\rightarrow\infty }PR_{n}^{2}=\eta^{2}/(1+\eta^{2})$, where \[ \eta^{2}=\mu_{\beta1}^{2}+\sigma_{\beta1}^{2}+\left( \mu_{\beta2}^{2} +\sigma_{\beta2}^{2}\right) \left[ 1+E\left( \boldsymbol{\gamma} _{xi}^{^{\prime}}\boldsymbol{\gamma}_{xi}\right) \right] +\frac{2\mu _{\beta2}E\left( \boldsymbol{\gamma}_{xi}^{^{\prime}}\boldsymbol{\gamma} _{i}\right) }{\sqrt{m_{0}}}+\frac{E\left( \boldsymbol{\gamma}_{i}^{^{\prime }}\boldsymbol{\gamma}_{i}\right) }{m_{0}}. \] By controlling the value of $\eta^{2}$ across the experiments we ensure that the pooled R$^{2}$ in large samples is the same for all values of $\sigma _{i}^{2}$. In particular, in the case of the pure latent model we have $\eta^{2}=m_{0}^{-1}E\left( \boldsymbol{\gamma}_{i}^{^{\prime}} \boldsymbol{\gamma}_{i}\right) =O\left( n^{\alpha-1}\right) ,$ where $\alpha=max(\alpha_{1},\alpha_{2})$.

DGP under alternative hypotheses

Under alternative hypotheses, using ((ref)), we consider a spatial alternative defined by

equation[equation omitted — 201 chars of source]

where $\boldsymbol{\varepsilon}_{\circ t}\left( \lambda\right) =\left( \varepsilon_{1t}\left( \lambda\right) ,\varepsilon_{2t}\left( \lambda\right) ,\ldots,\varepsilon_{nt}\left( \lambda\right) \right) ^{\prime},$ $\mathbf{W}=(w_{ij})$, and $\boldsymbol{\varepsilon}_{\circ t}=(\varepsilon_{1t},\varepsilon_{2t},\ldots,\varepsilon_{nt})^{\prime}$. The errors $\varepsilon_{it}$ are generated as described above. For the spatial weights $w_{ij}$, we first set $w_{ij}^{0}=1$ if $j=i-2,i-1,i+1,i+2,$ and zero otherwise. We then row normalize the weights such that $w_{ij}=\left( \sum_{j=1}^{n}w_{ij}^{0}\right) ^{-1}w_{ij}^{0}$. We also set $c\left( \lambda\right) ^{2}=n/$$tr$$\left[ \left( \mathbf{I}_{n} -\lambda\mathbf{W}\right) ^{-1}\left( \mathbf{I}_{n}-\lambda\mathbf{W} \right) ^{\prime-1}\right] $, which ensures that $n^{-1}\sum_{i=1} ^{n}Var(\varepsilon_{it}\left( \lambda\right) )=1$, for all values of $\lambda$. In practice, only positive values of $\lambda$ are of interest, and the power function need not be symmetric for all positive and negative values of $\lambda$.

CD, CD$^{\ast}$ and CD$_{W^{+}}$ tests

All experiments are carried out for $n=100,200,500,1000$ and $T=100,200,500$, and the number of replications is set to $2000$. Firstly we consider the DGPs with serially independent errors. For the pure latent factor models, we compute the filtered residuals as $\hat{v}_{it}=y_{it}-\hat{\text{a}}_{i}$, where $\hat{\text{a}}_{i}=T^{-1}\sum_{t=1}^{T}y_{it}$. For the panel regressions with latent factors, the filtered residuals are computed as

equation[equation omitted — 126 chars of source]

where $\left( \hat{\text{a}}_{CCE,i},\hat{\beta}_{CCE,i1},\hat{\beta }_{CCE,i2}\right) $ is the CCE estimator of a$_{i}$, $\beta_{i1}$ and $\beta_{i2}$, as set out in pesaran2006estimation. The residuals $\{\hat{v}_{it};$ $i=1,2,\ldots,n;$ and $t=1,2,\ldots,T\}$, together with their first $m$ PCs and the associated factor loadings, $(\boldsymbol{\hat {\gamma}}_{i},\hat{\mathbf{f}}_{t}),$ are then used to compute the filtered residuals, $\hat{u}_{it}=\hat{v}_{it}-\boldsymbol{\hat{\gamma}}_{i}^{\prime }\hat{\mathbf{f}}_{t}$, to compute the CD test statistics, $CD$ and $CD^{\ast }$, given by ((ref)) and ((ref)), respectively. For comparison, we also consider the power enhanced version of the randomized CD test statistic proposed by JR given by

equation[equation omitted — 56 chars of source]

where

equation[equation omitted — 228 chars of source]

The weights $w_{i}$, for $i=1,2,...,n$ are independently drawn from a Rademacher distribution and

equation[equation omitted — 209 chars of source]

where $\hat{\rho}_{ij,T}=T^{-1}\sum_{t=1}^{T}\tilde{\varepsilon}_{it,T} \tilde{\varepsilon}_{jt,T}$, and $\tilde{\varepsilon}_{it,T}$ is defined by ((ref)). As shown by JR, $CD_{W}$ has a zero mean by construction and avoids the over-rejection problem of the CD test, but it can also lack power by the very nature of the randomization process. JR further suggest $CD_{W+}$ by adding a screening component $\Delta_{nT}$ proposed by fan2015power, which enhances the power of the test since $\Delta_{nT}$ converges to zero as $n$ and $T\rightarrow\infty$ under the null hypotheses, but can diverge under alternatives with a sufficient number of $(i,j)$ pairs with non-zero correlations, $\rho_{ij}$.

As discussed in Section (ref), the CD$^{\text{*}}$ test is not valid when the errors are serially correlated. In the simulations, we apply the variance and ARDL adjustments to $CD$, $CD_{W+}$, and $CD^{\ast}$. The variance adjusted versions are computed by scaling the original statistics by the standard deviation of the CD statistics using the expression in ((ref)) with $\tilde{\varepsilon}_{it,T}=\hat{u}_{it}/\hat{\sigma}_{i,T}$, where $\hat{u}_{it}=\hat{v}_{it}-\boldsymbol{\hat{\gamma} }_{i}^{\prime}\hat{\mathbf{f}}_{t}$. The ARDL adjusted versions of $CD$, $CD^{\ast}$, and $CD_{W+}$, are computed using the residuals from the following dynamic panel data model with latent factors,

equation[equation omitted — 204 chars of source]

In the simulations we set $S=1$, but higher order values can also be considered. The number of latent factors in $\mathbf{h}_{t}$ depends on $S$ and is given by $m_{h}=(S+1)m_{0}$. Accordingly, the number of selected PCs, $\hat{m}$, should satisfy $\hat{m}\geq(S+1)m_{0}$. In the simulations if $S=0$, we consider $\hat{m}=1$ and $2$ if $m_{0}=1$, and $\hat{m}=2$ and $4$ if $m_{0}=2$. But if $S=1$ we consider $\hat{m}=2$ and $4$ if $m_{0}=1$, and $\hat{m}=4$ and $6$ if $m_{0}=2$. Seen from this perspective, the variance adjustment approach to dealing with error serial correlation seems preferable since it does not require specifying the lag order $S$.

Simulation results

We first report the simulation results for the DGPs with normally distributed errors, followed by the results based on DGPs with chi-squared distributed errors. Next, we report simulation results for the DGPs with serially correlated errors, using the variance and ARDL adjusted CD tests discussed in Section (ref). Finally, to investigate the power of the CD$^{\text{*}}$ test we consider the spatial $SAR(1)$ alternative with $\lambda=0.25$. As to be expected the power rises very quickly as $\lambda$ deviates from $0$.\footnote{Simulated power functions are provided in the supplement for $\lambda=\pm0.05$, $\pm0.1$, $\pm0.2$, $\pm0.3$, $\pm0.4$, $\pm0.5$, $\pm0.6$, $\pm0.7$, $\pm0.8$, $\pm0.9$, $\pm0.95$.}

Serially independent errors: normally distributed errors

The simulation results for the DGPs with the errors following Gaussian distribution are shown in Tables (ref) to (ref). Tables (ref) and (ref) report the test results for the latent factor model with one factor. Table (ref) gives the results for the case where the number of selected PCs, denoted by $\hat{m}$, is the same as the true number of factors ($m_{0}=1$), while Table (ref) reports the results when $\hat{m}=2$. As to be expected the standard CD test over-rejects when the factor is strong, namely when $\alpha=1$. By comparison, the rejection frequencies of both CD$^{\ast}$ and CD$_{W+}$ tests under null ($\lambda=0)$ are generally around the nominal size of $5$ per cent. Under the alternative (when $\lambda=0.25$), the CD$^{\ast}$ test has satisfactory power properties with significantly high rejection frequencies even when the sample size is small. But the CD$_{W+}$ test performs quite poorly under the spatial alternative, especially when $T$ is small.

Tables (ref) and (ref) summarize the size and power results for the latent factor model with $m_{0}=2,$ and reports the results when $\hat{m}$ (the selected number of PCs) is set to $2$ (Table (ref)) and $4$ (Table (ref)). The results are qualitatively similar to the ones reported for the one factor model. The CD test over-rejects if at least one of the factors is strong, and the empirical sizes of CD$^{\ast}$ and CD$_{W+}$ tests are close to their nominal value of $5$ per cent, although we now observe some mild over-rejection when $n=100$ and the selected number of PCs is 4. In terms of power, the CD$^{\ast}$ test performs well, although there is some loss of power as the numbers of factors and selected PCs rise. Similarly, the power of the CD$_{W^{+}}$ test is now even lower and quite close to $5$ per cent when $T<500$ even if the number of PCs is set to $m_{0}=2$.

Turning to panel regression models with latent factors estimated by CCE, the associated simulation results are summarized in Tables (ref) to (ref). As can be seen, the results are very close to the ones reported in Tables (ref) to (ref) for the latent factor model, and are in line with the asymptotic result in ((ref)) that underlies the use of CCE approach to filter out the effects of observed covariates, as well as latent factors.

table[table omitted — 3,523 chars of source]
table[table omitted — 2,957 chars of source]
table[table omitted — 3,244 chars of source]
table[table omitted — 3,057 chars of source]
table[table omitted — 3,202 chars of source]
table[table omitted — 2,952 chars of source]
table[table omitted — 3,310 chars of source]
table[table omitted — 3,069 chars of source]

Serially independent errors: chi-squared distributed errors

To save space, the simulation results for the DGPs with chi-squared errors are provided in Tables (ref) to (ref) in the supplement. For the standard CD test and its biased-corrected version, CD$^{\ast}$, as shown in Tables (ref) and (ref), the results are very similar to the ones with Gaussian errors, suggesting that the CD$^{\text{*}}$ test is likely to be robust to departures from Gaussianity. As with the experiments with Gaussian errors, the standard CD test continues to over-reject unless $\alpha<2/3$, and the CD$^{\text{*}}$ test has the correct size for all $n$ and $T$ combinations, except when the number of selected PCs is large relative to $m_{0}$, and $T=100$. The main difference between the results with and without Gaussian errors is the tendency for the CD$_{W^{+}}$ test to over-reject when $n>T$, which seems to be a universal feature of this test and holds for all choices of $m_{0}$ and the number of selected PCs, irrespective of whether the factors are strong or weak. This could be due to the screening component of the CD$_{W+}$ test not tending to zero sufficiently fast with $n$ and $T$. Furthermore, the CD$^{\text{*}}$ test continues to have satisfactory power, but the CD$_{W^{+}}$ test clearly lacks power against spatial or network alternatives that are of primary interest.

Similar results are obtained for panel regression models with latent factors, summarized in Tables (ref) and (ref) in the supplement.

Serially correlated errors

To save space, the results for the DGPs with serially correlated errors are summarized in Tables (ref) to (ref) in the supplement. Tables (ref) to (ref) give the simulation results for the variance adjusted CD tests, whilst Tables (ref) to (ref) provide the results for the ARDL adjusted tests. Overall, the results corroborate our earlier findings obtained for DGPs with serially independent errors. Both adjustments for serial error correlation work well, with size and power of the adjusted CD$^{\text{*}}$ tests being quite close to the results already reported for DGPs with serially independent errors. It is also clear that without adjustments for latent factors and error serial correlation, the standard CD test will lead to large size distortions when the latent factors are strong. But in line with our theoretical results, the standard CD test, when adjusted for error serial correlation if needed, tends to have the correct size when the latent factors are weak.

Comparing the two types of adjustments for error serial correlations (for pure latent factor models as well as for panel regression models with latent factors), the variance adjusted CD$^{\text{*}}$ test works particularly well, and only shows mild over-rejection in the case where $T=100$ and $n>T$. In contrast, the CD$_{W+}$ test with variance adjustment over-rejects for all combinations of $n$ and $T$.

The ARDL adjusted version of the CD$^{\text{*}}$ test also works well when the number of PCs is not too large, and tends to have the correct size for all $(n,T)$ combinations and only shows slight over-rejection when $n=100$. The CD$_{W+}$ test using ARDL adjustment does better in controlling for the size when the errors are Gaussian, but tends to over-reject when the errors are chi-squared distributed and $n>T$. Both adjusted versions of the CD$_{W+}$ test continue to lack power against spatial or network alternatives.

Empirical application

It is well known that house price changes are spatially correlated, but it is unclear if such correlations are mainly due to common factors (national or regional) or arise from spatial spillover effects not related to the common factors, a phenomenon also referred to as the ripple effect. See, for example, holly2011spatial, tsai2015spillover, chiang2016ripple, bailey2016two, and aquaro2021estimation. To test for the presence of ripple effects the influence of common factors must first be filtered out and this is often a challenging exercise due to the latent nature of regional and national factors. Therefore, to find if there exist local spillover effects, one needs to test for significant residual cross-sectional dependence once the effects of common factors are filtered out.

We consider quarterly data on real house prices at the level of Metropolitan Statistical Areas (MSAs) in the U.S. There are 381 MSAs, under the February 2013 definition provided by the U.S. Office of Management and Budget (OMB). We use quarterly data on real house price changes compiled by yang2021common which covers $n=377$ MSAs from the contiguous United States over the period 1975Q1-2014Q4 ($T=160$ quarters). To allow for possible regional factors, we also follow bailey2016two and start with the Bureau of Economic Analysis eight regional classification, namely New England, Mideast, Great Lakes, Plains, Southeast, Southwest, Rocky Mountain and Far West. But due to the low number of MSAs in New England and Rocky Mountain regions, we combine New England and Mideast, and Southwest and Rocky Mountain as two regions. We end up with a six region classification ($R=6$), each covering a reasonable number of MSAs.

We model house price changes and consider an extended factor model with deterministic seasonal dummies to allow for seasonal movements in house prices. bailey2016two find evidence of regional factors in U.S. house price changes which might not be picked up when using PCA. Given this finding, our model includes observed regional and national factors, as well as latent factors. Specifically, we suppose

equation[equation omitted — 224 chars of source]

where $\pi_{irt}$ is the real house price change in MSA $i$ located in region $r=1,2,\ldots,R$, l$\left\{ q_{t}=j\right\} $ is the index for quarter $j$, and $\mathbf{f}_{t}$ is the $m_{0}\times1$ vector of latent factors. $\bar{\pi }_{rt}=n_{r}^{-1}\sum_{i=1}^{n_{r}}\pi_{irt}$, where $n=\sum_{r=1}^Rn_r$, and $n_r$ is the number of MSAs in region $r$, and $\bar{\pi}_{t}=n^{-1} \sum_{r=1}^{R}\sum_{i=1}^{n_{r}}\pi_{irt}$ are proxies for the regional and national factors. To filter out the effects of seasonal dummies as well as observed factors, we first run the least squares regression of $\pi_{irt}$ on an intercept and $\left( \text{l}\left\{ q_{t}=j\right\} ,\bar{\pi} _{rt},\bar{\pi}_{t}\right) $ for each $i$ to generate the residuals

equation[equation omitted — 201 chars of source]

and then apply PCA to $\left\{ \hat{v}_{irt}:i=1,2,\ldots,n_r,r=1,2,\ldots ,R,t=1,2,\ldots,T\right\} $ to obtain $\boldsymbol{\hat{\gamma}}_{ir}$ and $\mathbf{\hat{f}}_{t}$, yielding the residuals

equation[equation omitted — 265 chars of source]

For the case without adjusting for error serial correlation, the above residuals are used to compute $CD$, $CD^{\ast}$ and $CD_{W+}$, given by ((ref)), ((ref)) and ((ref)). For the case with serially correlated errors, the variance adjusted versions of the three CD statistics are generated by scaling original test statistics using ((ref)), where $\tilde{\varepsilon}_{it,T}$ is replaced by the standardized residuals generated from ((ref)), while the ARDL adjusted versions are computed using the residuals from the following dynamic panel data model with latent factors, $\mathbf{h}_{t}$:

align[align omitted — 383 chars of source]

To estimate the number of latent factors, $m$, we consider the information criteria $IC_{P1}$ and $IC_{P2}$ proposed by bai2002determining, and the $ER$ and $GR$ criteria proposed by ahn2013eigenvalue. Given the spatial diversity of U.S. housing market, we set $m_{\max}=$ $10$, although once we allow for national and regional factors we would expect $m_{0}$ and its estimate, $\hat{m}$, to be relatively small. The estimated number of factors and the associated CD test statistics are summarized in Table (ref). The first four columns of the table report $\hat{m}$, $CD$, $CD^{\ast}$ and $CD_{W+}$ statistics that are not adjusted for error serial correlation, whilst the middle and the final four columns report the variance adjusted and ARDL adjusted versions of these statistics, respectively.

As can be seen in the case of no error serial correlations, there are large differences in the number of factors selected by the different criteria, with $IC_{p1}$ selecting the assumed maximum number of factors, $IC_{p2}$ selecting $4$, and $ER$ and $GR$ both selecting $2$ factors. These estimates are not affected when we allow for error serial correlations and consider variance adjustment.

CD, CD$^{\ast}$ and CD$_{W+}$ tests all reject the null hypothesis of cross-sectional independence, irrespective of the choices of $\hat{m}$ and whether we allow for error serial correlation. In view of the theoretical and finite sample results reported in this paper, it is advisable to focus on the CD$^{\text{*}}$ test results and recognize that the relatively large magnitudes obtained for CD and CD$_{W+}$ test statistics could be due to their tendencies to over-rejection in the presence of strong latent factors and non-Gaussian errors. Focusing on $CD^{\ast}$, we find that even with $\hat {m}=10$ the CD$^{\text{*}}$ test strongly rejects the null of cross-sectional independence with $CD^{\ast}$ statistic of $25.5$, compared to 95 per cent critical value of $1.96$. There is clear evidence that in addition to latent factors, spatial modeling of the type carried out in bailey2016two and aquaro2021estimation is likely to be necessary to account for the remaining error cross-sectional dependence.

table[table omitted — 2,070 chars of source]

Concluding remarks

This paper revisits the problem of testing error cross-sectional independence in panel data models with latent factors. Starting with a pure latent multi-factor model we show that the standard CD test proposed by pesaran2004general remains valid if the latent factors are weak, but over-reject when one or more of the latent factors are strong. The over-rejection of the CD test in the case of strong factors is also established by juodis2022incidental, who propose a randomized test statistic to correct for over-rejection and add a screening component to achieve power. However, as we show, JR's CD$_{W^{+}}$ test is not guaranteed to have the correct size and need not be powerful against spatial or network alternatives. Such alternatives are of particular interest in the analyses of ripple effects in housing markets, and clustering of firms within industries in capital or arbitrage asset pricing models. In fact, using Monte Carlo experiments we show that under non-Gaussian errors the JR test continues to over-reject when the cross section dimension ($n$) is larger than the time dimension ($T$), and often has power close to size against spatial alternatives. To overcome some of these shortcomings, we propose a simple bias-corrected CD test statistic, labeled $CD^{\ast}$, which is shown to be asymptotically $\mathcal{N}(0,1)$ under the null when $n$ and $T\rightarrow \infty$ such that $n/T\rightarrow\kappa$, for a fixed constant $\kappa$. In addition, the CD$^{\text{*}}$ test is shown to have power against network type dependence. These results hold for pure latent factor models as well as for panel regression models with latent factors. To deal with possible error serial dependence, following baltagi2016testing, we also consider a variance adjusted version of $CD^{\ast}$, as well as an alternative ARDL adjusted version that eliminates the error serial dependence before the application of the CD$^{\text{*}}$ test procedure. Both of these approaches are shown to perform well within the Monte Carlo set up of the paper.

\setcounter{equation}{0} \setcounter{section}{0}

center[center omitted — 40 chars of source]

In this appendix we provide proofs of the propositions and and theorems. The auxiliary lemmas and the associated proofs are given in the supplement.

Proof of Proposition (ref)

Here we provide a proof for part (b) of Proposition (ref). The proof for part (a) follows trivially by setting $\lambda_{T}=0$. To this end we first note that the $CD$ statistic given by ((ref)) can be written as (for a proof see Lemma (ref) of the supplement) \[ CD=\left( \sqrt{\frac{n}{n-1}}\right) \frac{1}{\sqrt{2T}}\sum_{t=1} ^{T}\left[ \left( \frac{1}{\sqrt{n}}\sum_{i=1}^{n}\frac{\hat{u}_{it}} {\hat{\sigma}_{i,T}}\right) ^{2}-1\right] , \] where $\hat{u}_{it}$ is defined by ((ref)). Using Lemma (ref) of the supplement we also note that

equation[equation omitted — 57 chars of source]

where

equation[equation omitted — 242 chars of source]

with $\omega_{i,T}=\left( T^{-1}\sigma_{i}^{2}\boldsymbol{\varepsilon }_{i\circ}^{\prime}\mathbf{M}_{F}\boldsymbol{\varepsilon}_{i\circ}\right) ^{1/2}$, $\boldsymbol{\varepsilon}_{i\circ}=\left( \varepsilon_{i1} ,\varepsilon_{i2},\ldots,\varepsilon_{iT}\right) ^{\prime}$ and $\mathbf{M}_{F}=\mathbf{I}_{T}-\mathbf{F}\left( \mathbf{F}^{\prime} \mathbf{F}\right) ^{-1}\mathbf{F}^{\prime}$. Also letting

equation[equation omitted — 163 chars of source]

where $\lambda_{T}=c_{\lambda}T^{-1/2}$ with $c_{\lambda}\neq0$, $\mathbf{w}_{i0}=\left( w_{i1},w_{i2},\ldots,w_{in}\right) ^{\prime}$ and $\boldsymbol{\varepsilon}_{\circ t}=\left( \varepsilon_{1t},\varepsilon _{2t},\ldots,\varepsilon_{nt}\right) ^{\prime}$, under ((ref)) $\hat {u}_{it}$ can now be expressed as

equation[equation omitted — 410 chars of source]

Let $\boldsymbol{\delta}_{i,T}=\boldsymbol{\gamma}_{i}/\omega_{i,T}$, and $\hat{\boldsymbol{\delta}}_{i,T}=\hat{\boldsymbol{\gamma}}_{i}/\omega_{i,T}$. Then

equation[equation omitted — 445 chars of source]

Also, subject to the normalization $n^{-1}\sum_{j=1}^{n}\boldsymbol{\hat {\gamma}}_{j}\boldsymbol{\hat{\gamma}}_{j}^{\prime}=\mathbf{I}_{m_{0}}$ and $n^{-1}\sum_{j=1}^{n}\boldsymbol{\gamma}_{j}\boldsymbol{\gamma}_{j}^{\prime }=\mathbf{I}_{m_{0}}$ we have \[ \mathbf{\hat{f}}_{t}=n^{-1}\sum_{j=1}^{n}\boldsymbol{\hat{\gamma}}_{j} y_{jt}=\left( n^{-1}\sum_{j=1}^{n}\boldsymbol{\hat{\gamma}}_{j} \boldsymbol{\gamma}_{j}^{\prime}\right) \mathbf{f}_{t}+n^{-1}\sum_{j=1} ^{n}\boldsymbol{\hat{\gamma}}_{j}\sigma_{j}\varepsilon_{jt}\left( \lambda _{T}\right) , \] and hence \[ \mathbf{\hat{f}}_{t}-\mathbf{f}_{t}=\left[ n^{-1}\sum_{j=1}^{n}\left( \boldsymbol{\hat{\gamma}}_{j}-\boldsymbol{\gamma}_{j}\right) \boldsymbol{\gamma}_{j}^{\prime}\right] \mathbf{f}_{t}+n^{-1}\sum_{j=1} ^{n}\left( \boldsymbol{\hat{\gamma}}_{j}-\boldsymbol{\gamma}_{j}\right) \sigma_{j}\varepsilon_{jt}\left( \lambda_{T}\right) +n^{-1}\sum_{j=1} ^{n}\boldsymbol{\gamma}_{j}\sigma_{j}\varepsilon_{jt}\left( \lambda _{T}\right) . \] Using this result in ((ref)) we obtain

align[align omitted — 919 chars of source]

and summing over $i$ yields

align*[align* omitted — 1,023 chars of source]

where $\boldsymbol{\varphi}_{nT}=n^{-1}\sum_{i=1}^{n}\boldsymbol{\delta} _{i,T}$. Written more compactly

equation[equation omitted — 191 chars of source]

where

align[align omitted — 818 chars of source]

Further, let

equation[equation omitted — 242 chars of source]

where $\boldsymbol{\varphi}_{n}=n^{-1}\sum_{i=1}^{n}\boldsymbol{\delta}_{i},$ and $\boldsymbol{\delta}_{i}=\boldsymbol{\gamma}_{i}/\sigma_{i}$. Then $\psi_{t,nT}\left( \lambda_{T}\right) ,$ given by ((ref)), can be written as

align*[align* omitted — 1,266 chars of source]

where

equation[equation omitted — 263 chars of source]

Writing $\psi_{t,nT}\left( \lambda_{T}\right) $ more compactly we have

equation[equation omitted — 284 chars of source]

where

align[align omitted — 568 chars of source]

Using ((ref)) in ((ref)) and after some algebra we have (where we have made the dependence of $\widetilde{CD}$ on $\lambda_{T}$ explicit) \[ \widetilde{CD}(\lambda_{T})=\left( \sqrt{\frac{n}{n-1}}\right) \frac {1}{\sqrt{T}}\sum_{t=1}^{T}\left( \frac{\psi_{t,nT}^{2}\left( \lambda _{T}\right) -1}{\sqrt{2}}\right) +\left( \sqrt{\frac{n}{n-1}}\right) \left( p_{nT}\left( \lambda_{T}\right) -q_{nT}\left( \lambda_{T}\right) \right) , \] where $\psi_{t,nT}\left( \lambda_{T}\right) $ is defined by ((ref)),

equation[equation omitted — 125 chars of source]

and

equation[equation omitted — 160 chars of source]

By Lemma (ref) of the supplement $p_{nT}\left( \lambda _{T}\right) =o_{p}(1)$, and $q_{nT}\left( \lambda_{T}\right) =o_{p}(1)$. Hence

equation[equation omitted — 202 chars of source]

Now consider $T^{-1/2}\sum_{t=1}^{T}\psi_{t,nT}^{2}\left( \lambda_{T}\right) $ and using ((ref)) note that

align[align omitted — 1,246 chars of source]

By Lemma (ref) of the supplement, it follows that

equation[equation omitted — 388 chars of source]

Consider now the bias-corrected version of $\widetilde{CD}\left( \lambda _{T}\right) $ defined by

equation[equation omitted — 172 chars of source]

where $\theta_{n}=1-\frac{1}{n}\sum_{i=1}^{n}a_{i,n}^{2},$ and $a_{i,n} =1-\sigma_{i}\boldsymbol{\varphi}_{n}^{\prime}\boldsymbol{\gamma}_{i}$. Using ((ref)) in ((ref)), we have

align*[align* omitted — 461 chars of source]

Now using ((ref)) in the above and after some re-arrangement of the terms we obtain \[ \widetilde{CD}^{\ast}\left( \lambda_{T}\right) =\frac{\left[ \frac{1} {\sqrt{T}}\sum_{t=1}^{T}\left( \frac{\xi_{t,n}^{2}\left( \lambda_{T}\right) -\left( 1-\theta_{n}\right) }{\sqrt{2}}\right) \right] \left( 1+\frac {2}{\sqrt{T}}w_{nT}\left( \lambda_{T}\right) \right) }{1-\theta_{n}} +\sqrt{2}w_{nT}\left( \lambda_{T}\right) +o_{p}\left( 1\right) , \] where \[ w_{nT}\left( \lambda_{T}\right) =\frac{T^{-1/2}\sum_{t=1}^{T}\xi _{t,n}\left( \lambda_{T}\right) \upsilon_{t,nT}\left( \lambda_{T}\right) }{T^{-1}\sum_{t=1}^{T}\xi_{t,n}^{2}\left( \lambda_{T}\right) }. \] By Lemma (ref) of the supplement, $w_{nT}\left( \lambda _{T}\right) =o_{p}\left( 1\right) $. Hence\ \[ \widetilde{CD}^{\ast}\left( \lambda_{T}\right) =\frac{\frac{1}{\sqrt{T}} \sum_{t=1}^{T}\left( \frac{\xi_{t,n}^{2}\left( \lambda_{T}\right) -\left( 1-\theta_{n}\right) }{\sqrt{2}}\right) }{1-\theta_{n}}+o_{p}\left( 1\right) . \] In particular, since $\varepsilon_{it}(\lambda_{T})=\varepsilon_{it} +\lambda_{T}\mathbf{w}_{i0}^{\prime}\boldsymbol{\varepsilon}_{\circ t}$, then

align*[align* omitted — 322 chars of source]

Using this result, $\widetilde{CD}^{\ast}(\lambda_{T})$ can be written as

align[align omitted — 912 chars of source]

The second term of ((ref)) can be written as

align*[align* omitted — 270 chars of source]

where $\phi_{n}=\frac{\sqrt{2}c_{\lambda}}{\left( 1-\theta_{n}\right) }T^{-1}\sum_{t=1}^{T}E\left( A_{nt}B_{nt}\right) $. Further

align*[align* omitted — 500 chars of source]

with $\mathbf{a}_{n}=(a_{1,n},a_{2,n},...,a_{n,n})^{\prime}$. Hence

equation[equation omitted — 267 chars of source]

Under part (a) of Assumption (ref), $\varepsilon_{it}\sim IID\left( 0,1\right) $ for all $i$ and $t$, with $E\left( \varepsilon _{it}^{8}\right) <C$, it then follows that $A_{nt}B_{nt}-E\left( A_{nt}B_{nt}\right) $ will be serially independent with zero means and finite variances and by weak law of large numbers $T^{-1}\sum_{t=1}^{T}\left[ A_{nt}B_{nt}-E\left( A_{nt}B_{nt}\right) \right] =o_{p}(1)$. Therefore

equation[equation omitted — 60 chars of source]

Similarly, for the third term of ((ref)) we first note that

align*[align* omitted — 543 chars of source]

and hence

equation[equation omitted — 273 chars of source]

Using ((ref)) and ((ref)) in ((ref)) now yields

equation[equation omitted — 152 chars of source]

Consider the first term of ((ref)), $\widetilde{CD}^{\ast}\left( 0\right) $, and note that $A_{nt}$ can be written as $\,A_{nt}=n^{-1/2} \mathbf{a}_{n}^{\prime}\boldsymbol{\varepsilon}_{\circ t}$, where $\boldsymbol{a}_{n}^{\prime}=(a_{1,n},a_{2,n},...,a_{n,n})$, and we have (recall that $a_{i}=1-\sigma_{i}\boldsymbol{\varphi}_{n}^{\prime }\boldsymbol{\gamma}_{i}$) \[ E\left( A_{nt}^{2}\right) =n^{-1}E\left( \boldsymbol{\varepsilon}_{\circ t}^{\prime}\mathbf{a}_{n}\mathbf{a}_{n}^{\prime}\boldsymbol{\varepsilon }_{\circ t}\right) =n^{-1}\mathbf{a}_{n}^{\prime}\mathbf{a}_{n}=n^{-1} \sum_{i=1}^{n}a_{i,n}^{2}=\frac{1}{n}\sum_{i=1}^{n}\left( 1-\sigma _{i}\boldsymbol{\varphi}_{n}^{\prime}\boldsymbol{\gamma}_{i}\right) ^{2}=1-\theta_{n}>0, \] and (using result (S.7) of Lemma 6 in pesaran2024testing)

align*[align* omitted — 474 chars of source]

where $\kappa_{2}=E\left( \varepsilon_{it}^{4}\right) -3$ and $\mathbf{A} =\mathbf{a}_{n}\mathbf{a}_{n}^{\prime}$. Hence \[ Var\left( A_{nt}^{2}\right) =2\left( \frac{1}{n}\sum_{i=1}^{n}a_{i,n} ^{2}\right) ^{2}-\frac{\kappa_{2}}{n}\left( \frac{1}{n}\sum_{i=1}^{n} a_{i,n}^{4}\right) . \] Furthermore, since $\sup_{i}\left\vert a_{i,n}\right\vert =\sup_{i}\left\vert 1-\sigma_{i}\boldsymbol{\varphi}_{n}^{\prime}\boldsymbol{\gamma} _{i}\right\vert <1+\left( \sup_{i}\sigma_{i}\right) \left( \sup _{i}\left\Vert \boldsymbol{\gamma}_{i}\right\Vert \right) \left\Vert \boldsymbol{\varphi}_{n}\right\Vert <C$, then $n^{-2}\sum_{i=1}^{n}a_{i,n} ^{4}=O(n^{-1})$, and $Var\left( A_{nt}^{2}\right) =2\left( 1-\theta _{n}\right) ^{2}+O(n^{-1})$. Using the above results it now readily follows that,

equation[equation omitted — 327 chars of source]

Since under part (a) of Assumption (ref), $\varepsilon_{it}\sim IID\left( 0,1\right) $ for all $i$ and $t$, then $A_{nt}=\frac{1}{\sqrt{n} }\sum_{i=1}^{n}a_{i,n}\varepsilon_{it}$, and $A_{nt}^{2}$ are also independently distributed over $t$ with finite second order moments. Then by Lindeberg-L\'{e}vy central limit theorem it follows that $\widetilde{CD} ^{\ast}(0)\rightarrow_{d}\mathcal{N}(0,1),$ as $n$ and $T\rightarrow\infty$. Using this result in ((ref)) we further have $\widetilde{CD}^{\ast }\left( \lambda_{T}\right) \rightarrow_{d}\mathcal{N}\left( \phi,1\right) $, where $\phi=\lim_{n\rightarrow\infty}\phi_{n}$. By Lemma (ref) of the supplement we have $CD=\widetilde{CD}+o_{p}(1)$, then it follows

align*[align* omitted — 235 chars of source]

where the final line holds by ((ref)). Now result ((ref)) is established as required.

Proof of Proposition (ref)

Note that $\theta_{n}$ define by ((ref)) can be written as $\theta _{n}=2g_{n}-\boldsymbol{\varphi}_{n}^{\prime}\mathbf{H}_{n}\boldsymbol{\varphi }_{n}$, where $g_{n}=n^{-1}\sum_{i=1}^{n}\sigma_{i}\boldsymbol{\varphi} _{n}^{\prime}\boldsymbol{\gamma}_{i}$, $\mathbf{H}_{n}=n^{-1}\sum_{i=1} ^{n}\sigma_{i}\left( \boldsymbol{\gamma}_{i}\boldsymbol{\gamma}_{i}^{\prime }\right) $, $\boldsymbol{\varphi}_{n}=n^{-1}\sum_{i=1}^{n}\boldsymbol{\delta }_{i},$ and $\boldsymbol{\delta}_{i}=\boldsymbol{\gamma}_{i}/\sigma_{i}$. Similarly using ((ref)) we have $\hat{\theta}_{nT}=2\hat{g} _{nT}-\boldsymbol{\hat{\varphi}}_{nT}^{\prime}\mathbf{\hat{H}}_{nT} \boldsymbol{\hat{\varphi}}_{nT}$, where $\hat{g}_{nT}=n^{-1}\sum_{i=1}^{n} \hat{\sigma}_{i,T}\boldsymbol{\hat{\varphi}}_{nT}^{\prime}\boldsymbol{\hat {\gamma}}_{i},$ $\mathbf{\hat{H}}_{nT}=n^{-1}\sum_{i=1}^{n}\hat{\sigma} _{i,T}^{2}\left( \boldsymbol{\hat{\gamma}}_{i}\boldsymbol{\hat{\gamma}} _{i}^{\prime}\right) ,$ $\boldsymbol{\hat{\varphi}}_{nT}=n^{-1}\sum_{i=1} ^{n}\boldsymbol{\hat{\delta}}_{i,nT},$ and\ $\boldsymbol{\hat{\delta}} _{i,nT}=\boldsymbol{\hat{\gamma}}_{i}/\hat{\sigma}_{i,T}$. Then

equation[equation omitted — 329 chars of source]

Consider the first term of the above

align[align omitted — 466 chars of source]

and since $\sigma_{i}$ and $\boldsymbol{\gamma}_{i}$ are bounded then $n^{-1}\sum_{i=1}^{n}\sigma_{i}\boldsymbol{\gamma}_{i}=O(1)$. Also by ((ref)) of Lemma (ref) in the supplement we have $\sqrt{T}\left( \boldsymbol{\hat{\varphi}}_{nT}-\boldsymbol{\varphi} _{n}\right) =o_{p}(1)$, and hence the first term of the above is $o_{p}(1)$. To establish the probability order of the second term of ((ref)), we first note that

align[align omitted — 737 chars of source]

But by ((ref)) and ((ref)) of the supplement, $\boldsymbol{\varphi}_{n}=O_{p}(1)$ and $n^{-1}\sum_{i=1}^{n}\left( \hat{\sigma}_{i,T}\boldsymbol{\hat{\gamma}}_{i}-\sigma_{i}\boldsymbol{\gamma }_{i}\right) =O_{p}\left( \ln\left( n\right) /T\right) ,$ which also establishes that the second term of ((ref)) is $o_{p}(1)$. Therefore overall we have

equation[equation omitted — 89 chars of source]

Consider now the second term of ((ref)) and note that

align[align omitted — 688 chars of source]

where $\mathbf{\hat{H}}_{nT}=\frac{1}{n}\sum_{i=1}^{n}\hat{\sigma}_{i,T} ^{2}\left( \boldsymbol{\hat{\gamma}}_{i}\boldsymbol{\hat{\gamma}}_{i} ^{\prime}\right) $, and

align[align omitted — 865 chars of source]

The first two terms of ((ref)) are $o_{p}(1)$, since $\left\Vert \boldsymbol{\varphi}_{n}\right\Vert <C$, $\sqrt{T}\left( \boldsymbol{\hat {\varphi}}_{nT}-\boldsymbol{\varphi}_{n}\right) =o_{p}(1)$, and $n^{-1} \sum_{i=1}^{n}\hat{\sigma}_{i,T}^{2}\left( \boldsymbol{\hat{\gamma}} _{i}\boldsymbol{\hat{\gamma}}_{i}^{\prime}\right) =O_{p}(1)$. To establish the probability order of the third term of ((ref)), since $\left\Vert \boldsymbol{\varphi}_{n}\right\Vert <C$ it is sufficient to consider the four terms of $\sqrt{T}\left( \mathbf{\hat{H}}_{nT} -\mathbf{H}_{n}\right) $. It is clear that $\mathbf{D}_{2,nT}$ is dominated by $\mathbf{D}_{1,nT\ }$ and by ((ref)) of Lemma (ref) of the supplement, \[ \mathbf{D}_{1,nT\ }=\frac{\sqrt{T}}{n}\sum_{i=1}^{n}\sigma_{i}^{2}\left( \boldsymbol{\hat{\gamma}}_{i}\boldsymbol{\hat{\gamma}}_{i}^{\prime }-\boldsymbol{\gamma}_{i}\boldsymbol{\gamma}_{i}^{\prime}\right) =O_{p}\left( \sqrt{\frac{\ln\left( n\right) }{n}}\right) =o_{p}(1). \] Using ((ref)) of Lemma (ref) of the supplement and replacing $b_{ni}$ with $\gamma_{ij}\gamma_{ij^{^{\prime}}}$ for $j,j^{\prime }=1,2,\ldots,m_{0}$, it then follows that \[ \mathbf{D}_{3,nT}=\frac{\sqrt{T}}{n}\sum_{i=1}^{n}\left( \hat{\sigma} _{i,T}^{2}-\omega_{i,T}^{2}\right) \boldsymbol{\gamma}_{i}\boldsymbol{\gamma }_{i}^{\prime}=O_{p}\left( \frac{\ln\left( n\right) }{\sqrt{T}}\right) =o_{p}(1). \] Finally, denote the $\left( j,j^{\prime}\right) $ element of $\mathbf{D} _{4,nT}$ by $d_{4,nT}(j,j^{\prime})$ and note that \[ d_{4,nT}(j,j^{\prime})=\frac{1}{n}\sum_{i=1}^{n}\left( \sigma_{i}^{2} \gamma_{ij}\gamma_{ij^{\prime}}\right) \sqrt{T}\left( \frac {\boldsymbol{\varepsilon}_{i\circ}^{^{\prime}}\mathbf{M}_{F} \boldsymbol{\varepsilon}_{i\circ}}{T}-1\right) \text{, for }j,j^{\prime }=1,2,...,m_{0}. \] But under Assumptions (ref) and (ref), $\left\vert \sigma_{i}^{2}\gamma_{ij}\gamma_{ij^{\prime}}\right\vert <C$, and $\sqrt {T}\left( T^{-1}\boldsymbol{\varepsilon}_{i\circ}^{^{\prime}}\mathbf{M} _{F}\boldsymbol{\varepsilon}_{i\circ}-1\right) $, for $i=1,2,...,n$ are identically and independently distributed across $i$, with mean $1/\sqrt{T}$ and a finite variance\footnote{The mean and variance of $\sqrt{T}\left( T^{-1}\boldsymbol{\varepsilon}_{i\circ}^{^{\prime}}\mathbf{M}_{F} \boldsymbol{\varepsilon}_{i\circ}-1\right) $ can be obtained using ((ref)) and ((ref)) in Lemma (ref) of the supplement.}. Then by standard law of large numbers, for each $(j,j^{\prime})$, $d_{4,nT}(j,j^{\prime})\rightarrow_{p}0$, as $n$ and $T\rightarrow\infty$, and hence we also have $\mathbf{D}_{4,nT}=o_{p}(1)$. Overall, $\mathbf{\hat{H} }_{nT}-\mathbf{H}_{n}=o_{p}(1)$, and we have $\sqrt{T}\left( \boldsymbol{\hat {\varphi}}_{nT}^{\prime}\mathbf{\hat{H}}_{nT}\boldsymbol{\hat{\varphi}} _{nT}-\boldsymbol{\varphi}_{n}^{\prime}\mathbf{H}_{n}\boldsymbol{\varphi} _{n}\right) =o_{p}(1)$. Using this result and ((ref)) in ((ref)) now yields $\sqrt{T}\left( \hat{\theta}_{nT}-\theta _{n}\right) =o_{p}(1)$, as required.

Proof of Theorem (ref)

Recall from ((ref)) that $CD^{\ast}$ is given by \[ CD^{\ast}=\frac{CD+\sqrt{\frac{T}{2}}\hat{\theta}_{nT}}{1-\hat{\theta}_{nT}}, \] where $\hat{\theta}_{nT}=1-\frac{1}{n}\sum_{i=1}^{n}\hat{a}_{i,n}^{2}$, $\hat{a}_{i,n}=1-\hat{\sigma}_{i,T}\left( \boldsymbol{\varphi}_{nT}^{\prime }\boldsymbol{\hat{\gamma}}_{i}\right) ,$ and $\boldsymbol{\hat{\varphi}} _{nT}=n^{-1}\sum_{i=1}^{n}\boldsymbol{\hat{\gamma}}_{i}/\hat{\sigma}_{i,T}$, subject to the normalization $n^{-1}\sum_{i=1}^{n}\boldsymbol{\hat{\gamma} }_{i}\boldsymbol{\hat{\gamma}}_{i}^{^{\prime}}=\mathbf{I}_{m_{0}}$. By result ((ref)) of Proposition (ref), $\sqrt{T}\left( \hat{\theta}_{nT}-\theta_{n}\right) =o_{p}(1)$, and hence

align*[align* omitted — 329 chars of source]

Theorem (ref) is then established by following Proposition (ref).

Proof of Theorem (ref)

Let $v_{it}=y_{it}-\boldsymbol{\alpha}_{i}^{\prime}\mathbf{d}_{t} -\boldsymbol{\beta}_{i}^{^{\prime}}\mathbf{x}_{it}$, and $u_{it} =y_{it}-\boldsymbol{\alpha}_{i}^{\prime}\mathbf{d}_{t}-\boldsymbol{\beta} _{i}^{^{\prime}}\mathbf{x}_{it}-\boldsymbol{\gamma}_{i}^{^{\prime}} \mathbf{f}_{t}=v_{it}-\boldsymbol{\gamma}_{i}^{^{\prime}}\mathbf{f}_{t}$, and consider the following two optimization problems

align[align omitted — 350 chars of source]

where

align[align omitted — 710 chars of source]

We need to show that solving problem ((ref)) is asymptotically equivalent to solving problem ((ref)). First, using the results in pesaran2011large and the fact that $\mathbf{d}_{t}$\ and $\mathbf{x} _{it}$ are (stochastically) bounded\footnote{See equation (31) in pesaran2011large.},

align[align omitted — 491 chars of source]

then rewrite the criterion for ((ref)) with ((ref)),

align[align omitted — 1,830 chars of source]

Therefore, using ((ref)) and ((ref)), then $A_{2,nT}=O_{p}\left( \frac{1}{\sqrt{T}}\right) +O_{p}\left( \frac{1} {n}\right) +O_{p}\left( \frac{1}{\sqrt{nT}}\right) $ and $A_{3,nT} =O_{p}\left( \frac{1}{\sqrt{T}}\right) +O_{p}\left( \frac{1}{n}\right) +O_{p}\left( \frac{1}{\sqrt{nT}}\right) $. Also, consider the fourth term of ((ref)) and note that by Cauchy-Schwarz inequality,

align*[align* omitted — 490 chars of source]

where $u_{it}=v_{it}-\boldsymbol{\gamma}_{i}^{^{\prime}}\mathbf{f}_{t} =y_{it}-\boldsymbol{\alpha}_{i}^{\prime}\mathbf{d}_{t}-\boldsymbol{\beta} _{i}^{^{\prime}}\mathbf{x}_{it}-\boldsymbol{\gamma}_{i}^{^{\prime}} \mathbf{f}_{t}$, so given ((ref)) we have $A_{4,nT}=O_{p}\left( \frac{1}{\sqrt{T}}\right) +O_{p}\left( \frac{1}{n}\right) +O_{p}\left( \frac{1}{\sqrt{nT}}\right) $. Similarly, we can show $A_{5,nT}$ and $A_{6,nT}$ share the same probability order as $A_{4,nT}$. Since in both optimization problems $\boldsymbol{\gamma}_{i}$ and $\mathbf{f}_{t}$ are only identified up to $m_{0}\times m_{0}$ rotation matrices, it follows that

align*[align* omitted — 444 chars of source]

Hence, PCs based on $\hat{v}_{it}$ are asymptotically equivalent to those based on $v_{it}$. The remaining proof of Theorem (ref) follows from the proof of Theorem (ref).

{

thebibliography\bibitem[Ahn and Horenstein, 2013]{ahn2013eigenvalue} Ahn, S. C. and Horenstein, A. R. (2013). \newblock Eigenvalue ratio test for the number of factors. \newblock {\em Econometrica}, 81(3):1203--1227. \bibitem[Aquaro et al., 2021]{aquaro2021estimation} Aquaro, M., Bailey, N., and Pesaran, M. H. (2021). \newblock Estimation and inference for spatial models with heterogeneous coefficients: an application to {U.S.} house prices. \newblock {\em Journal of Applied Econometrics}, 36(1):18--44. \bibitem[Bai, 2003]{bai2003inferential} Bai, J. (2003). \newblock Inferential theory for factor models of large dimensions. \newblock {\em Econometrica}, 71(1):135--171. \bibitem[Bai, 2009]{bai2009panel} Bai, J. (2009). \newblock Panel data models with interactive fixed effects. \newblock {\em Econometrica}, 77(4):1229--1279. \bibitem[Bai and Li, 2021]{bai2021dynamic} Bai, J. and Li, K. (2021). \newblock Dynamic spatial panel data models with common shocks. \newblock {\em Journal of Econometrics}, 224(1):134--160. \bibitem[Bai and Ng, 2002]{bai2002determining} Bai, J. and Ng, S. (2002). \newblock Determining the number of factors in approximate factor models. \newblock {\em Econometrica}, 70(1):191--221. \bibitem[Bai and Ng, 2006]{bai2006confidence} Bai, J. and Ng, S. (2006). \newblock Confidence intervals for diffusion index forecasts and inference for factor-augmented regressions. \newblock {\em Econometrica}, 74(4):1133--1150. \bibitem[Bai and Ng, 2023]{bai2023approximate} Bai, J. and Ng, S. (2023). \newblock Approximate factor models with weaker loadings. \newblock {\em Journal of Econometrics}, 235(2):1893--1916. \bibitem[Bailey et al., 2016]{bailey2016two} Bailey, N., Holly, S., and Pesaran, M. H. (2016). \newblock A two-stage approach to spatio-temporal analysis with strong and weak cross-sectional dependence. \newblock {\em Journal of Applied Econometrics}, 31(1):249--280. \bibitem[Bailey et al., 2021]{bailey2021measurement} Bailey, N., Kapetanios, G., and Pesaran, M. H. (2021). \newblock Measurement of factor strength: Theory and practice. \newblock {\em Journal of Applied Econometrics}, 36(2):431--453. \bibitem[Baltagi et al., 2016]{baltagi2016testing} Baltagi, B. H., Kao, C., and Peng, B. (2016). \newblock Testing cross-sectional correlation in large panel data models with serial correlation. \newblock {\em Econometrics}, 4(4):44. \bibitem[Chamberlain and Rothschild, 1983]{chamberlain1983arbitrage} Chamberlain, G. and Rothschild, M. (1983). \newblock Arbitrage, factor structure, and mean-variance analysis on large asset markets. \newblock {\em Econometrica}, 51(1):1281--1304. \bibitem[Chiang and Tsai, 2016]{chiang2016ripple} Chiang, M.-C. and Tsai, I.-C. (2016). \newblock Ripple effect and contagious effect in the {U.S.} regional housing markets. \newblock {\em The Annals of Regional Science}, 56(1):55--82. \bibitem[Chudik et al., 2018]{chudik2018one} Chudik, A., Kapetanios, G., and Pesaran, M. H. (2018). \newblock A one covariate at a time, multiple testing approach to variable selection in high-dimensional linear regression models. \newblock {\em Econometrica}, 86(4):1479--1512. \bibitem[Chudik et al., 2011]{chudik2011weak} Chudik, A., Pesaran, M., and Tosetti, E. (2011). \newblock Weak and strong cross-section dependence and estimation of large panels. \newblock {\em The Econometrics Journal}, 14(1):C45--C90. \bibitem[Durbin and Watson, 1950]{durbin1950testing} Durbin, J. and Watson, G. S. (1950). \newblock Testing for serial correlation in least squares regression. {I}. \newblock {\em Biometrika}, 37(3-4):409--428. \bibitem[Fan et al., 2011]{fan2011high} Fan, J., Liao, Y., and Mincheva, M. (2011). \newblock High dimensional covariance matrix estimation in approximate factor models. \newblock {\em Annals of Statistics}, 39(6):3320--3356. \bibitem[Fan et al., 2013]{fan2013large} Fan, J., Liao, Y., and Mincheva, M. (2013). \newblock Large covariance estimation by thresholding principal orthogonal complements. \newblock {\em Journal of the Royal Statistical Society Series B: Statistical Methodology}, 75(4):603--680. \bibitem[Fan et al., 2015]{fan2015power} Fan, J., Liao, Y., and Yao, J. (2015). \newblock Power enhancement in high-dimensional cross-sectional tests. \newblock {\em Econometrica}, 83(4):1497--1541. \bibitem[Gagliardini et al., 2019]{gagliardini2019diagnostic} Gagliardini, P., Ossola, E., and Scaillet, O. (2019). \newblock A diagnostic criterion for approximate factor structure. \newblock {\em Journal of Econometrics}, 212(2):503--521. \bibitem[Goldie and Kl{\"u}ppelberg, 1998]{goldie1998subexponential} Goldie, C. M. and Kl{\"u}ppelberg, C. (1998). \newblock Subexponential distributions. \newblock In R. J. Adler, R. E. Feldman, and M. S. Taqqu (Eds.), {\em A Practical Guide to Heavy Tails: Statistical Techniques and Applications}, pages 435--459. Birkh{\"a}user Boston Inc., U.S. \bibitem[Holly et al., 2011]{holly2011spatial} Holly, S., Pesaran, M. H., and Yamagata, T. (2011). \newblock The spatial and temporal diffusion of house prices in the {U.K.} \newblock {\em Journal of Urban Economics}, 69(1):2--23. \bibitem[Hsiao et al., 2012]{hsiao2012diagnostic} Hsiao, C., Pesaran, M. H., and Pick, A. (2012). \newblock Diagnostic tests of cross-section independence for limited dependent variable panel data models. \newblock {\em Oxford Bulletin of Economics and Statistics}, 74(2):253--277. \bibitem[Jiang et al., 2023]{jiang2023revisiting} Jiang, P., Uematsu, Y., and Yamagata, T. (2023). \newblock Revisiting asymptotic theory for principal component estimators of approximate factor models. \newblock {\em arXiv preprint arXiv:2311.00625}. \bibitem[Juodis and Reese, 2022]{juodis2022incidental} Juodis, A. and Reese, S. (2022). \newblock The incidental parameters problem in testing for remaining cross-section correlation. \newblock {\em Journal of Business & Economic Statistics}, 40(3):1191--1203. \bibitem[Pesaran, 2004]{pesaran2004general} Pesaran, M. H. (2004). \newblock General diagnostic tests for cross-sectional dependence in panels. \newblock {\em University of Cambridge, Cambridge Working Papers in Economics}, 435. \bibitem[Pesaran, 2006]{pesaran2006estimation} Pesaran, M. H. (2006). \newblock Estimation and inference in large heterogeneous panels with a multifactor error structure. \newblock {\em Econometrica}, 74(4):967--1012. \bibitem[Pesaran, 2015]{pesaran2015testing} Pesaran, M. H. (2015). \newblock Testing weak cross-sectional dependence in large panels. \newblock {\em Econometric Reviews}, 34(6-10):1089--1117. \bibitem[Pesaran, 2021]{pesaran2021general} Pesaran, M. H. (2021). \newblock General diagnostic tests for cross-sectional dependence in panels. \newblock {\em Empirical Economics}, 60(1):13--50. \bibitem[Pesaran and Tosetti, 2011]{pesaran2011large} Pesaran, M. H. and Tosetti, E. (2011). \newblock Large panels with common factors and spatial correlation. \newblock {\em Journal of Econometrics}, 161(2):182--202. \bibitem[Pesaran and Yamagata, 2024]{pesaran2024testing} Pesaran, M. H. and Yamagata, T. (2024). \newblock Testing for alpha in linear factor pricing models with a large number of securities. \newblock {\em Journal of Financial Econometrics}, 22(2):407--460. \bibitem[Phillips and Sul, 2003]{phillips2003dynamic} Phillips, P. C. B. and Sul, D. (2003). \newblock Dynamic panel estimation and homogeneity testing under cross section dependence. \newblock {\em The Econometrics Journal}, 6(1):217--259. \bibitem[Phillips and Sul, 2007]{phillips2007bias} Phillips, P. C. B. and Sul, D. (2007). \newblock Bias in dynamic panel estimation with fixed effects, incidental trends and cross section dependence. \newblock {\em Journal of Econometrics}, 137(1):162--188. \bibitem[Shi and Lee, 2017]{shi2017spatial} Shi, W. and Lee, L.-f. (2017). \newblock Spatial dynamic panel data models with interactive fixed effects. \newblock {\em Journal of Econometrics}, 197(2):323--347. \bibitem[Tsai, 2015]{tsai2015spillover} Tsai, I.-C. (2015). \newblock Spillover effect between the regional and the national housing markets in the {U.K.} \newblock {\em Regional Studies}, 49(12):1957--1976. \bibitem[Vershynin, 2018]{vershynin2018high} Vershynin, R. (2018). \newblock {\em High-dimensional probability: An introduction with applications in data science}, volume 47. \newblock Cambridge university press. \bibitem[Yang, 2021]{yang2021common} Yang, C. F. (2021). \newblock Common factors and spatial dependence: An application to {U.S.} house prices. \newblock {\em Econometric Reviews}, 40(1):14--50.

}

center[center omitted — 302 chars of source]

\setcounter{equation}{0} \setcounter{section}{0} \setcounter{page}{1} \setcounter{theorem}{1} \setcounter{footnote}{0} \setcounter{table}{0} \setcounter{figure}{0} \setcounter{lemma}{0} \setcounter{remark}{0}

This supplement is in four sections. Section (ref) states and establishes the auxiliary lemmas used in the proofs of propositions and theorems in the paper. Section (ref) derives the order of $\theta_{n}$, defined by ((ref)) in the paper, in terms of the factor strengths. Section (ref) considers the CD$_{W+}$ test proposed by Juodis and Reese (2022), and discusses some of its properties. Section (ref) reports simulation results for the experiments discussed in Section (ref) of the main paper.

Statement and proofs of the lemmas

This section provides auxiliary lemmas and the associated proofs, which are required to establish the main results of the paper.

lemmaThe CD statistic defined by ((ref)) can be written equivalently as, \begin{equation} CD=\left( \sqrt{\frac{n}{n-1}}\right) \frac{1}{\sqrt{2T}}\sum_{t=1} ^{T}\left[ \left( \frac{1}{\sqrt{n}}\sum_{i=1}^{n}\frac{\hat{u}_{it}} {\hat{\sigma}_{i,T}}\right) ^{2}-1\right] . \end{equation}
proofUsing $\hat{\rho}_{ij,T}=\left( \frac{1}{T}\sum_{t=1}^{T}\hat{u}_{it}\hat {u}_{jt}\right) /\hat{\sigma}_{i,T}\hat{\sigma}_{j,T}$ in ((ref)) we have: \begin{equation} CD=\sqrt{\frac{2T}{n(n-1)}}\sum_{i=1}^{n-1}\sum_{j=i+1}^{n}\frac{\frac{1} {T}\sum_{t=1}^{T}\hat{u}_{it}\hat{u}_{jt}}{\hat{\sigma}_{i,T}\hat{\sigma }_{j,T}}=\sqrt{\frac{2T}{n(n-1)}}\frac{1}{T}\sum_{t=1}^{T}\left( \sum _{i=1}^{n-1}\sum_{j=i+1}^{n}\left( \frac{\hat{u}_{it}}{\hat{\sigma}_{i,T} }\right) \left( \frac{\hat{u}_{jt}}{\hat{\sigma}_{j,T}}\right) \right) . \end{equation} Further, we note that \[ \frac{1}{n}\sum_{i=1}^{n-1}\sum_{j=i+1}^{n}\left( \frac{\hat{u}_{it}} {\hat{\sigma}_{i,T}}\right) \left( \frac{\hat{u}_{jt}}{\hat{\sigma}_{j,T} }\right) =\frac{1}{2}\left[ \left( \frac{1}{\sqrt{n}}\sum_{i=1}^{n} \frac{\hat{u}_{it}}{\hat{\sigma}_{i,T}}\right) ^{2}-\frac{1}{n}\sum_{i=1} ^{n}\left( \frac{\hat{u}_{it}}{\hat{\sigma}_{i,T}}\right) ^{2}\right] . \] Then using this result in ((ref)), and after some algebra, we have \begin{align*} CD & =\sqrt{\frac{2Tn^{2}}{n(n-1)}}\frac{1}{2T}\sum_{t=1}^{T}\left[ \left( \frac{1}{\sqrt{n}}\sum_{i=1}^{n}\frac{\hat{u}_{it}}{\hat{\sigma}_{i,T} }\right) ^{2}-\frac{1}{n}\sum_{i=1}^{n}\left( \frac{\hat{u}_{it}} {\hat{\sigma}_{i,T}}\right) ^{2}\right] \\ & =\sqrt{\frac{2Tn^{2}}{n(n-1)}}\frac{1}{2}\left[ \frac{1}{T}\sum_{t=1} ^{T}\left( \frac{1}{\sqrt{n}}\sum_{i=1}^{n}\frac{\hat{u}_{it}}{\hat{\sigma }_{i,T}}\right) ^{2}-\frac{1}{n}\sum_{i=1}^{n}\frac{1}{T}\sum_{t=1} ^{T}\left( \frac{\hat{u}_{it}}{\hat{\sigma}_{i,T}}\right) ^{2}\right] \\ & =\left( \sqrt{\frac{n}{n-1}}\right) \frac{1}{\sqrt{2T}}\sum_{t=1} ^{T}\left[ \left( \frac{1}{\sqrt{n}}\sum_{i=1}^{n}\frac{\hat{u}_{it}} {\hat{\sigma}_{i,T}}\right) ^{2}-1\right] , \end{align*} as required.
lemmaConsider the latent factor model given by ((ref)) and ((ref)). The latent factors, $\mathbf{f}_{t}$, and their loadings, $\boldsymbol{\gamma}_{i}$, are estimated by principal components, $\mathbf{\hat{f}}_{t}$ and $\boldsymbol{\hat{\gamma}}_{i}$, given by ((ref)). Suppose that Assumptions (ref)-(ref) hold and $\left( n,T\right) \rightarrow\infty$, such that $n/T\rightarrow\kappa\,$ for $0<\kappa<\infty$. Then \begin{align} \left\Vert \mathbf{\hat{F}}-\mathbf{F}\right\Vert _{F} & =O_{p}\left( \frac{\sqrt{T}}{\delta_{nT}}\right) ,\\ \left\Vert \boldsymbol{\hat{\Gamma}}-\boldsymbol{\Gamma}\right\Vert _{F} & =O_{p}\left( \frac{\sqrt{n}}{\delta_{nT}}\right) ,\\ \left\Vert \mathbf{U}\left( \lambda_{T}\right) ^{^{\prime}}\mathbf{(\hat {F}-F)}\right\Vert _{F} & =O_{p}\left( \frac{\sqrt{nT}}{\delta_{nT} }\right) ,\\ \left\Vert \boldsymbol{\Gamma}^{\prime}(\boldsymbol{\hat{\Gamma} }-\boldsymbol{\Gamma})\right\Vert _{F} & =O_{p}\left( \frac{n}{\delta_{nT} }\right) ,\\ \left\Vert \mathbf{F}^{\prime}(\mathbf{\hat{F}}-\mathbf{F})\right\Vert _{F} & =O_{p}\left( \frac{T}{\delta_{nT}}\right) ,\\ \left( \mathbf{\hat{F}-F}\right) ^{^{\prime}}\mathbf{F} & =O_{p}\left( \frac{T}{\delta_{nT}^{2}}\right) ,\\ \left( \mathbf{\hat{F}-F}\right) ^{^{\prime}}\mathbf{\hat{F}} & =O_{p}\left( \frac{T}{\delta_{nT}^{2}}\right) ,\\ (\boldsymbol{\hat{\Gamma}}-\boldsymbol{\Gamma})^{\prime}\mathbf{u}_{\circ t}\left( \lambda_{T}\right) & =O_{p}\left( \frac{n}{\delta_{nT}^{2} }\right) , \end{align} where $\mathbf{F}=\left( \mathbf{f}_{1},\mathbf{f}_{2},\ldots,\mathbf{f} _{T}\right) ^{\prime}$, $\mathbf{\hat{F}}=\left( \mathbf{\hat{f}} _{1},\mathbf{\hat{f}}_{2},\ldots,\mathbf{\hat{f}}_{T}\right) ^{\prime}$, $\boldsymbol{\Gamma}=\left( \boldsymbol{\gamma}_{1},\boldsymbol{\gamma} _{2},...,\boldsymbol{\gamma}_{n}\right) ^{\prime}$, $\boldsymbol{\hat{\Gamma }}=\left( \boldsymbol{\hat{\gamma}}_{1},\boldsymbol{\hat{\gamma}} _{2},...,\boldsymbol{\hat{\gamma}}_{n}\right) ^{\prime}$, $\mathbf{U}\left( \lambda_{T}\right) =$ $(\mathbf{u}_{\circ1}\left( \lambda_{T}\right) ,\mathbf{u}_{\circ2}\left( \lambda_{T}\right) ,...,\mathbf{u}_{\circ T}\left( \lambda_{T}\right) )^{\prime}$, $\mathbf{u}_{\circ t}\left( \lambda_{T}\right) =(\sigma_{1}\varepsilon_{1t}\left( \lambda_{T}\right) ,\sigma_{2}\varepsilon_{2t}\left( \lambda_{T}\right) ,...,\sigma _{n}\varepsilon_{nt}\left( \lambda_{T}\right) )^{\prime}$, and $\varepsilon_{it}\left( \lambda_{T}\right) =\varepsilon_{it}+\lambda_{T} \sum_{j=1}^{n}w_{ij}\varepsilon_{jt}$.
proofSince Assumptions (ref)-(ref) are a sub-set of assumptions made by Bai (2003), so results ((ref)) to ((ref)), ((ref)) and ((ref)) follow directly from Lemmas B.1, B.2 and B.3, and Theorems 1 and 2 of Bai (2003). Results ((ref)) and ((ref)) can be established analogously.
lemmaSuppose that Assumptions (ref)-(ref) hold and $\left( n,T\right) \rightarrow\infty$, such that $n/T\rightarrow \kappa\,,$ for $0<\kappa<\infty$. Then \begin{align} \sup_{i}\left( T^{-1}\left\Vert \boldsymbol{\varepsilon}_{i\circ}\right\Vert ^{2}\right) & =O_{p}\left( 1\right) ,\\ \sup_{i}\left\Vert \frac{\mathbf{F}^{\prime}\boldsymbol{\varepsilon}_{i\circ} }{T}\right\Vert & =O_{p}\left( \sqrt{\frac{\ln(n)}{T}}\right) ,\\ \sup_{t}\left\Vert \frac{\boldsymbol{\Gamma}^{\prime}\boldsymbol{\varepsilon }_{\circ t}}{n}\right\Vert & =O_{p}\left( \sqrt{\frac{\ln(T)}{n}}\right) ,\\ \sup_{i}\left\Vert \frac{1}{nT}\sum_{t=1}^{T}\sum_{j=1}^{n}\sigma _{j}\boldsymbol{\gamma}_{j}\varepsilon_{it}\varepsilon_{jt}\right\Vert & =O_{p}\left( \sqrt{\frac{\ln\left( n\right) }{nT}}\right) , \end{align} where $\boldsymbol{\varepsilon}_{i\circ}=\left( \varepsilon_{i1} ,\varepsilon_{i2},\ldots,\varepsilon_{iT}\right) ^{\prime}$ and $\boldsymbol{\varepsilon}_{\circ t}=\left( \varepsilon_{1t},\varepsilon _{2t},\ldots,\varepsilon_{nt}\right) ^{\prime}$.
proofConsider ((ref)) and note \[ T^{-1}\left\Vert \boldsymbol{\varepsilon}_{i\circ}\right\Vert ^{2}=\frac{1} {T}\sum_{t=1}^{T}\left[ \varepsilon_{it}^{2}-E\left( \varepsilon_{it} ^{2}\right) \right] +\frac{1}{T}\sum_{t=1}^{T}E\left( \varepsilon_{it} ^{2}\right) =\frac{1}{T}\sum_{t=1}^{T}z_{it}+1, \] where $z_{it}=\varepsilon_{it}^{2}-E\left( \varepsilon_{it}^{2}\right) $. Then \begin{equation} \sup_{i}\left( T^{-1}\left\Vert \boldsymbol{\varepsilon}_{i\circ}\right\Vert ^{2}\right) \leq\sup_{i}\left\vert \frac{1}{T}\sum_{t=1}^{T}z_{it}\right\vert +1. \end{equation} To establish the probability of the first term, consider the filtration $\mathcal{I}_{it}^{\left( 1\right) }=\{\varepsilon_{i\tau}:\tau =t-1,t-2,\ldots\}$ and, given the serial independence of $\varepsilon_{it}$, note that $E\left( z_{it}|\mathcal{I}_{i,t-1}^{\left( 1\right) }\right) =0$, so $z_{it}$ is a martingale difference process with respect to $\mathcal{I}_{i,t-1}^{\left( 1\right) }$. In addition, $Var\left( z_{it}\right) =Var\left( \varepsilon_{it}^{2}\right) =E\left( \varepsilon_{it}^{4}\right) -\left( E\varepsilon_{it}^{2}\right) ^{2}$ which is bounded by assumption. Also, as $\varepsilon_{it}$ is sub-exponential by part (a) of Assumption (ref), $\varepsilon_{it}^{2}$ (and hence $z_{it}$) is sub-exponential, and there exist positive constants $C_{4}$, $C_{5}$ and $r_{3}$ such that \[ \sup_{i}\Pr\left( \left\vert z_{it}\right\vert >a\right) \leq C_{4} \exp\left( -C_{5}a^{r_{3}}\right) ,\text{ for all }a>0.\text{ } \] Then by Lemma A3 in the online theory supplement of Chudik et al. (2018), for $\varsigma_{T}=\ominus\left( T^{\mu}\right) $ and $0<\mu<\left( r_{3}+1\right) /\left( r_{3}+2\right) $, there exists a positive constant $C_{6}$ such that$\Pr\left( \left\vert \sum_{t=1}^{T}z_{it}\right\vert >\varsigma_{T}\right) \leq\exp\left( -C_{6}T^{-1}\varsigma_{T}^{2}\right) $, and if $\mu>\left( r_{3}+1\right) /\left( r_{3}+2\right) $ there exists a positive constant $C_{7}$ such that $\Pr\left( \left\vert \sum_{t=1} ^{T}z_{it}\right\vert >\varsigma_{T}\right) \leq\exp\left( -C_{7}\left( \varsigma_{T}\right) ^{\frac{r_{3}}{r_{3}+1}}\right) . $ By Boole's inequality, we have \[ \Pr\left( \sup_{i}\left\vert \sum_{t=1}^{T}z_{it}\right\vert >\varsigma _{T}\right) \leq\exp\left( \ln\left( n\right) -C_{6}T^{-1}\varsigma _{T}^{2}\right) ,\text{ if }0<\mu<\left( r_{3}+1\right) /\left( r_{3}+2\right) , \] \[ \Pr\left( \sup_{i}\left\vert \sum_{t=1}^{T}z_{it}\right\vert >\varsigma _{T}\right) \leq\exp\left( \ln\left( n\right) -C_{7}\left( \varsigma _{T}\right) ^{\frac{r_{3}}{r_{3}+1}}\right) ,\text{ if }\mu>\left( r_{3}+1\right) /\left( r_{3}+2\right) . \] Let $\varsigma_{T}=C_{8}\sqrt{T\ln\left( n\right) }$ where $C_{8}$ is a finite but sufficiently large constant. Then for $0<\mu<\left( r_{3} +1\right) /\left( r_{3}+2\right) $, we have \begin{align*} \Pr\left( \sup_{i}\left\vert \frac{1}{T}\sum_{t=1}^{T}z_{it}\right\vert >C_{8}\sqrt{\frac{\ln\left( n\right) }{T}}\right) & =\Pr\left( \sup _{i}\left\vert \sum_{t=1}^{T}\varepsilon_{it}-E\left( \varepsilon_{it} ^{2}\right) \right\vert >C_{8}\sqrt{T\ln\left( n\right) }\right) \\ & \leq\exp\left[ \ln\left( n\right) -C_{6}T^{-1}C_{8}^{2}T\ln\left( n\right) \right] \\ & =\exp\left[ \ln\left( n\right) -C_{6}C_{8}^{2}\ln\left( n\right) \right] , \end{align*} which is $o\left( 1\right) $ given $C_{8}$ is sufficiently large. Also for $\mu\geq\left( r_{3}+1\right) /\left( r_{3}+2\right) $, we have \begin{align*} \Pr\left( \sup_{i}\left\vert \frac{1}{T}\sum_{t=1}^{T}\varepsilon _{it}\right\vert >C_{8}\sqrt{\frac{\ln\left( n\right) }{T}}\right) & =\Pr\left( \sup_{i}\left\vert \sum_{t=1}^{T}\varepsilon_{it}\right\vert >C_{8}\sqrt{T\ln\left( n\right) }\right) \\ & \leq\exp\left[ \ln\left( n\right) -C_{7}\left( C_{8}\sqrt{T\ln\left( n\right) }\right) ^{\frac{r_{3}}{r_{3}+1}}\right] , \end{align*} which is also $o\left( 1\right) $ as$\ n$ and $T$ are of the same order of magnitude and sufficiently we have \[ \frac{\ln\left( n\right) }{\left[ \sqrt{n\ln\left( n\right) }\right] ^{\frac{r_{3}}{r_{3}+1}}}=\frac{\left[ \ln\left( n\right) \right] ^{\frac{r_{3}+2}{2\left( r_{3}+1\right) }}}{n^{\frac{r_{3}}{2\left( r_{3}+1\right) }}}=\left[ \frac{\left( \ln\left( n\right) \right) ^{\frac{r_{3}+2}{r_{3}}}}{n}\right] ^{\frac{r_{3}}{2\left( r_{3}+1\right) }}\rightarrow0. \] Therefore, $\sup_{i}\left\vert T^{-1}\sum_{t=1}^{T}z_{it}\right\vert =O_{p}\left( \sqrt{\frac{\ln\left( n\right) }{T}}\right) $ and ((ref)) follows from ((ref)). Next, consider ((ref)) and note by Assumptions (ref) and (ref), $\mathbf{f} _{t}$ is independent from $\varepsilon_{it^{\prime}}$ for all $t,t^{\prime }=1,2,\ldots,T$, also $\varepsilon_{it}$ is serially independent, then for a suitable choice of $\mathcal{I}_{i,t-1}^{\left( 2\right) }=\{\mathbf{f} _{\tau}\varepsilon_{i\tau}:$ $\tau=t-1,t-2,\ldots\}$ and $i=1,2,\ldots,n$, $E\left( \mathbf{f}_{t}\varepsilon_{it}|\mathcal{I}_{i,t-1}^{\left( 2\right) }\right) =E\left( \mathbf{f}_{t}|\mathcal{I}_{i,t-1}^{\left( 2\right) }\right) E\left( \varepsilon_{it}\right) =\mathbf{0}$ so $\mathbf{f}_{t}\varepsilon_{it}$ is a martingale difference sequence with respect to the filtration $\mathcal{I}_{i,t-1}^{\left( 2\right) }$. In addition, $E\left( \mathbf{f}_{t}\varepsilon_{it}\right) =E\left( \mathbf{f}_{t}\right) E\left( \varepsilon_{it}\right) =\mathbf{0}$ and $Var\left( \mathbf{f}_{t}\varepsilon_{it}\right) =E\left( \mathbf{f} _{t}\mathbf{f}_{t}^{\prime}\right) E\left( \varepsilon_{it}^{2}\right) $, which is bounded by Assumptions (ref) and (ref). Also by assumptions both $\mathbf{f}_{t}$ and $\varepsilon_{it}$ are sub-exponential, then it also follows that $\mathbf{f}_{t}\varepsilon_{it}$ is sub-exponential. Hence, the method of proof used above can also be applied to $\left\Vert \sum_{t=1}^{T}\mathbf{f}_{t}\varepsilon_{it}\right\Vert $, and result ((ref)) follows. Similarly, ((ref)) can be established by the symmetry of the standard factor models in $\boldsymbol{\gamma}_{i}$ and $\mathbf{f}_{t}$. Now consider ((ref)), and note that we have the following decomposition, \begin{align*} \mathbf{q}_{i,nT} & =\frac{1}{nT}\sum_{t=1}^{T}\sum_{j=1}^{n}\sigma _{j}\boldsymbol{\gamma}_{j}\varepsilon_{it}\varepsilon_{jt}\\ & =\frac{1}{nT}\sum_{t=1}^{T}\sum_{j=1}^{n}\left[ \sigma_{j} \boldsymbol{\gamma}_{j}\varepsilon_{it}\varepsilon_{jt}-\sigma_{j} \boldsymbol{\gamma}_{j}E\left( \varepsilon_{it}\varepsilon_{jt}\right) \right] +\frac{1}{nT}\sum_{t=1}^{T}\sum_{j=1}^{n}\sigma_{j}\boldsymbol{\gamma }_{j}E\left( \varepsilon_{it}\varepsilon_{jt}\right) =\mathbf{q} _{i,nT}^{(a)}+\mathbf{q}_{i,nT}^{(b)}. \end{align*} Since $E(\varepsilon_{it}\varepsilon_{jt})=0$ if $i\neq j$, and $E(\varepsilon _{it}\varepsilon_{jt})=1$, if $i=j$, then $\mathbf{q}_{i,nT}^{(b)}=\frac {1}{nT}\sum_{t=1}^{T}\sum_{j=1}^{n}\sigma_{j}\boldsymbol{\gamma}_{j}E\left( \varepsilon_{it}\varepsilon_{jt}\right) =n^{-1}\sigma_{j}\boldsymbol{\gamma }_{j}, $ and $\sup_{i}\left\Vert \mathbf{q}_{i,nT}^{(b)}\right\Vert =O(n^{-1})$. Consider now the first term and note that $\mathbf{q} _{i,nT}^{(a)}=\frac{1}{nT}\sum_{t=1}^{T}\sum_{j=1}^{n}\mathbf{s}_{i,jt}, $ where $\mathbf{s}_{i,jt}=\sigma_{j}\boldsymbol{\gamma}_{j}\left[ \varepsilon_{it}\varepsilon_{jt}-E\left( \varepsilon_{it}\varepsilon _{jt}\right) \right] $. Since by assumption $\varepsilon_{it}$ are independently distributed over all $i$ and $t$, then $E\left( \mathbf{s} _{i,jt}\left\vert \mathcal{I}_{i,t-1}^{\left( 3\right) }\right. \right) =\mathbf{0}$, where $\mathcal{I}_{i,t-1}^{\left( 3\right) }=\{\varepsilon _{i\tau}\varepsilon_{j\tau}:$ $j=1,2,...,n$ and $\tau=t-1,t-2,...\}$. Hence $\mathbf{s}_{i,jt}$ is a martingale difference process with respect to the filtration, $\mathcal{I}_{i,t-1}^{\left( 3\right) }$. The variance of $\mathbf{s}_{i,jt}$ is $\left( \sigma_{j}^{2}\boldsymbol{\gamma} _{j}\boldsymbol{\gamma}_{j}^{\prime}\right) Var(\varepsilon_{it} \varepsilon_{jt})$ where $Var(\varepsilon_{it}\varepsilon_{jt})=1$ if $i\neq j$, and $Var(\varepsilon_{it}\varepsilon_{jt})=Var(\varepsilon_{it} ^{2})=E(\varepsilon_{it}^{4})-1$ if $i=j$, so that by Assumption (ref), $\left\Vert Var(\mathbf{s}_{i,jt})\right\Vert <C$. Also, since by assumption $\varepsilon_{it}$ is sub-exponential, then it follows that $\mathbf{s}_{i,jt}$ is also sub-exponential, and the above method of proof can be applied to all elements of $\mathbf{q}_{i,nT}^{(a)}$. Specifically $\sup_{i}\left\Vert \mathbf{q}_{i,nT}^{(a)}\right\Vert =O_{p}\left( \sqrt{\frac{\ln(n)}{nT}}\right) $, $\sup_{i}\left\Vert \mathbf{q}_{i,nT}\right\Vert =O_{p}\left( \sqrt{\frac{\ln(n)}{nT}}\right) +O(n^{-1})=O_{p}\left( \sqrt{\frac{\ln(n)}{nT}}\right) $, and result ((ref)) follows, as required.
lemmaConsider the latent factor model given by ((ref)) and ((ref)). The latent factors, $\mathbf{f}_{t}$, and their loadings, $\boldsymbol{\gamma}_{i}$, are estimated by principal components, $\mathbf{\hat{f}}_{t}$ and $\boldsymbol{\hat{\gamma}}_{i}$, given by ((ref)). Suppose that Assumptions (ref)-(ref) hold and $\left( n,T\right) \rightarrow\infty$, such that $n/T\rightarrow\kappa\,,$ for $0<\kappa<\infty$. Then \begin{align} \sup_{i}\left( T^{-1}\left\Vert \boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) \right\Vert ^{2}\right) & =O_{p}\left( 1\right) ,\\ \sup_{i}\left\Vert \frac{\mathbf{F}^{\prime}\boldsymbol{\varepsilon}_{i\circ }\left( \lambda_{T}\right) }{T}\right\Vert & =O_{p}\left( \sqrt{\frac {\ln(n)}{T}}\right) ,\\ \sup_{t}\left\Vert \frac{\boldsymbol{\Gamma}^{\prime}\boldsymbol{\varepsilon }_{\circ t}\left( \lambda_{T}\right) }{n}\right\Vert & =O_{p}\left( \sqrt{\frac{\ln(T)}{n}}\right) ,\\ \sup_{i}\left\Vert \frac{1}{nT}\sum_{t=1}^{T}\sum_{j=1}^{n}\sigma _{j}\boldsymbol{\gamma}_{j}\varepsilon_{it}\left( \lambda_{T}\right) \varepsilon_{jt}\left( \lambda_{T}\right) \right\Vert & =O_{p}\left( \sqrt{\frac{\ln\left( n\right) }{nT}}\right) ,\\ \sup_{i}\left\Vert \boldsymbol{\hat{\gamma}}_{i}-\boldsymbol{\gamma} _{i}\right\Vert & =O_{p}\left( \sqrt{\frac{\ln(n)}{T}}\right) ,\\ \sup_{t}\left\Vert \mathbf{\hat{f}}_{t}-\mathbf{f}_{t}\right\Vert & =O_{p}\left( \sqrt{\frac{\ln(T)}{n}}\right) , \end{align} where $\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) =\left( \varepsilon_{i1}\left( \lambda_{T}\right) ,\varepsilon_{i2}\left( \lambda_{T}\right) ,\ldots,\varepsilon_{iT}\left( \lambda_{T}\right) \right) ^{\prime}$ and $\boldsymbol{\varepsilon}_{\circ t}\left( \lambda _{T}\right) =\left( \varepsilon_{1t}\left( \lambda_{T}\right) ,\varepsilon_{2t}\left( \lambda_{T}\right) ,\ldots,\varepsilon_{nt}\left( \lambda_{T}\right) \right) ^{\prime}$.
proofConsider ((ref)) and note by definition \begin{equation} \varepsilon_{it}\left( \lambda_{T}\right) =\varepsilon_{it}+\lambda _{T}\mathbf{w}_{i0}^{\prime}\boldsymbol{\varepsilon}_{\circ t}, \end{equation} where $\mathbf{w}_{i0}=\left( w_{i1},w_{i2},\ldots,w_{in}\right) ^{\prime}$ and $\boldsymbol{\varepsilon}_{\circ t}=\left( \varepsilon_{1t} ,\varepsilon_{2t},\ldots,\varepsilon_{nt}\right) ^{\prime}$. Then \[ \varepsilon_{it}^{2}\left( \lambda_{T}\right) =\varepsilon_{it}^{2} +2\lambda_{T}\mathbf{w}_{i0}^{\prime}\boldsymbol{\varepsilon}_{\circ t}\varepsilon_{it}+\lambda_{T}^{2}\mathbf{w}_{i0}^{\prime} \boldsymbol{\varepsilon}_{\circ t}\boldsymbol{\varepsilon}_{\circ t}^{\prime }\mathbf{w}_{i0}, \] and hence \[ \frac{1}{T}\sum_{t=1}^{T}\varepsilon_{it}^{2}\left( \lambda_{T}\right) =\frac{1}{T}\sum_{t=1}^{T}\varepsilon_{it}^{2}+2\lambda_{T}\mathbf{w} _{i0}^{\prime}\left( \frac{1}{T}\sum_{t=1}^{T}\boldsymbol{\varepsilon}_{\circ t}\varepsilon_{it}\right) +\lambda_{T}^{2}\mathbf{w}_{i0}^{\prime} \mathbf{V}_{\varepsilon T}\mathbf{w}_{i0}, \] where $\mathbf{V}_{\varepsilon T}=T^{-1}\sum_{t=1}^{T}\boldsymbol{\varepsilon }_{\circ t}\boldsymbol{\varepsilon}_{\circ t}^{\prime}$ and $\left\Vert \mathbf{V}_{\varepsilon T}\right\Vert =O_{p}\left( 1\right) $ by part (b) of Assumption (ref). It follows \[ \left\vert \frac{1}{T}\sum_{t=1}^{T}\varepsilon_{it}^{2}\left( \lambda _{T}\right) \right\vert \leq\left\vert \frac{1}{T}\sum_{t=1}^{T} \varepsilon_{it}^{2}\right\vert +2\left\vert \lambda_{T}\right\vert \left\Vert \mathbf{w}_{i0}\right\Vert \left\Vert \frac{1}{T}\sum_{t=1}^{T} \boldsymbol{\varepsilon}_{\circ t}\varepsilon_{it}\right\Vert +\lambda_{T} ^{2}\left\Vert \mathbf{w}_{i0}\right\Vert ^{2}\left\Vert \mathbf{V} _{\varepsilon T}\right\Vert . \] Denote $\mathbf{e}_{i}$ as $n\times1$ selection vector with $1$ on its $i^{th}$ element and zeros elsewhere, and note that \[ \sup_{i}\left\Vert \frac{1}{T}\sum_{t=1}^{T}\boldsymbol{\varepsilon}_{\circ t}\varepsilon_{it}\right\Vert =\sup_{i}\left\Vert \frac{1}{T}\sum_{t=1} ^{T}\boldsymbol{\varepsilon}_{\circ t}\boldsymbol{\varepsilon}_{\circ t}^{\prime}\mathbf{e}_{i}\right\Vert \leq\left\Vert \frac{1}{T}\sum_{t=1} ^{T}\boldsymbol{\varepsilon}_{\circ t}\boldsymbol{\varepsilon}_{\circ t}^{\prime}\right\Vert \left( \sup_{i}\left\Vert \mathbf{e}_{i}\right\Vert \right) =\left\Vert \mathbf{V}_{\varepsilon T}\right\Vert . \] Using this result we now have (recalling that $\lambda_{T}=c_{\lambda} T^{-1/2})$ \begin{align*} \sup_{i}\left\vert \frac{1}{T}\sum_{t=1}^{T}\varepsilon_{it}^{2}\left( \lambda_{T}\right) \right\vert & \leq\sup_{i}\left\vert \frac{1}{T} \sum_{t=1}^{T}\varepsilon_{it}^{2}\right\vert +2c_{\lambda}T^{-1/2}\left\Vert \mathbf{V}_{\varepsilon T}\right\Vert \left( \sup_{i}\left\Vert \mathbf{w}_{i0}\right\Vert \right) +\\ & c_{\lambda}^{2}T^{-1}\left\Vert \mathbf{V}_{\varepsilon T}\right\Vert \left( \sup_{i}\left\Vert \mathbf{w}_{i0}\right\Vert ^{2}\right) . \end{align*} Therefore, since $\sup_{i}\left\Vert \mathbf{w}_{i0}\right\Vert <C$ and by assumption $\left\Vert \mathbf{V}_{\varepsilon T}\right\Vert =O_{p}(1)$, then \begin{equation} \sup_{i}\left\vert \frac{1}{T}\sum_{t=1}^{T}\varepsilon_{it}^{2}\left( \lambda_{T}\right) \right\vert =\sup_{i}\left\vert \frac{1}{T}\sum_{t=1} ^{T}\varepsilon_{it}^{2}\right\vert +O_{p}\left( \frac{1}{\sqrt{T}}\right) . \end{equation} Result ((ref)) now follows from ((ref)). Similarly, to establish ((ref)) note that \[ \frac{\mathbf{F}^{\prime}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda _{T}\right) }{T}=\frac{1}{T}\sum_{t=1}^{T}\mathbf{f}_{t}\varepsilon _{it}\left( \lambda_{T}\right) =\frac{1}{T}\sum_{t=1}^{T}\mathbf{f} _{t}\varepsilon_{it}+\frac{\lambda_{T}}{T}\sum_{t=1}^{T}\mathbf{f} _{t}\boldsymbol{\varepsilon}_{\circ t}^{\prime}\mathbf{w}_{i0}, \] and \[ \left\Vert \frac{\mathbf{F}^{\prime}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) }{T}\right\Vert \leq\left\Vert \frac{1}{T}\sum_{t=1} ^{T}\mathbf{f}_{t}\varepsilon_{it}\right\Vert +\left\vert \lambda _{T}\right\vert \left\Vert \mathbf{w}_{i0}\right\Vert \left\Vert \frac{1} {T}\sum_{t=1}^{T}\mathbf{f}_{t}\boldsymbol{\varepsilon}_{\circ t}^{\prime }\right\Vert . \] Applying the supremum operator to both sides yields \[ \sup_{i}\left\Vert \frac{\mathbf{F}^{\prime}\boldsymbol{\varepsilon}_{i\circ }\left( \lambda_{T}\right) }{T}\right\Vert \leq\sup_{i}\left\Vert \frac {1}{T}\sum_{t=1}^{T}\mathbf{f}_{t}\varepsilon_{it}\right\Vert +\left\vert \lambda_{T}\right\vert \left( \sup_{i}\left\Vert \mathbf{w}_{i0}\right\Vert \right) \left\Vert \frac{1}{T}\sum_{t=1}^{T}\mathbf{f}_{t} \boldsymbol{\varepsilon}_{\circ t}^{\prime}\right\Vert . \] Also \begin{align*} E\left\Vert \frac{1}{T}\sum_{t=1}^{T}\mathbf{f}_{t}\boldsymbol{\varepsilon }_{\circ t}^{\prime}\right\Vert ^{2} & \leq E\left\Vert \frac{1}{T} \sum_{t=1}^{T}\mathbf{f}_{t}\boldsymbol{\varepsilon}_{\circ t}^{\prime }\right\Vert _{F}^{2}=\mathrm{tr}\left[ E\left[ \left( \frac{1}{T} \sum_{t=1}^{T}\mathbf{f}_{t}\boldsymbol{\varepsilon}_{\circ t}^{\prime }\right) \left( \frac{1}{T}\sum_{t=1}^{T}\mathbf{f}_{t} \boldsymbol{\varepsilon}_{\circ t}^{\prime}\right) ^{\prime}\right] \right] \\ & =\mathrm{tr}\left( \frac{1}{T^{2}}\sum_{t=1}^{T}\sum_{t^{\prime}=1} ^{T}E\left( \mathbf{f}_{t}\boldsymbol{\varepsilon}_{\circ t}^{\prime }\boldsymbol{\varepsilon}_{\circ t^{\prime}}\mathbf{f}_{t^{\prime}}^{\prime }\right) \right) =\frac{1}{T^{2}}\mathrm{tr}\left( \sum_{t=1}^{T}E\left( \mathbf{f}_{t}\mathbf{f}_{t}^{\prime}\right) E\left( \boldsymbol{\varepsilon }_{\circ t}^{\prime}\boldsymbol{\varepsilon}_{\circ t}\right) \right) \\ & =\frac{n}{T}\mathrm{tr}\left( \mathbf{\Sigma}_{ff}\right) =O\left( \frac{nm_{0}}{T}\right) , \end{align*} which establishes that $T^{-1}\sum_{t=1}^{T}\mathbf{f}_{t} \boldsymbol{\varepsilon}_{\circ t}^{\prime}=O_{p}\left( 1\right) $ since by assumption $n$ and $T$ have the same orders of magnitudes. Given $\lambda _{T}=c_{\lambda}T^{-1/2}$ and $\sup_{i}\left\Vert \mathbf{w}_{i0}\right\Vert <C$, we now have \[ \sup_{i}\left\Vert \frac{\mathbf{F}^{\prime}\boldsymbol{\varepsilon}_{i\circ }\left( \lambda_{T}\right) }{T}\right\Vert =\sup_{i}\left\Vert \frac{1} {T}\sum_{t=1}^{T}\mathbf{f}_{t}\varepsilon_{it}\right\Vert +O_{p}\left( \frac{1}{\sqrt{T}}\right) , \] and ((ref)) follows using ((ref)). Similarly, ((ref)) can be established using result ((ref)). Next, consider ((ref)) and using the definition of $\varepsilon_{it}\left( \lambda_{T}\right) $ in ((ref)) yields \begin{align*} \frac{1}{nT}\sum_{t=1}^{T}\sum_{j=1}^{n}\sigma_{j}\boldsymbol{\gamma} _{j}\varepsilon_{it}\left( \lambda_{T}\right) \varepsilon_{jt}\left( \lambda_{T}\right) & =\frac{1}{nT}\sum_{t=1}^{T}\sum_{j=1}^{n}\sigma _{j}\boldsymbol{\gamma}_{j}\varepsilon_{it}\varepsilon_{jt}+\frac{\lambda_{T} }{nT}\sum_{t=1}^{T}\sum_{j=1}^{n}\sigma_{j}\boldsymbol{\gamma}_{j} \mathbf{w}_{j0}^{\prime}\boldsymbol{\varepsilon}_{\circ t}\varepsilon_{it}+\\ & \frac{\lambda_{T}}{nT}\sum_{t=1}^{T}\sum_{j=1}^{n}\sigma_{j} \boldsymbol{\gamma}_{j}\mathbf{w}_{i0}^{\prime}\boldsymbol{\varepsilon}_{\circ t}\varepsilon_{jt}+\frac{\lambda_{T}^{2}}{n}\sum_{j=1}^{n}\sigma _{j}\boldsymbol{\gamma}_{j}\mathbf{w}_{i0}^{\prime}\left( \frac{1}{T} \sum_{t=1}^{T}\boldsymbol{\varepsilon}_{\circ t}\boldsymbol{\varepsilon }_{\circ t}^{\prime}\right) \mathbf{w}_{j0}, \end{align*} which implies \begin{align*} \left\Vert \frac{1}{nT}\sum_{t=1}^{T}\sum_{j=1}^{n}\sigma_{j} \boldsymbol{\gamma}_{j}\varepsilon_{it}\left( \lambda_{T}\right) \varepsilon_{jt}\left( \lambda_{T}\right) \right\Vert & \leq\left\Vert \frac{1}{nT}\sum_{t=1}^{T}\sum_{j=1}^{n}\sigma_{j}\boldsymbol{\gamma} _{j}\varepsilon_{it}\varepsilon_{jt}\right\Vert +\left\Vert \frac{\lambda_{T} }{nT}\sum_{t=1}^{T}\sum_{j=1}^{n}\sigma_{j}\boldsymbol{\gamma}_{j} \mathbf{w}_{j0}^{\prime}\boldsymbol{\varepsilon}_{\circ t}\varepsilon _{it}\right\Vert +\\ & \left\Vert \frac{\lambda_{T}}{nT}\sum_{t=1}^{T}\sum_{j=1}^{n}\sigma _{j}\boldsymbol{\gamma}_{j}\mathbf{w}_{i0}^{\prime}\boldsymbol{\varepsilon }_{\circ t}\varepsilon_{jt}\right\Vert +\left\Vert \frac{\lambda_{T}^{2}} {n}\sum_{j=1}^{n}\sigma_{j}\boldsymbol{\gamma}_{j}\mathbf{w}_{i0}^{\prime }\left( \frac{1}{T}\sum_{t=1}^{T}\boldsymbol{\varepsilon}_{\circ t}\boldsymbol{\varepsilon}_{\circ t}^{\prime}\right) \mathbf{w} _{j0}\right\Vert . \end{align*} Taking the supremum on both sides of this inequality yields \begin{align} & \sup_{i}\left\Vert \frac{1}{nT}\sum_{t=1}^{T}\sum_{j=1}^{n}\sigma _{j}\boldsymbol{\gamma}_{j}\varepsilon_{it}\left( \lambda_{T}\right) \varepsilon_{jt}\left( \lambda_{T}\right) \right\Vert \nonumber\\ & \leq\sup_{i}\left\Vert \frac{1}{nT}\sum_{t=1}^{T}\sum_{j=1}^{n}\sigma _{j}\boldsymbol{\gamma}_{j}\varepsilon_{it}\varepsilon_{jt}\right\Vert +\sup_{i}\left\Vert \frac{\lambda_{T}}{nT}\sum_{t=1}^{T}\sum_{j=1}^{n} \sigma_{j}\boldsymbol{\gamma}_{j}\mathbf{w}_{j0}^{\prime} \boldsymbol{\varepsilon}_{\circ t}\varepsilon_{it}\right\Vert +\nonumber\\ & \sup_{i}\left\Vert \frac{\lambda_{T}}{nT}\sum_{t=1}^{T}\sum_{j=1}^{n} \sigma_{j}\boldsymbol{\gamma}_{j}\mathbf{w}_{i0}^{\prime} \boldsymbol{\varepsilon}_{\circ t}\varepsilon_{jt}\right\Vert +\sup _{i}\left\Vert \frac{\lambda_{T}^{2}}{n}\sum_{j=1}^{n}\sigma_{j} \boldsymbol{\gamma}_{j}\mathbf{w}_{i0}^{\prime}\left( \frac{1}{T}\sum _{t=1}^{T}\boldsymbol{\varepsilon}_{\circ t}\boldsymbol{\varepsilon}_{\circ t}^{\prime}\right) \mathbf{w}_{j0}\right\Vert . \end{align} Note that $\varepsilon_{it}=\boldsymbol{\varepsilon}_{\circ t}^{\prime }\mathbf{e}_{i}$ where $\mathbf{e}_{i}$ is an $n\times1$ selection vector with $1$ on its $i^{th}$ element and zero elsewhere. Then the second term of the above can be bounded as \begin{align*} \sup_{i}\left\Vert \frac{\lambda_{T}}{nT}\sum_{t=1}^{T}\sum_{j=1}^{n} \sigma_{j}\boldsymbol{\gamma}_{j}\mathbf{w}_{j0}^{\prime} \boldsymbol{\varepsilon}_{\circ t}\varepsilon_{it}\right\Vert & =\sup _{i}\left\Vert \frac{\lambda_{T}}{n}\sum_{j=1}^{n}\sigma_{j}\boldsymbol{\gamma }_{j}\mathbf{w}_{j0}^{\prime}\left( \frac{1}{T}\sum_{t=1}^{T} \boldsymbol{\varepsilon}_{\circ t}\boldsymbol{\varepsilon}_{\circ t}^{\prime }\right) \mathbf{e}_{i}\right\Vert \\ & \leq\left\vert \lambda_{T}\right\vert \left\Vert \frac{1}{n}\sum_{j=1} ^{n}\sigma_{j}\boldsymbol{\gamma}_{j}\mathbf{w}_{j0}^{\prime}\left( \frac {1}{T}\sum_{t=1}^{T}\boldsymbol{\varepsilon}_{\circ t}\boldsymbol{\varepsilon }_{\circ t}^{\prime}\right) \right\Vert \left( \sup_{i}\left\Vert \mathbf{e}_{i}\right\Vert \right) \\ & \leq\left\vert \lambda_{T}\right\vert \left\Vert \frac{1}{n}\sum_{j=1} ^{n}\sigma_{j}\boldsymbol{\gamma}_{j}\mathbf{w}_{j0}^{\prime}\right\Vert \left\Vert \mathbf{V}_{\varepsilon T}\right\Vert . \end{align*} Since $\sup_{i}\sigma_{i}^{2}<C$, and by Assumptions (ref) -(ref), $\sup_{s,i}\gamma_{si}^{2}<C$ and $\sup_{i}\sum_{j=1} ^{n}\left\vert w_{ji}\right\vert <C$, then it follows \begin{align*} \left\Vert \frac{1}{n}\sum_{j=1}^{n}\sigma_{j}\boldsymbol{\gamma} _{j}\mathbf{w}_{j0}^{\prime}\right\Vert ^{2} & \leq\left\Vert \frac{1} {n}\sum_{j=1}^{n}\sigma_{j}\boldsymbol{\gamma}_{j}\mathbf{w}_{j0}^{\prime }\right\Vert _{F}^{2}=\frac{1}{n^{2}}\sum_{s=1}^{m_{0}}\sum_{i=1}^{n}\left( \sum_{j=1}^{n}\sigma_{j}\gamma_{sj}w_{ji}\right) ^{2}\\ & \leq\frac{1}{n^{2}}\sum_{s=1}^{m_{0}}\sum_{i=1}^{n}\left( \sum_{j=1} ^{n}\left\vert \sigma_{j}\gamma_{sj}\right\vert \left\vert w_{ji}\right\vert \right) ^{2}\\ & \leq\left( \sup_{i}\sigma_{i}^{2}\right) \left( \sup_{s,i}\gamma _{si}^{2}\right) \left[ \frac{1}{n}\sum_{s=1}^{m_{0}}\left( \frac{1}{n} \sum_{i=1}^{n}\left( \sum_{j=1}^{n}\left\vert w_{ji}\right\vert \right) ^{2}\right) \right] \\ & =O\left( \frac{1}{n}\right) . \end{align*} In addition, $\lambda_{T}=c_{\lambda}T^{-1/2}$ and $\left\Vert \mathbf{V} _{\varepsilon T}\right\Vert =O_{p}\left( n/T\right) $ with $0<n/T<C$. Hence, we have $\sup_{i}\left\Vert \frac{\lambda_{T}}{nT}\sum_{t=1}^{T}\sum_{j=1} ^{n}\sigma_{j}\boldsymbol{\gamma}_{j}\mathbf{w}_{j0}^{\prime} \boldsymbol{\varepsilon}_{\circ t}\varepsilon_{it}\right\Vert =O_{p}\left( \left( nT\right) ^{-1/2}\right) $. Similarly the third term of ((ref)) is also $O_{p}\left( \left( nT\right) ^{-1/2}\right) $. For the fourth term of ((ref)), \begin{align*} \sup_{i}\left\Vert \frac{\lambda_{T}^{2}}{n}\sum_{j=1}^{n}\sigma _{j}\boldsymbol{\gamma}_{j}\mathbf{w}_{i0}^{\prime}\left( \frac{1}{T} \sum_{t=1}^{T}\boldsymbol{\varepsilon}_{\circ t}\boldsymbol{\varepsilon }_{\circ t}^{\prime}\right) \mathbf{w}_{j0}\right\Vert & \leq\lambda _{T}^{2}\left( \sup_{i}\left\Vert \mathbf{w}_{i0}\right\Vert \right) \left( \frac{1}{n}\sum_{j=1}^{n}\sigma_{j}\left\Vert \boldsymbol{\gamma} _{j}\right\Vert \left\Vert \mathbf{w}_{j0}\right\Vert \right) \left\Vert \mathbf{V}_{\varepsilon T}\right\Vert \\ & =O_{p}\left( \frac{1}{T}\right) . \end{align*} Using the above results in ((ref)) we now have \[ \sup_{i}\left\Vert \frac{1}{nT}\sum_{t=1}^{T}\sum_{j=1}^{n}\sigma _{j}\boldsymbol{\gamma}_{j}\varepsilon_{it}\left( \lambda_{T}\right) \varepsilon_{jt}\left( \lambda_{T}\right) \right\Vert \leq\sup_{i}\left\Vert \frac{1}{nT}\sum_{t=1}^{T}\sum_{j=1}^{n}\sigma_{j}\boldsymbol{\gamma} _{j}\varepsilon_{it}\varepsilon_{jt}\right\Vert +O_{p}\left( \sqrt{\frac {\ln\left( n\right) }{nT}}\right) , \] and ((ref)) follows using ((ref)) to establish the order of the first term of the above. To establish (ref)) note that by definition of $\boldsymbol{\hat{\gamma}}_{i}$, \begin{align*} \boldsymbol{\hat{\gamma}}_{i}-\boldsymbol{\gamma}_{i} & =\left( \mathbf{\hat{F}}^{\prime}\mathbf{\hat{F}}\right) ^{-1}\mathbf{\hat{F} }^{\prime}\left( \mathbf{F}\boldsymbol{\gamma}_{i}+\sigma_{i} \boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) \right) -\boldsymbol{\gamma}_{i}=\left( \mathbf{\hat{F}}^{\prime}\mathbf{\hat{F} }\right) ^{-1}\mathbf{\hat{F}}^{\prime}\left[ \left( \mathbf{F-\hat{F} }\right) \boldsymbol{\gamma}_{i}+\mathbf{\hat{F}}\boldsymbol{\gamma} _{i}+\sigma_{i}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) \right] -\boldsymbol{\gamma}_{i}\\ & =\left( \frac{\mathbf{\hat{F}}^{\prime}\mathbf{\hat{F}}}{T}\right) ^{-1}\left( \frac{\mathbf{\hat{F}}-\mathbf{F+F}}{T}\right) ^{\prime}\left[ \left( \mathbf{F-\hat{F}}\right) \boldsymbol{\gamma}_{i}+\sigma _{i}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) \right] \\ & =\left( \frac{\mathbf{\hat{F}}^{\prime}\mathbf{\hat{F}}}{T}\right) ^{-1}\left[ \frac{\left( \mathbf{\hat{F}}-\mathbf{F}\right) ^{\prime }\left( \mathbf{F-\hat{F}}\right) }{T}\right] \boldsymbol{\gamma} _{i}+\left( \frac{\mathbf{\hat{F}}^{\prime}\mathbf{\hat{F}}}{T}\right) ^{-1}\left[ \frac{\mathbf{F}^{\prime}\left( \mathbf{F-\hat{F}}\right) } {T}\right] \boldsymbol{\gamma}_{i}\\ & +\left( \frac{\mathbf{\hat{F}}^{\prime}\mathbf{\hat{F}}}{T}\right) ^{-1}\left[ \frac{\sigma_{i}\left( \mathbf{\hat{F}}-\mathbf{F}\right) ^{\prime}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) } {T}\right] +\left( \frac{\mathbf{\hat{F}}^{\prime}\mathbf{\hat{F}}} {T}\right) ^{-1}\left( \frac{\sigma_{i}\mathbf{F}^{\prime} \boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) }{T}\right) =\sum_{j=1}^{4}\mathbf{a}_{j,iT}, \end{align*} and \begin{equation} \left\Vert \boldsymbol{\hat{\gamma}}_{i}-\boldsymbol{\gamma}_{i}\right\Vert \leq\sum_{j=1}^{4}\left\Vert \mathbf{a}_{j,iT}\right\Vert . \end{equation} Firstly we have \begin{align*} \left\Vert \mathbf{a}_{1,iT}\right\Vert & \leq\left\Vert \left( \frac{\mathbf{\hat{F}}^{\prime}\mathbf{\hat{F}}}{T}\right) ^{-1}\right\Vert \left\Vert \left[ \frac{\left( \mathbf{\hat{F}}-\mathbf{F}\right) ^{\prime }\left( \mathbf{F-\hat{F}}\right) }{T}\right] \right\Vert \left\Vert \boldsymbol{\gamma}_{i}\right\Vert \\ & \leq\left\Vert \left( \frac{\mathbf{\hat{F}}^{\prime}\mathbf{\hat{F}}} {T}\right) ^{-1}\right\Vert \left( \frac{\left\Vert \mathbf{\hat{F} }-\mathbf{F}\right\Vert ^{2}}{T}\right) \left\Vert \boldsymbol{\gamma} _{i}\right\Vert , \end{align*} which implies \[ \sup_{i}\left\Vert \mathbf{a}_{1,iT}\right\Vert \leq\left\Vert \left( \frac{\mathbf{\hat{F}}^{\prime}\mathbf{\hat{F}}}{T}\right) ^{-1}\right\Vert \left( \frac{\left\Vert \mathbf{\hat{F}}-\mathbf{F}\right\Vert ^{2}} {T}\right) \sup_{i}\left\Vert \boldsymbol{\gamma}_{i}\right\Vert . \] Using ((ref)) in Lemma (ref) and ((ref)) in Lemma (ref), we note that \begin{equation} \frac{\mathbf{\hat{F}}^{^{\prime}}\mathbf{\hat{F}}}{T}=O_{p}(1), \left( \frac{\mathbf{\hat{F}}^{^{\prime}}\mathbf{\hat{F}}}{T}\right) ^{-1}=O_{p}(1), and T^{-1}\left\Vert \mathbf{\hat{F}}-\mathbf{F} \right\Vert _{F}^{2}=O_{p}\left( \frac{1}{\delta_{nT}^{2}}\right) . \end{equation} Using this result and $\sup_{i}\left\Vert \boldsymbol{\gamma}_{i}\right\Vert <C$, we obtain \begin{equation} \sup_{i}\left\Vert \mathbf{a}_{1,iT}\right\Vert =O_{p}\left( \frac{1} {\delta_{nT}^{2}}\right) . \end{equation} Similarly, $\left\Vert \mathbf{a}_{2,iT}\right\Vert \leq\left\Vert \left( \frac{\mathbf{\hat{F}}^{\prime}\mathbf{\hat{F}}}{T}\right) ^{-1}\right\Vert \left\Vert \frac{\mathbf{F}^{\prime}\left( \mathbf{F-\hat{F}}\right) } {T}\right\Vert \left\Vert \boldsymbol{\gamma}_{i}\right\Vert ,$ so using ((ref)) and ((ref)) it yields \begin{equation} \sup_{i}\left\Vert \mathbf{a}_{2,iT}\right\Vert \leq\left\Vert \left( \frac{\mathbf{\hat{F}}^{\prime}\mathbf{\hat{F}}}{T}\right) ^{-1}\right\Vert \left( \frac{\left\Vert \mathbf{F}^{\prime}\left( \mathbf{F-\hat{F}}\right) \right\Vert }{T}\right) \sup_{i}\left\Vert \boldsymbol{\gamma}_{i}\right\Vert =O_{p}\left( \frac{1}{\delta_{nT}^{2}}\right) . \end{equation} Regarding $\mathbf{a}_{3,iT}$, by Cauchy-Schwarz inequality we have \[ \left\Vert \mathbf{a}_{3,iT}\right\Vert \leq\left\Vert \left( \frac {\mathbf{\hat{F}}^{\prime}\mathbf{\hat{F}}}{T}\right) ^{-1}\right\Vert \left\Vert \frac{\sigma_{i}\left( \mathbf{\hat{F}}-\mathbf{F}\right) ^{\prime}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) } {T}\right\Vert \leq\left\Vert \left( \frac{\mathbf{\hat{F}}^{\prime }\mathbf{\hat{F}}}{T}\right) ^{-1}\right\Vert \left( \frac{\left\Vert \mathbf{\hat{F}}-\mathbf{F}\right\Vert ^{2}}{T}\right) ^{1/2}\left( \frac{\sigma_{i}^{2}\left\Vert \boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) \right\Vert ^{2}}{T}\right) ^{1/2}, \] and therefore \[ \sup_{i}\left\Vert \mathbf{a}_{3,iT}\right\Vert \leq\left( \sup_{i}\sigma _{i}\right) \left\Vert \left( \frac{\mathbf{\hat{F}}^{\prime}\mathbf{\hat {F}}}{T}\right) ^{-1}\right\Vert \left( \frac{\left\Vert \mathbf{\hat{F} }-\mathbf{F}\right\Vert ^{2}}{T}\right) ^{1/2}\left[ \sup_{i}\left( \frac{\left\Vert \boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) \right\Vert ^{2}}{T}\right) \right] ^{1/2}. \] Now using ((ref)) and ((ref)) it follows that \begin{equation} \sup_{i}\left\Vert \mathbf{a}_{3,iT}\right\Vert =O_{p}\left( \frac{1} {\delta_{nT}}\right) . \end{equation} Next, note \[ \left\Vert \mathbf{a}_{4,iT}\right\Vert \leq\left\Vert \left( \frac {\mathbf{\hat{F}}^{\prime}\mathbf{\hat{F}}}{T}\right) ^{-1}\right\Vert \left\Vert \frac{\sigma_{i}\mathbf{F}^{\prime}\boldsymbol{\varepsilon} _{i\circ}\left( \lambda_{T}\right) }{T}\right\Vert \] then by ((ref)) we also have \begin{equation} \sup_{i}\left\Vert \mathbf{a}_{4,iT}\right\Vert \leq\left( \sup_{i}\sigma _{i}\right) \left\Vert \left( \frac{\mathbf{\hat{F}}^{\prime}\mathbf{\hat {F}}}{T}\right) ^{-1}\right\Vert \left( \sup_{i}\left\Vert \frac {\mathbf{F}^{\prime}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda _{T}\right) }{T}\right\Vert \right) =O_{p}\left( \sqrt{\frac{\ln(n)}{T} }\right) . \end{equation} Hence using ((ref))-((ref)) in ((ref)) we have \[ \sup_{i}\left\Vert \boldsymbol{\hat{\gamma}}_{i}-\boldsymbol{\gamma} _{i}\right\Vert \leq\sum_{j=1}^{4}\sup_{i}\left\Vert \mathbf{a}_{j,iT} \right\Vert =O_{p}\left( \sqrt{\frac{\ln(n)}{T}}\right) , \] as required. Result ((ref)) follows by symmetry.
lemmaConsider $\varepsilon_{it}\left( \lambda_{T}\right) =\varepsilon_{it}+\lambda_{T}\mathbf{w}_{i0}^{\prime}\boldsymbol{\varepsilon }_{\circ t}$, where $\boldsymbol{\varepsilon}_{\circ t}=(\varepsilon _{1t},\varepsilon_{2t},...,\varepsilon_{nt})^{\prime}$, $\varepsilon_{it}\sim IID\left( 0,1\right) $ for all $i$ and $t$, $\mathbf{w}_{i0}=\left( w_{i1},w_{i2},\ldots,w_{in}\right) ^{\prime}$, and $\mathbf{W}=\left( w_{ij}\right) $ satisfy the bounded conditions $\left\Vert \mathbf{W} \right\Vert _{1}=\sup_{j}\sum_{i=1}^{n}\left\vert w_{ij}\right\vert <C,$ and $\left\Vert \mathbf{W}\right\Vert _{\infty}=\sup_{i}\sum_{j=1}^{n}\left\vert w_{ij}\right\vert <C$. Then for all $\left\vert \lambda_{T}\right\vert <C$ we have \[ \sup_{j}\sum_{i=1}^{n}\left\vert E\left[ \varepsilon_{it}\left( \lambda _{T}\right) \varepsilon_{jt}\left( \lambda_{T}\right) \right] \right\vert <C,\text{ }\sup_{i}\sum_{j=1}^{n}\left\vert E\left[ \varepsilon_{it}\left( \lambda_{T}\right) \varepsilon_{jt}\left( \lambda_{T}\right) \right] \right\vert <C, \] and \begin{equation} n^{-1}\sum_{i=1}^{n}\sum_{j=1}^{n}\left\vert E\left[ \varepsilon_{it}\left( \lambda_{T}\right) \varepsilon_{jt}\left( \lambda_{T}\right) \right] \right\vert <C. \end{equation}
proofLet $\boldsymbol{\varepsilon}_{\circ t}(\lambda_{T})=$ $\boldsymbol{\varepsilon}_{\circ t}+\lambda_{T}\mathbf{W} \boldsymbol{\varepsilon}_{\circ t}$, where $\mathbf{W}^{\prime}=(\mathbf{w} _{10},\mathbf{w}_{20},...,\mathbf{w}_{n0})$. Then \[ \boldsymbol{\varepsilon}_{\circ t}(\lambda_{T})\boldsymbol{\varepsilon}_{\circ t}^{\prime}(\lambda_{T})=\boldsymbol{\varepsilon}_{\circ t} \boldsymbol{\varepsilon}_{\circ t}^{\prime}+\lambda_{T}^{2}\mathbf{W} \boldsymbol{\varepsilon}_{\circ t}\boldsymbol{\varepsilon}_{\circ t}^{\prime }\mathbf{W}^{\prime}+\lambda_{T}\mathbf{W}\boldsymbol{\varepsilon}_{\circ t}\boldsymbol{\varepsilon}_{\circ t}^{\prime}+\lambda_{T} \boldsymbol{\varepsilon}_{\circ t}\boldsymbol{\varepsilon}_{\circ t}^{\prime }\mathbf{W}^{\prime}, \] and \[ \mathbf{V}_{\varepsilon}(\lambda_{T})=E\left[ \boldsymbol{\varepsilon}_{\circ t}(\lambda_{T})\boldsymbol{\varepsilon}_{\circ t}^{\prime}(\lambda _{T})\right] =\mathbf{I}_{n}+\lambda_{T}\left( \mathbf{W+W}^{\prime}\right) +\lambda_{T}^{2}\mathbf{WW}^{\prime}. \] Consider the maximum absolute column sum norm of $\mathbf{V}_{\varepsilon }(\lambda_{T})$ and note that \[ \left\Vert \mathbf{V}_{\varepsilon}(\lambda_{T})\right\Vert _{1}=\sup_{j} \sum_{i=1}^{n}\left\vert E\left[ \varepsilon_{it}\left( \lambda_{T}\right) \varepsilon_{jt}\left( \lambda_{T}\right) \right] \right\vert <1+\left\vert \lambda_{T}\right\vert \left( \left\Vert \mathbf{W}\right\Vert _{1} +\left\Vert \mathbf{W}\right\Vert _{\infty}\right) +\lambda_{T}^{2}\left\Vert \mathbf{W}\right\Vert _{1}\left\Vert \mathbf{W}\right\Vert _{\infty}<C\text{. } \] Similarly for the maximum absolute row sum norm of $\mathbf{V}_{\varepsilon }(\lambda_{T})$ \[ \left\Vert \mathbf{V}_{\varepsilon}(\lambda_{T})\right\Vert _{\infty}=\sup _{i}\sum_{j=1}^{n}\left\vert E\left[ \varepsilon_{it}\left( \lambda _{T}\right) \varepsilon_{jt}\left( \lambda_{T}\right) \right] \right\vert <1+\left\vert \lambda_{T}\right\vert \left( \left\Vert \mathbf{W}\right\Vert _{\infty}+\left\Vert \mathbf{W}\right\Vert _{1}\right) +\lambda_{T} ^{2}\left\Vert \mathbf{W}\right\Vert _{\infty}\left\Vert \mathbf{W}\right\Vert _{1}<C\text{,} \] and result ((ref)) follows.
lemmaConsider the latent factor model given by ((ref)) and ((ref)). Suppose that Assumptions (ref)-(ref) hold and $\left( n,T\right) \rightarrow\infty$, such that $n/T\rightarrow \kappa\,,$ for $0<\kappa<\infty$. Then for the estimator of factors, we have \begin{equation} \sup_{i}\left\Vert \frac{\left( \mathbf{\hat{F}-F}\right) ^{\prime }\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) }{T}\right\Vert =O_{p}\left( \sqrt{\frac{\ln\left( n\right) }{nT}}\right) . \end{equation}
proofBy (A.1) of Bai (2003) we note that \[ \mathbf{\hat{f}}_{t}-\mathbf{f}_{t}=\frac{1}{T}\sum_{t^{\prime}=1} ^{T}\mathbf{\hat{f}}_{t^{\prime}}\eta_{tt^{\prime}}+\frac{1}{T}\sum _{t^{\prime}=1}^{T}\mathbf{\hat{f}}_{t^{\prime}}\zeta_{tt^{\prime}}+\frac {1}{T}\sum_{t^{\prime}=1}^{T}\mathbf{\hat{f}}_{t^{\prime}}\varkappa _{tt^{\prime}}+\frac{1}{T}\sum_{t^{\prime}=1}^{T}\mathbf{\hat{f}}_{t^{\prime} }\xi_{tt^{\prime}} \] where $\eta_{tt^{\prime}}=n^{-1}\sum_{i=1}^{n}\sigma_{i}^{2}E\left( \varepsilon_{it}\left( \lambda_{T}\right) \varepsilon_{it^{\prime}}\left( \lambda_{T}\right) \right) $, $\zeta_{tt^{\prime}}=n^{-1}\sum_{i=1} ^{n}\sigma_{i}^{2}\varepsilon_{it}^{2}\left( \lambda_{T}\right) -\eta_{tt^{\prime}}$, $\varkappa_{tt^{\prime}}=n^{-1}\sum_{i=1}^{n}\sigma _{i}\mathbf{f}_{t^{\prime}}^{\prime}\boldsymbol{\gamma}_{i}^{\prime }\varepsilon_{it}\left( \lambda_{T}\right) $, and $\xi_{tt^{\prime}} =n^{-1}\sum_{i=1}^{n}\sigma_{i}\mathbf{f}_{t}^{\prime}\boldsymbol{\gamma} _{i}^{\prime}\varepsilon_{it^{\prime}}\left( \lambda_{T}\right) $. Hence \begin{align*} \frac{1}{T}\left( \mathbf{\hat{F}-F}\right) ^{\prime}\boldsymbol{\varepsilon }_{i\circ}\left( \lambda_{T}\right) & =\frac{1}{T}\sum_{t=1}^{T}\left( \mathbf{\hat{f}}_{t}-\mathbf{f}_{t}\right) \varepsilon_{it}\left( \lambda_{T}\right) \\ & =\frac{1}{T^{2}}\sum_{t=1}^{T}\sum_{t^{\prime}=1}^{T}\mathbf{\hat{f} }_{t^{\prime}}\eta_{tt^{\prime}}\varepsilon_{it}\left( \lambda_{T}\right) +\frac{1}{T^{2}}\sum_{t=1}^{T}\sum_{t^{\prime}=1}^{T}\mathbf{\hat{f} }_{t^{\prime}}\zeta_{tt^{\prime}}\varepsilon_{it}\left( \lambda_{T}\right) \\ & +\frac{1}{T^{2}}\sum_{t=1}^{T}\sum_{t^{\prime}=1}^{T}\mathbf{\hat{f} }_{t^{\prime}}\varkappa_{tt^{\prime}}\varepsilon_{it}\left( \lambda _{T}\right) +\frac{1}{T^{2}}\sum_{t=1}^{T}\sum_{t^{\prime}=1}^{T} \mathbf{\hat{f}}_{t^{\prime}}\xi_{tt^{\prime}}\varepsilon_{it}\left( \lambda_{T}\right) \\ & =\sum_{j=1}^{4}\mathbf{b}_{j,iT}, \end{align*} and \begin{equation} \sup_{i}\left\Vert \frac{\left( \mathbf{\hat{F}-F}\right) ^{\prime }\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) }{T}\right\Vert \leq\sum_{j=1}^{4}\sup_{i}\left\Vert \mathbf{b}_{j,iT}\right\Vert . \end{equation} Firstly, consider $\mathbf{b}_{1,iT}$ and note that \begin{align*} \mathbf{b}_{1,iT} & =\frac{1}{T^{2}}\sum_{t=1}^{T}\sum_{t^{\prime}=1} ^{T}\left( \mathbf{\hat{f}}_{t^{\prime}}-\mathbf{f}_{t^{\prime}}\right) \eta_{tt^{\prime}}\varepsilon_{it}\left( \lambda_{T}\right) +\frac{1}{T^{2} }\sum_{t=1}^{T}\sum_{t^{\prime}=1}^{T}\mathbf{f}_{t^{\prime}}\eta_{tt^{\prime }}\varepsilon_{it}\left( \lambda_{T}\right) =\mathbf{b}_{1,1,iT} +\mathbf{b}_{1,2,iT}, \end{align*} so $\left\Vert \mathbf{b}_{1,iT}\right\Vert \leq\left\Vert \mathbf{b} _{1,1,iT}\right\Vert +\left\Vert \mathbf{b}_{1,2,iT}\right\Vert . $ Note for the first term on the right hand side, \begin{align*} \left\Vert \mathbf{b}_{1,1,iT}\right\Vert & =\left\Vert \frac{1}{T} \sum_{t^{\prime}=1}^{T}\left( \mathbf{\hat{f}}_{t^{\prime}}-\mathbf{f} _{t^{\prime}}\right) \left( \frac{1}{T}\sum_{t=1}^{T}\eta_{tt^{\prime} }\varepsilon_{it}\left( \lambda_{T}\right) \right) \right\Vert \\ & \leq\left( \frac{1}{T}\sum_{t^{\prime}=1}^{T}\left\Vert \mathbf{\hat{f} }_{t^{\prime}}-\mathbf{f}_{t^{\prime}}\right\Vert ^{2}\right) ^{1/2}\left[ \frac{1}{T}\sum_{t^{\prime}=1}^{T}\left( \frac{1}{T}\sum_{t=1}^{T} \eta_{tt^{\prime}}\varepsilon_{it}\left( \lambda_{T}\right) \right) ^{2}\right] ^{1/2}\\ & \leq\frac{1}{\sqrt{T}}\left( \frac{1}{T}\sum_{t^{\prime}=1}^{T}\left\Vert \mathbf{\hat{f}}_{t^{\prime}}-\mathbf{f}_{t^{\prime}}\right\Vert ^{2}\right) ^{1/2}\left( \frac{1}{T}\sum_{t^{\prime}=1}^{T}\sum_{t=1}^{T}\eta _{tt^{\prime}}^{2}\right) ^{1/2}\left( \frac{1}{T}\sum_{t=1}^{T} \varepsilon_{it}^{2}\left( \lambda_{T}\right) \right) ^{1/2}, \end{align*} where the second and third lines hold by Cauchy-Schwarz inequality. By definition of $\eta_{tt^{\prime}}$ and serial independence of $\varepsilon _{it}\left( \lambda_{T}\right) ,$ $\eta_{tt^{\prime}}=\bar{\sigma}_{n}^{2}$ for $t=t^{\prime}$ but $0$ otherwise, where $\bar{\sigma}_{n}^{2}=n^{-1} \sum_{i=1}^{n}\sigma_{i}^{2}E\left( \varepsilon_{it}^{2}\left( \lambda _{T}\right) \right) $. Under assumption on weight $\left\{ w_{ij}\right\} $, $E\left( \varepsilon_{it}^{2}\left( \lambda_{T}\right) \right) =1+\lambda_{T}^{2}\left( \sum_{j=1}^{n}w_{ij}^{2}\right) <C $, so it follows that \begin{equation} \frac{1}{T}\sum_{t^{\prime}=1}^{T}\sum_{t=1}^{T}\eta_{tt^{\prime}}^{2}=\left( \bar{\sigma}_{n}^{2}\right) ^{2}<C. \end{equation} Given results ((ref)), ((ref)) and ((ref)), we further obtain \begin{align*} \sup_{i}\left\Vert \mathbf{b}_{1,1,iT}\right\Vert & \leq\frac{1}{\sqrt{T} }\left( \frac{1}{T}\sum_{t^{\prime}=1}^{T}\left\Vert \mathbf{\hat{f} }_{t^{\prime}}-\mathbf{f}_{t^{\prime}}\right\Vert ^{2}\right) ^{1/2}\left( \frac{1}{T}\sum_{t^{\prime}=1}^{T}\sum_{t=1}^{T}\eta_{tt^{\prime}}^{2}\right) ^{1/2}\left( \sup_{i}\frac{1}{T}\sum_{t=1}^{T}\varepsilon_{it}^{2}\left( \lambda_{T}\right) \right) ^{1/2}\\ & =O_{p}\left( T^{-1/2}\delta_{nT}^{-1}\right) . \end{align*} Now consider $\mathbf{b}_{1,2,iT}$. Using properties of $\eta_{tt^{\prime}}$ we have \[ \frac{1}{T^{2}}\sum_{t=1}^{T}\sum_{t^{\prime}=1}^{T}\mathbf{f}_{t^{\prime} }\eta_{tt^{\prime}}\varepsilon_{it}\left( \lambda_{T}\right) =\frac{1} {T^{2}}\sum_{t=1}^{T}\mathbf{f}_{t}\eta_{tt}\varepsilon_{it}\left( \lambda_{T}\right) +\frac{1}{T^{2}}\sum_{t=1}^{T}\sum_{t^{\prime}\neq t^{\prime}}^{T}\mathbf{f}_{t^{\prime}}\eta_{tt^{\prime}}\varepsilon _{it}\left( \lambda_{T}\right) =\bar{\sigma}_{n}^{2}\left( \frac{1}{T^{2} }\sum_{t=1}^{T}\mathbf{f}_{t}\varepsilon_{it}\left( \lambda_{T}\right) \right) , \] and therefore \[ \left\Vert \mathbf{b}_{1,2,iT}\right\Vert =\left\Vert \frac{1}{T^{2}} \sum_{t=1}^{T}\sum_{t^{\prime}=1}^{T}\mathbf{f}_{t^{\prime}}\eta_{tt^{\prime} }\varepsilon_{it}\left( \lambda_{T}\right) \right\Vert \leq\frac{1}{T} \bar{\sigma}_{n}^{2}\left( \left\Vert \frac{1}{T}\sum_{t=1}^{T}\mathbf{f} _{t}\varepsilon_{it}\left( \lambda_{T}\right) \right\Vert \right) . \] Then using ((ref)) we have \[ \sup_{i}\left\Vert \mathbf{b}_{1,2,iT}\right\Vert \leq\frac{1}{T}\bar{\sigma }_{n}^{2}\sup_{i}\left\Vert \frac{1}{T}\sum_{t=1}^{T}\mathbf{f}_{t} \varepsilon_{it}\left( \lambda_{T}\right) \right\Vert =O_{p}\left( \frac {1}{T}\sqrt{\frac{\ln\left( n\right) }{T}}\right) . \] Hence, combining the probability orders of $\sup_{i}\left\Vert \mathbf{b} _{1,1,iT}\right\Vert $ and $\sup_{i}\left\Vert \mathbf{b}_{1,2,iT}\right\Vert $, we have \begin{equation} \sup_{i}\left\Vert \mathbf{b}_{1,nT}\right\Vert \leq\sup_{i}\left\Vert \mathbf{b}_{1,1,iT}\right\Vert +\sup_{i}\left\Vert \mathbf{b}_{1,2,iT} \right\Vert =O_{p}\left( \frac{1}{T^{1/2}\delta_{nT}}\right) . \end{equation} Next, consider $\mathbf{b}_{2,iT}$ in ((ref)), which can be written as \begin{align*} \mathbf{b}_{2,iT} & =\frac{1}{T^{2}}\sum_{t=1}^{T}\sum_{t^{\prime}=1} ^{T}\left( \mathbf{\hat{f}}_{t^{\prime}}-\mathbf{f}_{t^{\prime}}\right) \zeta_{tt^{\prime}}\varepsilon_{it}\left( \lambda_{T}\right) +\frac{1} {T^{2}}\sum_{t=1}^{T}\sum_{t^{\prime}=1}^{T}\mathbf{f}_{t^{\prime}} \zeta_{tt^{\prime}}\varepsilon_{it}\left( \lambda_{T}\right) =\mathbf{b} _{2,1,iT}+\mathbf{b}_{2,2,iT}. \end{align*} For the first term, we can apply Cauchy-Schwarz inequality to obtain \begin{align*} \left\Vert \mathbf{b}_{2,1,iT}\right\Vert & =\left\Vert \frac{1}{T^{2}} \sum_{t^{\prime}=1}^{T}\left( \mathbf{\hat{f}}_{t^{\prime}}-\mathbf{f} _{t^{\prime}}\right) \left( \sum_{t=1}^{T}\zeta_{tt^{\prime}}\varepsilon _{it}\left( \lambda_{T}\right) \right) \right\Vert \\ & \leq\left( \frac{1}{T}\sum_{t^{\prime}=1}^{T}\left\Vert \mathbf{\hat{f} }_{t^{\prime}}-\mathbf{f}_{t^{\prime}}\right\Vert ^{2}\right) ^{1/2}\left[ \frac{1}{T^{3}}\sum_{t^{\prime}=1}^{T}\left( \sum_{t=1}^{T}\zeta_{tt^{\prime }}\varepsilon_{it}\left( \lambda_{T}\right) \right) ^{2}\right] ^{1/2}\\ & \leq\left( \frac{1}{T}\sum_{t^{\prime}=1}^{T}\left\Vert \mathbf{\hat{f} }_{t^{\prime}}-\mathbf{f}_{t^{\prime}}\right\Vert ^{2}\right) ^{1/2}\left( \frac{1}{T^{2}}\sum_{t^{\prime}=1}^{T}\sum_{t=1}^{T}\zeta_{tt^{\prime}} ^{2}\right) ^{1/2}\left( \frac{1}{T}\sum_{t=1}^{T}\varepsilon_{it} ^{2}\left( \lambda_{T}\right) \right) ^{1/2}. \end{align*} Since \[ \zeta_{tt^{\prime}}=n^{-1}\sum_{i=1}^{n}\sigma_{i}^{2}\varepsilon_{it} ^{2}\left( \lambda_{T}\right) -\eta_{tt^{\prime}}=\frac{1}{n}\sum_{j=1} ^{n}\sigma_{j}^{2}\left[ \varepsilon_{jt}\left( \lambda_{T}\right) \varepsilon_{jt^{\prime}}\left( \lambda_{T}\right) -E\left( \varepsilon _{jt}\left( \lambda_{T}\right) \varepsilon_{jt^{\prime}}\left( \lambda _{T}\right) \right) \right] , \] then it follows that \begin{align} E\left( \frac{1}{T^{2}}\sum_{t^{\prime}=1}^{T}\sum_{t=1}^{T}\zeta _{tt^{\prime}}^{2}\right) & =\frac{1}{T^{2}}\sum_{t^{\prime}=1}^{T} \sum_{t=1}^{T}E\left( \zeta_{tt^{\prime}}^{2}\right) \nonumber\\ & =\frac{1}{T^{2}n^{2}}\sum_{t^{\prime}=1}^{T}\sum_{t=1}^{T}E\left( \sum_{j=1}^{n}\sigma_{j}^{2}\left[ \varepsilon_{jt}\left( \lambda _{T}\right) \varepsilon_{jt^{\prime}}\left( \lambda_{T}\right) -E\left( \varepsilon_{jt}\left( \lambda_{T}\right) \varepsilon_{jt^{\prime}}\left( \lambda_{T}\right) \right) \right] \right) ^{2}\nonumber\\ & =\frac{1}{T^{2}n^{2}}\sum_{t=1}^{T}\sum_{t^{\prime}=1}^{T}\sum_{j=1} ^{n}\sum_{j^{\prime}=1}^{n}\sigma_{j}^{2}\sigma_{j^{\prime}}^{2}E\left( \varepsilon_{jt}\left( \lambda_{T}\right) \varepsilon_{jt^{\prime}}\left( \lambda_{T}\right) \varepsilon_{j^{\prime}t}\left( \lambda_{T}\right) \varepsilon_{j^{\prime}t^{\prime}}\left( \lambda_{T}\right) \right) \nonumber\\ & -\frac{1}{T^{2}n^{2}}\sum_{t=1}^{T}\sum_{t^{\prime}=1}^{T}\sum_{j=1} ^{n}\sum_{j^{\prime}=1}^{n}\sigma_{j}^{2}\sigma_{j^{\prime}}^{2}E\left( \varepsilon_{jt}\left( \lambda_{T}\right) \varepsilon_{jt^{\prime}}\left( \lambda_{T}\right) \right) E\left( \varepsilon_{j^{\prime}t}\left( \lambda_{T}\right) \varepsilon_{j^{\prime}t^{\prime}}\left( \lambda _{T}\right) \right) . \end{align} For the first term of ((ref)), given the serial independence of $\varepsilon_{it}\left( \lambda_{T}\right) $ and note $E\left( \varepsilon_{it}\left( \lambda_{T}\right) \right) =E\left( \varepsilon _{it}+\lambda_{T}\mathbf{w}_{i0}^{\prime}\boldsymbol{\varepsilon}_{\circ t}\right) =0$, some algebra yields \begin{align} & \frac{1}{T^{2}n^{2}}\sum_{t=1}^{T}\sum_{t^{\prime}=1}^{T}\sum_{j=1}^{n} \sum_{j^{\prime}=1}^{n}\sigma_{j}^{2}\sigma_{j^{\prime}}^{2}E\left( \varepsilon_{jt}\left( \lambda_{T}\right) \varepsilon_{jt^{\prime}}\left( \lambda_{T}\right) \varepsilon_{j^{\prime}t}\left( \lambda_{T}\right) \varepsilon_{j^{\prime}t^{\prime}}\left( \lambda_{T}\right) \right) \nonumber\\ & =\frac{1}{T^{2}n^{2}}\sum_{t=1}^{T}\sum_{j=1}^{n}\sigma_{j}^{4}E\left( \varepsilon_{jt}^{4}\left( \lambda_{T}\right) \right) +\frac{1}{T^{2}n^{2} }\sum_{t=1}^{T}\sum_{t^{\prime}\neq t}^{T}\sum_{j=1}^{n}\sigma_{j}^{4}E\left( \varepsilon_{jt}^{2}\left( \lambda_{T}\right) \right) E\left( \varepsilon_{jt^{\prime}}^{2}\left( \lambda_{T}\right) \right) +\nonumber\\ & \frac{1}{T^{2}n^{2}}\sum_{t=1}^{T}\sum_{j=1}^{n}\sum_{j^{\prime}\neq j} ^{n}\sigma_{j}^{2}\sigma_{j^{\prime}}^{2}E\left( \varepsilon_{jt}^{2}\left( \lambda_{T}\right) \varepsilon_{j^{\prime}t}^{2}\left( \lambda_{T}\right) \right) +\nonumber\\ & \frac{1}{T^{2}n^{2}}\sum_{t=1}^{T}\sum_{t^{\prime}\neq t}^{T}\sum_{j=1} ^{n}\sum_{j^{\prime}\neq j}^{n}\sigma_{j}^{2}\sigma_{j^{\prime}}^{2}E\left( \varepsilon_{jt}\left( \lambda_{T}\right) \varepsilon_{jt^{\prime}}\left( \lambda_{T}\right) \varepsilon_{j^{\prime}t}\left( \lambda_{T}\right) \varepsilon_{j^{\prime}t^{\prime}}\left( \lambda_{T}\right) \right) . \end{align} To show the order of ((ref)), we note $\varepsilon_{it}=\mathbf{e} _{i}^{\prime}\boldsymbol{\varepsilon}_{\circ t}$ where $\mathbf{e}_{i}$ is an $n\times1$ selection vector with $1$ on its $i^{th}$ element and zero elsewhere, then \begin{align*} \varepsilon_{jt}^{4}\left( \lambda_{T}\right) & =\left( \varepsilon _{it}+\lambda_{T}\mathbf{w}_{i0}^{\prime}\boldsymbol{\varepsilon}_{\circ t}\right) ^{4}=\left[ \left( \mathbf{e}_{i}+\lambda_{T}\mathbf{w} _{i0}\right) ^{\prime}\boldsymbol{\varepsilon}_{\circ t}\right] ^{4}\\ & =\left[ \boldsymbol{\varepsilon}_{\circ t}^{\prime}\left( \mathbf{e} _{i}+\lambda_{T}\mathbf{w}_{i0}\right) \left( \mathbf{e}_{i}+\lambda _{T}\mathbf{w}_{i0}\right) ^{\prime}\boldsymbol{\varepsilon}_{\circ t}\right] ^{2}=\left( \boldsymbol{\varepsilon}_{\circ t}^{\prime} \mathbf{A}_{i}\boldsymbol{\varepsilon}_{\circ t}\right) ^{2}, \end{align*} where $\mathbf{A}_{i}=\left( \mathbf{e}_{i}+\lambda_{T}\mathbf{w} _{i0}\right) \left( \mathbf{e}_{i}+\lambda_{T}\mathbf{w}_{i0}\right) ^{\prime}.$ Using result (S.7) of Lemma 6 in Pesaran and Yamagata (2024) and noting $w_{ii}=0$ for all $i$, we have \begin{align} E\left( \varepsilon_{jt}^{4}\left( \lambda_{T}\right) \right) & =E\left( \boldsymbol{\varepsilon}_{\circ t}^{\prime}\mathbf{A}_{i} \boldsymbol{\varepsilon}_{\circ t}\right) ^{2} =\kappa_{2}\mathrm{tr}\left[ \left( \mathbf{A}_{i}\odot\mathbf{A}_{i}\right) \right] +\left[ \mathrm{tr}\left( \mathbf{A}_{i}\right) \right] ^{2}+2\mathrm{tr}\left( \mathbf{A}_{i}^{2}\right) ,\nonumber \end{align} where $\kappa_{2}=E(\varepsilon_{it}^{4})-3$. Also, by condition ((ref)) we have \[ \sum_{j=1}^{n}w_{ij}^{2}\leq\left( \sum_{j=1}^{n}\left\vert w_{ij}\right\vert \right) ^{2}<C,\sum_{j=1}^{n}w_{ij}^{4}\leq\left( \sum_{j=1}^{n}\left\vert w_{ij}\right\vert \right) ^{4}<C, \] using which yields \begin{equation} E\left( \varepsilon_{jt}^{4}\left( \lambda_{T}\right) \right) =\kappa _{2}\left( 1+\lambda_{T}^{4}\sum_{j=1}^{n}w_{ij}^{4}\right) +3\left( 1+\lambda_{T}^{2}\sum_{j=1}^{n}w_{ij}^{2}\right) =O\left( 1\right) . \end{equation} It therefore follows that $E\left( \varepsilon_{jt}^{2}\left( \lambda _{T}\right) \right) $ and $E\left( \varepsilon_{jt}^{4}\left( \lambda _{T}\right) \right) $ are bounded so that the first two terms of ((ref)) satisfy \begin{align*} \frac{1}{T^{2}n^{2}}\sum_{t=1}^{T}\sum_{j=1}^{n}\sigma_{j}^{4}E\left( \varepsilon_{jt}^{4}\left( \lambda_{T}\right) \right) & =O\left( \frac{1}{nT}\right) ,\\ \frac{1}{T^{2}n^{2}}\sum_{t=1}^{T}\sum_{t^{\prime}\neq t}^{T}\sum_{j=1} ^{n}\sigma_{j}^{4}E\left( \varepsilon_{jt}^{2}\left( \lambda_{T}\right) \right) E\left( \varepsilon_{jt^{\prime}}^{2}\left( \lambda_{T}\right) \right) & =O\left( \frac{1}{n}\right) . \end{align*} In addition, the third term of ((ref)) can be bounded using Cauchy-Schwarz inequality, \begin{align*} & \frac{1}{T^{2}n^{2}}\sum_{t=1}^{T}\sum_{j=1}^{n}\sum_{j^{\prime}\neq j} ^{n}\sigma_{j}^{2}\sigma_{j^{\prime}}^{2}E\left( \varepsilon_{jt}^{2}\left( \lambda_{T}\right) \varepsilon_{j^{\prime}t}^{2}\left( \lambda_{T}\right) \right) \\ & =\frac{1}{T^{2}n^{2}}\sum_{t=1}^{T}\sum_{j=1}^{n}\sum_{j^{\prime}\neq j}^{n}\sigma_{j}^{2}\sigma_{j^{\prime}}^{2}\left[ E\left( \varepsilon _{jt}^{4}\left( \lambda_{T}\right) \right) \right] ^{1/2}\times\left[ E\left( \varepsilon_{j^{\prime}t}^{4}\left( \lambda_{T}\right) \right) \right] ^{1/2}=O\left( \frac{1}{T}\right) . \end{align*} The fourth term of ((ref)) can be expanded based on the serial independence of $\varepsilon_{it}\left( \lambda_{T}\right) $, so that \begin{align*} & \frac{1}{T^{2}n^{2}}\sum_{t=1}^{T}\sum_{t^{\prime}\neq t}^{T}\sum_{j=1} ^{n}\sum_{j^{\prime}\neq j}^{n}\sigma_{j}^{2}\sigma_{j^{\prime}}^{2}E\left( \varepsilon_{jt}\left( \lambda_{T}\right) \varepsilon_{jt^{\prime}}\left( \lambda_{T}\right) \varepsilon_{j^{\prime}t}\left( \lambda_{T}\right) \varepsilon_{j^{\prime}t^{\prime}}\left( \lambda_{T}\right) \right) \\ & =\frac{1}{T^{2}n^{2}}\sum_{j=1}^{n}\sum_{j^{\prime}\neq j}^{n}\sigma _{j}^{2}\sigma_{j^{\prime}}^{2}\left[ \sum_{t=1}^{T}E\left( \varepsilon _{jt}\left( \lambda_{T}\right) \varepsilon_{j^{\prime}t}\left( \lambda _{T}\right) \right) \right] \left[ \sum_{t^{\prime}\neq t}^{T}E\left( \varepsilon_{jt^{\prime}}\left( \lambda_{T}\right) \varepsilon_{j^{\prime }t^{\prime}}\left( \lambda_{T}\right) \right) \right] . \end{align*} Also note by definition of $\varepsilon_{jt}\left( \lambda_{T}\right) $, for $j\neq j^{\prime}$, \[ E\left( \varepsilon_{jt}\left( \lambda_{T}\right) \varepsilon_{j^{\prime} t}\left( \lambda_{T}\right) \right) =\lambda_{T}g_{jj^{\prime}}+\lambda _{T}^{2}\mathbf{w}_{j0}^{\prime}\mathbf{w}_{j^{\prime}0}, \] where $g_{jj^{\prime}}=w_{jj^{\prime}}+w_{j^{\prime}j}$ and $\mathbf{w} _{j0}=\left( w_{j1},w_{j2},\ldots,w_{jn}\right) ^{\prime}$, so that \[ \left[ \sum_{t=1}^{T}E\left( \varepsilon_{jt}\left( \lambda_{T}\right) \varepsilon_{j^{\prime}t}\left( \lambda_{T}\right) \right) \right] \left[ \sum_{t^{\prime}\neq t}^{T}E\left( \varepsilon_{jt^{\prime}}\left( \lambda_{T}\right) \varepsilon_{j^{\prime}t^{\prime}}\left( \lambda _{T}\right) \right) \right] =T\left( T-1\right) \left( \lambda _{T}g_{jj^{\prime}}+\lambda_{T}^{2}\mathbf{w}_{j0}^{\prime}\mathbf{w} _{j^{\prime}0}\right) ^{2}. \] Then it follows \begin{align*} & \frac{1}{T^{2}n^{2}}\sum_{t=1}^{T}\sum_{t^{\prime}\neq t}^{T}\sum_{j=1} ^{n}\sum_{j^{\prime}\neq j}^{n}\sigma_{j}^{2}\sigma_{j^{\prime}}^{2}E\left( \varepsilon_{jt}\left( \lambda_{T}\right) \varepsilon_{jt^{\prime}}\left( \lambda_{T}\right) \varepsilon_{j^{\prime}t}\left( \lambda_{T}\right) \varepsilon_{j^{\prime}t^{\prime}}\left( \lambda_{T}\right) \right) \\ & =\frac{T\left( T-1\right) }{T^{2}n^{2}}\sum_{j=1}^{n}\sum_{j^{\prime}\neq j}^{n}\sigma_{j}^{2}\sigma_{j^{\prime}}^{2}\left( \lambda_{T}g_{jj^{\prime} }+\lambda_{T}^{2}\mathbf{w}_{j0}^{\prime}\mathbf{w}_{j^{\prime}0}\right) ^{2}\\ & \leq\frac{2T\left( T-1\right) }{T^{2}n^{2}}\sum_{j=1}^{n}\sum_{j^{\prime }\neq j}^{n}\sigma_{j}^{2}\sigma_{j^{\prime}}^{2}\left( \lambda _{T}g_{jj^{\prime}}\right) ^{2}+\frac{2T\left( T-1\right) }{T^{2}n^{2}} \sum_{j=1}^{n}\sum_{j^{\prime}\neq j}^{n}\sigma_{j}^{2}\sigma_{j^{\prime}} ^{2}\left( \lambda_{T}^{2}\mathbf{w}_{j0}^{\prime}\mathbf{w}_{j^{\prime} 0}\right) ^{2}. \end{align*} Since $\sum_{j^{\prime}=1}^{n}\left\vert w_{jj^{\prime}}\right\vert ^{2} \leq\left( \sum_{j^{\prime}=1}^{n}\left\vert w_{jj^{\prime}}\right\vert \right) ^{2}\leq\left\Vert \mathbf{w}_{j0}\right\Vert ^{2}<C$ \ by ((ref)), $\lambda_{T}=c_{\lambda}T^{-1/2}$ and $\left\vert \mathbf{w}_{j0}^{\prime}\mathbf{w}_{j^{\prime}0}\right\vert \leq\left\Vert \mathbf{w}_{j0}\right\Vert \left\Vert \mathbf{w}_{j^{\prime}0}\right\Vert <C$, we further have \begin{align*} \frac{2T\left( T-1\right) }{T^{2}n^{2}}\sum_{j=1}^{n}\sum_{j^{\prime}\neq j}^{n}\sigma_{j}^{2}\sigma_{j^{\prime}}^{2}\left( \lambda_{T}g_{jj^{\prime} }\right) ^{2} & =\frac{2T\left( T-1\right) \lambda_{T}^{2}}{T^{2}n^{2} }\sum_{j=1}^{n}\sum_{j^{\prime}\neq j}^{n}\sigma_{j}^{2}\sigma_{j^{\prime} }^{2}\left( w_{jj^{\prime}}+w_{j^{\prime}j}\right) ^{2}\\ & \leq\frac{4T\left( T-1\right) \lambda_{T}^{2}}{T^{2}n^{2}}\sum_{j=1} ^{n}\sum_{j^{\prime}\neq j}^{n}\sigma_{j}^{2}\sigma_{j^{\prime}}^{2}\left( \left\vert w_{jj^{\prime}}\right\vert ^{2}+\left\vert w_{j^{\prime} j}\right\vert ^{2}\right) \\ & \leq\frac{4T\left( T-1\right) \lambda_{T}^{2}}{T^{2}n^{2}}\left( \sup_{i}\sigma_{i}^{4}\right) \sum_{j=1}^{n}\sum_{j^{\prime}\neq j} ^{n}\left( \left\vert w_{jj^{\prime}}\right\vert ^{2}+\left\vert w_{j^{\prime}j}\right\vert ^{2}\right) \\ & =O\left( \frac{1}{nT}\right) , \end{align*} and \begin{align*} \frac{2T\left( T-1\right) }{T^{2}n^{2}}\sum_{j=1}^{n}\sum_{j^{\prime}\neq j}^{n}\sigma_{j}^{2}\sigma_{j^{\prime}}^{2}\left( \lambda_{T}^{2} \mathbf{w}_{j0}^{\prime}\mathbf{w}_{j^{\prime}0}\right) ^{2} & =\frac{2T\left( T-1\right) \lambda_{T}^{4}}{T^{2}n^{2}}\sum_{j=1}^{n} \sum_{j^{\prime}\neq j}^{n}\sigma_{j}^{2}\sigma_{j^{\prime}}^{2}\left( \mathbf{w}_{j0}^{\prime}\mathbf{w}_{j^{\prime}0}\right) ^{2}\\ & \leq\frac{2T\left( T-1\right) \lambda_{T}^{4}}{T^{2}}\left( \sup _{i}\sigma_{i}^{4}\right) \left( \sup_{j,j^{\prime}}\left\vert \mathbf{w}_{j0}^{\prime}\mathbf{w}_{j^{\prime}0}\right\vert \right) ^{2}\\ & =O_{p}\left( \frac{1}{T^{2}}\right) . \end{align*} Overall, using ((ref)) we are able to show the first term of ((ref)) is $O\left( T^{-1}\right) $ as $n,T\rightarrow\infty$ such that $n/T=\kappa$ where $0<\kappa<\infty$. Compared to that, the second term of ((ref)) satisfies \begin{align*} & \frac{1}{T^{2}n^{2}}\sum_{t=1}^{T}\sum_{t^{\prime}=1}^{T}\sum_{j=1}^{n} \sum_{j^{\prime}=1}^{n}\sigma_{j}^{2}\sigma_{j^{\prime}}^{2}E\left( \varepsilon_{jt}\left( \lambda_{T}\right) \varepsilon_{jt^{\prime}}\left( \lambda_{T}\right) \right) E\left( \varepsilon_{j^{\prime}t}\left( \lambda_{T}\right) \varepsilon_{j^{\prime}t^{\prime}}\left( \lambda _{T}\right) \right) \\ & =\frac{1}{T^{2}n^{2}}\sum_{t=1}^{T}\sum_{j=1}^{n}\sum_{j^{\prime}=1} ^{n}\sigma_{j}^{2}\sigma_{j^{\prime}}^{2}E\left( \varepsilon_{jt}^{2}\left( \lambda_{T}\right) \right) E\left( \varepsilon_{j^{\prime}t}^{2}\left( \lambda_{T}\right) \right) =O\left( \frac{1}{T}\right) . \end{align*} As we have shown the orders of the two terms in ((ref)), it follows that $E\left( \frac{1}{T^{2}}\sum_{t^{\prime}=1}^{T}\sum_{t=1}^{T} \zeta_{tt^{\prime}}^{2}\right) =O\left( \frac{1}{T}\right) $, so by Markov inequality $T^{-2}\sum_{t^{\prime}=1}^{T}\sum_{t=1}^{T}\zeta_{tt^{\prime}} ^{2}=O_{p}\left( \frac{1}{T}\right) $. Using this result, ((ref)), and ((ref)), it follows that \begin{align*} \sup_{i}\left\Vert \mathbf{b}_{2,1,iT}\right\Vert & \leq\left( \frac{1} {T}\sum_{t^{\prime}=1}^{T}\left\Vert \mathbf{\hat{f}}_{t^{\prime}} -\mathbf{f}_{t^{\prime}}\right\Vert ^{2}\right) ^{1/2}\left( \frac{1}{T^{2} }\sum_{t^{\prime}=1}^{T}\sum_{t=1}^{T}\zeta_{tt^{\prime}}^{2}\right) ^{1/2}\left( \sup_{i}\frac{1}{T}\sum_{t=1}^{T}\varepsilon_{it}^{2}\left( \lambda_{T}\right) \right) ^{1/2}\\ & =O_{p}\left( \frac{1}{\delta_{nT}}\right) \times O_{p}\left( \frac {1}{\sqrt{n}}\right) =O_{p}\left( \frac{1}{\sqrt{n}\delta_{nT}}\right) . \end{align*} Note that $\mathbf{b}_{2,2,iT}$ can be written as \[ \mathbf{b}_{2,2,iT}=\frac{1}{\sqrt{nT}}\frac{1}{T}\sum_{t=1}^{T}\mathbf{z} _{t}\varepsilon_{it}\left( \lambda_{T}\right) , \] where \[ \mathbf{z}_{t}=\frac{1}{\sqrt{nT}}\sum_{t^{\prime}=1}^{T}\sum_{k=1}^{n} \sigma_{k}\mathbf{f}_{t^{\prime}}\left[ \varepsilon_{kt^{\prime}}\left( \lambda_{T}\right) \varepsilon_{kt}\left( \lambda_{T}\right) -E\left( \varepsilon_{kt^{\prime}}\left( \lambda_{T}\right) \varepsilon_{kt}\left( \lambda_{T}\right) \right) \right] , \] and $\left\Vert \mathbf{z}_{t}\right\Vert ^{2}=O_{p}\left( 1\right) $. Then by Cauchy-Schwarz inequality, \[ \left\Vert \mathbf{b}_{2,2,iT}\right\Vert =\frac{1}{\sqrt{nT}}\left\Vert \frac{1}{T}\sum_{t=1}^{T}\mathbf{z}_{t}\varepsilon_{it}\left( \lambda _{T}\right) \right\Vert \leq\frac{1}{\sqrt{nT}}\left( \frac{1}{T}\sum _{t=1}^{T}\left\Vert \mathbf{z}_{t}\right\Vert ^{2}\right) ^{1/2}\left( \frac{1}{T}\sum_{t=1}^{T}\varepsilon_{it}^{2}\left( \lambda_{T}\right) \right) ^{1/2}, \] and in view of ((ref)) we have \[ \sup_{i}\left\Vert \mathbf{b}_{2,2,iT}\right\Vert \leq\frac{1}{\sqrt{nT} }\left( \frac{1}{T}\sum_{t=1}^{T}\left\Vert \mathbf{z}_{t}\right\Vert ^{2}\right) ^{1/2}\sup_{i}\left( \frac{1}{T}\sum_{t=1}^{T}\varepsilon _{it}^{2}\left( \lambda_{T}\right) \right) ^{1/2}=O_{p}\left( \frac {1}{\sqrt{nT}}\right) . \] Hence, \[ \sup_{i}\left\Vert \mathbf{b}_{2,iT}\right\Vert \leq\sup_{i}\left\Vert \mathbf{b}_{2,1,iT}\right\Vert +\sup_{i}\left\Vert \mathbf{b}_{2,2,iT} \right\Vert =O_{p}\left( \frac{1}{\sqrt{n}\delta_{nT}}\right) +O_{p}\left( \frac{1}{\sqrt{nT}}\right) . \] Now consider $\mathbf{b}_{3,iT}$ in ((ref)) and note that \begin{align*} \mathbf{b}_{3,iT} & =\frac{1}{T^{2}}\sum_{t=1}^{T}\sum_{t^{\prime}=1} ^{T}\left( \mathbf{\hat{f}}_{t^{\prime}}-\mathbf{f}_{t^{\prime}}\right) \varkappa_{tt^{\prime}}\varepsilon_{it}\left( \lambda_{T}\right) +\frac {1}{T^{2}}\sum_{t=1}^{T}\sum_{t^{\prime}=1}^{T}\mathbf{f}_{t^{\prime} }\varkappa_{tt^{\prime}}\varepsilon_{it}\left( \lambda_{T}\right) =\mathbf{b}_{3,1,iT}+\mathbf{b}_{3,2,iT}. \end{align*} To bound the first term, note that \begin{align*} \left\Vert \mathbf{b}_{3,1,iT}\right\Vert & =\left\Vert \frac{1}{T^{2}} \sum_{t^{\prime}=1}^{T}\left( \mathbf{\hat{f}}_{t^{\prime}}-\mathbf{f} _{t^{\prime}}\right) \left( \sum_{t=1}^{T}\varkappa_{tt^{\prime}} \varepsilon_{it}\left( \lambda_{T}\right) \right) \right\Vert \\ & \leq\left( \frac{1}{T}\sum_{t^{\prime}=1}^{T}\left\Vert \mathbf{\hat{f} }_{t^{\prime}}-\mathbf{f}_{t^{\prime}}\right\Vert ^{2}\right) ^{1/2}\left[ \frac{1}{T^{3}}\sum_{t^{\prime}=1}^{T}\left( \sum_{t=1}^{T}\varkappa _{tt^{\prime}}\varepsilon_{it}\left( \lambda_{T}\right) \right) ^{2}\right] ^{1/2}\\ & \leq\left( \frac{1}{T}\sum_{t^{\prime}=1}^{T}\left\Vert \mathbf{\hat{f} }_{t^{\prime}}-\mathbf{f}_{t^{\prime}}\right\Vert ^{2}\right) ^{1/2}\left( \frac{1}{T^{2}}\sum_{t^{\prime}=1}^{T}\sum_{t=1}^{T}\varkappa_{tt^{\prime} }^{2}\right) ^{1/2}\left( \frac{1}{T}\sum_{t=1}^{T}\varepsilon_{it} ^{2}\left( \lambda_{T}\right) \right) ^{1/2}. \end{align*} By definition of $\varkappa_{tt^{\prime}}$ and given ((ref)), it follows that \begin{align*} E\left( \frac{1}{T^{2}}\sum_{t=1}^{T}\sum_{t^{\prime}=1}^{T}\varkappa _{tt^{\prime}}^{2}\right) & =\frac{1}{T^{2}}\sum_{t^{\prime}=1}^{T} \sum_{t=1}^{T}E\left( \frac{1}{n}\sum_{j=1}^{n}\sigma_{j}\mathbf{f} _{t^{\prime}}^{\prime}\boldsymbol{\gamma}_{j}\varepsilon_{jt}\left( \lambda_{T}\right) \right) ^{2}\\ & =\frac{1}{T^{2}n^{2}}\sum_{t^{\prime}=1}^{T}\sum_{t=1}^{T}\sum_{j=1} ^{n}\sum_{j^{\prime}=1}^{n}\sigma_{j}\sigma_{j^{\prime}}E\left( \mathbf{f}_{t^{\prime}}^{\prime}\boldsymbol{\gamma}_{j}\mathbf{f}_{t^{\prime} }\boldsymbol{\gamma}_{j^{\prime}}\right) E\left( \varepsilon_{jt}\left( \lambda_{T}\right) \varepsilon_{j^{\prime}t}\left( \lambda_{T}\right) \right) \\ & \leq\left( \sup_{i}\left\Vert \boldsymbol{\gamma}_{i}\right\Vert ^{2}\right) \left( \sup_{t}E\left\Vert \mathbf{f}_{t}\right\Vert ^{2}\right) \frac{1}{T^{2}n^{2}}\sum_{t^{\prime}=1}^{T}\sum_{t=1}^{T} \sum_{j=1}^{n}\sum_{j^{\prime}=1}^{n}\sigma_{j}\sigma_{j^{\prime}}E\left( \varepsilon_{jt}\left( \lambda_{T}\right) \varepsilon_{j^{\prime}t}\left( \lambda_{T}\right) \right) \\ & \leq\left( \sup_{i}\left\Vert \boldsymbol{\gamma}_{i}\right\Vert ^{2}\right) \left( \sup_{t}E\left\Vert \mathbf{f}_{t}\right\Vert ^{2}\right) \left( \sup_{i}\sigma_{i}^{2}\right) \frac{1}{nT}\sum_{t=1} ^{T}\left( \frac{1}{n}\sum_{j=1}^{n}\sum_{j^{\prime}=1}^{n}\left\vert E\left( \varepsilon_{jt}\left( \lambda_{T}\right) \varepsilon_{j^{\prime} t}\left( \lambda_{T}\right) \right) \right\vert \right) \\ & =O\left( \frac{1}{n}\right) , \end{align*} using which and Markov inequality yields $T^{-2}\sum_{t^{\prime}=1}^{T} \sum_{t=1}^{T}\varkappa_{tt^{\prime}}^{2}=O_{p}\left( n^{-1}\right) $. Then given this result and ((ref)), ((ref)), it follows \begin{align*} \sup_{i}\left\Vert \mathbf{b}_{3,1,iT}\right\Vert & =\left( \frac{1}{T} \sum_{t^{\prime}=1}^{T}\left\Vert \mathbf{\hat{f}}_{t^{\prime}}-\mathbf{f} _{t^{\prime}}\right\Vert ^{2}\right) ^{1/2}\left( \frac{1}{T^{2}} \sum_{t^{\prime}=1}^{T}\sum_{t=1}^{T}\varkappa_{tt^{\prime}}^{2}\right) ^{1/2}\left( \sup_{i}\frac{1}{T}\sum_{t=1}^{T}\varepsilon_{it}^{2}\left( \lambda_{T}\right) \right) ^{1/2}\\ & =O_{p}\left( \frac{1}{\delta_{nT}}\right) \times O_{p}\left( \frac {1}{\sqrt{n}}\right) =O_{p}\left( \frac{1}{\sqrt{n}\delta_{nT}}\right) . \end{align*} Next, we consider $\left\Vert \mathbf{b}_{3,2,iT}\right\Vert $ and observe that \[ \frac{1}{T^{2}}\sum_{t=1}^{T}\sum_{t^{\prime}=1}^{T}\mathbf{f}_{t^{\prime} }\varkappa_{tt^{\prime}}\varepsilon_{it}\left( \lambda_{T}\right) =\left( \frac{1}{T}\sum_{t^{\prime}=1}^{T}\mathbf{f}_{t^{\prime}}\mathbf{f} _{t^{\prime}}^{\prime}\right) \left( \frac{1}{Tn}\sum_{t=1}^{T}\sum _{j=1}^{n}\sigma_{j}\boldsymbol{\gamma}_{j}\varepsilon_{jt}\left( \lambda _{T}\right) \varepsilon_{it}\left( \lambda_{T}\right) \right) , \] using which yields \[ \left\Vert \mathbf{b}_{3,2,iT}\right\Vert =\left\Vert \frac{1}{T^{2}} \sum_{t=1}^{T}\sum_{t^{\prime}=1}^{T}\mathbf{f}_{t^{\prime}}\varkappa _{tt^{\prime}}\varepsilon_{it}\left( \lambda_{T}\right) \right\Vert \leq\left( \frac{1}{T}\sum_{t^{\prime}=1}^{T}\left\Vert \mathbf{f} _{t^{\prime}}\mathbf{f}_{t^{\prime}}^{\prime}\right\Vert \right) \left( \left\Vert \frac{1}{nT}\sum_{t=1}^{T}\sum_{j=1}^{n}\sigma_{j} \boldsymbol{\gamma}_{j}\varepsilon_{jt}\left( \lambda_{T}\right) \varepsilon_{it}\left( \lambda_{T}\right) \right\Vert \right) . \] Then given ((ref)), \[ \sup_{i}\left\Vert \mathbf{b}_{3,2,iT}\right\Vert \leq\left( \frac{1}{T} \sum_{t^{\prime}=1}^{T}\left\Vert \mathbf{f}_{t^{\prime}}\mathbf{f} _{t^{\prime}}^{\prime}\right\Vert \right) \left( \sup_{i}\left\Vert \frac {1}{nT}\sum_{t=1}^{T}\sum_{j=1}^{n}\sigma_{j}\boldsymbol{\gamma} _{j}\varepsilon_{jt}\left( \lambda_{T}\right) \varepsilon_{it}\left( \lambda_{T}\right) \right\Vert \right) =O_{p}\left( \sqrt{\frac{\ln\left( n\right) }{nT}}\right) . \] Hence, $\sup_{i}\left\Vert \mathbf{b}_{3,iT}\right\Vert \leq\sup_{i}\left\Vert \mathbf{b}_{3,1,iT}\right\Vert +\sup_{i}\left\Vert \mathbf{b}_{3,2,iT} \right\Vert =O_{p}\left( \sqrt{\frac{\ln\left( n\right) }{nT}}\right) . $ Similarly, the probability order of $\sup_{i}\left\Vert \mathbf{b} _{4,iT}\right\Vert $ can also be shown to be $O_{p}\left( \sqrt{\ln\left( n\right) /\left( nT\right) }\right) $. Overall, result ((ref)) follows as we can use ((ref)) to show \begin{align*} \sup_{i}\left\Vert \frac{\left( \mathbf{\hat{F}-F}\right) ^{\prime }\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) }{T}\right\Vert & \leq\sup_{i}\left( \sum_{j=1}^{4}\left\Vert \mathbf{b}_{j,iT}\right\Vert \right) \leq\sum_{j=1}^{4}\sup_{i}\left\Vert \mathbf{b}_{j,iT}\right\Vert \\ & =O_{p}\left( \frac{1}{\sqrt{T}\delta_{nT}}\right) +O_{p}\left( \frac {1}{\sqrt{n}\delta_{nT}}\right) +O_{p}\left( \frac{1}{\sqrt{nT}}\right) +O_{p}\left( \sqrt{\frac{\ln\left( n\right) }{nT}}\right) \\ & =O_{p}\left( \sqrt{\frac{\ln\left( n\right) }{nT}}\right) . \end{align*}
lemmaDenote $\boldsymbol{\varepsilon}_{i\circ}=\left( \varepsilon_{i1},\varepsilon_{i2},\ldots,\varepsilon_{iT}\right) ^{\prime}$ and $\mathbf{b}_{i}=\left( b_{i1},b_{i2},\ldots,b_{iT}\right) ^{\prime}$ with $b_{it}=\mathbf{w}_{i0}^{\prime}\boldsymbol{\varepsilon}_{\circ t}$, $\mathbf{w}_{i0}=\left( w_{i1},w_{is},\ldots,w_{in}\right) ^{\prime}$ and $\boldsymbol{\varepsilon}_{\circ t}=\left( \varepsilon_{1t},\varepsilon _{2t},\ldots,\varepsilon_{nt}\right) ^{\prime}$. Suppose that Assumptions (ref) and (ref) hold. Then as $\left( n,T\right) \rightarrow\infty$, such that $n/T\rightarrow\kappa\,,$ for $0<\kappa<\infty $, \begin{align} \sup_{i}\left\vert \frac{\boldsymbol{\varepsilon}_{i\circ}^{\prime} \mathbf{b}_{i}}{T}\right\vert & =O_{p}\left( \sqrt{\frac{\ln\left( n\right) }{T}}\right) ,\\ \sup_{i}\left\vert \frac{\mathbf{b}_{i}^{\prime}\mathbf{b}_{i}}{T}\right\vert & =O_{p}\left( 1\right) . \end{align}
proofConsider ((ref)) and denote $\mathcal{I}_{i,t-1}^{\left( \varepsilon b\right) }=\left\{ \varepsilon_{i\tau}b_{i\tau}:\tau =t-1,t-2,\ldots\right\} $ and $i=1,2,\ldots,n$. By assumption, $\varepsilon _{it}$ is cross-sectionally independent, and is independent from $b_{it}$ as $w_{ii}=0$ for all $i$. Then $E\left( \varepsilon_{it}b_{it}|\mathcal{I} _{i,t-1}^{\left( \varepsilon b\right) }\right) =$ $E\left( \varepsilon _{it}b_{it}\right) =0$. Also $Var\left( \varepsilon_{it}b_{it}\right) =Var\left( \varepsilon_{it}\right) Var\left( b_{it}\right) =E\left( b_{it}^{2}\right) , $ which is bounded as by condition ((ref)), \begin{equation} E\left( b_{it}^{2}\right) =E\left( \mathbf{w}_{i0}^{\prime} \boldsymbol{\varepsilon}_{\circ t}\right) ^{2}=\mathbf{w}_{i0}^{\prime }E\left( \boldsymbol{\varepsilon}_{\circ t}\boldsymbol{\varepsilon}_{\circ t}^{\prime}\right) \mathbf{w}_{i0}=\sum_{s=1}^{n}w_{is}^{2}<C. \end{equation} In addition, since $\varepsilon_{it}\left( \lambda_{T}\right) =$ $\varepsilon_{it}+\lambda_{T}b_{it}$ is sub-exponential for any $\left\vert \lambda_{T}\right\vert <C$, then it follows that $\varepsilon_{it}$ and $b_{it}$ are both sub-exponential, and hence $\varepsilon_{it}b_{it}$ is also sub-exponential. Hence, result ((ref)) can be established by applying the method of proof used for result ((ref)). Now consider ((ref)) and note \[ \frac{\mathbf{b}_{i}^{\prime}\mathbf{b}_{i}}{T}=\frac{1}{T}\sum_{t=1} ^{T}\left[ b_{it}^{2}-E\left( b_{it}^{2}\right) \right] +\frac{1}{T} \sum_{t=1}^{T}E\left( b_{it}^{2}\right) , \] which further implies \[ \sup_{i}\left\vert \frac{\mathbf{b}_{i}^{\prime}\mathbf{b}_{i}}{T}\right\vert \leq\sup_{i}\left\vert \frac{1}{T}\sum_{t=1}^{T}\left[ b_{it}^{2}-E\left( b_{it}^{2}\right) \right] \right\vert +\sup_{i}\left( \frac{1}{T}\sum _{t=1}^{T}E\left( b_{it}^{2}\right) \right) . \] Denote $\mathcal{I}_{i,t-1}^{\left( b\right) }=\left\{ b_{i\tau} :\tau=t-1,t-2,\ldots\right\} $ and $i=1,2,\ldots,n$, then $E\left( b_{it}^{2}-E\left( b_{it}^{2}\right) |\mathcal{I}_{i,t-1}^{\left( b\right) }\right) =$ $E\left( b_{it}^{2}-E\left( b_{it}^{2}\right) \right) =0$. Besides, $Var\left( b_{it}^{2}\right) =E\left( b_{it}^{4}\right) -\left[ E\left( b_{it}^{2}\right) \right] ^{2}$ is bounded as ((ref)) shows $E\left( b_{it}^{4}\right) =E\left( \mathbf{w}_{i0}^{\prime} \boldsymbol{\varepsilon}_{\circ t}\right) ^{4}=O\left( 1\right) $. Also $b_{it}^{2}-E\left( b_{it}^{2}\right) $ is sub-exponential given that it is already established that $b_{it}$ is sub-exponential. Therefore, it follows that $\sup_{i}\left\vert \frac{1}{T}\sum_{t=1}^{T}\left[ b_{it}^{2}-E\left( b_{it}^{2}\right) \right] \right\vert =O_{p}\left( \sqrt{\frac{\ln\left( n\right) }{T}}\right) $ by applying the method of proof used for result ((ref)). Also by ((ref)), we have $\sup_{i}\left( \frac{1} {T}\sum_{t=1}^{T}E\left( b_{it}^{2}\right) \right) =O\left( 1\right) $. Result ((ref)) now follows straightforwardly.
lemmaConsider the latent factor model given by ((ref)) and ((ref)). Let $\hat{\sigma}_{i,T}=\left( T^{-1}\mathbf{e}_{i}^{\prime }\mathbf{e}_{i}\right) ^{1/2}$, where $\mathbf{e}_{i}=\mathbf{M}_{\hat{F} }\mathbf{y}_{i}$, $\mathbf{M}_{\hat{F}}=\mathbf{I}_{T}-\mathbf{\hat{F}(\hat {F}}^{\prime}\mathbf{\hat{F})}^{-1}\mathbf{\hat{F}}^{\prime}$, $\mathbf{y} _{i}=(y_{i1},y_{i2},...,y_{iT})^{\prime}$, and $\mathbf{\hat{F}}$ is given by ((ref)). Also let $\omega_{i,T}=\left( T^{-1}\sigma_{i}^{2} \boldsymbol{\varepsilon}_{i\circ}^{^{\prime}}\mathbf{M}_{F} \boldsymbol{\varepsilon}_{i\circ}\right) ^{1/2},$ where $\mathbf{M} _{F}=\mathbf{I}_{T}-\mathbf{F(F}^{\prime}\mathbf{F)}^{-1}\mathbf{F}^{\prime}$. Suppose that Assumptions (ref)-(ref) hold and $\left( n,T\right) \rightarrow\infty$, such that $n/T\rightarrow\kappa\,,$ for $0<\kappa<\infty$. Then \begin{align} \sup_{i}\left\vert \hat{\sigma}_{i,T}^{2}-\omega_{i,T}^{2}\right\vert & =O_{p}\left( \frac{\ln\left( n\right) }{T}\right) , \\ \sup_{i}\left\vert \hat{\sigma}_{i,T}-\omega_{i,T}\right\vert & =O_{p}\left( \frac{\ln\left( n\right) }{T}\right) ,\\ \sup_{i}\left\vert \frac{1}{\hat{\sigma}_{i,T}}-\frac{1}{\omega_{i,T} }\right\vert & =O_{p}\left( \frac{\ln\left( n\right) }{T}\right) , \end{align} and \begin{align} \frac{1}{n}\sum_{i=1}^{n}\left\vert \hat{\sigma}_{i,T}^{2}-\omega_{i,T} ^{2}\right\vert & =O_{p}\left( \frac{\ln\left( n\right) }{T}\right) ,\\ \frac{1}{n}\sum_{i=1}^{n}\left\vert \hat{\sigma}_{i,T}-\omega_{i,T} \right\vert & =O_{p}\left( \frac{\ln\left( n\right) }{T}\right) ,\\ \frac{1}{n}\sum_{i=1}^{n}\left\vert \frac{1}{\hat{\sigma}_{i,T}}-\frac {1}{\omega_{i,T}}\right\vert & =O_{p}\left( \frac{\ln\left( n\right) } {T}\right) . \end{align}
proofNote that by ((ref)), $\mathbf{y}_{i}=\mathbf{F}\boldsymbol{\gamma} _{i}+\sigma_{i}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) $, where $\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) =\left( \varepsilon_{i1}\left( \lambda_{T}\right) ,\varepsilon_{i2}\left( \lambda_{T}\right) ,\ldots,\varepsilon_{iT}\left( \lambda_{T}\right) \right) ^{\prime}$, which in turn implies \begin{align*} \mathbf{e}_{i} & =\mathbf{M}_{\hat{F}}\mathbf{y}_{i}=\mathbf{M}_{\hat{F} }\left( \mathbf{F}\boldsymbol{\gamma}_{i}+\sigma_{i}\boldsymbol{\varepsilon }_{i\circ}\left( \lambda_{T}\right) \right) \\ & =\sigma_{i}\mathbf{M}_{F}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) +\sigma_{i}\left( \mathbf{M}_{\hat{F}}-\mathbf{M} _{F}\right) \boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) +\mathbf{M}_{\hat{F}}\mathbf{F}\boldsymbol{\gamma}_{i}. \end{align*} Then $\hat{\sigma}_{i,T}^{2}$ can be decomposed as \begin{align} \hat{\sigma}_{i,T}^{2}-\omega_{i,T}^{2} & =\frac{\boldsymbol{\gamma} _{i}^{^{\prime}}\mathbf{F}^{^{\prime}}\mathbf{M}_{\hat{F}}\mathbf{F} \boldsymbol{\gamma}_{i}}{T}+\frac{\sigma_{i}^{2}\boldsymbol{\varepsilon }_{i\circ}\left( \lambda_{T}\right) ^{\prime}\left( \mathbf{M}_{\hat{F} }-\mathbf{M}_{F}\right) \left( \mathbf{M}_{\hat{F}}-\mathbf{M}_{F}\right) \boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) }{T}\nonumber\\ & +\frac{2\sigma_{i}^{2}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda _{T}\right) ^{\prime}\mathbf{M}_{F}\left( \mathbf{M}_{\hat{F}} -\mathbf{M}_{F}\right) \boldsymbol{\varepsilon}_{i\circ}\left( \lambda _{T}\right) }{T}+\frac{2\sigma_{i}^{2}\boldsymbol{\varepsilon}_{i\circ }\left( \lambda_{T}\right) ^{\prime}\mathbf{M}_{F}\mathbf{M}_{\hat{F} }\mathbf{F}\boldsymbol{\gamma}_{i}}{T}\nonumber\\ & +\frac{2\sigma_{i}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda _{T}\right) ^{\prime}\left( \mathbf{M}_{\hat{F}}-\mathbf{M}_{F}\right) \mathbf{M}_{\hat{F}}\mathbf{F}\boldsymbol{\gamma}_{i}}{T}+\frac{\sigma_{i} ^{2}\left( \boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) +\boldsymbol{\varepsilon}_{i\circ}\right) ^{\prime}\mathbf{M}_{F}\left( \boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) -\boldsymbol{\varepsilon}_{i\circ}\right) }{T}\nonumber\\ & =\sum_{j=1}^{6}B_{j,iT}, \end{align} and $\left\vert \hat{\sigma}_{i,T}^{2}-\omega_{i,T}^{2}\right\vert \leq \sum_{j=1}^{6}\left\vert B_{j,iT}\right\vert $. Starting with $B_{1,iT}$, note that \[ \left\vert B_{1,iT}\right\vert =\left\vert \frac{\boldsymbol{\gamma} _{i}^{^{\prime}}\mathbf{F}^{^{\prime}}\mathbf{M}_{\hat{F}}\mathbf{F} \boldsymbol{\gamma}_{i}}{T}\right\vert =\left\vert \frac{\boldsymbol{\gamma }_{i}^{^{\prime}}\left( \mathbf{F-\hat{F}}\right) ^{^{\prime}} \mathbf{M}_{\hat{F}}\left( \mathbf{F-\hat{F}}\right) \boldsymbol{\gamma} _{i}}{T}\right\vert \leq\left\Vert \boldsymbol{\gamma}_{i}\right\Vert ^{2}\left\Vert \mathbf{M}_{\hat{F}}\right\Vert \left( \frac{\left\Vert \mathbf{\hat{F}-F}\right\Vert ^{2}}{T}\right) . \] Also $\sup_{i}\left\Vert \boldsymbol{\gamma}_{i}\right\Vert <C$ and $\left\Vert \mathbf{M}_{\hat{F}}\right\Vert =1$, and using ((ref)) of Lemma (ref) we have \[ \sup_{i}\left\vert B_{1,iT}\right\vert \leq\left( \sup_{i}\left\Vert \boldsymbol{\gamma}_{i}\right\Vert ^{2}\right) \left\Vert \mathbf{M}_{\hat {F}}\right\Vert \left( \frac{\left\Vert \mathbf{\hat{F}-F}\right\Vert ^{2} }{T}\right) =O_{p}\left( \frac{1}{\delta_{nT}^{2}}\right) . \] To establish the probability orders of the remaining terms of ((ref)), we first observe that \[ \frac{\mathbf{\hat{F}}^{^{\prime}}\mathbf{\hat{F}}}{T}-\frac{\mathbf{F} ^{^{\prime}}\mathbf{F}}{T}=\frac{\left( \mathbf{\hat{F}-F}\right) ^{^{\prime}}\left( \mathbf{\hat{F}-F}\right) }{T}+\frac{\left( \mathbf{\hat{F}-F}\right) ^{^{\prime}}\mathbf{F}}{T}+\frac{\mathbf{F} ^{^{\prime}}\left( \mathbf{\hat{F}-F}\right) }{T}. \] Using results ((ref)) and ((ref)) it follows that \begin{equation} \left\Vert \frac{\mathbf{\hat{F}}^{^{\prime}}\mathbf{\hat{F}}}{T} -\frac{\mathbf{F}^{^{\prime}}\mathbf{F}}{T}\right\Vert \leq\frac{\left\Vert \mathbf{\hat{F}-F}\right\Vert ^{2}}{T}+\frac{\left\Vert \left( \mathbf{\hat {F}-F}\right) ^{\prime}\mathbf{F}\right\Vert }{T}+\frac{\left\Vert \mathbf{F}^{\prime}\left( \mathbf{\hat{F}-F}\right) \right\Vert }{T} =O_{p}\left( \frac{1}{\delta_{nT}^{2}}\right) . \end{equation} By assumption $T^{-1}\mathbf{F}^{^{\prime}}\mathbf{F}$ is a positive definite matrix, then \[ \left\Vert \left( \frac{\mathbf{\hat{F}}^{^{\prime}}\mathbf{\hat{F}}} {T}\right) ^{-1}-\left( \frac{\mathbf{F}^{^{\prime}}\mathbf{F}}{T}\right) ^{-1}\right\Vert \leq\left\Vert \left( \frac{\mathbf{\hat{F}}^{^{\prime} }\mathbf{\hat{F}}}{T}\right) ^{-1}\right\Vert \left\Vert \frac{\mathbf{\hat {F}}^{^{\prime}}\mathbf{\hat{F}}}{T}-\frac{\mathbf{F}^{^{\prime}}\mathbf{F} }{T}\right\Vert \left\Vert \left( \frac{\mathbf{F}^{^{\prime}}\mathbf{F}} {T}\right) ^{-1}\right\Vert =O_{p}\left( \frac{1}{\delta_{nT}^{2}}\right) , \] so it follows that \begin{equation} \frac{\mathbf{\hat{F}}^{^{\prime}}\mathbf{\hat{F}}}{T}=\frac{\mathbf{F} ^{^{\prime}}\mathbf{F}}{T}+O_{p}\left( \frac{1}{\delta_{nT}^{2}}\right) , and \left( \frac{\mathbf{\hat{F}}^{^{\prime}}\mathbf{\hat{F}}} {T}\right) ^{-1}=\left( \frac{\mathbf{F}^{^{\prime}}\mathbf{F}}{T}\right) ^{-1}+O_{p}\left( \frac{1}{\delta_{nT}^{2}}\right) . \end{equation} Now we consider $B_{2,iT},$ and note that \begin{align} B_{2,iT} & =T^{-1}\sigma_{i}^{2}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) ^{\prime}\left[ \mathbf{\hat{F}}\left( \mathbf{\hat{F} }^{^{\prime}}\mathbf{\hat{F}}\right) ^{-1}\mathbf{\hat{F}}^{^{\prime} }-\mathbf{F}\left( \mathbf{F}^{^{\prime}}\mathbf{F}\right) ^{-1} \mathbf{F}^{^{\prime}}\right] \boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) +\nonumber\\ & T^{-1}\sigma_{i}^{2}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda _{T}\right) ^{\prime}\left( \mathbf{I}_{m_{0}}-\mathbf{\hat{F}}\left( \mathbf{\hat{F}}^{^{\prime}}\mathbf{\hat{F}}\right) ^{-1}\mathbf{\hat{F} }^{^{\prime}}\right) \mathbf{F}\left( \mathbf{F}^{^{\prime}}\mathbf{F} \right) ^{-1}\mathbf{F}^{^{\prime}}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) +\nonumber\\ & T^{-1}\sigma_{i}^{2}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda _{T}\right) ^{\prime}\mathbf{F}\left( \mathbf{F}^{^{\prime}}\mathbf{F} \right) ^{-1}\mathbf{F}^{^{\prime}}\left( \mathbf{I}_{m_{0}}-\mathbf{\hat {F}}\left( \mathbf{\hat{F}}^{^{\prime}}\mathbf{\hat{F}}\right) ^{-1}\mathbf{\hat{F}}^{^{\prime}}\right) \boldsymbol{\varepsilon}_{i\circ }\left( \lambda_{T}\right) . \end{align} We further note \begin{align*} \mathbf{\hat{F}}\left( \mathbf{\hat{F}}^{^{\prime}}\mathbf{\hat{F}}\right) ^{-1}\mathbf{\hat{F}}^{^{\prime}}-\mathbf{F}\left( \mathbf{F}^{^{\prime} }\mathbf{F}\right) ^{-1}\mathbf{F}^{^{\prime}} & =\left( \mathbf{\hat {F}-F}\right) \left( \mathbf{\hat{F}}^{^{\prime}}\mathbf{\hat{F}}\right) ^{-1}\left( \mathbf{\hat{F}-F}\right) ^{^{\prime}}+\left[ \mathbf{F}\left( \mathbf{\hat{F}}^{^{\prime}}\mathbf{\hat{F}}\right) ^{-1}\mathbf{F} ^{^{\prime}}-\mathbf{F}\left( \mathbf{F}^{^{\prime}}\mathbf{F}\right) ^{-1}\mathbf{F}^{^{\prime}}\right] \\ & +\mathbf{F}\left( \mathbf{\hat{F}}^{^{\prime}}\mathbf{\hat{F}}\right) ^{-1}\left( \mathbf{\hat{F}-F}\right) ^{^{\prime}}+\left( \mathbf{\hat {F}-F}\right) \left( \mathbf{\hat{F}}^{^{\prime}}\mathbf{\hat{F}}\right) ^{-1}\mathbf{F}^{\prime}. \end{align*} Using ((ref)), we have $B_{2,iT}=\sum_{j=1}^{6}B_{2,j,iT}$, where \begin{align*} B_{2,1,iT} & =T^{-1}\sigma_{i}^{2}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) ^{\prime}\left( \mathbf{\hat{F}-F}\right) \left( \mathbf{\hat{F}}^{^{\prime}}\mathbf{\hat{F}}\right) ^{-1}\left( \mathbf{\hat{F}-F}\right) ^{\prime}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) ,\\ B_{2,2,iT} & =T^{-1}\sigma_{i}^{2}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) ^{\prime}\left[ \mathbf{F}\left( \mathbf{\hat{F} }^{^{\prime}}\mathbf{\hat{F}}\right) ^{-1}\mathbf{F}^{^{\prime}} -\mathbf{F}\left( \mathbf{F}^{^{\prime}}\mathbf{F}\right) ^{-1} \mathbf{F}^{^{\prime}}\right] \boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) ,\\ B_{2,3,iT} & =B_{2,4,iT}=T^{-1}\sigma_{i}^{2}\boldsymbol{\varepsilon }_{i\circ}\left( \lambda_{T}\right) ^{\prime}\mathbf{F}\left( \mathbf{\hat{F}}^{^{\prime}}\mathbf{\hat{F}}\right) ^{-1}\left( \mathbf{\hat{F}-F}\right) ^{\prime}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) ,\\ B_{2,5,iT} & =B_{2,6,iT}=T^{-1}\sigma_{i}^{2}\boldsymbol{\varepsilon }_{i\circ}\left( \lambda_{T}\right) ^{\prime}\left( \mathbf{I}_{m_{0} }-\mathbf{\hat{F}}\left( \mathbf{\hat{F}}^{^{\prime}}\mathbf{\hat{F}}\right) ^{-1}\mathbf{\hat{F}}^{^{\prime}}\right) \mathbf{F}\left( \mathbf{F} ^{^{\prime}}\mathbf{F}\right) ^{-1}\mathbf{F}^{\prime}\boldsymbol{\varepsilon }_{i\circ}\left( \lambda_{T}\right) , \end{align*} and $B_{2,iT}\leq\sum_{j=1}^{6}\left\vert B_{2,j,iT}\right\vert $. Starting with the first term we note that \begin{align*} \left\vert B_{2,1,iT}\right\vert & =\left\vert T^{-1}\sigma_{i} ^{2}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) ^{\prime }\left( \mathbf{\hat{F}-F}\right) \left( \mathbf{\hat{F}}^{^{\prime} }\mathbf{\hat{F}}\right) ^{-1}\left( \mathbf{\hat{F}-F}\right) ^{\prime }\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) \right\vert \\ & \leq\sigma_{i}^{2}\left( \frac{1}{T}\left\Vert \boldsymbol{\varepsilon }_{i\circ}\left( \lambda_{T}\right) \right\Vert ^{2}\right) \left( \frac{\left\Vert \mathbf{\hat{F}-F}\right\Vert ^{2}}{T}\right) \left\Vert \left( \frac{\mathbf{\hat{F}}^{^{\prime}}\mathbf{\hat{F}}}{T}\right) ^{-1}\right\Vert . \end{align*} By ((ref)) and ((ref)) we have \[ \sup_{i}\left\vert B_{2,1,iT}\right\vert \leq\left( \sup_{i}\sigma_{i} ^{2}\right) \left( \sup_{i}\frac{1}{T}\left\Vert \boldsymbol{\varepsilon }_{i\circ}\left( \lambda_{T}\right) \right\Vert ^{2}\right) \left( \frac{\left\Vert \mathbf{\hat{F}-F}\right\Vert ^{2}}{T}\right) \left\Vert \left( \frac{\mathbf{\hat{F}}^{^{\prime}}\mathbf{\hat{F}}}{T}\right) ^{-1}\right\Vert =O_{p}\left( \frac{1}{\delta_{nT}^{2}}\right) . \] Next, consider $B_{2,2,iT}$ and note that \begin{align*} \left\vert B_{2,2,iT}\right\vert & =\left\vert \frac{1}{T}\sigma_{i} ^{2}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) ^{\prime }\left[ \mathbf{F}\left( \mathbf{\hat{F}}^{^{\prime}}\mathbf{\hat{F} }\right) ^{-1}\mathbf{F}^{^{\prime}}-\mathbf{F}\left( \mathbf{F}^{^{\prime} }\mathbf{F}\right) ^{-1}\mathbf{F}^{^{\prime}}\right] \boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) \right\vert \\ & =\sigma_{i}^{2}\left\vert \left( \frac{\boldsymbol{\varepsilon}_{i\circ }\left( \lambda_{T}\right) ^{\prime}\mathbf{F}}{T}\right) \left[ \left( \frac{\mathbf{\hat{F}}^{^{\prime}}\mathbf{\hat{F}}}{T}\right) ^{-1}-\left( \frac{\mathbf{F}^{^{\prime}}\mathbf{F}}{T}\right) ^{-1}\right] \left( \frac{\mathbf{F}^{^{\prime}}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) }{T}\right) \right\vert \\ & \leq\sigma_{i}^{2}\left( \left\Vert \frac{\boldsymbol{\varepsilon} _{i\circ}\left( \lambda_{T}\right) ^{\prime}\mathbf{F}}{T}\right\Vert ^{2}\right) \left\Vert \left( \frac{\mathbf{\hat{F}}^{^{\prime}} \mathbf{\hat{F}}}{T}\right) ^{-1}-\left( \frac{\mathbf{F}^{^{\prime} }\mathbf{F}}{T}\right) ^{-1}\right\Vert . \end{align*} Using ((ref)) (from Lemma (ref)) and ((ref))), we further have \begin{align*} \sup_{i}\left\vert B_{2,2,iT}\right\vert & \leq\left( \sup_{i}\sigma _{i}^{2}\right) \left( \sup_{i}\left\Vert \frac{\boldsymbol{\varepsilon }_{i\circ}\left( \lambda_{T}\right) ^{\prime}\mathbf{F}}{T}\right\Vert ^{2}\right) \left\Vert \left( \frac{\mathbf{\hat{F}}^{^{\prime}} \mathbf{\hat{F}}}{T}\right) ^{-1}-\left( \frac{\mathbf{F}^{^{\prime} }\mathbf{F}}{T}\right) ^{-1}\right\Vert \\ & =O_{p}\left( \frac{\ln\left( n\right) }{T}\right) \times O_{p}\left( \frac{1}{\delta_{nT}^{2}}\right) =O_{p}\left( \frac{\ln\left( n\right) }{T\delta_{nT}^{2}}\right) . \end{align*} Next, consider \begin{align*} \left\vert B_{2,3,iT}\right\vert & =\sigma_{i}^{2}\left\vert \left( \frac{\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) ^{\prime }\mathbf{F}}{T}\right) \left( \frac{\mathbf{\hat{F}}^{^{\prime}} \mathbf{\hat{F}}}{T}\right) ^{-1}\left[ \frac{\left( \mathbf{\hat{F} -F}\right) ^{\prime}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda _{T}\right) }{T}\right] \right\vert \\ & \leq\sigma_{i}^{2}\left\Vert \frac{\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) ^{\prime}\mathbf{F}}{T}\right\Vert \left\Vert \left( \frac{\mathbf{\hat{F}}^{^{\prime}}\mathbf{\hat{F}}}{T}\right) ^{-1} \right\Vert \left\Vert \frac{\left( \mathbf{\hat{F}-F}\right) ^{\prime }\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) }{T}\right\Vert , \end{align*} and using ((ref)), ((ref)) and ((ref)) now yields \begin{align*} \sup_{i}\left\vert B_{2,3,iT}\right\vert & \leq\left\Vert \left( \frac{\mathbf{\hat{F}}^{^{\prime}}\mathbf{\hat{F}}}{T}\right) ^{-1} \right\Vert \left( \sup_{i}\left\Vert \frac{\boldsymbol{\varepsilon}_{i\circ }\left( \lambda_{T}\right) ^{\prime}\mathbf{F}}{T}\right\Vert \right) \left( \sup_{i}\left\Vert \frac{\left( \mathbf{\hat{F}-F}\right) ^{\prime }\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) }{T}\right\Vert \right) \\ & =O_{p}\left( \sqrt{\frac{\ln\left( n\right) }{T}}\right) \times O_{p}\left( \sqrt{\frac{\ln\left( n\right) }{nT}}\right) . \end{align*} Similarly, \begin{align} \left\vert B_{2,5,iT}\right\vert & =\left\vert T^{-1}\sigma_{i} ^{2}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) ^{\prime }\left( \mathbf{I}_{m_{0}}-\mathbf{\hat{F}}\left( \mathbf{\hat{F}} ^{^{\prime}}\mathbf{\hat{F}}\right) ^{-1}\mathbf{\hat{F}}^{^{\prime}}\right) \mathbf{F}\left( \mathbf{F}^{^{\prime}}\mathbf{F}\right) ^{-1} \mathbf{F}^{\prime}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda _{T}\right) \right\vert \nonumber\\ & \leq\sigma_{i}^{2}\left\vert \frac{\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) ^{\prime}\mathbf{P}_{F}\boldsymbol{\varepsilon}_{i\circ }\left( \lambda_{T}\right) }{T}\right\vert +\sigma_{i}^{2}\left\vert \frac{\boldsymbol{\varepsilon}_{i\circ}^{\prime}\left( \lambda_{T}\right) \mathbf{P}_{F}\left( \mathbf{P}_{\hat{F}}-\mathbf{P}_{F}\right) \boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) }{T}\right\vert , \end{align} where $\mathbf{P}_{F}=\mathbf{F}\left( \mathbf{F}^{^{\prime}}\mathbf{F} \right) ^{-1}\mathbf{F}^{^{\prime}}$ and $\mathbf{P}_{\hat{F}}=\mathbf{\hat {F}}\left( \mathbf{\hat{F}}^{^{\prime}}\mathbf{\hat{F}}\right) ^{-1}\mathbf{\hat{F}}^{^{\prime}}$. Further \[ \left\vert \frac{\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) ^{\prime}\mathbf{P}_{F}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda \right) }{T}\right\vert =\left( \frac{\boldsymbol{\varepsilon}_{i\circ }\left( \lambda_{T}\right) ^{\prime}\mathbf{F}}{T}\right) \left( \frac{\mathbf{F}^{^{\prime}}\mathbf{F}}{T}\right) ^{-1}\left( \frac {\mathbf{F}^{^{\prime}}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda _{T}\right) }{T}\right) \leq\left\Vert \left( \frac{\mathbf{F}^{^{\prime} }\mathbf{F}}{T}\right) ^{-1}\right\Vert \left\Vert \frac {\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) ^{\prime }\mathbf{F}}{T}\right\Vert ^{2}, \] and using ((ref)) it follows that \[ \sup_{i}\left\vert \frac{\boldsymbol{\varepsilon}_{i\circ}\left( \lambda _{T}\right) ^{\prime}\mathbf{P}_{F}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) }{T}\right\vert \leq\left\Vert \left( \frac {\mathbf{F}^{^{\prime}}\mathbf{F}}{T}\right) ^{-1}\right\Vert \left( \sup_{i}\left\Vert \frac{\boldsymbol{\varepsilon}_{i\circ}\left( \lambda _{T}\right) ^{\prime}\mathbf{F}}{T}\right\Vert ^{2}\right) =O_{p}\left( \frac{\ln\left( n\right) }{T}\right) . \] Also, the probability order of the first term in ((ref)) dominates the second term, and we have \[ \sup_{i}\left\vert B_{2,5,iT}\right\vert \leq\left( \sup_{i}\sigma_{i} ^{2}\right) \left( \sup_{i}\left\vert \frac{\boldsymbol{\varepsilon} _{i\circ}\left( \lambda_{T}\right) ^{\prime}\mathbf{P}_{F} \boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) }{T}\right\vert \right) =O_{p}\left( \frac{\ln\left( n\right) }{T}\right) . \] Overall, using the above results we obtain \[ \sup_{i}\left\vert B_{2,iT}\right\vert \leq\sup_{i}\left( \sum_{j=1} ^{6}\left\vert B_{2,j,iT}\right\vert \right) \leq\sum_{j=1}^{6}\sup _{i}\left\vert B_{2,j,iT}\right\vert =O_{p}\left( \frac{\ln\left( n\right) }{T}\right) . \] Next, consider $B_{3,iT}$ which can be rewritten as \begin{align*} B_{3,iT} & =\frac{2\sigma_{i}^{2}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) ^{\prime}\left( \mathbf{I}_{m_{0}}-\mathbf{P}_{F}\right) \left( \mathbf{P}_{F}-\mathbf{P}_{\hat{F}}\right) \boldsymbol{\varepsilon }_{i\circ}\left( \lambda_{T}\right) }{T}=\frac{2\sigma_{i}^{2} \boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) ^{\prime}\left( \mathbf{P}_{F}-\mathbf{P}_{\hat{F}}-\mathbf{P}_{F}+\mathbf{P}_{F} \mathbf{P}_{\hat{F}}\right) \boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) }{T}\\ & =\frac{2\sigma_{i}^{2}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda _{T}\right) ^{\prime}\mathbf{P}_{\hat{F}}\boldsymbol{\varepsilon}_{i\circ }\left( \lambda_{T}\right) }{T}+\frac{2\sigma_{i}^{2}\boldsymbol{\varepsilon }_{i\circ}\left( \lambda_{T}\right) ^{\prime}\mathbf{P}_{F}\mathbf{P} _{\hat{F}}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) } {T}=B_{3,1,nT}+B_{3,2,nT}. \end{align*} Note that \begin{align*} \left\vert B_{3,1,iT}\right\vert & =2\left\vert \sigma_{i}^{2}\left( \frac{\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) ^{\prime }\mathbf{\hat{F}}}{T}\right) \left( \frac{\mathbf{\hat{F}}^{\prime }\mathbf{\hat{F}}}{T}\right) ^{-1}\left( \frac{\mathbf{\hat{F}}^{\prime }\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) }{T}\right) \right\vert \leq2\sigma_{i}^{2}\left\Vert \frac{\boldsymbol{\varepsilon }_{i\circ}\left( \lambda_{T}\right) ^{\prime}\mathbf{\hat{F}}}{T}\right\Vert ^{2}\left\Vert \left( \frac{\mathbf{\hat{F}}^{\prime}\mathbf{\hat{F}}} {T}\right) ^{-1}\right\Vert ^{2}\\ & \leq4\sigma_{i}^{2}\left( \left\Vert \frac{\boldsymbol{\varepsilon }_{i\circ}\left( \lambda_{T}\right) ^{\prime}\mathbf{F}}{T}\right\Vert ^{2}+\left\Vert \frac{\boldsymbol{\varepsilon}_{i\circ}\left( \lambda _{T}\right) ^{\prime}\left( \mathbf{\hat{F}}-\mathbf{F}\right) } {T}\right\Vert ^{2}\right) \left\Vert \left( \frac{\mathbf{\hat{F}}^{\prime }\mathbf{\hat{F}}}{T}\right) ^{-1}\right\Vert ^{2}. \end{align*} Then using ((ref)) and ((ref)), we have \begin{align*} \sup_{i}\left\vert B_{3,1,iT}\right\vert & \leq4\left( \sup_{i}\sigma _{i}^{2}\right) \sup_{i}\left( \left\Vert \frac{\boldsymbol{\varepsilon }_{i\circ}\left( \lambda_{T}\right) ^{\prime}\mathbf{F}}{T}\right\Vert ^{2}+\left\Vert \frac{\boldsymbol{\varepsilon}_{i\circ}\left( \lambda _{T}\right) ^{\prime}\left( \mathbf{\hat{F}}-\mathbf{F}\right) } {T}\right\Vert ^{2}\right) \left\Vert \left( \frac{\mathbf{\hat{F}}^{\prime }\mathbf{\hat{F}}}{T}\right) ^{-1}\right\Vert ^{2}\\ & \leq4\left( \sup_{i}\sigma_{i}^{2}\right) \left( \sup_{i}\left\Vert \frac{\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) ^{\prime }\mathbf{F}}{T}\right\Vert ^{2}+\sup_{i}\left\Vert \frac {\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) ^{\prime}\left( \mathbf{\hat{F}}-\mathbf{F}\right) }{T}\right\Vert ^{2}\right) \times\left\Vert \left( \frac{\mathbf{\hat{F}}^{\prime}\mathbf{\hat{F}}} {T}\right) ^{-1}\right\Vert ^{2}\\ & =O_{p}\left( \frac{\ln\left( n\right) }{T}\right) . \end{align*} Also, \begin{align*} \left\vert B_{3,2,iT}\right\vert & =2\sigma_{i}^{2}\left\vert \left( \frac{\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) ^{\prime }\mathbf{F}}{T}\right) \left( \frac{\mathbf{F}^{\prime}\mathbf{F}} {T}\right) ^{-1}\left( \frac{\mathbf{F}^{\prime}\mathbf{\hat{F}}}{T}\right) \left( \frac{\mathbf{\hat{F}}^{\prime}\mathbf{\hat{F}}}{T}\right) ^{-1}\left( \frac{\mathbf{\hat{F}}^{\prime}\boldsymbol{\varepsilon}_{i\circ }\left( \lambda_{T}\right) }{T}\right) \right\vert \\ & \leq2\sigma_{i}^{2}\left\vert \left( \frac{\boldsymbol{\varepsilon }_{i\circ}\left( \lambda_{T}\right) ^{\prime}\mathbf{F}}{T}\right) \left( \frac{\mathbf{F}^{\prime}\mathbf{F}}{T}\right) ^{-1}\left( \frac {\mathbf{F}^{\prime}\mathbf{\hat{F}}}{T}\right) \left( \frac{\mathbf{\hat {F}}^{\prime}\mathbf{\hat{F}}}{T}\right) ^{-1}\left[ \frac{\mathbf{F} ^{\prime}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) } {T}+\frac{\left( \mathbf{\hat{F}-F}\right) ^{\prime}\boldsymbol{\varepsilon }_{i\circ}\left( \lambda_{T}\right) }{T}\right] \right\vert \\ & \leq2\sigma_{i}^{2}\left\Vert \frac{\boldsymbol{\varepsilon}_{i\circ }\left( \lambda_{T}\right) ^{\prime}\mathbf{F}}{T}\right\Vert \left( \left\Vert \frac{\mathbf{F}^{\prime}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) }{T}\right\Vert +\left\Vert \frac{\left( \mathbf{\hat {F}-F}\right) ^{\prime}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda _{T}\right) }{T}\right\Vert \right) \left\Vert \left( \frac{\mathbf{F} ^{\prime}\mathbf{F}}{T}\right) ^{-1}\right\Vert \times\\ & \left\Vert \left( \frac{\mathbf{\hat{F}}^{\prime}\mathbf{\hat{F}}} {T}\right) ^{-1}\right\Vert \left\Vert \frac{\mathbf{F}^{\prime} \mathbf{\hat{F}}}{T}\right\Vert . \end{align*} By ((ref)) and ((ref)) we obtain \begin{align} & \sup_{i}\sigma_{i}^{2}\left\Vert \frac{\boldsymbol{\varepsilon}_{i\circ }\left( \lambda_{T}\right) ^{\prime}\mathbf{F}}{T}\right\Vert \left( \left\Vert \frac{\mathbf{F}^{\prime}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) }{T}\right\Vert +\left\Vert \frac{\left( \mathbf{\hat {F}-F}\right) ^{\prime}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda _{T}\right) }{T}\right\Vert \right) \nonumber\\ & \leq\left( \sup_{i}\sigma_{i}^{2}\right) \left[ \sup_{i}\left\Vert \frac{\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) ^{\prime }\mathbf{F}}{T}\right\Vert ^{2}+\sup_{i}\left\Vert \frac {\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) ^{\prime }\mathbf{F}}{T}\right\Vert \left( \sup_{i}\left\Vert \frac{\left( \mathbf{\hat{F}-F}\right) ^{\prime}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) }{T}\right\Vert \right) \right] \nonumber\\ & =O_{p}\left( \frac{\ln\left( n\right) }{T}\right) +O_{p}\left( \sqrt{\frac{\ln\left( n\right) }{T}}\right) \times O_{p}\left( \sqrt {\frac{\ln\left( n\right) }{nT}}\right) =O_{p}\left( \frac{\ln\left( n\right) }{T}\right) . \end{align} Further, by ((ref)), \begin{equation} \left\Vert \frac{\mathbf{F}^{\prime}\mathbf{\hat{F}}}{T}\right\Vert \leq\left\Vert \frac{\mathbf{F}^{\prime}\mathbf{F}}{T}\right\Vert +\left\Vert \frac{\mathbf{F}^{\prime}\left( \mathbf{\hat{F}-F}\right) }{T}\right\Vert =O_{p}\left( 1\right) . \end{equation} Using ((ref)), ((ref)), and ((ref)) now yields \begin{align*} \sup_{i}\left\vert B_{3,2,iT}\right\vert & \leq2\left[ \sup_{i}\sigma _{i}^{2}\left\Vert \frac{\boldsymbol{\varepsilon}_{i\circ}\left( \lambda _{T}\right) ^{\prime}\mathbf{F}}{T}\right\Vert \left( \left\Vert \frac{\mathbf{F}^{\prime}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda _{T}\right) }{T}\right\Vert +\left\Vert \frac{\left( \mathbf{\hat{F} -F}\right) ^{\prime}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda _{T}\right) }{T}\right\Vert \right) \right] \times\\ & \left\Vert \left( \frac{\mathbf{F}^{\prime}\mathbf{F}}{T}\right) ^{-1}\right\Vert \left\Vert \left( \frac{\mathbf{\hat{F}}^{\prime }\mathbf{\hat{F}}}{T}\right) ^{-1}\right\Vert \left\Vert \frac{\mathbf{F} ^{\prime}\mathbf{\hat{F}}}{T}\right\Vert =O_{p}\left( \frac{\ln\left( n\right) }{T}\right) . \end{align*} Hence, \[ \sup_{i}\left\vert B_{3,iT}\right\vert \leq\sup_{i}\left( \left\vert B_{3,1,iT}\right\vert +\left\vert B_{3,2,iT}\right\vert \right) \leq\sup _{i}\left\vert B_{3,1,iT}\right\vert +\sup_{i}\left\vert B_{3,2,iT}\right\vert =O_{p}\left( \frac{\ln\left( n\right) }{T}\right) . \] Now consider $B_{4,iT}$, and note that \begin{align*} B_{4,iT} & =\frac{2\sigma_{i}\boldsymbol{\gamma}_{i}^{\prime}\left( \mathbf{F-\hat{F}}\right) ^{\prime}\mathbf{M}_{\hat{F}}\mathbf{M} _{F}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) }{T}\\ & =-\frac{2\sigma_{i}\boldsymbol{\gamma}_{i}^{\prime}\left( \mathbf{\hat {F}-F}\right) ^{\prime}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda _{T}\right) }{T}+\frac{2\sigma_{i}\boldsymbol{\gamma}_{i}^{\prime}\left( \mathbf{\hat{F}-F}\right) ^{\prime}\mathbf{\hat{F}}\left( \mathbf{\hat{F} }^{\prime}\mathbf{\hat{F}}\right) ^{-1}\mathbf{\hat{F}}^{\prime }\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) }{T}\\ & -\frac{2\sigma_{i}\boldsymbol{\gamma}_{i}^{\prime}\left( \mathbf{\hat {F}-F}\right) ^{\prime}\mathbf{M}_{\hat{F}}\left( \mathbf{\hat{F}-F}\right) \left( \mathbf{F}^{\prime}\mathbf{F}\right) ^{-1}\mathbf{F}^{\prime }\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) }{T}=\sum _{j=1}^{3}B_{4,j,1T}. \end{align*} For the first term of the above equation, we have \[ \left\vert B_{4,1,iT}\right\vert \leq2\left\vert \frac{\sigma_{i} \boldsymbol{\gamma}_{i}^{\prime}\left( \mathbf{\hat{F}-F}\right) ^{\prime }\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) }{T}\right\vert \leq2\sigma_{i}\left\Vert \boldsymbol{\gamma}_{i}\right\Vert \left\Vert \frac{\left( \mathbf{\hat{F}-F}\right) ^{\prime}\boldsymbol{\varepsilon }_{i\circ}\left( \lambda_{T}\right) }{T}\right\Vert , \] where $\sigma_{i}$\ and $\boldsymbol{\gamma}_{i}$ are bounded. Then using ((ref)) it follows that \[ \sup_{i}\left\vert B_{4,1,iT}\right\vert \leq2\left( \sup_{i}\sigma _{i}\right) \left( \sup_{i}\left\Vert \boldsymbol{\gamma}_{i}\right\Vert \right) \left( \sup_{i}\left\Vert \frac{\left( \mathbf{\hat{F}-F}\right) ^{\prime}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) } {T}\right\Vert \right) =O_{p}\left( \sqrt{\frac{\ln\left( n\right) }{nT} }\right) . \] Similarly, \begin{align*} \left\vert B_{4,2,iT}\right\vert & =\left\vert \frac{2\sigma_{i} \boldsymbol{\gamma}_{i}^{\prime}\left( \mathbf{\hat{F}-F}\right) ^{\prime }\mathbf{\hat{F}}\left( \mathbf{\hat{F}}^{\prime}\mathbf{\hat{F}}\right) ^{-1}\mathbf{\hat{F}}^{\prime}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) }{T}\right\vert \\ & \leq2\sigma_{i}\left\Vert \boldsymbol{\gamma}_{i}\right\Vert \left\Vert \frac{\left( \mathbf{\hat{F}-F}\right) ^{\prime}\mathbf{\hat{F}}} {T}\right\Vert \left\Vert \left( \frac{\mathbf{\hat{F}}^{^{\prime} }\mathbf{\hat{F}}}{T}\right) ^{-1}\right\Vert \left( \frac{\left\Vert \mathbf{F}^{\prime}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda _{T}\right) \right\Vert }{T}+\frac{\left\Vert \left( \mathbf{\hat{F} -F}\right) ^{\prime}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda _{T}\right) \right\Vert }{T}\right) , \end{align*} and using ((ref)) and ((ref)) we have \begin{align*} \sup_{i}\left\vert B_{4,2,iT}\right\vert & \leq2\left( \sup_{i}\sigma _{i}\right) \left( \sup_{i}\left\Vert \boldsymbol{\gamma}_{i}\right\Vert \right) \left( \sup_{i}\frac{\left\Vert \mathbf{F}^{\prime} \boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) \right\Vert } {T}+\sup_{i}\frac{\left\Vert \left( \mathbf{\hat{F}-F}\right) ^{\prime }\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) \right\Vert } {T}\right) \\ & \times\left\Vert \frac{\left( \mathbf{\hat{F}-F}\right) ^{\prime }\mathbf{\hat{F}}}{T}\right\Vert \left\Vert \left( \frac{\mathbf{\hat{F} }^{^{\prime}}\mathbf{\hat{F}}}{T}\right) ^{-1}\right\Vert =O_{p}\left( \sqrt{\frac{\ln\left( n\right) }{T}}\right) \times O_{p}\left( \frac {1}{\delta_{nT}^{2}}\right) . \end{align*} Moreover, \begin{align*} \left\vert B_{4,3,iT}\right\vert & =\left\vert \frac{2\sigma_{i} \boldsymbol{\gamma}_{i}^{\prime}\left( \mathbf{\hat{F}-F}\right) ^{\prime }\mathbf{M}_{\hat{F}}\left( \mathbf{\hat{F}-F}\right) \left( \mathbf{F} ^{\prime}\mathbf{F}\right) ^{-1}\mathbf{F}^{\prime}\boldsymbol{\varepsilon }_{i\circ}\left( \lambda_{T}\right) }{T}\right\vert \\ & \leq2\sigma_{i}\left\Vert \boldsymbol{\gamma}_{i}\right\Vert \left( \frac{\left\Vert \mathbf{\hat{F}-F}\right\Vert ^{2}}{T}\right) \left\Vert \left( \frac{\mathbf{F}^{\prime}\mathbf{F}}{T}\right) ^{-1}\right\Vert \left\Vert \frac{\mathbf{F}^{\prime}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) }{T}\right\Vert , \end{align*} then taking the supremum, \begin{align*} \sup_{i}\left\vert B_{4,3,iT}\right\vert & \leq2\left( \sup_{i}\sigma _{i}\right) \left( \sup_{i}\left\Vert \boldsymbol{\gamma}_{i}\right\Vert \right) \left( \sup_{i}\left\Vert \frac{\mathbf{F}^{\prime} \boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) }{T}\right\Vert \right) \left( \frac{\left\Vert \mathbf{\hat{F}-F}\right\Vert ^{2}} {T}\right) \left\Vert \left( \frac{\mathbf{F}^{\prime}\mathbf{F}}{T}\right) ^{-1}\right\Vert \\ & =O_{p}\left( \sqrt{\frac{\ln\left( n\right) }{T}}\right) \times O_{p}\left( \frac{1}{\delta_{nT}^{2}}\right) . \end{align*} Hence, it follows that \[ \sup_{i}\left\vert B_{4,iT}\right\vert \leq\sup_{i}\sum_{j=1}^{3}\left\vert B_{4,j,iT}\right\vert \leq\sum_{j=1}^{3}\sup_{i}\left\vert B_{4,j,iT} \right\vert =O_{p}\left( \sqrt{\frac{\ln\left( n\right) }{nT}}\right) . \] Similarly, $\sup_{i}\left\vert B_{5,iT}\right\vert =O_{p}\left( \sqrt {\ln\left( n\right) /\left( nT\right) }\right) $. For the final term $B_{6,iT}$, note that \begin{align*} B_{6,iT} & =\frac{\sigma_{i}^{2}\left( \boldsymbol{\varepsilon}_{i\circ }\left( \lambda_{T}\right) +\boldsymbol{\varepsilon}_{i\circ}\right) ^{\prime}\left( \boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) -\boldsymbol{\varepsilon}_{i\circ}\right) }{T}-\frac{\sigma_{i}^{2}\left( \boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) +\boldsymbol{\varepsilon}_{i\circ}\right) ^{\prime}\mathbf{F}}{T}\left( \frac{\mathbf{F}^{\prime}\mathbf{F}}{T}\right) ^{-1}\frac{\mathbf{F}^{\prime }\left( \boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) -\boldsymbol{\varepsilon}_{i\circ}\right) }{T}\\ & =B_{6,1,iT}+B_{6,2,iT}. \end{align*} Since $\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) =\boldsymbol{\varepsilon}_{i\circ}+\lambda_{T}\mathbf{b}_{i}$ where $\boldsymbol{\varepsilon}_{i\circ}=\left( \varepsilon_{i1},\varepsilon _{i2},\ldots,\varepsilon_{iT}\right) ^{\prime}$, $\mathbf{b}_{i}=\left( b_{i1},b_{i2},\ldots,b_{iT}\right) ^{\prime}$, $b_{it}=\mathbf{w} _{i0}^{\prime}\boldsymbol{\varepsilon}_{\circ t}$, $\mathbf{w}_{i0}=\left( w_{i1},w_{is},\ldots,w_{in}\right) ^{\prime}$ and $\boldsymbol{\varepsilon }_{\circ t}=\left( \varepsilon_{1t},\varepsilon_{2t},\ldots,\varepsilon _{nt}\right) ^{\prime}$, then \[ B_{6,1,iT}=\sigma_{i}^{2}\left( \frac{2\lambda_{T}\boldsymbol{\varepsilon }_{i\circ}^{\prime}\mathbf{b}_{i}}{T}+\frac{\lambda_{T}^{2}\mathbf{b} _{i}^{\prime}\mathbf{b}_{i}}{T}\right) . \] Further, using results ((ref)) and ((ref)) we have \[ \sup_{i}\left\vert B_{6,1,iT}\right\vert \leq\left( \sup_{i}\sigma_{i} ^{2}\right) \left( 2\left\vert \lambda_{T}\right\vert \sup_{i}\left\vert \frac{\boldsymbol{\varepsilon}_{i\circ}^{\prime}\mathbf{b}_{i}}{T}\right\vert +\lambda_{T}^{2}\sup_{i}\left\vert \frac{\mathbf{b}_{i}^{\prime}\mathbf{b} _{i}}{T}\right\vert \right) =O_{p}\left( \frac{\sqrt{\ln\left( n\right) } }{T}\right) +O_{p}\left( \frac{1}{T}\right) . \] Now consider $B_{6,2,iT}$ and note that \[ \left\vert B_{6,2,iT}\right\vert \leq\sigma_{i}^{2}\left\Vert \frac{\left( \boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) +\boldsymbol{\varepsilon}_{i\circ}\right) ^{\prime}\mathbf{F}}{T}\right\Vert \left\Vert \left( \frac{\mathbf{F}^{\prime}\mathbf{F}}{T}\right) ^{-1}\right\Vert \left\Vert \frac{\mathbf{F}^{\prime}\left( \boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) -\boldsymbol{\varepsilon}_{i\circ}\right) }{T}\right\Vert . \] Further, using ((ref)) we have \begin{align*} \sup_{i}\left\Vert \frac{\left( \boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) +\boldsymbol{\varepsilon}_{i\circ}\right) ^{\prime }\mathbf{F}}{T}\right\Vert & \leq\sup_{i}\left\Vert \frac {\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) ^{\prime }\mathbf{F}}{T}\right\Vert +\sup_{i}\left\Vert \frac{\boldsymbol{\varepsilon }_{i\circ}^{\prime}\mathbf{F}}{T}\right\Vert =O_{p}\left( \sqrt{\ln\left( n\right) /T}\right) ,\\ \sup_{i}\left\Vert \frac{\mathbf{F}^{\prime}\left( \boldsymbol{\varepsilon }_{i\circ}\left( \lambda_{T}\right) -\boldsymbol{\varepsilon}_{i\circ }\right) }{T}\right\Vert & \leq\sup_{i}\left\Vert \frac {\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) ^{\prime }\mathbf{F}}{T}\right\Vert +\sup_{i}\left\Vert \frac{\boldsymbol{\varepsilon }_{i\circ}^{\prime}\mathbf{F}}{T}\right\Vert =O_{p}\left( \sqrt{\ln\left( n\right) /T}\right) . \end{align*} Hence, it follows that \begin{align*} \sup_{i}\left\vert B_{6,2,iT}\right\vert & \leq\left\Vert \left( \frac{\mathbf{F}^{\prime}\mathbf{F}}{T}\right) ^{-1}\right\Vert \left( \sup_{i}\sigma_{i}^{2}\right) \left( \sup_{i}\left\Vert \frac{\left( \boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) +\boldsymbol{\varepsilon}_{i\circ}\right) ^{\prime}\mathbf{F}}{T}\right\Vert \right) \\ & \times\left( \sup_{i}\left\Vert \frac{\mathbf{F}^{\prime}\left( \boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) -\boldsymbol{\varepsilon}_{i\circ}\right) }{T}\right\Vert \right) =O_{p}\left( \frac{\ln\left( n\right) }{T}\right) . \end{align*} Overall, \[ \sup_{i}\left\vert B_{6,iT}\right\vert \leq\sup_{i}\left\vert B_{6,1,iT} \right\vert +\sup_{i}\left\vert B_{6,2,iT}\right\vert =O_{p}\left( \frac {\ln\left( n\right) }{T}\right) . \] Using the above results of $B_{1,iT}$ to $B_{6,iT}$, and noting that $n$ and $T$ are assumed to be of the same order of magnitude, we obtain \[ \sup_{i}\left\vert \hat{\sigma}_{i,T}^{2}-\omega_{i,T}^{2}\right\vert \leq \sup_{i}\sum_{j=1}^{6}\left\vert B_{j,iT}\right\vert \leq\sum_{j=1}^{6} \sup_{i}\left\vert B_{j,iT}\right\vert =O_{p}\left( \frac{\ln\left( n\right) }{T}\right) , \] so ((ref)) is established. To prove ((ref)), note that \[ \sup_{i}\left\vert \hat{\sigma}_{i,T}-\omega_{i,T}\right\vert =\sup_{i} \frac{\left\vert \hat{\sigma}_{i,T}^{2}-\omega_{i,T}^{2}\right\vert } {\hat{\sigma}_{i,T}+\omega_{i,T}}=\sup_{i}\left( \frac{1}{\hat{\sigma} _{i,T}+\omega_{i,T}}\right) \sup_{i}\left\vert \hat{\sigma}_{i,T}^{2} -\omega_{i,T}^{2}\right\vert . \] Since $\hat{\sigma}_{i,T}^{2}>0$ (by construction) \[ \sup_{i}\left\vert \hat{\sigma}_{i,T}-\omega_{i,T}\right\vert \leq\sup _{i}\left( \frac{1}{\omega_{i,T}}\right) \sup_{i}\left\vert \hat{\sigma }_{i,T}^{2}-\omega_{i,T}^{2}\right\vert . \] By definition of $\omega_{i,T}$, we have $\omega_{i,T}^{-2}=\sigma_{i} ^{-2}\left( T^{-1}\boldsymbol{\varepsilon}_{i\circ}^{\prime}\mathbf{M} _{F}\boldsymbol{\varepsilon}_{i\circ}\right) ^{-1}$, so that \begin{align*} \sup_{i}\left( \frac{1}{\omega_{i,T}}\right) & =\sup_{i}\left( \frac {1}{\omega_{i,T}^{2}}\right) ^{1/2}\leq\left( \sup_{i}\frac{1}{\sigma _{i}^{2}}\right) ^{1/2}\left( \sup_{i}\frac{1}{T^{-1}\boldsymbol{\varepsilon }_{i\circ}^{\prime}\mathbf{M}_{F}\boldsymbol{\varepsilon}_{i\circ}}\right) ^{1/2}\\ & =\left( \frac{1}{\inf_{i}\sigma_{i}^{2}}\right) ^{1/2}\left( \frac {1}{\inf_{i}\left( T^{-1}\boldsymbol{\varepsilon}_{i\circ}^{\prime} \mathbf{M}_{F}\boldsymbol{\varepsilon}_{i\circ}\right) }\right) ^{1/2}. \end{align*} Since $\inf_{i}\sigma_{i}^{2}>c>0$ and by condition ((ref)) in Assumption (ref), $\inf_{i}T^{-1}\boldsymbol{\varepsilon}_{i\circ }^{\prime}\mathbf{M}_{F}\boldsymbol{\varepsilon}_{i\circ}>c>0$ as $T\rightarrow\infty$, then it readily follows $\sup_{i}\omega_{i,T} ^{-1}<C<\infty$, which together with ((ref)) now establishes result ((ref)). Similarly, note that \[ \sup_{i}\left\vert \frac{1}{\hat{\sigma}_{i,T}}-\frac{1}{\omega_{i,T} }\right\vert \leq\sup_{i}\left( \frac{1}{\hat{\sigma}_{i,T}}\right) \sup _{i}\left( \frac{1}{\omega_{i,T}}\right) \sup_{i}\left\vert \hat{\sigma }_{i,T}-\omega_{i,T}\right\vert , \] where $\sup_{i}\left( \frac{1}{\hat{\sigma}_{i,T}}\right) <C$, by construction. Hence, result ((ref)) can be established using ((ref)). Finally, results ((ref)),((ref)) and ((ref)) follow using ((ref)), ((ref)) and ((ref)), respectively.
lemmaConsider the latent factor model given by ((ref)) and ((ref)). The latent factors, $\mathbf{f}_{t}$, and their loadings, $\boldsymbol{\gamma}_{i}$, are estimated by principal components, $\mathbf{\hat{f}}_{t}$ and $\boldsymbol{\hat{\gamma}}_{i}$, given by ((ref)). Suppose that Assumptions (ref)-(ref) hold and $\left( n,T\right) \rightarrow\infty$, such that $n/T\rightarrow\kappa\,,$ for $0<\kappa<\infty$. Then \begin{align} \mathbf{d}_{1,nT} & =\frac{1}{n}\sum_{i=1}^{n}b_{in}(\boldsymbol{\hat {\gamma}}_{i}-\boldsymbol{\gamma}_{i})=O_{p}\left( \sqrt{\frac{\ln\left( n\right) }{nT}}\right) ,\\ \mathbf{d}_{2,nT} & =\frac{1}{n}\sum_{i=1}^{n}(\boldsymbol{\hat{\delta} }_{i,T}-\boldsymbol{\delta}_{i,T})=O_{p}\left( \sqrt{\frac{\ln\left( n\right) }{nT}}\right) ,\\ \mathbf{d}_{3,nT} & =\frac{1}{n}\sum_{i=1}^{n}b_{in}\left( \boldsymbol{\hat {\gamma}}_{i}-\boldsymbol{\gamma}_{i}\right) \boldsymbol{\gamma}_{i}^{\prime }=O_{p}\left( \sqrt{\frac{\ln\left( n\right) }{nT}}\right) ,\\ \mathbf{d}_{4,nT} & =\frac{1}{n}\sum_{i=1}^{n}\left( \omega_{i,T} -\sigma_{i}\right) (\boldsymbol{\hat{\gamma}}_{i}-\boldsymbol{\gamma} _{i})=O_{p}\left( \frac{1}{\delta_{nT}^{2}}\right) ,\\ \mathbf{d}_{5,nT} & =\frac{1}{n}\sum_{i=1}^{n}\left( \frac{1}{\omega_{i,T} }-\frac{1}{\sigma_{i}}\right) (\boldsymbol{\hat{\gamma}}_{i} -\boldsymbol{\gamma}_{i})=O_{p}\left( \frac{1}{\delta_{nT}^{2}}\right) ,\\ \mathbf{d}_{6,nT} & =\frac{1}{n}\sum_{i=1}^{n}\left( \hat{\sigma} _{i,T}-\omega_{i,T}\right) (\boldsymbol{\hat{\gamma}}_{i}-\boldsymbol{\gamma }_{i})=O_{p}\left[ \left( \frac{\ln\left( n\right) }{T}\right) ^{3/2}\right] ,\\ \mathbf{d}_{7,nT} & =\frac{1}{n}\sum_{i=1}^{n}\left( \frac{1}{\hat{\sigma }_{i,T}}-\frac{1}{\omega_{i,T}}\right) (\boldsymbol{\hat{\gamma}} _{i}-\boldsymbol{\gamma}_{i})=O_{p}\left[ \left( \frac{\ln\left( n\right) }{T}\right) ^{3/2}\right] , \end{align} where $\left\{ b_{in}\right\} _{i=1}^{n}$ is a sequence of fixed values bounded in $n$, such that $n^{-1}\sum_{i=1}^{n}b_{in}^{2}=O(1)$, $\boldsymbol{\delta}_{i,T}=\boldsymbol{\gamma}_{i}/\omega_{i,T},$ $\boldsymbol{\hat{\delta}}_{i,T}=\hat{\boldsymbol{\gamma}}_{i}/\omega_{i,T},$ and $\omega_{i,T}=\left( T^{-1}\sigma_{i}^{2}\boldsymbol{\varepsilon} _{i\circ}^{\prime}\mathbf{M}_{F}\boldsymbol{\varepsilon}_{i\circ}\right) ^{1/2}.$
proofNote that in general \begin{align} \boldsymbol{\hat{\gamma}}_{i}-\boldsymbol{\gamma}_{i} & =\left( \frac{\mathbf{\hat{F}}^{\prime}\mathbf{\hat{F}}}{T}\right) ^{-1}\left( \frac{\mathbf{\hat{F}}^{\prime}\mathbf{F}\boldsymbol{\gamma}_{i}}{T} +\frac{\sigma_{i}\mathbf{\hat{F}}^{\prime}\boldsymbol{\varepsilon}_{i\circ }\left( \lambda_{T}\right) }{T}\right) -\left( \frac{\mathbf{\hat{F} }^{\prime}\mathbf{\hat{F}}}{T}\right) ^{-1}\left( \frac{\mathbf{\hat{F} }^{\prime}\mathbf{\hat{F}}}{T}\right) \boldsymbol{\gamma}_{i}\nonumber\\ & =-\left( \frac{\mathbf{\hat{F}}^{\prime}\mathbf{\hat{F}}}{T}\right) ^{-1}\left[ \frac{\mathbf{\hat{F}}^{\prime}\left( \mathbf{\hat{F}-F}\right) \boldsymbol{\gamma}_{i}}{T}\right] +\left( \frac{\mathbf{\hat{F}}^{\prime }\mathbf{\hat{F}}}{T}\right) ^{-1}\left( \frac{\sigma_{i}\mathbf{\hat{F} }^{\prime}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) } {T}\right) , \end{align} and we have \begin{align} \mathbf{d}_{1,nT} & =\frac{1}{n}\sum_{i=1}^{n}b_{in}\left( \boldsymbol{\hat {\gamma}}_{i}-\boldsymbol{\gamma}_{i}\right) \nonumber\\ & =-\left( \frac{\mathbf{\hat{F}}^{\prime}\mathbf{\hat{F}}}{T}\right) ^{-1}\left[ \frac{\mathbf{\hat{F}}^{\prime}\left( \mathbf{\hat{F}-F}\right) }{T}\right] \left( \frac{1}{n}\sum_{i=1}^{n}b_{in}\boldsymbol{\gamma} _{i}\right) +\left( \frac{\mathbf{\hat{F}}^{\prime}\mathbf{\hat{F}}} {T}\right) ^{-1}T^{-1}\left( \frac{1}{n}\sum_{i=1}^{n}b_{in}\sigma _{i}\mathbf{\hat{F}}^{\prime}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) \right) . \end{align} Since by assumption $\left\Vert \boldsymbol{\gamma}_{i}\right\Vert <C$, we have \[ \left\Vert \frac{1}{n}\sum_{i=1}^{n}b_{in}\boldsymbol{\gamma}_{i}\right\Vert \leq\left( \frac{1}{n}\sum_{i=1}^{n}b_{in}^{2}\right) ^{1/2}\left\Vert \frac{1}{n}\sum_{i=1}^{n}\boldsymbol{\gamma}_{i}\boldsymbol{\gamma} _{i}^{^{\prime}}\right\Vert ^{1/2}\leq\left( \frac{1}{n}\sum_{i=1}^{n} b_{in}^{2}\right) ^{1/2}\left( \frac{1}{n}\sum_{i=1}^{n}\left\Vert \boldsymbol{\gamma}_{i}\right\Vert ^{2}\right) ^{1/2}<C. \] Also by results ((ref)) and ((ref)), the first term of ((ref)) is $O_{p}\left( \delta_{nT}^{-2}\right) $. For the second term of ((ref)), since $\left( T^{-1}\mathbf{\hat{F}}^{\prime}\mathbf{\hat {F}}\right) ^{-1}=O_{p}(1)$, we note that \[ T^{-1}\left( \mathbf{\hat{F}}-\mathbf{F+F}\right) ^{\prime}\left( \frac {1}{n}\sum_{i=1}^{n}b_{in}\sigma_{i}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) \right) =\frac{1}{n}\sum_{i=1}^{n}b_{in}\left( \frac{\mathbf{\hat{F}}-\mathbf{F}}{T}\right) ^{\prime}\sigma_{i} \boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) +\frac{1}{Tn} \sum_{i=1}^{n}b_{in}\mathbf{F}^{^{\prime}}\sigma_{i}\boldsymbol{\varepsilon }_{i\circ}\left( \lambda_{T}\right) . \] Using result ((ref)), we have \begin{align} & \left\Vert \frac{1}{n}\sum_{i=1}^{n}b_{in}\left( \frac{\mathbf{\hat{F} }-\mathbf{F}}{T}\right) ^{\prime}\sigma_{i}\boldsymbol{\varepsilon}_{i\circ }\left( \lambda_{T}\right) \right\Vert \nonumber\\ & \leq\frac{1}{n}\sum_{i=1}^{n}\left\vert b_{in}\sigma_{i}\right\vert \left\Vert \left( \frac{\mathbf{\hat{F}}-\mathbf{F}}{T}\right) ^{\prime }\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) \right\Vert \leq\left( \sup_{i}\left\vert b_{in}\right\vert \right) \left( \sup _{i}\sigma_{i}\right) \left( \sup_{i}\left\Vert \frac{\left( \mathbf{\hat {F}}-\mathbf{F}\right) ^{\prime}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) }{T}\right\Vert \right) \nonumber\\ & =O_{p}\left( \sqrt{\frac{\ln\left( n\right) }{nT}}\right) . \end{align} Under part (a) of Assumption (ref) and by the serial independence of $\varepsilon_{it}$, \begin{align*} E\left\Vert \frac{1}{\sqrt{nT}}\sum_{i=1}^{n}\sum_{t=1}^{T}b_{in} \mathbf{f}_{t}\sigma_{i}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda _{T}\right) \right\Vert ^{2} & =\frac{1}{nT}\sum_{i=1}^{n}\sum_{j=1} ^{n}\sum_{t=1}^{T}b_{in}b_{jn}\sigma_{i}\sigma_{j}E\left\Vert \mathbf{f} _{t}\right\Vert ^{2}E\left( \varepsilon_{it}\left( \lambda_{T}\right) \varepsilon_{jt}\left( \lambda_{T}\right) \right) \\ & \leq E\left\Vert \mathbf{f}_{t}\right\Vert ^{2}\left( \sup_{i}b_{in} ^{2}\right) \left( \sup_{i}\sigma_{i}^{2}\right) \left[ \frac{1}{nT} \sum_{t=1}^{T}\sum_{i=1}^{n}\sum_{j=1}^{n}\left\vert E\left( \varepsilon _{it}\left( \lambda_{T}\right) \varepsilon_{jt}\left( \lambda_{T}\right) \right) \right\vert \right] \end{align*} which is $O\left( 1\right) $ based on ((ref)) and the boundedness of $E\left\Vert \mathbf{f}_{t}\right\Vert ^{2}$, $b_{in}^{2}$ and $\sigma_{i} ^{2}$ required by assumptions. So it follows \begin{equation} \frac{1}{nT}\sum_{i=1}^{n}b_{in}\mathbf{F}^{^{\prime}}\sigma_{i} \boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) =\frac{1} {\sqrt{nT}}\left( \frac{1}{\sqrt{nT}}\sum_{i=1}^{n}\sum_{t=1}^{T} b_{in}\mathbf{f}_{t}\sigma_{i}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) \right) =O_{p}\left( \frac{1}{\sqrt{nT}}\right) . \end{equation} Result ((ref)) now follows using ((ref)) and ((ref)) in ((ref)), and noting that by assumption $n$ and $T$ are of the same order. Consider now ((ref)), which can be written as \begin{align*} \mathbf{d}_{2,nT} & =\frac{1}{n}\sum_{i=1}^{n}\left( \frac{\boldsymbol{\hat {\gamma}}_{i}}{\omega_{i,T}}-\frac{\boldsymbol{\gamma}_{i}}{\omega_{i,T} }\right) \\ & =\frac{1}{n}\sum_{i=1}^{n}\left( \frac{\boldsymbol{\hat{\gamma}} _{i}-\boldsymbol{\gamma}_{i}}{\sigma_{i}}\right) \left( 1-\frac{\omega _{i,T}-\sigma_{i}}{\omega_{i,T}}\right) \\ & =\frac{1}{n}\sum_{i=1}^{n}\left( \frac{\boldsymbol{\hat{\gamma}} _{i}-\boldsymbol{\gamma}_{i}}{\sigma_{i}}\right) -\frac{1}{n}\sum_{i=1} ^{n}\left( \frac{\boldsymbol{\hat{\gamma}}_{i}-\boldsymbol{\gamma}_{i} }{\sigma_{i}}\right) \left( 1-\frac{T}{\boldsymbol{\varepsilon}_{i\circ }^{\prime}\mathbf{M}_{F}\boldsymbol{\varepsilon}_{i\circ}}\right) . \end{align*} The first term of the above has the same form as ((ref)), and becomes identical to it if we replace $a_{i}$ in ((ref)) with $1/\sigma_{i}$, since by assumption $\inf_{i}(\sigma_{i})>c$. Hence, the order of the first term is $O_{p}(\sqrt{\ln\left( n\right) /\left( nT\right) })$. Also the second term is dominated by the first term, since $1-\left( T^{-1} \boldsymbol{\varepsilon}_{i\circ}^{\prime}\mathbf{M}_{F} \boldsymbol{\varepsilon}_{i\circ}\right) ^{-1}=O_{p}(T^{-1})$ based on result ((ref)). Therefore, ((ref)) is established as required. For ((ref)), note that \begin{align} \mathbf{d}_{3,nT} & =\frac{1}{n}\sum_{i=1}^{n}b_{in}\left[ -\left( \mathbf{\hat{F}}^{\prime}\mathbf{\hat{F}}\right) ^{-1}\mathbf{\hat{F} }^{\prime}\left( \mathbf{\hat{F}-F}\right) \boldsymbol{\gamma}_{i} +\sigma_{i}\left( \mathbf{\hat{F}}^{\prime}\mathbf{\hat{F}}\right) ^{-1}\mathbf{\hat{F}}^{\prime}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) \right] \boldsymbol{\gamma}_{i}^{\prime}\nonumber\\ & =-\left( \frac{\mathbf{\hat{F}}^{\prime}\mathbf{\hat{F}}}{T}\right) ^{-1}\frac{\mathbf{\hat{F}}^{\prime}\left( \mathbf{\hat{F}-F}\right) } {T}\left( \frac{1}{n}\sum_{i=1}^{n}b_{in}\boldsymbol{\gamma}_{i} \boldsymbol{\gamma}_{i}^{\prime}\right) +\left( \frac{\mathbf{\hat{F} }^{\prime}\mathbf{\hat{F}}}{T}\right) ^{-1}T^{-1}\mathbf{\hat{F}}^{\prime }\left( \frac{1}{n}\sum_{i=1}^{n}b_{in}\sigma_{i}\boldsymbol{\varepsilon }_{i\circ}\left( \lambda_{T}\right) \boldsymbol{\gamma}_{i}^{\prime}\right) \nonumber\\ & =-\left( \frac{\mathbf{\hat{F}}^{\prime}\mathbf{\hat{F}}}{T}\right) ^{-1}\frac{\mathbf{\hat{F}}^{\prime}\left( \mathbf{\hat{F}-F}\right) } {T}\left( \frac{1}{n}\sum_{i=1}^{n}b_{in}\boldsymbol{\gamma}_{i} \boldsymbol{\gamma}_{i}^{\prime}\right) +\left( \frac{\mathbf{\hat{F} }^{\prime}\mathbf{\hat{F}}}{T}\right) ^{-1}\left( \frac{1}{n}\sum_{i=1} ^{n}T^{-1}b_{in}\sigma_{i}\left( \mathbf{\hat{F}-F}\right) ^{\prime }\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) \boldsymbol{\gamma}_{i}^{\prime}\right) \nonumber\\ & +\left( \frac{\mathbf{\hat{F}}^{\prime}\mathbf{\hat{F}}}{T}\right) ^{-1}\left( \frac{1}{n}\sum_{i=1}^{n}T^{-1}b_{in}\sigma_{i}\mathbf{F} ^{\prime}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) \boldsymbol{\gamma}_{i}^{\prime}\right) . \end{align} Recall that $\left( T^{-1}\mathbf{\hat{F}}^{\prime}\mathbf{\hat{F}}\right) ^{-1}=O_{p}(1)$, and $n^{-1}\sum_{i=1}^{n}\boldsymbol{\gamma}_{i} \boldsymbol{\gamma}_{i}^{\prime}=O_{p}(1)$. Also note that $b_{in}$ is bounded in $n$. Then using ((ref)) it follows that ($n$ and $T$ being of the same order) \[ \left( \frac{\mathbf{\hat{F}}^{\prime}\mathbf{\hat{F}}}{T}\right) ^{-1} \frac{\mathbf{\hat{F}}^{\prime}\left( \mathbf{\hat{F}-F}\right) }{T}\left( n^{-1}\sum_{i=1}^{n}b_{in}\boldsymbol{\gamma}_{i}\boldsymbol{\gamma} _{i}^{\prime}\right) =O_{p}\left( \frac{1}{\min(n,T)}\right) =O_{p} (\delta_{nT\ }^{-2}). \] Similarly, using ((ref)) \[ \left( \frac{\mathbf{\hat{F}}^{\prime}\mathbf{\hat{F}}}{T}\right) ^{-1}\left( n^{-1}\sum_{i=1}^{n}T^{-1}\left( \mathbf{\hat{F}-F}\right) ^{\prime}b_{in}\sigma_{i}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda _{T}\right) \boldsymbol{\gamma}_{i}^{\prime}\right) =O_{p}\left( \sqrt{\frac{\ln\left( n\right) }{nT}}\right) . \] The last term of ( (ref))\ can be written as $\left( nT\right) ^{-1/2}\left( T^{-1}\mathbf{\hat{F}}^{\prime}\mathbf{\hat{F}}\right) ^{-1}\left( n^{-1/2}T^{-1/2}\sum_{i=1}^{n}b_{in}\sigma_{i}\mathbf{F}^{\prime }\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) \boldsymbol{\gamma}_{i}^{\prime}\right) ,$ where $n^{-1/2}T^{-1/2}\sum _{i=1}^{n}b_{in}\sigma_{i}\mathbf{F}^{\prime}\boldsymbol{\varepsilon}_{i\circ }\left( \lambda_{T}\right) \boldsymbol{\gamma}_{i}^{\prime}$ is an $m_{0}\times m_{0}$ matrix with its $(j,j^{\prime})$ element given by $n^{-1/2}T^{-1/2}\sum_{i=1}^{n}\sum_{t=1}^{T}$ $b_{in}\sigma_{i} f_{jt}\varepsilon_{it}\left( \lambda_{T}\right) \gamma_{ij^{\prime}}$ for $j,j^{\prime}=1,2,\ldots,m_{0}$. It can be further shown \begin{align*} & E\left( n^{-1/2}T^{-1/2}\sum_{i=1}^{n}\sum_{t=1}^{T}b_{in}\sigma_{i} f_{jt}\gamma_{ij^{\prime}}\varepsilon_{it}\left( \lambda_{T}\right) \right) ^{2}\\ & =\frac{1}{nT}\sum_{i=1}^{n}\sum_{i^{\prime}=1}^{n}\sum_{t=1}^{T} b_{in}b_{i^{\prime}n}\sigma_{i}\sigma_{i^{\prime}}\gamma_{ij^{\prime}} \gamma_{i^{\prime}j^{\prime}}E\left( f_{jt}^{2}\right) E\left( \varepsilon_{it}\left( \lambda_{T}\right) \varepsilon_{i^{\prime}t}\left( \lambda_{T}\right) \right) \\ & \leq\frac{1}{nT}\sum_{i=1}^{n}\sum_{i^{\prime}=1}^{n}\sum_{t=1} ^{T}\left\vert b_{in}b_{i^{\prime}n}\sigma_{i}\sigma_{i^{\prime}} \gamma_{ij^{\prime}}\gamma_{i^{\prime}j^{\prime}}\right\vert E\left( f_{jt}^{2}\right) \left\vert E\left( \varepsilon_{it}\left( \lambda _{T}\right) \varepsilon_{i^{\prime}t}\left( \lambda_{T}\right) \right) \right\vert \\ & \leq\left( \sup_{i}b_{in}^{2}\right) \left( \sup_{i}\gamma_{ij^{\prime} }^{2}\right) \left( \sup_{i}\sigma_{i}^{2}\right) E\left( f_{jt} ^{2}\right) \left[ \frac{1}{nT}\sum_{i=1}^{n}\sum_{i^{\prime}=1}^{n} \sum_{t=1}^{T}\left\vert E\left( \varepsilon_{it}\left( \lambda_{T}\right) \varepsilon_{i^{\prime}t}\left( \lambda_{T}\right) \right) \right\vert \right] \end{align*} which is $O\left( 1\right) $ based on ((ref)) and the boundedness of $b_{in}^{2}$, $\gamma_{ij^{\prime}}^{2}$, $\sigma_{i}^{2}$ and $E\left( f_{jt}^{2}\right) $. Consequentially, $n^{-1/2}T^{-1/2}\sum_{i=1}^{n} b_{in}\sigma_{i}\mathbf{F}^{\prime}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) \boldsymbol{\gamma}_{i}^{\prime}=O_{p}\left( 1\right) $ and the last term of ((ref)) are also $O_{p}(\delta_{nT\ }^{-2})$. Thus result ((ref)) is established, as required. To prove ((ref)) we first write it as \[ \mathbf{d}_{4,nT}=\left( \sqrt{\frac{1}{T}}\right) \frac{1}{n}\sum_{i=1} ^{n}q_{iT}\left( \boldsymbol{\hat{\gamma}}_{i}-\boldsymbol{\gamma} _{i}\right) , \] where \[ q_{iT}=\sqrt{T}\left( \omega_{i,T}-\sigma_{i}\right) =\sigma_{i}\sqrt {T}\left[ \left( \frac{\boldsymbol{\varepsilon}_{i\circ}^{\prime} \mathbf{M}_{F}\boldsymbol{\varepsilon}_{i\circ}}{T}\right) ^{1/2}-1\right] , \] and conditional on $\mathbf{F}$ and $\sigma_{i}$, $q_{iT}$ are independently distributed across $i$. Using results in Lemma (ref) it is easily seen that $E\left( q_{iT}\right) =O(T^{-1/2})$ and $Var\left( q_{iT}\right) =O(1)$, and hence $n^{-1}\sum_{i=1}^{n}q_{iT}^{2}=O_{p}(1)$. Also by Cauchy-Schwarz inequality we have \[ \left\Vert \mathbf{d}_{4,nT}\right\Vert \leq\left( \sqrt{\frac{1}{T}}\right) \left( n^{-1}\sum_{i=1}^{n}q_{iT}^{2}\right) ^{1/2}\left( n^{-1/2} \left\Vert \mathbf{\hat{\Gamma}}-\mathbf{\Gamma}\right\Vert \right) , \] where $T^{-1}n=\ominus(1)$, and by ((ref)) $n^{-1/2}\left\Vert \mathbf{\hat{\Gamma}}-\mathbf{\Gamma}\right\Vert =O_{p}(\delta_{nT}^{-1})$, and ((ref)) is established. Result ((ref)) follows similarly, with $q_{iT}$ re-defined as $q_{iT}=\sigma_{i}^{-1}\sqrt{T}\left[ \left( \frac{\boldsymbol{\varepsilon}_{i\circ}^{\prime}\mathbf{M}_{F} \boldsymbol{\varepsilon}_{i\circ}}{T}\right) ^{-1/2}-1\right] $, and noting that $\sup_{i}(1/\sigma_{i}^{2})<C$, and using results in Lemma (ref). Result ((ref)) is established as \begin{align*} \left\Vert \mathbf{d}_{6,nT}\right\Vert & =\left\Vert \frac{1}{n}\sum _{i=1}^{n}\left( \hat{\sigma}_{i,T}-\omega_{i,T}\right) \left( \boldsymbol{\hat{\gamma}}_{i}-\boldsymbol{\gamma}_{i}\right) \right\Vert \leq\frac{1}{n}\sum_{i=1}^{n}\left\Vert \left( \hat{\sigma}_{i,T} -\omega_{i,T}\right) \left( \boldsymbol{\hat{\gamma}}_{i}-\boldsymbol{\gamma }_{i}\right) \right\Vert \\ & \leq\left( \frac{1}{n}\sum_{i=1}^{n}\left\vert \hat{\sigma}_{i,T} -\omega_{i,T}\right\vert \right) \left( \sup_{i}\left\Vert \boldsymbol{\hat {\gamma}}_{i}-\boldsymbol{\gamma}_{i}\right\Vert \right) =O_{p}\left[ \left( \frac{\ln\left( n\right) }{T}\right) ^{3/2}\right] \end{align*} where by ((ref)) $\sup_{i}\left\Vert \boldsymbol{\hat{\gamma} }_{i}-\boldsymbol{\gamma}_{i}\right\Vert =O_{p}\left( \sqrt{\ln\left( n\right) /T}\right) $, and by ((ref)) $n^{-1}\sum_{i=1}^{n}\left\vert \hat{\sigma}_{i,T}-\omega_{i,T}\right\vert =O_{p}(\ln\left( n\right) /T)$. Similarly by ((ref)) and ((ref)) we have \begin{align*} \left\Vert \mathbf{d}_{7,nT}\right\Vert & =\left\Vert \frac{1}{n}\sum _{i=1}^{n}\left( \frac{1}{\hat{\sigma}_{i,T}}-\frac{1}{\omega_{i,T}}\right) \left( \boldsymbol{\hat{\gamma}}_{i}-\boldsymbol{\gamma}_{i}\right) \right\Vert \leq\frac{1}{n}\sum_{i=1}^{n}\left\Vert \left( \frac{1} {\hat{\sigma}_{i,T}}-\frac{1}{\omega_{i,T}}\right) \left( \boldsymbol{\hat {\gamma}}_{i}-\boldsymbol{\gamma}_{i}\right) \right\Vert \\ & \leq\left( \frac{1}{n}\sum_{i=1}^{n}\left\vert \frac{1}{\hat{\sigma} _{i,T}}-\frac{1}{\omega_{i,T}}\right\vert \right) \left( \sup_{i}\left\Vert \boldsymbol{\hat{\gamma}}_{i}-\boldsymbol{\gamma}_{i}\right\Vert \right) =O_{p}\left[ \left( \frac{\ln\left( n\right) }{T}\right) ^{3/2}\right] . \end{align*}
lemmaConsider the latent factor model given by ((ref)) and ((ref)). Suppose that Assumptions (ref)-(ref) hold and $\left( n,T\right) \rightarrow\infty$, such that $n/T\rightarrow \kappa\,,$ for $0<\kappa<\infty$. Then \begin{align} p_{nT}\left( \lambda_{T}\right) & =\frac{1}{\sqrt{T}}\sum_{t=1} ^{T}s_{t,nT}^{2}\left( \lambda_{T}\right) =o_{p}(1),\\ q_{nT}\left( \lambda_{T}\right) & =\frac{1}{\sqrt{T}}\sum_{t=1}^{T} \psi_{t,nT}\left( \lambda_{T}\right) s_{t,nT}\left( \lambda_{T}\right) =o_{p}(1), \end{align} where $\psi_{t,nT}\left( \lambda_{T}\right) $ and $s_{t,nT}\left( \lambda_{T}\right) $ are defined by ((ref)) and ((ref)), respectively.
proofUsing ((ref)), recall that \begin{align} s_{t,nT}\left( \lambda_{T}\right) & =\boldsymbol{\varphi}_{nT}^{\prime }\left[ n^{-1/2}\sum_{i=1}^{n}\left( \boldsymbol{\hat{\gamma}} _{i}-\boldsymbol{\gamma}_{i}\right) \sigma_{i}\varepsilon_{it}\left( \lambda_{T}\right) \right] +\boldsymbol{\varphi}_{nT}^{\prime}\left[ n^{-1/2}\sum_{i=1}^{n}\left( \boldsymbol{\hat{\gamma}}_{i}-\boldsymbol{\gamma }_{i}\right) \boldsymbol{\gamma}_{i}^{\prime}\right] \mathbf{f} _{t}\nonumber\\ & +\left[ n^{-1/2}\sum_{i=1}^{n}\left( \boldsymbol{\hat{\delta}} _{i,T}-\boldsymbol{\delta}_{i,T}\right) ^{\prime}\right] \mathbf{f} _{t}+\left[ n^{-1/2}\sum_{i=1}^{n}\left( \boldsymbol{\hat{\delta}} _{i,T}-\boldsymbol{\delta}_{i,T}\right) ^{^{\prime}}\right] \left( \mathbf{\hat{f}}_{t}-\mathbf{f}_{t}\right) .\nonumber\\ & \end{align} We also note that using ((ref)), $\psi_{t,nT}\left( \lambda _{T}\right) $ can be written as \begin{equation} \psi_{t,nT}\left( \lambda_{T}\right) =\xi_{t,n}\left( \lambda_{T}\right) -\left( \boldsymbol{\varphi}_{nT}-\boldsymbol{\varphi}_{n}\right) ^{\prime }\boldsymbol{\kappa}_{t,n}\left( \lambda_{T}\right) +\upsilon_{t,nT}\left( \lambda_{T}\right) \end{equation} where \begin{align} \xi_{t,n}\left( \lambda_{T}\right) & =\frac{1}{\sqrt{n}}\sum_{i=1} ^{n}a_{i,n}\varepsilon_{it}\left( \lambda_{T}\right) , a_{i,n}=1-\sigma_{i}\boldsymbol{\varphi}_{n}^{\prime}\boldsymbol{\gamma} _{i},\\ \boldsymbol{\kappa}_{t,n}\left( \lambda_{T}\right) & =\frac{1}{\sqrt{n} }\sum_{i=1}^{n}\boldsymbol{\gamma}_{i}\sigma_{i}\varepsilon_{it}\left( \lambda_{T}\right) ,\\ \upsilon_{t,nT}\left( \lambda_{T}\right) & =\frac{1}{\sqrt{n}}\sum _{i=1}^{n}\left[ \frac{1}{\left( \boldsymbol{\varepsilon}_{i\circ}^{\prime }\mathbf{M}_{F}\boldsymbol{\varepsilon}_{i\circ}/T\right) ^{1/2}}-1\right] \varepsilon_{it}\left( \lambda_{T}\right) . \end{align} After squaring $s_{t,nT}\left( \lambda_{T}\right) $, we end up with $p_{nT}\left( \lambda_{T}\right) =\sum_{j=1}^{10}A_{j,nT}\left( \lambda _{T}\right) $, composed of four squared terms and six cross product terms. For the first square term we have \[ A_{1,nT}\left( \lambda_{T}\right) =\sqrt{T}\boldsymbol{\varphi}_{nT} ^{\prime}\left( \frac{1}{T}\sum_{t=1}^{T}\mathbf{b}_{t,n}\left( \lambda _{T}\right) \mathbf{b}_{t,n}^{^{\prime}}\left( \lambda_{T}\right) \right) \boldsymbol{\varphi}_{nT}, \] where $\mathbf{b}_{t,n}\left( \lambda_{T}\right) =n^{-1/2}\sum_{i=1} ^{n}\left( \boldsymbol{\hat{\gamma}}_{i}-\boldsymbol{\gamma}_{i}\right) \sigma_{i}\varepsilon_{it}\left( \lambda_{T}\right) $. Let $\mathbf{u} _{\circ t}\left( \lambda_{T}\right) =\left( \sigma_{1}\varepsilon _{1t}\left( \lambda_{T}\right) ,\sigma_{2}\varepsilon_{2t}\left( \lambda_{T}\right) ,\ldots,\sigma_{n}\varepsilon_{nt}\left( \lambda _{T}\right) \right) ^{\prime}$ so that $\mathbf{b}_{t,n}\left( \lambda _{T}\right) =n^{-1/2}\left( \boldsymbol{\hat{\Gamma}}-\boldsymbol{\Gamma }\right) ^{^{\prime}}\mathbf{u}_{\circ t}\left( \lambda_{T}\right) $. Then \begin{equation} \left\vert A_{1,nT}\left( \lambda_{T}\right) \right\vert \leq\frac{\sqrt{T} }{n}\left\Vert \boldsymbol{\varphi}_{nT}\right\Vert ^{2}\left\Vert \boldsymbol{\hat{\Gamma}}-\boldsymbol{\Gamma}\right\Vert ^{2}\left\Vert \mathbf{V}_{T}\left( \lambda_{T}\right) \right\Vert , \end{equation} where $\mathbf{V}_{T}\left( \lambda_{T}\right) =T^{-1}\sum_{t=1} ^{T}\mathbf{u}_{\circ t}\left( \lambda_{T}\right) \mathbf{u}_{\circ t}^{^{\prime}}\left( \lambda_{T}\right) .$ Since $\boldsymbol{\varphi} _{n}=n^{-1}\sum_{i=1}^{n}\boldsymbol{\gamma}_{i}/\sigma_{i}=O(1)$ by Assumption (ref), and $\boldsymbol{\varphi}_{nT} =\boldsymbol{\varphi}_{n}+o_{p}\left( 1\right) $ by result ((ref)) in Lemma (ref), then \begin{equation} \boldsymbol{\varphi}_{nT}=O_{p}\left( 1\right) . \end{equation} Further, using ((ref)) in the paper, $\mathbf{u}_{\circ t}\left( \lambda_{T}\right) =\mathbf{D}_{0}\left( \boldsymbol{\varepsilon}_{\circ t}+\lambda_{T}\mathbf{W}\boldsymbol{\varepsilon}_{\circ t}\right) $, where $\mathbf{D}_{0}=diag\left( \sigma_{1},\sigma_{2},\ldots,\sigma_{n}\right) $, $\boldsymbol{\varepsilon}_{\circ t}=\left( \varepsilon_{1t},\varepsilon _{2t},\ldots,\varepsilon_{nT}\right) ^{\prime}$, and $\mathbf{W}=\left( w_{ij}\right) $. Therefore \[ \mathbf{u}_{\circ t}\left( \lambda_{T}\right) \mathbf{u}_{\circ t}^{\prime }\left( \lambda_{T}\right) =\mathbf{D}_{0}\left( \boldsymbol{\varepsilon }_{\circ t}\boldsymbol{\varepsilon}_{\circ t}^{\prime}+\lambda_{T} ^{2}\mathbf{W}\boldsymbol{\varepsilon}_{\circ t}\boldsymbol{\varepsilon }_{\circ t}^{\prime}\mathbf{W}^{\prime}+\lambda_{T}\boldsymbol{\varepsilon }_{\circ t}\boldsymbol{\varepsilon}_{\circ t}^{\prime}\mathbf{W}^{\prime }+\lambda_{T}\mathbf{W}\boldsymbol{\varepsilon}_{\circ t} \boldsymbol{\varepsilon}_{\circ t}^{\prime}\right) \mathbf{D}_{0}, \] and \[ \mathbf{V}_{T}\left( \lambda_{T}\right) =\mathbf{D}_{0}\left( \mathbf{V}_{\varepsilon T}+\lambda_{T}^{2}\mathbf{W\mathbf{V}}_{\varepsilon T}\mathbf{W}^{\prime}+\lambda_{T}\mathbf{V}_{\varepsilon T}\mathbf{W}^{\prime }+\lambda_{T}\mathbf{W\mathbf{V}}_{\varepsilon T}\right) \mathbf{D}_{0}, \] where $\mathbf{V}_{\varepsilon T}=T^{-1}\sum_{t=1}^{T}\boldsymbol{\varepsilon }_{\circ t}\boldsymbol{\varepsilon}_{\circ t}^{\prime}$. It follows that \[ \left\Vert \mathbf{V}_{T}\left( \lambda_{T}\right) \right\Vert \leq\left\Vert \mathbf{\mathbf{V}}_{\varepsilon T}\right\Vert \left\Vert \mathbf{D}_{0}\right\Vert ^{2}\left( \left\Vert \mathbf{I}_{n}\right\Vert +\lambda_{T}^{2}\left\Vert \mathbf{W}\right\Vert \left\Vert \mathbf{W} ^{\prime}\right\Vert +\left\vert \lambda_{T}\right\vert \left\Vert \mathbf{W}^{\prime}\right\Vert +\left\vert \lambda_{T}\right\vert \left\Vert \mathbf{W}\right\Vert \right) . \] Note that $\left\Vert \mathbf{D}_{0}\right\Vert $ and $\left\Vert \mathbf{W}\right\Vert $ are both bounded, and by part (b) of Assumption (ref) we have $\left\Vert \mathbf{V}_{\varepsilon T}\right\Vert =\mu_{max}\left( \mathbf{V}_{\varepsilon T}\right) =O_{p}\left( \frac{n} {T}\right) $. Hence, \begin{equation} \left\Vert \mathbf{V}_{T}\left( \lambda_{T}\right) \right\Vert =O_{p}\left( \frac{n}{T}\right) . \end{equation} Since $n$ and $T$ are of the same order of magnitude, then using results ((ref)), ((ref)), and ((ref)) in ((ref)) yields \[ \left\vert A_{1,nT}\left( \lambda_{T}\right) \right\vert =\frac{\sqrt{T}} {n}O_{p}\left( \frac{n}{\delta_{nT}^{2}}\right) =O_{p}\left( \frac {1}{\delta_{nT}}\right) . \] For the second squared term we have \[ A_{2,nT}\left( \lambda_{T}\right) =\sqrt{T}\left[ n^{-1/2}\sum_{i=1} ^{n}\left( \boldsymbol{\hat{\delta}}_{i,T}-\boldsymbol{\delta}_{i,T}\right) ^{\prime}\right] \left( T^{-1}\mathbf{F}^{\prime}\mathbf{F}\right) \left[ n^{-1/2}\sum_{i=1}^{n}\left( \boldsymbol{\hat{\delta}}_{i,T} -\boldsymbol{\delta}_{i,T}\right) \right] . \] where $T^{-1}\mathbf{F}^{\prime}\mathbf{F}=O_{p}(1)$ and using ((ref)) $n^{-1/2}\sum_{i=1}^{n}\left( \boldsymbol{\hat{\delta}}_{i,T} -\boldsymbol{\delta}_{i,T}\right) =O_{p}\left( \sqrt{\ln\left( n\right) /T}\right) $. Hence, $A_{2,nT}\left( \lambda_{T}\right) =$ $O_{p}\left( \ln\left( n\right) /\sqrt{T}\right) =o_{p}(1)$. Similarly, \begin{align*} A_{3,nT}\left( \lambda_{T}\right) & =\sqrt{T}\boldsymbol{\varphi} _{nT}^{\prime}\left[ n^{-1/2}\sum_{i=1}^{n}\left( \boldsymbol{\hat{\gamma} }_{i}-\boldsymbol{\gamma}_{i}\right) \boldsymbol{\gamma}_{i}^{\prime}\right] \left( T^{-1}\sum_{t=1}^{T}\mathbf{f}_{t}\mathbf{f}_{t}^{\prime}\right) \left[ n^{-1/2}\sum_{i=1}^{n}\boldsymbol{\gamma}_{i}\left( \boldsymbol{\hat {\gamma}}-\boldsymbol{\gamma}_{i}\right) ^{\prime}\right] \boldsymbol{\varphi}_{nT}\\ & =\sqrt{T}\boldsymbol{\varphi}_{nT}^{\prime}\left[ n^{-1/2}\sum_{i=1} ^{n}\left( \boldsymbol{\hat{\gamma}}_{i}-\boldsymbol{\gamma}_{i}\right) \boldsymbol{\gamma}_{i}^{\prime}\right] \left( T^{-1}\mathbf{F}^{\prime }\mathbf{F}\right) \left[ n^{-1/2}\sum_{i=1}^{n}\boldsymbol{\gamma} _{i}\left( \boldsymbol{\hat{\gamma}}-\boldsymbol{\gamma}_{i}\right) ^{\prime}\right] \boldsymbol{\varphi}_{nT}, \end{align*} where $\Vert\boldsymbol{\varphi}_{nT}\Vert$ is bounded by ((ref)), and by ((ref)) $n^{-1/2}\sum_{i=1}^{n}\left( \boldsymbol{\hat{\gamma} }_{i}-\boldsymbol{\gamma}_{i}\right) \boldsymbol{\gamma}_{i}^{\prime} =O_{p}\left( \sqrt{\ln\left( n\right) /T}\right) $. Hence, $A_{3,nT} \left( \lambda_{T}\right) =$ $O_{p}\left( \ln\left( n\right) /\sqrt {T}\right) =o_{p}(1)$. For the final squared term, \[ A_{4,nT}\left( \lambda_{T}\right) =\sqrt{T}\left[ \frac{1}{\sqrt{n}} \sum_{i=1}^{n}\left( \boldsymbol{\hat{\delta}}_{i,T}-\boldsymbol{\delta }_{i,T}\right) \right] ^{\prime}\left( \frac{\left\Vert \hat{\mathbf{F} }-\mathbf{F}\right\Vert ^{2}}{T}\right) \left[ \frac{1}{\sqrt{n}}\sum _{i=1}^{n}\left( \boldsymbol{\hat{\delta}}_{i,T}-\boldsymbol{\delta} _{i,T}\right) \right] , \] By results ((ref)) and ((ref)), it follows $A_{4,nT}\left( \lambda_{T}\right) =\sqrt{T}O_{p}\left( \delta_{nT\ }^{-2}\right) \times O_{p}\left( \ln\left( n\right) /T\right) =o_{p}(1)$. The probability orders of the cross product terms of $p_{nT}\left( \lambda_{T}\right) $, namely $A_{5,NT}\left( \lambda_{T}\right) ,\ldots,A_{10,NT}\left( \lambda_{T}\right) $, are also easily seen to be $o_{p}(1)$, by application of the Cauchy-Schwarz inequality to the product pairs of the terms $A_{1,NT}\left( \lambda_{T}\right) ,A_{2,NT}\left( \lambda_{T}\right) ,A_{3,NT}\left( \lambda_{T}\right) ,$ and $A_{4,NT}\left( \lambda _{T}\right) $. Thus, overall $p_{nT}\left( \lambda_{T}\right) =o_{p}(1)$, as required. Consider now $q_{nT}\left( \lambda_{T}\right) $ and note that it can be written as (using ((ref)) in ((ref))) \begin{align*} q_{nT}\left( \lambda_{T}\right) & =\frac{1}{\sqrt{T}}\sum_{t=1} ^{T}s_{t,nT}\left( \lambda_{T}\right) \xi_{t,n}\left( \lambda_{T}\right) -\left( \boldsymbol{\varphi}_{nT}-\boldsymbol{\varphi}_{n}\right) ^{\prime }\frac{1}{\sqrt{T}}\sum_{t=1}^{T}s_{t,nT}\left( \lambda_{T}\right) \boldsymbol{\kappa}_{t,n}\left( \lambda_{T}\right) \\ & +\frac{1}{\sqrt{T}}\sum_{t=1}^{T}s_{t,nT}\left( \lambda_{T}\right) \upsilon_{t,nT}\left( \lambda_{T}\right) , \end{align*} where $\xi_{t,n}\left( \lambda_{T}\right) $, $\upsilon_{t,nT}\left( \lambda_{T}\right) $ and $\boldsymbol{\kappa}_{t,n}\left( \lambda _{T}\right) \mathbf{\ }$are given by ((ref)), ((ref)) and ((ref)), respectively. The first term of the above can be written as \begin{align*} \frac{1}{\sqrt{T}}\sum_{t=1}^{T}s_{t,nT}\left( \lambda_{T}\right) \xi _{t,n}\left( \lambda_{T}\right) & =\boldsymbol{\varphi}_{nT}^{\prime }\left[ \frac{1}{\sqrt{n}}\sum_{i=1}^{n}\left( \boldsymbol{\hat{\gamma}} _{i}-\boldsymbol{\gamma}_{i}\right) \frac{1}{\sqrt{T}}\sum_{t=1}^{T}\xi _{t,n}\left( \lambda_{T}\right) \sigma_{i}\varepsilon_{it}\left( \lambda_{T}\right) \right] \\ & +\boldsymbol{\varphi}_{nT}^{\prime}\left[ n^{-1/2}\sum_{i=1}^{n}\left( \boldsymbol{\hat{\gamma}}_{i}-\boldsymbol{\gamma}_{i}\right) \boldsymbol{\gamma}_{i}^{\prime}\right] \left( \frac{1}{\sqrt{T}}\sum _{t=1}^{T}\xi_{t,n}\left( \lambda_{T}\right) \mathbf{f}_{t}\right) \\ & +\left[ n^{-1/2}\sum_{i=1}^{n}\left( \boldsymbol{\hat{\delta}} _{i,T}-\boldsymbol{\delta}_{i,T}\right) ^{\prime}\right] \left( \frac {1}{\sqrt{T}}\sum_{t=1}^{T}\xi_{t,n}\left( \lambda_{T}\right) \mathbf{f} _{t}\right) \\ & +\left[ n^{-1/2}\sum_{i=1}^{n}\left( \boldsymbol{\hat{\delta}} _{i,T}-\boldsymbol{\delta}_{i,T}\right) ^{^{\prime}}\right] \left[ \frac {1}{\sqrt{T}}\sum_{t=1}^{T}\xi_{t,n}\left( \lambda_{T}\right) \left( \mathbf{\hat{f}}_{t}-\mathbf{f}_{t}\right) \right] \\ & =\sum_{j=1}^{4}B_{j,nT}\left( \lambda_{T}\right) . \end{align*} Using ((ref)), $B_{1,nT}\left( \lambda_{T}\right) $ can be written as \[ B_{1,nT}\left( \lambda_{T}\right) =\boldsymbol{\varphi}_{nT}^{\prime}\left[ \frac{\sqrt{T}}{\sqrt{n}}\sum_{i=1}^{n}\left( \boldsymbol{\hat{\gamma}} _{i}-\boldsymbol{\gamma}_{i}\right) \frac{1}{\sqrt{n}}\sum_{j=1}^{n} a_{j,n}\left( \frac{1}{T}\sum_{t=1}^{T}\sigma_{i}\varepsilon_{jt}\left( \lambda_{T}\right) \varepsilon_{it}\left( \lambda_{T}\right) \right) , \right] \] where $a_{i,n}=1-\sigma_{i}\boldsymbol{\varphi}_{n}^{\prime}\boldsymbol{\gamma }_{i}$ and $\boldsymbol{\varphi}_{nT}=O_{p}(1)$. Also since $\varepsilon _{it}\left( \lambda_{T}\right) $ are independently distributed over $t$ and weakly cross-sectionally dependent, and $n$ and $T$ are of the same order, then \[ B_{1,nT}\left( \lambda_{T}\right) =O_{p}\left( \frac{1}{\sqrt{n}}\sum _{i=1}^{n}a_{i,n}\sigma_{i}\left( \boldsymbol{\hat{\gamma}}_{i} -\boldsymbol{\gamma}_{i}\right) \right) . \] Further, letting $b_{in}=a_{i,n}\sigma_{i}$, it follows from ((ref)) that $n^{-1/2}\sum_{i=1}^{n}a_{i,n}\sigma_{i}\left( \boldsymbol{\hat{\gamma}} _{i}-\boldsymbol{\gamma}_{i}\right) =O_{p}(\sqrt{\ln\left( n\right) /T})$, which in turn establishes that $B_{1,nT}\left( \lambda_{T}\right) =o_{p} (1)$. Similarly, using ((ref)), $B_{2,nT}\left( \lambda_{T}\right) $ can be written as \[ B_{2,nT}\left( \lambda_{T}\right) =\boldsymbol{\varphi}_{nT}^{\prime}\left[ n^{-1/2}\sum_{i=1}^{n}\left( \boldsymbol{\hat{\gamma}}_{i}-\boldsymbol{\gamma }_{i}\right) \boldsymbol{\gamma}_{i}^{\prime}\right] \left( \frac{1} {\sqrt{nT}}\sum_{j=1}^{n}\sum_{t=1}^{T}a_{j,n}\mathbf{f}_{t}\varepsilon _{jt}\left( \lambda_{T}\right) \right) , \] where $\boldsymbol{\varphi}_{nT}=O_{p}(1)$. Under parts (a) and (c) of Assumption (ref) $ \frac{1}{\sqrt{nT}}\sum_{j=1}^{n}\sum_{t=1}^{T}a_{j,n}\mathbf{f} _{t}\varepsilon_{jt}\left( \lambda_{T}\right) =O_{p}(1). $ Using this result together with ((ref)) it follows that $B_{2,nT}\left( \lambda_{T}\right) =o_{p}(1)$. Similarly, using ((ref)) we can establish that $B_{3,nT}\left( \lambda_{T}\right) =o_{p}(1)$. The final term, $B_{4,nT}\left( \lambda_{T}\right) $, is dominated by the third term and is also $o_{p}(1)$. Thus overall, $T^{-1/2}\sum_{t=1}^{T}s_{t,nT}\left( \lambda_{T}\right) \xi_{t,n}\left( \lambda_{T}\right) =o_{p}(1)$. Using the same line of reasoning, it is also readily established that $T^{-1/2} \sum_{t=1}^{T}s_{t,nT}\left( \lambda_{T}\right) \boldsymbol{\kappa} _{t,n}\left( \lambda_{T}\right) =o_{p}(1)$, considering that, $\boldsymbol{\kappa}_{t,n}\left( \lambda_{T}\right) $ $=n^{-1/2}\sum _{i=1}^{n}\boldsymbol{\gamma}_{i}\sigma_{i}\varepsilon_{it}\left( \lambda _{T}\right) $ has the same format as $\xi_{t,n}\left( \lambda_{T}\right) $, and in addition by ((ref)) $\boldsymbol{\varphi}_{nT} -\boldsymbol{\varphi}_{n}=O_{p}(n^{-1/2}T^{-1/2})+O_{p}(T^{-1})$. Finally, the last term of $q_{nT}$ is given by \begin{align*} \frac{1}{\sqrt{T}}\sum_{t=1}^{T}s_{t,nT}\left( \lambda_{T}\right) \upsilon_{t,nT}\left( \lambda_{T}\right) & =\frac{1}{\sqrt{T}}\sum _{t=1}^{T}\upsilon_{t,nT}\left( \lambda_{T}\right) \boldsymbol{\varphi} _{nT}^{\prime}\left[ n^{-1/2}\sum_{i=1}^{n}\left( \boldsymbol{\hat{\gamma} }_{i}-\boldsymbol{\gamma}_{i}\right) \sigma_{i}\varepsilon_{it}\left( \lambda_{T}\right) \right] \\ & +\frac{1}{\sqrt{T}}\sum_{t=1}^{T}\upsilon_{t,nT}\left( \lambda_{T}\right) \boldsymbol{\varphi}_{nT}^{\prime}\left[ n^{-1/2}\sum_{i=1}^{n}\left( \boldsymbol{\hat{\gamma}}_{i}-\boldsymbol{\gamma}_{i}\right) \boldsymbol{\gamma}_{i}^{\prime}\right] \mathbf{f}_{t}\\ & +\frac{1}{\sqrt{T}}\sum_{t=1}^{T}\upsilon_{t,nT}\left( \lambda_{T}\right) \left[ n^{-1/2}\sum_{i=1}^{n}\left( \boldsymbol{\hat{\delta}}_{i,T} -\boldsymbol{\delta}_{i,T}\right) ^{\prime}\right] \mathbf{f}_{t}\\ & +\frac{1}{\sqrt{T}}\sum_{t=1}^{T}\upsilon_{t,nT}\left( \lambda_{T}\right) \left[ n^{-1/2}\sum_{i=1}^{n}\left( \boldsymbol{\hat{\delta}}_{i,T} -\boldsymbol{\delta}_{i,T}\right) ^{^{\prime}}\right] \left( \mathbf{\hat {f}}_{t}-\mathbf{f}_{t}\right) \\ & =\sum_{j=1}^{4}C_{j,nT}\left( \lambda_{T}\right) . \end{align*} Using ((ref)) we have \begin{align*} C_{1,nT}\left( \lambda_{T}\right) & =\frac{1}{\sqrt{T}}\sum_{t=1}^{T} \frac{1}{\sqrt{n}}\sum_{j=1}^{n}\left( \frac{1}{\left( \boldsymbol{\varepsilon}_{j\circ}^{\prime}\mathbf{M}_{F} \boldsymbol{\varepsilon}_{j\circ}/T\right) ^{1/2}}-1\right) \varepsilon _{jt}\left( \lambda_{T}\right) \boldsymbol{\varphi}_{nT}^{\prime}\left[ n^{-1/2}\sum_{i=1}^{n}\left( \boldsymbol{\hat{\gamma}}_{i}-\boldsymbol{\gamma }_{i}\right) \sigma_{i}\varepsilon_{it}\left( \lambda_{T}\right) \right] \\ & =\sqrt{\frac{T}{n}}\boldsymbol{\varphi}_{nT}^{\prime}\sum_{j=1}^{n}\left[ n^{-1/2}\sum_{i=1}^{n}\left( \boldsymbol{\hat{\gamma}}_{i}-\boldsymbol{\gamma }_{i}\right) \frac{1}{T}\sum_{t=1}^{T}\sigma_{i}\left( \frac{1}{\left( \boldsymbol{\varepsilon}_{j\circ}^{\prime}\mathbf{M}_{F} \boldsymbol{\varepsilon}_{j\circ}/T\right) ^{1/2}}-1\right) \varepsilon _{it}\left( \lambda_{T}\right) \varepsilon_{jt}\left( \lambda_{T}\right) \right] . \end{align*} Since $\varepsilon_{it}\left( \lambda_{T}\right) $ is distributed independently over $t$ and weakly cross-sectionally dependent, then \[ \frac{1}{T}\sum_{t=1}^{T}\sigma_{i}\left( \frac{1}{\left( \boldsymbol{\varepsilon}_{j\circ}^{\prime}\mathbf{M}_{F} \boldsymbol{\varepsilon}_{j\circ}/T\right) ^{1/2}}-1\right) \varepsilon _{it}\left( \lambda_{T}\right) \varepsilon_{jt}\left( \lambda_{T}\right) \rightarrow_{p}0\text{, if }i\neq j\text{, } \] and \begin{align*} & \frac{1}{T}\sum_{t=1}^{T}\sigma_{i}\left( \frac{1}{\left( \boldsymbol{\varepsilon}_{j\circ}^{\prime}\mathbf{M}_{F} \boldsymbol{\varepsilon}_{j\circ}/T\right) ^{1/2}}-1\right) \varepsilon _{it}\left( \lambda_{T}\right) \varepsilon_{jt}\left( \lambda_{T}\right) \\ & \rightarrow_{p}\lim_{T\rightarrow\infty}\frac{1}{T}\sum_{t=1}^{T}\sigma _{i}E\left\{ \left[ \left( \frac{\boldsymbol{\varepsilon}_{i\circ}^{\prime }\mathbf{M}_{F}\boldsymbol{\varepsilon}_{i\circ}}{T}\right) ^{-1/2}-1\right] \varepsilon_{it}^{2}\left( \lambda_{T}\right) \right\} , if i=j. \end{align*} Also, by definition $\varepsilon_{it}\left( \lambda_{T}\right) =\varepsilon_{it}+\lambda_{T}b_{it}$, where $b_{it}=\mathbf{w}_{i0}^{\prime }\boldsymbol{\varepsilon}_{\circ t}$, $w_{ii}=0$, and $\varepsilon_{it}$ is independent from $b_{it}$, and therefore \begin{align*} & E\left\{ \left[ \left( \frac{\boldsymbol{\varepsilon}_{i\circ}^{\prime }\mathbf{M}_{F}\boldsymbol{\varepsilon}_{i\circ}}{T}\right) ^{-1/2}-1\right] \varepsilon_{it}^{2}\left( \lambda_{T}\right) \right\} \\ & =E\left\{ \left[ \left( \frac{\boldsymbol{\varepsilon}_{i\circ}^{\prime }\mathbf{M}_{F}\boldsymbol{\varepsilon}_{i\circ}}{T}\right) ^{-1/2}-1\right] \varepsilon_{it}^{2}\right\} +E\left[ \left( \frac{\boldsymbol{\varepsilon }_{i\circ}^{\prime}\mathbf{M}_{F}\boldsymbol{\varepsilon}_{i\circ}}{T}\right) ^{-1/2}-1\right] \lambda_{T}^{2}E\left( b_{it}^{2}\right) \\ & =O\left( \frac{1}{T}\right) +O\left( \frac{1}{T^{2}}\right) , \end{align*} where the last line holds by ((ref)), ((ref)) and ((ref)). Moreover, by ((ref)) $n^{-1/2}\sum_{i=1}^{n}\left( \boldsymbol{\hat {\gamma}}_{i}-\boldsymbol{\gamma}_{i}\right) =O_{p}(\sqrt{\ln\left( n\right) /T})$. As $n$ and $T$ being of the same order, it then follows that $C_{1,nT}\left( \lambda_{T}\right) =o_{p}(1)$. Similarly to $B_{2,nT}\left( \lambda_{T}\right) $, we have \begin{align*} C_{2,nT}\left( \lambda_{T}\right) & =\boldsymbol{\varphi}_{nT}^{\prime }\left[ n^{-1/2}\sum_{i=1}^{n}\left( \boldsymbol{\hat{\gamma}} _{i}-\boldsymbol{\gamma}_{i}\right) \boldsymbol{\gamma}_{i}^{\prime}\right] \left[ \frac{1}{\sqrt{nT}}\sum_{j=1}^{n}\sum_{t=1}^{T}\left( \frac {1}{\left( \boldsymbol{\varepsilon}_{j\circ}^{\prime}\mathbf{M} _{F}\boldsymbol{\varepsilon}_{j\circ}/T\right) ^{1/2}}-1\right) \varepsilon_{jt}\left( \lambda_{T}\right) \mathbf{f}_{t}\right] \\ & =O_{p}\left( \sqrt{\frac{\ln\left( n\right) }{T}}\right) O_{p} (1)=o_{p}(1). \end{align*} The same line of reasoning as used for $B_{3,nT}\left( \lambda_{T}\right) $ and $B_{4,nT}\left( \lambda_{T}\right) $ can be used to establish $C_{j,nT}\left( \lambda_{T}\right) =o_{p}(1)$ for $j=3$ and $4.$ Hence, $T^{-1/2}\sum_{t=1}^{T}s_{t,nT}\left( \lambda_{T}\right) \upsilon _{t,nT}\left( \lambda_{T}\right) =o_{p}(1)$, and overall we have $q_{nT}\left( \lambda_{T}\right) =o_{p}(1)$, as required.
lemmaConsider the latent factor model given by ((ref)) and ((ref)). Suppose that Assumptions (ref)-(ref) hold and $\left( n,T\right) \rightarrow\infty$, such that $n/T\rightarrow \kappa\,,$ for $0<\kappa<\infty$. Then \begin{equation} \sqrt{T}\left( \boldsymbol{\varphi}_{n}-\boldsymbol{\varphi}_{nT}\right) =O_{p}\left( n^{-1/2}\right) +O_{p}\left( T^{-1/2}\right) , \end{equation} \begin{equation} \sqrt{T}\left( \boldsymbol{\hat{\varphi}}_{nT}-\boldsymbol{\varphi} _{n}\right) =o_{p}(1), \end{equation} where $\boldsymbol{\varphi}_{n}=n^{-1}\sum_{i=1}^{n}\boldsymbol{\gamma} _{i}/\sigma_{i},$ $\boldsymbol{\varphi}_{nT}=n^{-1}\sum_{i=1}^{n} \boldsymbol{\gamma}_{i}/\omega_{i,T}$ with $\omega_{i,T}=\left( T^{-1} \sigma_{i}^{2}\boldsymbol{\varepsilon}_{i\circ}^{\prime}\mathbf{M} _{F}\boldsymbol{\varepsilon}_{i\circ}\right) ^{1/2}$, $\boldsymbol{\hat {\varphi}}_{nT}=n^{-1}\sum_{i=1}^{n}\boldsymbol{\hat{\gamma}}_{i}/\hat{\sigma }_{i,T}$, $\hat{\sigma}_{i,T}=\left( T^{-1}\mathbf{y}_{i}^{\prime} \mathbf{M}_{\hat{F}}\mathbf{y}_{i}\right) ^{1/2}$ and $\boldsymbol{\hat {\gamma}}_{i}$ and $\mathbf{\hat{F}}$ are the principal component estimators of $\boldsymbol{\gamma}_{i}$ and $\mathbf{F}$.
proofFirst note that \begin{align*} \sqrt{T}\left( \boldsymbol{\varphi}_{n}-\boldsymbol{\varphi}_{nT}\right) & =\frac{\sqrt{T}}{n}\sum_{i=1}^{n}\frac{\boldsymbol{\gamma}_{i}}{\sigma_{i} }\left\{ \left( 1-\frac{\sigma_{i}}{\omega_{i,T}}\right) -\left[ 1-E\left( \frac{\sigma_{i}}{\omega_{i,T}}\right) \right] \right\} +\frac{\sqrt{T}}{n}\sum_{i=1}^{n}\frac{\boldsymbol{\gamma}_{i}}{\sigma_{i} }\left[ 1-E\left( \frac{\sigma_{i}}{\omega_{i,T}}\right) \right] ,\\ & =\mathbf{g}_{1,nT}+\mathbf{g}_{2,nT}, \end{align*} where \begin{align*} \mathbf{g}_{1,nT} & =-\frac{1}{n}\sum_{i=1}^{n}\sqrt{T}\left[ \frac {\sigma_{i}}{\omega_{i,T}}-E\left( \frac{\sigma_{i}}{\omega_{i,T}}\right) \right] \frac{\boldsymbol{\gamma}_{i}}{\sigma_{i}},\\ \mathbf{g}_{2,nT} & =\frac{\sqrt{T}}{n}\sum_{i=1}^{n}\left[ 1-E\left( \frac{\sigma_{i}}{\omega_{i,T}}\right) \right] \frac{\boldsymbol{\gamma} _{i}}{\sigma_{i}}. \end{align*} Since $\sigma_{i}/\omega_{i,T}=\left( T^{-1}\boldsymbol{\varepsilon}_{i\circ }^{\prime}\mathbf{M}_{F}\boldsymbol{\varepsilon}_{i\circ}\right) ^{-1/2}$, $\left\Vert \boldsymbol{\gamma}_{i}\right\Vert <C$, $\sigma_{i}<C$, then using result ((ref)) we have $E\left( \frac{\sigma_{i}}{\omega_{i,T}}\right) =1+O\left( T^{-1}\right) $, and $\mathbf{g}_{2,nT}=$ $O\left( T^{-1/2}\right) $. The first term can be written as $\mathbf{g}_{1,nT} =n^{-1}\sum_{i=1}^{n}\boldsymbol{\gamma}_{i}\sigma_{i}^{-1}\chi_{i,T}$, where $\chi_{i,T}=-\sqrt{T}\left[ \sigma_{i}/\omega_{i,T}-E\left( \sigma _{i}/\omega_{i,T}\right) \right] $. Conditional on $\mathbf{F}$ and $\sigma_{i}$, $\chi_{i,T}$ are distributed independently over $i$ with mean zero and bounded variances:\footnote{When $\varepsilon_{it}$ are normally distributed we have the exact result $E\left( \frac{T} {\boldsymbol{\varepsilon}_{i\circ}^{\prime}\mathbf{M}_{F} \boldsymbol{\varepsilon}_{i\circ}}\right) =T/(T-m_{0}-2).$} \[ Var\left( \chi_{i,T}\right) =T\left[ E\left( \frac{T} {\boldsymbol{\varepsilon}_{i\circ}^{\prime}\mathbf{M}_{F} \boldsymbol{\varepsilon}_{i\circ}}\right) -\left[ E\left( \frac{\sigma_{i} }{\omega_{i,T}}\right) \right] ^{2}\right] =T\left[ 1+O\left( \frac{1} {T}\right) -\left[ 1+O\left( \frac{1}{T}\right) \right] ^{2}\right] =O(1). \] Hence, $\mathbf{g}_{1,nT}=O_{p}\left( n^{-1/2}\right) $, and the desired result ((ref)) follows. Consider now ((ref)) and note that it can be decomposed as \begin{equation} \sqrt{T}\left( \boldsymbol{\hat{\varphi}}_{nT}-\boldsymbol{\varphi} _{n}\right) =\sqrt{T}\left( \boldsymbol{\varphi}_{nT}-\boldsymbol{\varphi }_{n}\right) +\sqrt{T}\left( \boldsymbol{\hat{\varphi}}_{nT} -\boldsymbol{\varphi}_{nT}\right) , \end{equation} where it is already established that the first term is $o_{p}(1)$. Consider now the second term of ((ref)) and note that it can be written as \begin{align*} \sqrt{T}\left( \boldsymbol{\hat{\varphi}}_{nT}-\boldsymbol{\varphi} _{nT}\right) & =\frac{\sqrt{T}}{n}\sum_{i=1}^{n}\left( \frac {\boldsymbol{\hat{\gamma}}_{i}}{\hat{\sigma}_{i,T}}\boldsymbol{-} \frac{\boldsymbol{\gamma}_{i}}{\omega_{i,T}}\right) \\ & =\frac{\sqrt{T}}{n}\sum_{i=1}^{n}\boldsymbol{\gamma}_{i}\left( \frac {1}{\hat{\sigma}_{i,T}}\boldsymbol{-}\frac{1}{\omega_{i,T}}\right) +\frac{\sqrt{T}}{n}\sum_{i=1}^{n}\frac{1}{\sigma_{i}}\left( \boldsymbol{\hat {\gamma}}_{i}-\boldsymbol{\gamma}_{i}\right) +\\ & \frac{\sqrt{T}}{n}\sum_{i=1}^{n}\left( \frac{1}{\omega_{i,T}}-\frac {1}{\sigma_{i}}\right) \left( \boldsymbol{\hat{\gamma}}_{i} -\boldsymbol{\gamma}_{i}\right) +\frac{\sqrt{T}}{n}\sum_{i=1}^{n}\left( \frac{1}{\hat{\sigma}_{i,T}}-\frac{1}{\omega_{i,T}}\right) \left( \boldsymbol{\hat{\gamma}}_{i}-\boldsymbol{\gamma}_{i}\right) . \end{align*} Now using ((ref)) of Lemma (ref)\ we have \[ \frac{\sqrt{T}}{n}\sum_{i=1}^{n}\boldsymbol{\gamma}_{i}\left( \frac{1} {\hat{\sigma}_{i,T}}-\frac{1}{\omega_{i,T}}\right) =O_{p}\left( \frac {\ln\left( n\right) }{\sqrt{T}}\right) . \] Using this result as well as ((ref)), ((ref)) and ((ref)), we have $\sqrt{T}\left( \boldsymbol{\hat{\varphi}}_{nT}-\boldsymbol{\varphi }_{nT}\right) =o_{p}(1)$ as required.
lemmaSuppose $\mathbf{M}_{F}=\mathbf{I}_{T}-\mathbf{F(F}^{\prime }\mathbf{F)}^{-1}\mathbf{F}^{\prime}$, where $\mathbf{F}$ is a $T\times m_{0}$ matrix, and $\boldsymbol{\tau}_{T}$ is a $T\times1$ vector of ones. Then \begin{align} \mathrm{tr}\left( \mathbf{M}_{F}\right) & =v, \mathrm{tr}\left( \mathbf{M}_{F}\mathbf{\odot M}_{F}\right) =O\left( v\right) ,\\ \mathrm{tr}\left( \mathbf{M}_{F}\mathbf{\odot M}_{F}\mathbf{\odot M} _{F}\right) & =O\left( v\right) ,\mathrm{tr}\left( \mathbf{M} _{F}\mathbf{\odot M}_{F}\mathbf{\odot M}_{F}\mathbf{\odot M}_{F}\right) =O\left( v\right) ,\\ \mathrm{tr}\left[ \left( \mathbf{I}_{T}\mathbf{\odot M}_{F}\right) \mathbf{M}_{F}\right] & =O\left( v\right) ,\mathrm{tr}\left[ \left( \mathbf{M}_{F}\mathbf{\odot M}_{F}\right) \mathbf{M}_{F}\right] =O\left( v\right) , \end{align} \begin{align} \boldsymbol{\tau}_{T}^{\prime}\left( \mathbf{M}_{F}\mathbf{\odot M} _{F}\right) \boldsymbol{\tau}_{T} & =O\left( v\right) ,\boldsymbol{\tau }_{T}^{\prime}\left( \mathbf{M}_{F}\mathbf{\odot M}_{F}\mathbf{\odot M} _{F}\right) \boldsymbol{\tau}_{T}=O\left( v\right) ,\\ \boldsymbol{\tau}_{T}^{\prime}\left( \mathbf{M}_{F}\mathbf{\odot M} _{F}\mathbf{\odot M}_{F}\mathbf{\odot M}_{F}\right) \boldsymbol{\tau}_{T} & =O\left( v\right) ,\boldsymbol{\tau}_{T}^{\prime}\left( \mathbf{I} _{T}\mathbf{\odot M}_{F}\right) \left( \mathbf{M}_{F}\mathbf{\odot M} _{F}\mathbf{\odot M}_{F}\right) \boldsymbol{\tau}_{T}=O\left( v\right) ,\\ \boldsymbol{\tau}_{T}^{\prime}\left( \mathbf{I}_{T}\mathbf{\odot M} _{F}\right) \left( \mathbf{M}_{F}\mathbf{\odot M}_{F}\right) \left( \mathbf{I}_{T}\mathbf{\odot M}_{F}\right) \boldsymbol{\tau}_{T} & =O\left( v\right) ,\boldsymbol{\tau}_{T}^{\prime}\left( \mathbf{I}_{T}\mathbf{\odot M}_{F}\right) \left( \mathbf{I}_{T}\mathbf{\odot M}_{F}\right) \boldsymbol{\tau}_{T}=O\left( v\right) , \end{align} \begin{align} \boldsymbol{\tau}_{T}^{\prime}\left( \mathbf{I}_{T}\mathbf{\odot M} _{F}\right) \mathbf{M}_{F}\left( \mathbf{I}_{T}\mathbf{\odot M}_{F}\right) \boldsymbol{\tau}_{T} & =O\left( v^{3/2}\right) ,\boldsymbol{\tau} _{T}^{\prime}\left( \mathbf{M}_{F}\mathbf{\odot M}_{F}\right) \mathbf{M} _{F}\left( \mathbf{I}_{T}\mathbf{\odot M}_{F}\right) \boldsymbol{\tau} _{T}=O\left( v^{3/2}\right) ,\\ \boldsymbol{\tau}_{T}^{\prime}\left( \mathbf{I}_{T}\mathbf{\odot M} _{F}\right) \mathbf{M}_{F}\left( \mathbf{I}_{T}\mathbf{\odot M} _{F}\mathbf{\odot M}_{F}\right) \boldsymbol{\tau}_{T} & =O\left( v^{3/2}\right) , \end{align} where $v=T-m_{0}$.
proofSee Lemma 10 of Pesaran and Yamagata (2024).
lemmaSuppose the $T\times1$ vector $\boldsymbol{\varepsilon} =\mathbf{(}\varepsilon_{1},\varepsilon_{2},...,\varepsilon_{T})^{\prime}$ is $\boldsymbol{\varepsilon}\thicksim IID(\mathbf{0},\mathbf{I}_{T})$, $\sup _{t}E\left( \left\vert \varepsilon_{t}\right\vert ^{8+s}\right) <C$ for some small $s>0$, and $\mathbf{M}_{F}=\mathbf{I}_{T}-\mathbf{F(F}^{\prime }\mathbf{F)}^{-1}\mathbf{F}^{\prime}$, where the $T\times m_{0}$ matrix $\mathbf{F}$ is distributed independently of $\boldsymbol{\varepsilon}$, then \begin{equation} E\left( \frac{\boldsymbol{\varepsilon}^{\prime}\mathbf{M}_{F} \boldsymbol{\varepsilon}}{v}\right) =1, \end{equation} \begin{equation} E\left[ \left( \frac{\boldsymbol{\varepsilon}^{\prime}\mathbf{M} _{F}\boldsymbol{\varepsilon}}{v}\right) ^{2}\right] =1+O\left( \frac{1} {v}\right) , \end{equation} \begin{equation} E\left( q_{v}\right) =0\mathbf{,} E\left( q_{v}^{4}\right) =O(1)\mathbf{,} \end{equation} where $v=T-m_{0}$, and $q_{v}=\sqrt{v}\left( \frac{\boldsymbol{\varepsilon }^{\prime}\mathbf{M}_{F}\boldsymbol{\varepsilon}}{v}-1\right) $.
proofDenote $\boldsymbol{\tau}_{T}$ as a $T\times1$ vector of ones and suppose $\gamma_{1}=E\left( \varepsilon_{t}^{3}\right) $, $\gamma_{2}=E\left( \varepsilon_{t}^{4}\right) -3$, $\gamma_{3}=E\left( \varepsilon_{t} ^{5}\right) -10\gamma_{1}$, $\gamma_{4}=E\left( \varepsilon_{t}^{6}\right) -15\gamma_{2}-10\gamma_{1}^{2}-15$, $\gamma_{6}=E\left( \varepsilon_{t} ^{8}\right) -28\gamma_{4}-56\gamma_{3}\gamma_{1}-35\gamma_{2}^{2} -210\gamma_{2}-280\gamma_{1}^{2}-105$, which are all bounded as it is assumed $\sup_{t}E\left( \left\vert \varepsilon_{t}\right\vert ^{8+\epsilon}\right) <C$. Since $\boldsymbol{\varepsilon}\thicksim IID(\mathbf{0},\mathbf{I}_{T})$ and $\mathbf{M}_{F}=\left( m_{tt^{\prime}}\right) $ is an idempotent matrix then results (S.6) to (S.9) of Lemma 6 in Pesaran and Yamagata (2024) apply and we have (since $\mathrm{tr}\left( \mathbf{M}_{F}^{s}\right) =\mathrm{tr}\left( \mathbf{M}_{F}\right) =v$, for $s=1,2,\ldots$) \begin{equation} E\left( \boldsymbol{\varepsilon}^{\prime}\mathbf{M}_{F} \boldsymbol{\varepsilon}\right) =v, \end{equation} \begin{equation} E\left[ \left( \boldsymbol{\varepsilon}^{\prime}\mathbf{M}_{F} \boldsymbol{\varepsilon}\right) ^{2}\right] =v^{2}+2v+\gamma_{2} \mathrm{tr}\left( \mathbf{M}_{F}\mathbf{\odot M}_{F}\right) , \end{equation} \begin{align} E\left[ \left( \boldsymbol{\varepsilon}^{\prime}\mathbf{M}_{F} \boldsymbol{\varepsilon}\right) ^{3}\right] & =v^{3}+6v^{2}+8v+\gamma _{4}\mathrm{tr}\left( \mathbf{M}_{F}\mathbf{\odot M}_{F}\mathbf{\odot M} _{F}\right) +3\gamma_{2}\left( v+4\right) \mathrm{tr}\left( \mathbf{M} _{F}\mathbf{\odot M}_{F}\right) \nonumber\\ & +6\gamma_{1}^{2}\left[ \boldsymbol{\tau}_{T}^{\prime}\left( \mathbf{I}_{T}\mathbf{\odot M}_{F}\right) \mathbf{M}_{F}\left( \mathbf{I}_{T}\mathbf{\odot M}_{F}\right) \boldsymbol{\tau}_{T}\right] +4\gamma_{1}^{2}\left[ \boldsymbol{\tau}_{T}^{\prime}\left( \mathbf{M} _{F}\mathbf{\odot M}_{F}\mathbf{\odot M}_{F}\right) \boldsymbol{\tau} _{T}\right] \end{align} \begin{equation} E\left[ \left( \boldsymbol{\varepsilon}^{\prime}\mathbf{M}_{F} \boldsymbol{\varepsilon}\right) ^{4}\right] =v^{4}+12v^{3}+44v^{2} +48v+\gamma_{2}f_{\gamma_{2}}+\gamma_{4}f_{\gamma_{4}}+\gamma_{6}f_{\gamma _{6}}+\gamma_{1}^{2}f_{\gamma_{1}^{2}}+\gamma_{2}^{2}f_{\gamma_{2}^{2}} +\gamma_{1}\gamma_{3}f_{\gamma_{1}\gamma_{3}}, \end{equation} where \begin{align*} f_{\gamma_{2}} & =\left( 6v^{2}+48v\right) \mathrm{tr}\left( \mathbf{M}_{F}\mathbf{\odot M}_{F}\right) +12\left[ \boldsymbol{\tau} _{T}^{\prime}\left( \mathbf{M}_{F}\mathbf{\odot M}_{F}\right) \boldsymbol{\tau}_{T}\mathrm{tr}\left( \mathbf{M}_{F}\mathbf{\odot M} _{F}\right) \right] \\ & +96\mathrm{tr}\left[ \left( \mathbf{I}_{T}\mathbf{\odot M}_{F}\right) \mathbf{M}_{F}\right] +48\boldsymbol{\tau}_{T}^{\prime}\left( \mathbf{I} _{T}\mathbf{\odot M}_{F}\right) \left( \mathbf{I}_{T}\mathbf{\odot M} _{F}\right) \boldsymbol{\tau}_{T}, \end{align*} \[ f_{\gamma_{4}}=\left( 4v+24\right) \mathrm{tr}\left( \mathbf{M} _{F}\mathbf{\odot M}_{F}\mathbf{\odot M}_{F}\right) , \] \[ f_{\gamma_{6}}=\mathrm{tr}\left( \mathbf{M}_{F}\mathbf{\odot M} _{F}\mathbf{\odot M}_{F}\mathbf{\odot M}_{F}\right) , \] \begin{align*} f_{\gamma_{1}^{2}} & =24v\boldsymbol{\tau}_{T}^{\prime}\left( \mathbf{I}_{T}\mathbf{\odot M}_{F}\right) \mathbf{M}_{F}\left( \mathbf{I}_{T}\mathbf{\odot M}_{F}\right) \boldsymbol{\tau}_{T} +48\boldsymbol{\tau}_{T}^{\prime}\left( \mathbf{I}_{T}\mathbf{\odot M} _{F}\right) \mathbf{M}_{F}\left( \mathbf{I}_{T}\mathbf{\odot M}_{F}\right) \boldsymbol{\tau}_{T}\\ & +16v\boldsymbol{\tau}_{T}^{\prime}\left( \mathbf{M}_{F}\mathbf{\odot M}_{F}\mathbf{\odot M}_{F}\right) \boldsymbol{\tau}_{T}+96\boldsymbol{\tau }_{T}^{\prime}\left( \mathbf{M}_{F}\mathbf{\odot M}_{F}\right) \mathbf{M}_{F}\left( \mathbf{I}_{T}\mathbf{\odot M}_{F}\right) \boldsymbol{\tau}_{T}\\ & +96\mathrm{tr}\left[ \mathbf{M}_{F}\left( \mathbf{M}_{F}\mathbf{\odot M}_{F}\right) \mathbf{M}_{F}\right] , \end{align*} \begin{align*} f_{\gamma_{2}^{2}} & =3\left[ \mathrm{tr}\left( \mathbf{M}_{F} \mathbf{\odot M}_{F}\right) \right] ^{2}+24\boldsymbol{\tau}_{T}^{\prime }\left( \mathbf{I}_{T}\mathbf{\odot M}_{F}\right) \left( \mathbf{M} _{F}\mathbf{\odot M}_{F}\right) \left( \mathbf{I}_{T}\mathbf{\odot M} _{F}\right) \boldsymbol{\tau}_{T}\\ & +8\boldsymbol{\tau}_{T}^{\prime}\left( \mathbf{M}_{F}\mathbf{\odot M} _{F}\mathbf{\odot M}_{F}\mathbf{\odot M}_{F}\right) \boldsymbol{\tau}_{T}, \end{align*} \[ f_{\gamma_{1}\gamma_{3}}=24\boldsymbol{\tau}_{T}^{\prime}\left( \mathbf{I}_{T}\mathbf{\odot M}_{F}\right) \mathbf{M}_{F}\left( \mathbf{I}_{T}\mathbf{\odot M}_{F}\mathbf{\odot M}_{F}\right) \boldsymbol{\tau}_{T}+32\boldsymbol{\tau}_{T}^{\prime}\left( \mathbf{I} _{T}\mathbf{\odot M}_{F}\right) \left( \mathbf{M}_{F}\mathbf{\odot M} _{F}\mathbf{\odot M}_{F}\right) \boldsymbol{\tau}_{T}. \] Result ((ref)) follows from ((ref)). To establish ((ref)), using ((ref)), we first note that \[ E\left( \frac{\boldsymbol{\varepsilon}^{\prime}\mathbf{M}_{F} \boldsymbol{\varepsilon}}{v}\right) ^{2}=1+\frac{2}{v}+\gamma_{2} \frac{\mathrm{tr}\left( \mathbf{M}_{F}\mathbf{\odot M}_{F}\right) }{v^{2}}. \] But by ((ref)) $\mathrm{tr}\left( \mathbf{M}_{F}\mathbf{\odot M} _{F}\right) =O\left( v\right) $ and by assumption $\gamma_{2}$ is bounded. Hence $E\left( \frac{\boldsymbol{\varepsilon}^{\prime}\mathbf{M} _{F}\boldsymbol{\varepsilon}}{v}\right) ^{2}=1+O(\frac{1}{v})$, as required. To prove ((ref)), noting that $E\left( \boldsymbol{\varepsilon}^{\prime }\mathbf{M}_{F}\boldsymbol{\varepsilon}\right) =v$ then $E(q_{v})=0$. Also \[ q_{v}^{4}=v^{2}\left( \frac{\boldsymbol{\varepsilon}^{\prime}\mathbf{M} _{F}\boldsymbol{\varepsilon}}{v}-1\right) ^{4}=v^{2}\left[ \left( \frac{\boldsymbol{\varepsilon}^{\prime}\mathbf{M}_{F}\boldsymbol{\varepsilon} }{v}\right) ^{4}-4\left( \frac{\boldsymbol{\varepsilon}^{\prime} \mathbf{M}_{F}\boldsymbol{\varepsilon}}{v}\right) ^{3}+6\left( \frac{\boldsymbol{\varepsilon}^{\prime}\mathbf{M}_{F}\boldsymbol{\varepsilon} }{v}\right) ^{2}-4\left( \frac{\boldsymbol{\varepsilon}^{\prime} \mathbf{M}_{F}\boldsymbol{\varepsilon}}{v}\right) +1\right] , \] and taking expectation yields \begin{align*} E\left( q_{v}^{4}\right) & =v^{2}\left[ E\left[ \left( \frac {\boldsymbol{\varepsilon}^{\prime}\mathbf{M}_{F}\boldsymbol{\varepsilon}} {v}\right) ^{4}\right] -4E\left[ \left( \frac{\boldsymbol{\varepsilon }^{\prime}\mathbf{M}_{F}\boldsymbol{\varepsilon}}{v}\right) ^{3}\right] +6E\left[ \left( \frac{\boldsymbol{\varepsilon}^{\prime}\mathbf{M} _{F}\boldsymbol{\varepsilon}}{v}\right) ^{2}\right] -4E\left( \frac{\boldsymbol{\varepsilon}^{\prime}\mathbf{M}_{F}\boldsymbol{\varepsilon} }{v}\right) +1\right] \\ & =\frac{1}{v^{2}}E\left[ \left( \boldsymbol{\varepsilon}^{\prime }\mathbf{M}_{F}\boldsymbol{\varepsilon}\right) ^{4}\right] -\frac{4} {v}E\left[ \left( \boldsymbol{\varepsilon}^{\prime}\mathbf{M}_{F} \boldsymbol{\varepsilon}\right) ^{3}\right] +6E\left[ \left( \boldsymbol{\varepsilon}^{\prime}\mathbf{M}_{F}\boldsymbol{\varepsilon }\right) ^{2}\right] -4vE\left( \boldsymbol{\varepsilon}^{\prime} \mathbf{M}_{F}\boldsymbol{\varepsilon}\right) +v^{2}. \end{align*} Now using the results in ((ref))-((ref)), and after some algebra, we obtain \begin{align} E\left( q_{v}^{4}\right) & =\frac{1}{v^{2}}\left[ \begin{array} [c]{c} 12v^{2}+48v+12\gamma_{2}\left[ \boldsymbol{\tau}_{T}^{\prime}\left( \mathbf{M}_{F}\mathbf{\odot M}_{F}\right) \boldsymbol{\tau}_{T}\right] \mathrm{tr}\left( \mathbf{M}_{F}\mathbf{\odot M}_{F}\right) \\ +96\gamma_{2}\mathrm{tr}\left[ \left( \mathbf{I}_{T}\mathbf{\odot M} _{F}\right) \mathbf{M}_{F}\right] +48\gamma_{2}\left[ \boldsymbol{\tau} _{T}^{\prime}\left( \mathbf{I}_{T}\mathbf{\odot M}_{F}\right) \left( \mathbf{I}_{T}\mathbf{\odot M}_{F}\right) \boldsymbol{\tau}_{T}\right] \\ +\left( 4\gamma_{4}v+24\gamma_{4}\right) \mathrm{tr}\left( \mathbf{M} _{F}\mathbf{\odot M}_{F}\mathbf{\odot M}_{F}\right) +\gamma_{6} \mathrm{tr}\left( \mathbf{M}_{F}\mathbf{\odot M}_{F}\mathbf{\odot M} _{F}\mathbf{\odot M}_{F}\right) \\ +48\gamma_{1}^{2}\left[ \boldsymbol{\tau}_{T}^{\prime}\left( \mathbf{I} _{T}\mathbf{\odot M}_{F}\right) \mathbf{M}_{F}\left( \mathbf{I} _{T}\mathbf{\odot M}_{F}\right) \boldsymbol{\tau}_{T}\right] +96\gamma _{1}^{2}\left[ \boldsymbol{\tau}_{T}^{\prime}\left( \mathbf{M} _{F}\mathbf{\odot M}_{F}\right) \mathbf{M}_{F}\left( \mathbf{I} _{T}\mathbf{\odot M}_{F}\right) \boldsymbol{\tau}_{T}\right] \\ +96\gamma_{1}^{2}\mathrm{tr}\left[ \left( \mathbf{M}_{F}\mathbf{\odot M} _{F}\right) \mathbf{M}_{F}\right] +3\gamma_{2}^{2}\left[ \mathrm{tr}\left( \mathbf{M}_{F}\mathbf{\odot M}_{F}\right) \right] ^{2}\\ +24\gamma_{2}^{2}\left[ \boldsymbol{\tau}_{T}^{\prime}\left( \mathbf{I} _{T}\mathbf{\odot M}_{F}\right) \left( \mathbf{M}_{F}\mathbf{\odot M} _{F}\right) \left( \mathbf{I}_{T}\mathbf{\odot M}_{F}\right) \boldsymbol{\tau}_{T}\right] +8\gamma_{2}^{2}\left[ \boldsymbol{\tau} _{T}^{\prime}\left( \mathbf{M}_{F}\mathbf{\odot M}_{F}\mathbf{\odot M} _{F}\mathbf{\odot M}_{F}\right) \boldsymbol{\tau}_{T}\right] \\ +24\gamma_{1}\gamma_{3}\left[ \boldsymbol{\tau}_{T}^{\prime}\left( \mathbf{I}_{T}\mathbf{\odot M}_{F}\right) \mathbf{M}_{F}\left( \mathbf{I}_{T}\mathbf{\odot M}_{F}\mathbf{\odot M}_{F}\right) \boldsymbol{\tau}_{T}\right] \\ +32\gamma_{1}\gamma_{3}\left[ \boldsymbol{\tau}_{T}^{\prime}\left( \mathbf{I}_{T}\mathbf{\odot M}_{F}\right) \left( \mathbf{M}_{F}\mathbf{\odot M}_{F}\mathbf{\odot M}_{F}\right) \boldsymbol{\tau}_{T}\right] \end{array} \right] \nonumber\\ & =\sum_{s=1}^{15}a_{s,v}. \end{align} Further noting that $\gamma_{1},\gamma_{2}$,$\gamma_{3}$,$\gamma_{4}$ ,$\gamma_{5}$, and $\gamma_{6}$ are all bounded, then using the results ((ref))-((ref)) we have \[ a_{1,v}=12,a_{2,v}=\frac{48}{v}, \] \[ a_{3,v}=\frac{12\gamma_{2}\left[ \boldsymbol{\tau}_{T}^{\prime}\left( \mathbf{M}_{F}\mathbf{\odot M}_{F}\right) \boldsymbol{\tau}_{T}\right] \mathrm{tr}\left( \mathbf{M}_{F}\mathbf{\odot M}_{F}\right) }{v^{2} }=O\left( 1\right) , \] \[ a_{4,v}=\frac{96\gamma_{2}\mathrm{tr}\left[ \left( \mathbf{I}_{T} \mathbf{\odot M}_{F}\right) \mathbf{M}_{F}\right] }{v^{2}}=O\left( \frac {1}{v}\right) , \] \[ a_{5,v}=\frac{48\gamma_{2}\left[ \boldsymbol{\tau}_{T}^{\prime}\left( \mathbf{I}_{T}\mathbf{\odot M}_{F}\right) \left( \mathbf{I}_{T}\mathbf{\odot M}_{F}\right) \boldsymbol{\tau}_{T}\right] }{v^{2}}=O\left( \frac{1} {v}\right) , \] \[ a_{6,v}=\frac{\left( 4\gamma_{4}v+24\gamma_{4}\right) \mathrm{tr}\left( \mathbf{M}_{F}\mathbf{\odot M}_{F}\mathbf{\odot M}_{F}\right) }{v^{2} }=O\left( 1\right) , \] \[ a_{7,v}=\frac{\gamma_{6}\mathrm{tr}\left( \mathbf{M}_{F}\mathbf{\odot M} _{F}\mathbf{\odot M}_{F}\mathbf{\odot M}_{F}\right) }{v^{2}}=O\left( \frac{1}{v}\right) , \] \[ a_{8,v}=\frac{48\gamma_{1}^{2}\left[ \boldsymbol{\tau}_{T}^{\prime}\left( \mathbf{I}_{T}\mathbf{\odot M}_{F}\right) \mathbf{M}_{F}\left( \mathbf{I}_{T}\mathbf{\odot M}_{F}\right) \boldsymbol{\tau}_{T}\right] }{v^{2}}=O\left( \frac{1}{\sqrt{v}}\right) , \] \[ a_{9,v}=\frac{96\gamma_{1}^{2}\boldsymbol{\tau}_{T}^{\prime}\left( \mathbf{M}_{F}\mathbf{\odot M}_{F}\right) \mathbf{M}_{F}\left( \mathbf{I}_{T}\mathbf{\odot M}_{F}\right) \boldsymbol{\tau}_{T}}{v^{2} }=O\left( \frac{1}{\sqrt{v}}\right) , \] \[ a_{10,v}=\frac{96\gamma_{1}^{2}\mathrm{tr}\left[ \left( \mathbf{M} _{F}\mathbf{\odot M}_{F}\right) \mathbf{M}_{F}\right] }{v^{2}}=O\left( \frac{1}{v}\right) , \] \[ a_{11,v}=\frac{3\gamma_{2}^{2}\left[ \mathrm{tr}\left( \mathbf{M} _{F}\mathbf{\odot M}_{F}\right) \right] ^{2}}{v^{2}}=O\left( 1\right) , \] \[ a_{12,v}=\frac{24\gamma_{2}^{2}\left[ \boldsymbol{\tau}_{T}^{\prime}\left( \mathbf{I}_{T}\mathbf{\odot M}_{F}\right) \left( \mathbf{M}_{F}\mathbf{\odot M}_{F}\right) \left( \mathbf{I}_{T}\mathbf{\odot M}_{F}\right) \boldsymbol{\tau}_{T}\right] }{v^{2}}=O\left( \frac{1}{v}\right) , \] \[ a_{13,v}=\frac{8\gamma_{2}^{2}\left[ \boldsymbol{\tau}_{T}^{\prime}\left( \mathbf{M}_{F}\mathbf{\odot M}_{F}\mathbf{\odot M}_{F}\mathbf{\odot M} _{F}\right) \boldsymbol{\tau}_{T}\right] }{v^{2}}=O\left( \frac{1} {v}\right) , \] \[ a_{14,v}=\frac{24\gamma_{1}\gamma_{3}\left[ \boldsymbol{\tau}_{T}^{\prime }\left( \mathbf{I}_{T}\mathbf{\odot M}_{F}\right) \mathbf{M}_{F}\left( \mathbf{I}_{T}\mathbf{\odot M}_{F}\mathbf{\odot M}_{F}\right) \boldsymbol{\tau}_{T}\right] }{v^{2}}=O\left( \frac{1}{\sqrt{v}}\right) , \] \[ a_{15,v}=\frac{32\gamma_{1}\gamma_{3}\left[ \boldsymbol{\tau}_{T}^{\prime }\left( \mathbf{I}_{T}\mathbf{\odot M}_{F}\right) \left( \mathbf{M} _{F}\mathbf{\odot M}_{F}\mathbf{\odot M}_{F}\right) \boldsymbol{\tau} _{T}\right] }{v^{2}}=O\left( \frac{1}{v}\right) . \] Using these results in ((ref)) it now follows that $E\left( q_{v} ^{4}\right) =O\left( 1\right) $, as required.
lemmaSuppose the $T\times1$ vector $\boldsymbol{\varepsilon }=\mathbf{(}\varepsilon_{1},\varepsilon_{2},...,\varepsilon_{T})^{\prime}$ is $\boldsymbol{\varepsilon}\thicksim IID(\mathbf{0},\mathbf{I}_{T})$, $\sup _{t}E\left( \left\vert \varepsilon_{t}\right\vert ^{8+s}\right) <C$ for some small $s>0$, and $\mathbf{M}_{F}=\mathbf{I}_{T}-\mathbf{F(F}^{\prime }\mathbf{F)}^{-1}\mathbf{F}^{\prime}$, where the $T\times m_{0}$ matrix $\mathbf{F}$ is distributed independently of $\boldsymbol{\varepsilon}$. Suppose there exists a finite integer $v_{0}$ such that for all $v>v_{0}$, \begin{equation} \frac{\boldsymbol{\varepsilon}^{\prime}\mathbf{M}_{F}\boldsymbol{\varepsilon} }{v}>c>0. \end{equation} Then for $v>v_{0}$, \begin{equation} E\left[ \left( \frac{\boldsymbol{\varepsilon}^{\prime}\mathbf{M} _{F}\boldsymbol{\varepsilon}}{v}\right) ^{1/2}\right] =1+O\left( \frac {1}{v}\right) , E\left[ \left( \frac{\boldsymbol{\varepsilon }^{\prime}\mathbf{M}_{F}\boldsymbol{\varepsilon}}{v}\right) ^{-s/2}\right] =1+O\left( \frac{1}{v}\right) , for s=1,2,3,4, \end{equation} \begin{equation} E\left[ \varepsilon_{t}^{2}\left( \frac{\boldsymbol{\varepsilon}^{\prime }\mathbf{M}_{F}\boldsymbol{\varepsilon}}{v}\right) \right] =1+O\left( \frac{1}{v}\right) , \end{equation} \begin{equation} E\left[ \varepsilon_{t}^{2}\left( \frac{\boldsymbol{\varepsilon}^{\prime }\mathbf{M}_{F}\boldsymbol{\varepsilon}}{v}\right) ^{-1}\right] =1+O\left( \frac{1}{v}\right) , \end{equation} \begin{equation} E\left[ \varepsilon_{t}^{2}\left( \frac{\boldsymbol{\varepsilon}^{\prime }\mathbf{M}_{F}\boldsymbol{\varepsilon}}{v}\right) ^{-1/2}\right] =1+O\left( \frac{1}{v}\right) , \end{equation} \begin{equation} E\left[ \varepsilon_{t}\varepsilon_{t^{\prime}}\left( \frac {\boldsymbol{\varepsilon}^{\prime}\mathbf{M}_{F}\boldsymbol{\varepsilon}} {v}\right) ^{-1/2}\right] =O\left( \frac{1}{v}\right) , for t\neq t^{\prime}, \end{equation} \begin{equation} E\left[ \varepsilon_{t}\left( \frac{\boldsymbol{\varepsilon}^{\prime }\mathbf{M}_{F}\boldsymbol{\varepsilon}}{v}\right) ^{-1/2}\right] =O\left( \frac{1}{v}\right) . \end{equation}
proofTo establish the results in ((ref)) we first note that \begin{equation} \frac{\boldsymbol{\varepsilon}^{\prime}\mathbf{M}_{F}\boldsymbol{\varepsilon} }{v}=1+\frac{1}{\sqrt{v}}q_{v} \end{equation} where \begin{equation} q_{v}=\sqrt{v}\left( \frac{\boldsymbol{\varepsilon}^{\prime}\mathbf{M} _{F}\boldsymbol{\varepsilon}}{v}-1\right) . \end{equation} Applying the Taylor Theorem to $\left( v^{-1}\boldsymbol{\varepsilon} ^{\prime}\mathbf{M}_{F}\boldsymbol{\varepsilon}\right) ^{1/2}$ we have \begin{equation} \left( \frac{\boldsymbol{\varepsilon}^{\prime}\mathbf{M}_{F} \boldsymbol{\varepsilon}}{v}\right) ^{1/2}=1+\frac{1}{2}\frac{q_{v}}{\sqrt {v}}-\frac{1}{8v}R_{v}, \end{equation} where \[ R_{v}=\left( 1+\frac{\bar{q}_{v}}{\sqrt{v}}\right) ^{^{-3/2}}q_{v}^{2}, \] and $\bar{q}_{v}$ lies on the interval between $0$ and $q_{v}$. Since $E\left( q_{v}\right) =0$ as shown by result ((ref)) of Lemma (ref), taking expectations of both sides of ((ref)) yields \ \ \begin{equation} E\left( \frac{\boldsymbol{\varepsilon}^{\prime}\mathbf{M}_{F} \boldsymbol{\varepsilon}}{v}\right) ^{1/2}=1-\frac{1}{8v}E\left( R_{v}\right) . \end{equation} It is, therefore, sufficient to show that $E\left( \left\vert R_{v} \right\vert \right) <C$. By Cauchy-Schwarz inequality we have \begin{equation} E\left( \left\vert R_{v}\right\vert \right) \leq\left[ E\left( \left\vert 1+\frac{\bar{q}_{v}}{\sqrt{v}}\right\vert ^{-3}\right) \right] ^{1/2}\left[ E\left( q_{v}^{4}\right) \right] ^{1/2}. \end{equation} Consider $\left\vert 1+\frac{\bar{q}_{v}}{\sqrt{v}}\right\vert $\ and distinguish the cases (a) $q_{v}\geq0$ or equivalently if $\frac {\boldsymbol{\varepsilon}^{\prime}\mathbf{M}_{F}\boldsymbol{\varepsilon}} {v}\geq1$ and (b) $q_{v}<0$ or equivalently if $\frac{\boldsymbol{\varepsilon }^{\prime}\mathbf{M}_{F}\boldsymbol{\varepsilon}}{v}<1$. Under (a)$\ 0\leq$ $\bar{q}_{v}<q_{v}$, we have$\left\vert 1+\frac{\bar{q}_{v}}{\sqrt{v} }\right\vert \geq1$. Under (b) $q_{v}<\bar{q}_{v}<0$, we have $\left\vert 1+\frac{\bar{q}_{v}}{\sqrt{v}}\right\vert >\left\vert 1+\frac{q_{v}}{\sqrt{v} }\right\vert =\frac{\boldsymbol{\varepsilon}^{\prime}\mathbf{M}_{F} \boldsymbol{\varepsilon}}{v}$, and under condition ((ref)) $\left\vert 1+\frac{\bar{q}_{v}}{\sqrt{v}}\right\vert >c>0$. Hence, irrespective of whether $q_{v}\geq0$ or not, \begin{equation} \left\vert 1+\frac{\bar{q}_{v}}{\sqrt{v}}\right\vert >c>0, \end{equation} and we have $E\left( \left\vert 1+\frac{\bar{q}_{v}}{\sqrt{v}}\right\vert ^{-3}\right) <C$. Also it is established that $E\left( q_{v}^{4}\right) =O(1)$ by result ((ref)) of Lemma (ref). Using these results in ((ref)) it follows that $E\left( \left\vert R_{v}\right\vert \right) <C$, and given ((ref)) we can show$E\left[ \left( \boldsymbol{\varepsilon }^{\prime}\mathbf{M}_{F}\boldsymbol{\varepsilon}/v\right) ^{1/2}\right] =1+O\left( v^{-1}\right) . $ The other results in ((ref)) can also be established similarly. Result ((ref)) follows (S.7) in Lemma 6 of Pesaran and Yamagata (2024) by setting\ $\varepsilon_{t}^{2} =\boldsymbol{\varepsilon}^{\prime}\mathbf{A}_{1}\boldsymbol{\varepsilon}$, where $\mathbf{A}_{1}$ has only one none-zero element on its diagonal. Result ((ref)) can be established using a result due to Lieberman (1994) (see Lemmas 5 and 21 in the online supplement of Pesaran and Yamagata (2024)). To establish ((ref)), note that\ by applying the Taylor Theorem to $\left( v^{-1}\boldsymbol{\varepsilon}^{\prime}\mathbf{M}_{F}\boldsymbol{\varepsilon }\right) ^{-1/2}$, \[ \varepsilon_{t}^{2}\left( \frac{\boldsymbol{\varepsilon}^{\prime} \mathbf{M}_{F}\boldsymbol{\varepsilon}}{v}\right) ^{^{-1/2}}=\varepsilon _{t}^{2}-\frac{1}{2}\frac{q_{v}\varepsilon_{t}^{2}}{\sqrt{v}}+\frac{3} {8v}R_{e,v} \] where $R_{e,v}=\left( 1+\frac{\bar{q}_{v}}{\sqrt{v}}\right) ^{-5/2}q_{v} ^{2}\varepsilon_{t}^{2}, $ and taking expectations yields \begin{equation} E\left[ \varepsilon_{t}^{2}\left( \frac{\boldsymbol{\varepsilon}^{\prime }\mathbf{M}_{F}\boldsymbol{\varepsilon}}{v}\right) ^{^{-1/2}}\right] =E\left( \varepsilon_{t}^{2}\right) -\frac{1}{2}\frac{E\left( q_{v}\varepsilon_{t}^{2}\right) }{\sqrt{v}}+\frac{3}{8v}E\left( R_{e,v}\right) . \end{equation} $E(\varepsilon_{t}^{2})=1$, and using ((ref)) we have \begin{equation} E\left[ \varepsilon_{t}^{2}\left( \frac{q_{v}}{\sqrt{v}}\right) \right] =E\left[ \varepsilon_{t}^{2}\left( \frac{\boldsymbol{\varepsilon}^{\prime }\mathbf{M}_{F}\boldsymbol{\varepsilon}}{v}-1\right) \right] =O\left( \frac{1}{v}\right) . \end{equation} By Cauchy-Schwarz inequality \begin{align} E\left( \left\vert R_{e,v}\right\vert \right) & \leq\left[ E\left[ \varepsilon_{t}^{4}\left( \left\vert 1+\frac{\bar{q}_{v}}{\sqrt{v} }\right\vert ^{-5}\right) \right] \right] ^{1/2}\left[ E\left( q_{v} ^{4}\right) \right] ^{1/2}\nonumber\\ & \leq\left[ E\left( \left\vert 1+\frac{\bar{q}_{v}}{\sqrt{v}}\right\vert ^{-10}\right) \right] ^{1/4}\left[ E\left( \varepsilon_{t}^{8}\right) \right] ^{1/4}\left[ E\left( q_{v}^{4}\right) \right] ^{1/2}. \end{align} Given ((ref)), it is easily seen that $E\left( \left\vert 1+\frac{\bar {q}_{v}}{\sqrt{v}}\right\vert ^{-10}\right) <C$. Also $E\left( \varepsilon_{t}^{8}\right) <C$ by assumption and $E\left( q_{v}^{4}\right) <C$ by ((ref)). Hence, using these results in ((ref)) it follows $E\left( \left\vert R_{e,v}\right\vert \right) <C$, which completes the proof of ((ref)). To establish ((ref)), using ((ref)) note that for $t\neq t^{\prime}$, \[ E\left[ \varepsilon_{t}\varepsilon_{t^{\prime}}\left( \frac {\boldsymbol{\varepsilon}^{\prime}\mathbf{M}_{F}\boldsymbol{\varepsilon}} {v}\right) ^{-1/2}\right] =E\left( \varepsilon_{t}\varepsilon_{t^{\prime} }\right) -\frac{1}{2}\frac{E\left( q_{v}\varepsilon_{t}\varepsilon _{t^{\prime}}\right) }{\sqrt{v}}+\frac{3}{8v}E\left( R_{tt^{\prime} ,v}\right) \] where $R_{tt^{\prime},v}=\left( 1+\frac{\bar{q}_{v}}{\sqrt{v}}\right) ^{-5/2}q_{v}^{2}\varepsilon_{t}\varepsilon_{t^{\prime}}. $ Note that $E\left( \varepsilon_{t}\varepsilon_{t^{\prime}}\right) =0$ for $t\neq t^{\prime}$ by serial independence of $\varepsilon_{t}$. In addition, using definition of $q_{v}$ in ((ref)) yields \begin{align*} \frac{E\left( q_{v}\varepsilon_{t}\varepsilon_{t^{\prime}}\right) }{\sqrt {v}} & =E\left[ \left( \frac{\boldsymbol{\varepsilon}^{\prime} \mathbf{M}_{F}\boldsymbol{\varepsilon}}{v}-1\right) \varepsilon _{t}\varepsilon_{t^{\prime}}\right] =E\left[ \left( \frac {\boldsymbol{\varepsilon}^{\prime}\mathbf{M}_{F}\boldsymbol{\varepsilon}} {v}\right) \varepsilon_{t}\varepsilon_{t^{\prime}}\right] =\frac{1} {2v}E\left[ \left( \boldsymbol{\varepsilon}^{\prime}\mathbf{A} \boldsymbol{\varepsilon}\right) \left( \boldsymbol{\varepsilon}^{\prime }\mathbf{B}\boldsymbol{\varepsilon}\right) \right] , \end{align*} where $\mathbf{A}=\mathbf{M}_{F}$ and $\mathbf{B}=\mathbf{(}b_{tt^{\prime}})$ with $b_{tt^{\prime}}$ and $b_{t^{\prime}t}$ ($t\neq t^{\prime}$) being the only non-zero elements. Now using (S.7) of Lemma 6 in Pesaran and Yamagata (2024) it follows that $E\left( q_{v}\varepsilon_{t}\varepsilon_{t^{\prime} }\right) =0$. Also by Cauchy-Schwarz inequality , \begin{align*} E\left( \left\vert R_{tt^{\prime},v}\right\vert \right) & \leq\left[ E\left[ \varepsilon_{t}^{2}\varepsilon_{t^{\prime}}^{2}\left( \left\vert 1+\frac{\bar{q}_{v}}{\sqrt{v}}\right\vert ^{-5}\right) \right] \right] ^{1/2}\left[ E\left( q_{v}^{4}\right) \right] ^{1/2}\\ & \leq\left[ E\left( \left\vert 1+\frac{\bar{q}_{v}}{\sqrt{v}}\right\vert ^{-10}\right) \right] ^{1/4}\left[ E\left( \varepsilon_{t}^{4}\right) E\left( \varepsilon_{t^{\prime}}^{4}\right) \right] ^{1/4}\left[ E\left( q_{v}^{4}\right) \right] ^{1/2}<C, \end{align*} where the final inequality follows using the same line of argument used to bound $E\left( \left\vert R_{e,v}\right\vert \right) $ in ((ref)). Overall, result ((ref)) is established. Finally consider ((ref)) and note that\ \[ E\left[ \varepsilon_{t}\left( \frac{\boldsymbol{\varepsilon}^{\prime }\mathbf{M}_{F}\boldsymbol{\varepsilon}}{v}\right) ^{-1/2}\right] =E\left( \varepsilon_{t}\right) -\frac{1}{2}\frac{E\left( q_{v}\varepsilon _{t}\right) }{\sqrt{v}}+\frac{3}{8v}E\left( R_{t,v}\right) , \] where $R_{t,v}=\left( 1+\frac{\bar{q}_{v}}{\sqrt{v}}\right) ^{-5/2}q_{v} ^{2}\varepsilon_{t}. $ We have $E\left( \varepsilon_{t}\right) =0$ and \[ \frac{E\left( q_{v}\varepsilon_{t}\right) }{\sqrt{v}}=E\left( \frac{\boldsymbol{\varepsilon}^{\prime}\mathbf{M}_{F}\boldsymbol{\varepsilon }\varepsilon_{t}}{v}-\varepsilon_{t}\right) =E\left( \frac {\boldsymbol{\varepsilon}^{\prime}\mathbf{M}_{F}\boldsymbol{\varepsilon }\varepsilon_{t}}{v}\right) . \] Denote $\left\{ m_{jj^{\prime}}:j,j^{\prime}=1,2,\ldots,T\right\} $ as the element of $\mathbf{M}_{F}$, such that based on part (c) of Assumption (ref) $m_{jj^{\prime}}$ is independent from $\varepsilon_{t}$ and $E\left( m_{jj^{\prime}}\right) =O\left( 1\right) $. Then it follows that \begin{align*} E\left( \frac{\boldsymbol{\varepsilon}^{\prime}\mathbf{M}_{F} \boldsymbol{\varepsilon}\varepsilon_{t}}{v}\right) & =\frac{1}{v}\sum _{j=1}^{v}\sum_{j^{\prime}=1}^{v}E\left( \varepsilon_{j}m_{jj^{\prime} }\varepsilon_{j^{\prime}}\varepsilon_{t}\right) =\frac{1}{v}\sum_{j=1} ^{v}\sum_{j^{\prime}=1}^{v}E\left( \varepsilon_{j}\varepsilon_{j^{\prime} }\varepsilon_{t}\right) E\left( m_{jj^{\prime}}\right) =\frac{1}{v}E\left( \varepsilon_{t}^{3}\right) E\left( m_{tt}\right) \end{align*} which is $O\left( v^{-1}\right) $ where the second equation holds due to the independence of $\varepsilon_{t}$ and $m_{jj^{\prime}},$ while the third equation holds due to the serial independence of $\varepsilon_{t}$. Besides by Cauchy-Schwarz inequality, \ \begin{align*} E\left( \left\vert R_{t,v}\right\vert \right) & \leq\left[ E\left[ \varepsilon_{t}^{2}\left( \left\vert 1+\frac{\bar{q}_{v}}{\sqrt{v} }\right\vert ^{-5}\right) \right] \right] ^{1/2}\left[ E\left( q_{v} ^{4}\right) \right] ^{1/2}\\ & \leq\left[ E\left( \left\vert 1+\frac{\bar{q}_{v}}{\sqrt{v}}\right\vert ^{-10}\right) \right] ^{1/4}\left[ E\left( \varepsilon_{t}^{4}\right) \right] ^{1/4}\left[ E\left( q_{v}^{4}\right) \right] ^{1/2}<C, \end{align*} where the last inequality holds again using the same line of argument used to bound $E\left( \left\vert R_{e,v}\right\vert \right) $ in ((ref)). Overall, result ((ref)) is established.
lemmaConsider the latent factor model given by ((ref)) and ((ref)). $\widetilde{CD}$ and $CD$ statistics are defined by ((ref)) and ((ref)). Suppose that Assumptions (ref)-(ref) hold and $\left( n,T\right) \rightarrow \infty$, such that $n/T\rightarrow\kappa\,,$ for $0<\kappa<\infty$. Then \begin{equation} CD=\widetilde{CD}+o_{p}\left( 1\right) . \end{equation}
proofUsing ((ref)) and ((ref)) we first note that \begin{equation} \left( \sqrt{\frac{2\left( n-1\right) }{n}}\right) \left( CD-\widetilde{CD}\right) =\frac{1}{\sqrt{T}}\sum_{t=1}^{T}\left[ \left( \frac{1}{\sqrt{n}}\sum_{i=1}^{n}\frac{\hat{u}_{it}}{\hat{\sigma}_{i,T} }\right) ^{2}-\left( \frac{1}{\sqrt{n}}\sum_{i=1}^{n}\frac{\hat{u}_{it} }{\omega_{i,T}}\right) ^{2}\right] . \end{equation} Also note that \begin{equation} \frac{1}{\sqrt{n}}\sum_{i=1}^{n}\frac{\hat{u}_{it}}{\hat{\sigma}_{i,T} }=h_{t,nT\ }+g_{t,nT} \end{equation} where (also see ((ref))) \[ h_{t,nT\ }=\frac{1}{\sqrt{n}}\sum_{i=1}^{n}\frac{\hat{u}_{it}}{\omega_{i,T} }=\frac{\mathbf{c}_{nT}^{\prime}\mathbf{\hat{u}}_{\circ t}}{\sqrt{n}}\text{, and }g_{t,nT}=\frac{1}{\sqrt{n}}\sum_{i=1}^{n}\hat{u}_{it}\left( \frac {1}{\hat{\sigma}_{i,T}}-\frac{1}{\omega_{i,T}}\right) =\frac{\mathbf{d} _{nT}^{\prime}\mathbf{\hat{u}}_{\circ t}}{\sqrt{n}} \] $\mathbf{\hat{u}}_{\circ t}=(\hat{u}_{1t},\hat{u}_{2t},...,\hat{u} _{nt})^{\prime}$, $\mathbf{c}_{nT}=(\omega_{1,T\ }^{-1},\omega_{2,T} ^{-1},...,\omega_{n,T}^{-1})^{\prime}$, $\mathbf{d}_{nT}=(d_{1T} ,d_{2T},...,d_{nT})^{\prime}\,$, and $d_{iT}=\hat{\sigma}_{i,T}^{-1} -\omega_{i,T}^{-1}$. Then squaring both sides of ((ref)) and using the result in ((ref)) we have \begin{align} \left( \sqrt{\frac{2\left( n-1\right) }{n}}\right) \left( CD-\widetilde{CD}\right) & =\frac{1}{\sqrt{T}}\sum_{t=1}^{T}g_{t,nT} ^{2}+\frac{2}{\sqrt{T}}\sum_{t=1}^{T}h_{t,nT\ }g_{t,nT}\nonumber\\ & =\sqrt{\frac{T}{n}}\left( \frac{1}{\sqrt{n}}\mathbf{d}_{nT}^{\prime }\mathbf{\hat{V}}_{T}\mathbf{d}_{nT}+\frac{2}{\sqrt{n}}\mathbf{c}_{nT} ^{\prime}\mathbf{\hat{V}}_{T}\mathbf{d}_{nT}\right) , \end{align} where $\mathbf{\hat{V}}_{T}=T^{-1}\sum_{t=1}^{T}\mathbf{\hat{u}}_{\circ t}\mathbf{\hat{u}}_{\circ t}^{\prime}$. Now using ((ref)), the error vector $\mathbf{\hat{u}}_{\circ t}$ can be written as \[ \mathbf{\hat{u}}_{\circ t}=\mathbf{u}_{\circ t}\left( \lambda_{T}\right) -\mathbf{\Gamma}\left( \mathbf{\hat{f}}_{t}-\mathbf{f}_{t}\right) -\left( \mathbf{\hat{\Gamma}}-\mathbf{\Gamma}\right) \mathbf{f}_{t}-\left( \mathbf{\hat{\Gamma}}-\mathbf{\Gamma}\right) \left( \mathbf{\hat{f}} _{t}-\mathbf{f}_{t}\right) . \] where $\mathbf{u}_{\circ t}\left( \lambda_{T}\right) =\left( \sigma _{1}\varepsilon_{1t}\left( \lambda_{T}\right) ,\sigma_{2}\varepsilon _{2t}\left( \lambda_{T}\right) ,\ldots,\sigma_{n}\varepsilon_{nt}\left( \lambda_{T}\right) \right) ^{\prime}$. Using this expression we now have \begin{align*} \mathbf{\hat{V}}_{T} & =T^{-1}\sum_{t=1}^{T}\mathbf{u}_{\circ t}\left( \lambda_{T}\right) \mathbf{u}_{\circ t}^{\prime}\left( \lambda_{T}\right) +\mathbf{\Gamma}\left[ T^{-1}\sum_{t=1}^{T}\left( \mathbf{\hat{f}} _{t}-\mathbf{f}_{t}\right) \left( \mathbf{\hat{f}}_{t}-\mathbf{f} _{t}\right) ^{\prime}\right] \mathbf{\Gamma}^{\prime}\\ & +\left( \mathbf{\hat{\Gamma}-\Gamma}\right) \left( T^{-1}\sum_{t=1} ^{T}\mathbf{f}_{t}\mathbf{f}_{t}^{\prime}\right) \left( \mathbf{\hat{\Gamma }-\Gamma}\right) ^{\prime}+\left( \mathbf{\hat{\Gamma}-\Gamma}\right) \left[ T^{-1}\sum_{t=1}^{T}\left( \mathbf{\hat{f}}_{t}-\mathbf{f} _{t}\right) \left( \mathbf{\hat{f}}_{t}-\mathbf{f}_{t}\right) ^{\prime }\right] \left( \mathbf{\hat{\Gamma}-\Gamma}\right) ^{\prime} \mathbf{\Gamma}^{\prime}\\ & -\left[ T^{-1}\sum_{t=1}^{T}\mathbf{u}_{\circ t}\left( \lambda _{T}\right) \left( \mathbf{\hat{f}}_{t}-\mathbf{f}_{t}\right) ^{\prime }\right] -\left[ T^{-1}\sum_{t=1}^{T}\mathbf{u}_{\circ t}\left( \lambda _{T}\right) \mathbf{f}_{t}^{\prime}\right] \left( \mathbf{\hat{\Gamma }-\Gamma}\right) ^{\prime}\\ & -\left[ T^{-1}\sum_{t=1}^{T}\mathbf{u}_{\circ t}\left( \lambda _{T}\right) \left( \mathbf{\hat{f}}_{t}-\mathbf{f}_{t}\right) ^{\prime }\right] \left( \mathbf{\hat{\Gamma}-\Gamma}\right) ^{\prime} +\mathbf{\Gamma}\left[ T^{-1}\sum_{t=1}^{T}\left( \mathbf{\hat{f}} _{t}-\mathbf{f}_{t}\right) \mathbf{f}_{t}^{\prime}\right] \left( \mathbf{\hat{\Gamma}-\Gamma}\right) ^{\prime}\\ & +\mathbf{\Gamma}\left[ T^{-1}\sum_{t=1}^{T}\left( \mathbf{\hat{f}} _{t}-\mathbf{f}_{t}\right) \left( \mathbf{\hat{f}}_{t}-\mathbf{f} _{t}\right) ^{\prime}\right] \left( \mathbf{\hat{\Gamma}-\Gamma}\right) ^{\prime}+\left( \mathbf{\hat{\Gamma}-\Gamma}\right) \left[ T^{-1} \sum_{t=1}^{T}\mathbf{f}_{t}\left( \mathbf{\hat{f}}_{t}-\mathbf{f} _{t}\right) ^{\prime}\right] \left( \mathbf{\hat{\Gamma}-\Gamma}\right) ^{\prime}, \end{align*} or in matrix forms \begin{align*} \mathbf{\hat{V}}_{T} & =\mathbf{V}_{T}\left( \lambda_{T}\right) +\mathbf{\Gamma}\left[ T^{-1}\left( \mathbf{\hat{F}-F}\right) ^{\prime }\left( \mathbf{\hat{F}-F}\right) \right] \mathbf{\Gamma}^{\prime}+\left( \mathbf{\hat{\Gamma}-\Gamma}\right) \mathbf{\Sigma}_{T,ff}\left( \mathbf{\hat{\Gamma}-\Gamma}\right) ^{\prime}\\ & +\left( \mathbf{\hat{\Gamma}-\Gamma}\right) \left[ T^{-1}\left( \mathbf{\hat{F}-F}\right) ^{\prime}\left( \mathbf{\hat{F}-F}\right) \right] \left( \mathbf{\hat{\Gamma}-\Gamma}\right) ^{\prime}-T^{-1} \mathbf{U}^{\prime}\left( \lambda_{T}\right) \left( \mathbf{\hat{F} -F}\right) \mathbf{\Gamma}^{\prime}\\ & -T^{-1}\mathbf{U}^{\prime}\left( \lambda_{T}\right) \mathbf{F}\left( \mathbf{\hat{\Gamma}-\Gamma}\right) ^{\prime}-T^{-1}\mathbf{U}^{\prime }\left( \lambda_{T}\right) \left( \mathbf{\hat{F}-F}\right) \left( \mathbf{\hat{\Gamma}-\Gamma}\right) ^{\prime}+\mathbf{\Gamma}\left[ T^{-1}\mathbf{F}^{\prime}\left( \mathbf{\hat{F}-F}\right) \right] \left( \mathbf{\hat{\Gamma}-\Gamma}\right) ^{\prime}\\ & +\mathbf{\Gamma}\left[ T^{-1}\left( \mathbf{\hat{F}-F}\right) ^{\prime }\left( \mathbf{\hat{F}-F}\right) \right] \left( \mathbf{\hat{\Gamma }-\Gamma}\right) ^{\prime}+\left( \mathbf{\hat{\Gamma}-\Gamma}\right) \left[ T^{-1}\mathbf{F}^{\prime}\left( \mathbf{\hat{F}-F}\right) \right] \left( \mathbf{\hat{\Gamma}-\Gamma}\right) ^{\prime}, \end{align*} where $\mathbf{V}_{T}\left( \lambda_{T}\right) =T^{-1}\sum_{t=1} ^{T}\mathbf{u}_{\circ t}\left( \lambda_{T}\right) \mathbf{u}_{\circ t}^{^{\prime}}\left( \lambda_{T}\right) $ and $\mathbf{U}\left( \lambda _{T}\right) =\left( \mathbf{u}_{\circ1}\left( \lambda_{T}\right) ,\mathbf{u}_{\circ2}\left( \lambda_{T}\right) ,\ldots,\mathbf{u}_{\circ T}\left( \lambda_{T}\right) \right) ^{\prime}$. By ((ref)) we have $\left\Vert \mathbf{V}_{T}\left( \lambda_{T}\right) \right\Vert =\mu_{\max }\left( \mathbf{V}_{T}\left( \lambda_{T}\right) \right) =O_{p}(\frac{n} {T})$. By results in Lemma (ref) all other terms of the $\mathbf{\hat{V} }_{T}$ are either $O_{p}(1)$ or of lower order, and we also have $\mathbf{\hat{V}}_{T}=O_{p}\left( 1\right) $ since $n$ and $T$ are of the same order. Consider now the terms in ((ref)) and note that \[ \left( \sqrt{\frac{2\left( n-1\right) }{n}}\right) \left\vert CD-\widetilde{CD}\right\vert <C\left\Vert \mathbf{\hat{V}}_{T}\right\Vert \left[ \left( \frac{1}{\sqrt{n}}\left\Vert \mathbf{d}_{nT}\right\Vert ^{2}\right) +\left( \frac{2}{\sqrt{n}}\left\Vert \mathbf{c}_{nT}\right\Vert \right) \left\Vert \mathbf{d}_{nT}\right\Vert \right] . \] where $\frac{1}{\sqrt{n}}\left\Vert \mathbf{c}_{nT}\right\Vert =\left( n^{-1}\sum_{i=1}^{n}\omega_{i,T}^{-2}\right) ^{1/2}$ and $\left\Vert \mathbf{d}_{nT}\right\Vert =\left( \sum_{i=1}^{n}\left( \hat{\sigma} _{i,T}^{-1}-\omega_{i,T}^{-1}\right) ^{2}\right) ^{1/2}$. By part (c) of Assumption (ref) $E\left( \omega_{i,T}^{-2}\right) <C<\infty$, such that $\frac{1}{n}E\left\Vert \mathbf{c}_{nT}\right\Vert ^{2}=n^{-1} \sum_{i=1}^{n}E\left( \omega_{i,T}^{-2}\right) <C$ and hence $n^{-1/2} \left\Vert \mathbf{c}_{nT}\right\Vert =O_{p}\left( 1\right) $. Also by ((ref)) we have \[ \sup_{i}\left\vert \hat{\sigma}_{i,T}^{-1}-\omega_{i,T}^{-1}\right\vert ^{2}=\left( \sup_{i}\left\vert \hat{\sigma}_{i,T}^{-1}-\omega_{i,T} ^{-1}\right\vert \right) ^{2}=O_{p}\left[ \left( \frac{\ln\left( n\right) }{T}\right) ^{2}\right] , \] therefore \[ \left\Vert \mathbf{d}_{nT}\right\Vert ^{2}=\sum_{i=1}^{n}\left( \hat{\sigma }_{i,T}^{-1}-\omega_{i,T}^{-1}\right) ^{2}\leq n\sup_{i}\left( \hat{\sigma }_{i,T}^{-1}-\omega_{i,T}^{-1}\right) ^{2}=O_{p}\left[ n\left( \frac {\ln\left( n\right) }{T}\right) ^{2}\right] =o_{p}\left( 1\right) , \] recalling that $n$ and $T$ are of the same order. Hence, $\left\vert CD-\widetilde{CD}\right\vert =o_{p}(1)$, as required.
lemmaSuppose the data are generated by the latent factor model given by ((ref)) and ((ref)). The latent factors, $\mathbf{f}_{t}$, and their loadings, $\boldsymbol{\gamma}_{i}$, are estimated by principal components, $\mathbf{\hat{f}}_{t}$ and $\boldsymbol{\hat{\gamma}}_{i}$, given by ((ref)). Suppose that Assumptions (ref)-(ref) hold and $\left( n,T\right) \rightarrow\infty$, such that $n/T\rightarrow \kappa\,,$ for $0<\kappa<\infty$. Then \begin{align} \frac{1}{n}\sum_{i=1}^{n}\frac{\boldsymbol{\hat{\gamma}}_{i} -\boldsymbol{\gamma}_{i}}{\sigma_{i}} & =O_{p}\left( \sqrt{\frac{\ln\left( n\right) }{nT}}\right) ,\\ \frac{1}{n}\sum_{i=1}^{n}\left( \boldsymbol{\hat{\gamma}}_{i} -\boldsymbol{\gamma}_{i}\right) \sigma_{i} & =O_{p}\left( \sqrt{\frac {\ln\left( n\right) }{nT}}\right) ,\\ \frac{1}{n}\sum_{i=1}^{n}\sigma_{i}^{2}\left( \boldsymbol{\hat{\gamma}} _{i}\boldsymbol{\hat{\gamma}}_{i}^{\prime}-\boldsymbol{\gamma}_{i} \boldsymbol{\gamma}_{i}^{\prime}\right) & =O_{p}\left( \sqrt{\frac {\ln\left( n\right) }{nT}}\right) ,\\ \frac{1}{n}\sum_{i=1}^{n}\hat{\sigma}_{i,T}\boldsymbol{\hat{\gamma}}_{i} -\frac{1}{n}\sum_{i=1}^{n}\sigma_{i}\boldsymbol{\gamma}_{i} & =O_{p}\left( \frac{\ln\left( n\right) }{T}\right) . \end{align}
proofResults ((ref)) and ((ref)) follow directly from ((ref)) by setting $b_{in}=\sigma_{i}^{-1}$ and $b_{in}=\sigma_{i}$, respectively. To prove ((ref)), note that \begin{align} \frac{1}{n}\sum_{i=1}^{n}\sigma_{i}^{2}\left( \boldsymbol{\hat{\gamma}} _{i}\boldsymbol{\hat{\gamma}}_{i}^{\prime}-\boldsymbol{\gamma}_{i} \boldsymbol{\gamma}_{i}^{\prime}\right) & =\frac{1}{n}\sum_{i=1}^{n} \sigma_{i}^{2}\left( \boldsymbol{\hat{\gamma}}_{i}-\boldsymbol{\gamma} _{i}\right) \left( \boldsymbol{\hat{\gamma}}_{i}-\boldsymbol{\gamma} _{i}\right) ^{\prime}\nonumber\\ & +\frac{1}{n}\sum_{i=1}^{n}\sigma_{i}^{2}\left( \boldsymbol{\hat{\gamma} }_{i}-\boldsymbol{\gamma}_{i}\right) \boldsymbol{\gamma}_{i}^{\prime} +\frac{1}{n}\sum_{i=1}^{n}\sigma_{i}^{2}\boldsymbol{\gamma}_{i}\left( \boldsymbol{\hat{\gamma}}_{i}-\boldsymbol{\gamma}_{i}\right) ^{\prime}. \end{align} Since $\sigma_{i}^{2}$ is bounded, then \begin{align*} \left\Vert \frac{1}{n}\sum_{i=1}^{n}\sigma_{i}^{2}\left( \boldsymbol{\hat {\gamma}}_{i}-\boldsymbol{\gamma}_{i}\right) \left( \boldsymbol{\hat{\gamma }}_{i}-\boldsymbol{\gamma}_{i}\right) ^{\prime}\right\Vert & \leq\left( \sup_{i}\sigma_{i}^{2}\right) \left\Vert \frac{1}{n}\sum_{i=1}^{n}\left( \boldsymbol{\hat{\gamma}}_{i}-\boldsymbol{\gamma}_{i}\right) \left( \boldsymbol{\hat{\gamma}}_{i}-\boldsymbol{\gamma}_{i}\right) ^{\prime }\right\Vert \\ & \leq\left( \sup_{i}\sigma_{i}^{2}\right) \left( \frac{1}{n}\sum _{i=1}^{n}\left\Vert \boldsymbol{\hat{\gamma}}_{i}-\boldsymbol{\gamma} _{i}\right\Vert ^{2}\right) , \end{align*} and using ((ref)) it follows that \[ \frac{1}{n}\sum_{i=1}^{n}\sigma_{i}^{2}\left( \boldsymbol{\hat{\gamma}} _{i}-\boldsymbol{\gamma}_{i}\right) \left( \boldsymbol{\hat{\gamma}} _{i}-\boldsymbol{\gamma}_{i}\right) ^{\prime}=O_{p}\left( \frac{1} {\delta_{nT}^{2}}\right) . \] Also, using ((ref)) and setting $b_{in}=\sigma_{i}$, we have \[ \left\Vert \frac{1}{n}\sum_{i=1}^{n}\sigma_{i}^{2}\left( \boldsymbol{\hat {\gamma}}_{i}-\boldsymbol{\gamma}_{i}\right) \boldsymbol{\gamma}_{i}^{\prime }\right\Vert =\left\Vert \frac{1}{n}\sum_{i=1}^{n}\sigma_{i}^{2} \boldsymbol{\gamma}_{i}\left( \boldsymbol{\hat{\gamma}}_{i} -\boldsymbol{\gamma}_{i}\right) ^{\prime}\right\Vert =O_{p}\left( \sqrt{\frac{\ln\left( n\right) }{nT}}\right) , \] then ((ref)) is established based on ((ref)). Finally, consider ((ref)) and note that \begin{align} & \frac{1}{n}\sum_{i=1}^{n}\hat{\sigma}_{i,T}\boldsymbol{\hat{\gamma}} _{i}-\frac{1}{n}\sum_{i=1}^{n}\sigma_{i}\boldsymbol{\gamma}_{i}\nonumber\\ & =\frac{1}{n}\sum_{i=1}^{n}\left[ \left( \hat{\sigma}_{i,T}-\omega _{i,T}\right) +\omega_{i,T}\right] \left( \boldsymbol{\hat{\gamma}} _{i}-\boldsymbol{\gamma}_{i}+\boldsymbol{\gamma}_{i}\right) -\frac{1}{n} \sum_{i=1}^{n}\sigma_{i}\boldsymbol{\gamma}_{i}\nonumber\\ & =\frac{1}{n}\sum_{i=1}^{n}\boldsymbol{\gamma}_{i}\left( \omega _{i,T}-\sigma_{i}\right) +\frac{1}{n}\sum_{i=1}^{n}\boldsymbol{\gamma} _{i}\left( \hat{\sigma}_{i,T}-\omega_{i,T}\right) +\frac{1}{n}\sum_{i=1} ^{n}\sigma_{i}\left( \boldsymbol{\hat{\gamma}}_{i}-\boldsymbol{\gamma} _{i}\right) \nonumber\\ & +\frac{1}{n}\sum_{i=1}^{n}\left( \omega_{i,T}-\sigma_{i}\right) \left( \boldsymbol{\hat{\gamma}}_{i}-\boldsymbol{\gamma}_{i}\right) +\frac{1}{n} \sum_{i=1}^{n}\left( \hat{\sigma}_{i,T}-\omega_{i,T}\right) \left( \boldsymbol{\hat{\gamma}}_{i}-\boldsymbol{\gamma}_{i}\right) \nonumber\\ & =\mathbf{A}_{1,nT}+\mathbf{A}_{2,nT}+\mathbf{A}_{3,nT}+\mathbf{A} _{4,nT}+\mathbf{A}_{5,nT}. \end{align} Recall also that under Assumptions (ref) and (ref) $\sigma_{i}$ and $\boldsymbol{\gamma}_{i}$ are bounded and $\omega_{i,T} ^{2}=T^{-1}\sigma_{i}^{2}\boldsymbol{\varepsilon}_{i\circ}^{^{\prime} }\mathbf{M}_{F}\boldsymbol{\varepsilon}_{i\circ}$, for $i=1,2,...,n$ are distributed independently across $i,$ and from $\sigma_{i}$ and $\boldsymbol{\gamma}_{i}$. Starting with $\mathbf{A}_{1,nT}$, by ((ref)) we have \[ E\left( \sqrt{nT}\mathbf{A}_{1,nT}\right) =\frac{\sqrt{nT}}{n}\sum_{i=1} ^{n}(\boldsymbol{\gamma}_{i}\sigma_{i})E\left[ \left( \frac {\boldsymbol{\varepsilon}_{i\circ}^{^{\prime}}\mathbf{M}_{F} \boldsymbol{\varepsilon}_{i\circ}}{T}\right) ^{1/2}-1\right] =O\left( \frac{\sqrt{nT}}{T}\right) . \] Since $n$ and $T$ are assumed to be of the same order then $E\left( \sqrt {nT}\mathbf{A}_{1,nT}\right) =O\left( 1\right) $. Also, using result ((ref)) \[ Var\left( \sqrt{nT}\mathbf{A}_{1,nT}\right) =\frac{nT}{n^{2}}\sum_{i=1} ^{n}\left( \sigma_{i}^{2}\boldsymbol{\gamma}_{i}\boldsymbol{\gamma} _{i}^{\prime}\right) Var\left[ \left( \frac{\boldsymbol{\varepsilon }_{i\circ}^{^{\prime}}\mathbf{M}_{F}\boldsymbol{\varepsilon}_{i\circ}} {T}\right) ^{1/2}\right] =O\left( 1\right) . \] Therefore, $\sqrt{nT}\mathbf{A}_{1,nT}=O_{p}(1)$ and it follows that $\mathbf{A}_{1,nT}=O_{p}\left[ \left( nT\right) ^{-1/2}\right] $. Further, using ((ref)) and setting $b_{in}=\gamma_{ij}$, for $j=1,2,...,m_{0}$, it follows that \[ \mathbf{A}_{2,nT}=\frac{1}{n}\sum_{i=1}^{n}\boldsymbol{\gamma}_{i}\left( \hat{\sigma}_{i,T}-\omega_{i,T}\right) =O_{p}\left( \frac{\ln\left( n\right) }{T}\right) . \] Since $\mathbf{A}_{3,nT}$ is the same as the result in ((ref)), which is already established, then $\mathbf{A}_{3,nT}=O_{p}\left( \sqrt{\ln\left( n\right) /\left( nT\right) }\right) $. Using result ((ref)) it follows that \[ \mathbf{A}_{4,nT}=\frac{1}{n}\sum_{i=1}^{n}\left( \omega_{i,T}-\sigma _{i}\right) \left( \boldsymbol{\hat{\gamma}}_{i}-\boldsymbol{\gamma} _{i}\right) =O_{p}\left( \frac{1}{\delta_{nT}^{2}}\right) . \] Using result ((ref)) we have \[ \mathbf{A}_{5,nT}=\frac{1}{n}\sum_{i=1}^{n}\left( \hat{\sigma}_{i,T} -\omega_{i,T}\right) \left( \boldsymbol{\hat{\gamma}}_{i}-\boldsymbol{\gamma }_{i}\right) =O_{p}\left[ \left( \frac{\ln\left( n\right) }{T}\right) ^{3/2}\right] . \] Result ((ref)) now follows straightforwardly based on ((ref)).
lemmaSuppose the data are generated by the latent factor model given by ((ref)) and ((ref)). Denote $\boldsymbol{\varphi }_{n}=n^{-1}\sum_{i=1}^{n}\boldsymbol{\gamma}_{i}/\sigma_{i}$, $\boldsymbol{\varphi}_{nT}=n^{-1}\sum_{i=1}^{n}\boldsymbol{\gamma}_{i} /\omega_{i,T}$ with $\omega_{i,T}=\left( T^{-1}\sigma_{i}^{2} \boldsymbol{\varepsilon}_{i\circ}^{\prime}\mathbf{M}_{F} \boldsymbol{\varepsilon}_{i\circ}\right) ^{1/2}$ where $\boldsymbol{\varepsilon}_{i\circ}=\left( \varepsilon_{i1},\varepsilon _{i2},\ldots,\varepsilon_{iT}\right) ^{\prime}$ and $\mathbf{M} _{F}=\mathbf{I}_{T}-\mathbf{F}\left( \mathbf{F}^{\prime}\mathbf{F}\right) ^{-1}\mathbf{F}^{\prime}$. Suppose that Assumptions (ref) -(ref) hold and $\left( n,T\right) \rightarrow\infty$, such that $n/T\rightarrow\kappa\,,$ for $0<\kappa<\infty$. Then \begin{align*} g_{2,nT}\left( \lambda_{T}\right) & =\frac{1}{\sqrt{T}}\sum_{t=1} ^{T}\upsilon_{t,nT}^{2}\left( \lambda_{T}\right) =o_{p}\left( 1\right) ,\\ g_{3,nT}\left( \lambda_{T}\right) & =\sqrt{T}\left( \boldsymbol{\varphi }_{nT}-\boldsymbol{\varphi}_{n}\right) ^{\prime}\left( \frac{\sum_{t=1} ^{T}\boldsymbol{\kappa}_{t,n}\left( \lambda_{T}\right) \boldsymbol{\kappa }_{t,n}^{\prime}\left( \lambda_{T}\right) }{T}\right) \left( \boldsymbol{\varphi}_{nT}-\boldsymbol{\varphi}_{n}\right) =o_{p}\left( 1\right) ,\\ g_{4,nT}\left( \lambda_{T}\right) & =\sqrt{T}\left( \boldsymbol{\varphi }_{nT}-\boldsymbol{\varphi}_{n}\right) ^{\prime}\left( \frac{1}{T}\sum _{t=1}^{T}\boldsymbol{\kappa}_{t,n}\left( \lambda_{T}\right) \upsilon _{t,nT}\left( \lambda_{T}\right) \right) =o_{p}\left( 1\right) ,\\ g_{5,nT}\left( \lambda_{T}\right) & =\sqrt{T}\left( \boldsymbol{\varphi }_{nT}-\boldsymbol{\varphi}_{n}\right) ^{\prime}\left( \frac{1}{T}\sum _{t=1}^{T}\boldsymbol{\kappa}_{t,n}\left( \lambda_{T}\right) \xi _{t,n}\left( \lambda_{T}\right) \right) =o_{p}\left( 1\right) , \end{align*} where \begin{align*} \upsilon_{t,nT}\left( \lambda_{T}\right) & =\frac{1}{\sqrt{n}}\sum _{i=1}^{n}\left( \frac{1}{\left( \boldsymbol{\varepsilon}_{i\circ}^{\prime }\mathbf{M}_{F}\boldsymbol{\varepsilon}_{i\circ}/T\right) ^{1/2}}-1\right) \varepsilon_{it}\left( \lambda_{T}\right) ,\\ \boldsymbol{\kappa}_{t,n}\left( \lambda_{T}\right) & =\frac{1}{\sqrt{n} }\sum_{i=1}^{n}\boldsymbol{\gamma}_{i}\sigma_{i}\varepsilon_{it}\left( \lambda_{T}\right) ,\\ \xi_{t,n}\left( \lambda_{T}\right) & =\frac{1}{\sqrt{n}}\sum_{i=1} ^{n}a_{i,n}\varepsilon_{it}\left( \lambda_{T}\right) , a_{i,n}=1-\sigma_{i}\boldsymbol{\varphi}_{n}^{\prime}\boldsymbol{\gamma}_{i}. \end{align*}
proofStarting with the $g_{2,nT}\left( \lambda_{T}\right) $, note that \begin{equation} g_{2,nT}\left( \lambda_{T}\right) =\sqrt{T}\left( \frac{1}{T}\sum_{t=1} ^{T}\upsilon_{t,nT}^{2}\left( \lambda_{T}\right) \right) \leq\sqrt {T}\left( \sup_{t}\upsilon_{t,nT}^{2}\left( \lambda_{T}\right) \right) . \end{equation} Meanwhile, we have \[ \left\vert \upsilon_{t,nT}\left( \lambda_{T}\right) \right\vert \leq\frac {1}{\sqrt{n}}\sum_{i=1}^{n}\left\vert \left( \frac{1}{\left( \boldsymbol{\varepsilon}_{i\circ}^{\prime}\mathbf{M}_{F} \boldsymbol{\varepsilon}_{i\circ}/T\right) ^{1/2}}-1\right) \right\vert \left\vert \varepsilon_{it}\left( \lambda_{T}\right) \right\vert , \] and \begin{equation} \sup_{t}\left\vert \upsilon_{t,nT}\left( \lambda_{T}\right) \right\vert \leq\left[ \frac{1}{\sqrt{n}}\sum_{i=1}^{n}\left\vert \left( \frac {1}{\left( \boldsymbol{\varepsilon}_{i\circ}^{\prime}\mathbf{M} _{F}\boldsymbol{\varepsilon}_{i\circ}/T\right) ^{1/2}}-1\right) \right\vert \right] \left( \sup_{i,t}\left\vert \varepsilon_{it}\left( \lambda _{T}\right) \right\vert \right) . \end{equation} Consider the first term of the product in ((ref)) and note since $\boldsymbol{\varepsilon}_{i\circ}$ is cross-sectionally independent conditional then \[ E\left[ \frac{1}{\sqrt{n}}\sum_{i=1}^{n}\left\vert \left( \frac{1}{\left( \boldsymbol{\varepsilon}_{i\circ}^{\prime}\mathbf{M}_{F} \boldsymbol{\varepsilon}_{i\circ}/T\right) ^{1/2}}-1\right) \right\vert \right] ^{2}=\frac{1}{n}\sum_{i=1}^{n}E\left( \frac{1}{\left( \boldsymbol{\varepsilon}_{i\circ}^{\prime}\mathbf{M}_{F} \boldsymbol{\varepsilon}_{i\circ}/T\right) ^{1/2}}-1\right) ^{2}. \] Also, using ((ref)) we obtain \begin{align*} E\left( \frac{1}{\left( \boldsymbol{\varepsilon}_{i\circ}^{\prime} \mathbf{M}_{F}\boldsymbol{\varepsilon}_{i\circ}/T\right) ^{1/2}}-1\right) ^{2} & =E\left( \frac{1}{\boldsymbol{\varepsilon}_{i\circ}^{\prime }\mathbf{M}_{F}\boldsymbol{\varepsilon}_{i\circ}/T}\right) -2E\left( \frac{1}{\left( \boldsymbol{\varepsilon}_{i\circ}^{\prime}\mathbf{M} _{F}\boldsymbol{\varepsilon}_{i\circ}/T\right) ^{1/2}}\right) +1\\ & =\left[ E\left( \frac{1}{\boldsymbol{\varepsilon}_{i\circ}^{\prime }\mathbf{M}_{F}\boldsymbol{\varepsilon}_{i\circ}/T}\right) -1\right] -2\left[ E\left( \frac{1}{\left( \boldsymbol{\varepsilon}_{i\circ}^{\prime }\mathbf{M}_{F}\boldsymbol{\varepsilon}_{i\circ}/T\right) ^{1/2}}\right) -1\right] \\ & =O\left( \frac{1}{T}\right) . \end{align*} Therefore by Markov inequality, we have \begin{equation} \frac{1}{\sqrt{n}}\sum_{i=1}^{n}\left\vert \left( \frac{1}{\left( \boldsymbol{\varepsilon}_{i\circ}^{\prime}\mathbf{M}_{F} \boldsymbol{\varepsilon}_{i\circ}/T\right) ^{1/2}}-1\right) \right\vert =O_{p}\left( \frac{1}{\sqrt{T}}\right) . \end{equation} Now consider the second term of of the product in ((ref)) and note there exist $C_{\varepsilon,1}$, $C_{\varepsilon,2}$ and $r_{\varepsilon}>0$ such that \[ \Pr\left( \sup_{i,t}\left\vert \varepsilon_{it}\left( \lambda_{T}\right) \right\vert >a_{\varepsilon}\right) \leq nT\sup_{i,t}\Pr\left( \left\vert \varepsilon_{it}\left( \lambda_{T}\right) \right\vert >a_{\varepsilon }\right) \leq nTC_{\varepsilon,1}\exp\left( -C_{\varepsilon,2}\left( a_{\varepsilon}\right) ^{r_{\varepsilon}}\right) . \] Hence, by letting $a_{\varepsilon}=\ominus\left( \ln\left( nT\right) \right) $, we have \[ \Pr\left( \sup_{i,t}\left\vert \varepsilon_{it}\left( \lambda_{T}\right) \right\vert >a_{\varepsilon}\right) \leq C_{\varepsilon,1}\exp\left( \ln\left( nT\right) -C_{\varepsilon,2}\left( a_{\varepsilon}\right) ^{r_{\varepsilon}}\right) =O\left( 1\right) , \] which further implies $\sup_{i,t}\left\vert \varepsilon_{it}\left( \lambda_{T}\right) \right\vert =O_{p}\left( \ln\left( nT\right) \right) $. Invoking this result and ((ref)) in ((ref)) now yields \begin{equation} \sup_{t}\left\vert \upsilon_{t,nT}\left( \lambda_{T}\right) \right\vert =O_{p}\left( \frac{\ln\left( nT\right) }{\sqrt{T}}\right) . \end{equation} Then consider ((ref)) and it follows \begin{equation} g_{2,nT}\left( \lambda_{T}\right) \leq\sqrt{T}\left( \sup_{t} \upsilon_{t,nT}^{2}\left( \lambda_{T}\right) \right) \leq\sqrt{T}\left( \sup_{t}\left\vert \upsilon_{t,nT}\left( \lambda_{T}\right) \right\vert \right) ^{2}=O_{p}\left( \frac{\left[ \ln\left( nT\right) \right] ^{2} }{\sqrt{T}}\right) =o_{p}\left( 1\right) , \end{equation} as $n$ and $T$ are of the same order of magnitude. Consider $g_{3,nT}\left( \lambda_{T}\right) $ and note that the result of Lemma (ref) holds such that \begin{equation} \sqrt{T}\left( \boldsymbol{\varphi}_{n}-\boldsymbol{\varphi}_{nT}\right) =O_{p}\left( n^{-1/2}\right) +O_{p}\left( T^{-1/2}\right) . \end{equation} Furthermore, we have \begin{align*} E\left\Vert \boldsymbol{\kappa}_{t,n}\left( \lambda_{T}\right) \right\Vert ^{2} & =\frac{1}{n}\sum_{i=1}^{n}\sum_{j=1}^{n}\boldsymbol{\gamma} _{i}^{\prime}\boldsymbol{\gamma}_{j}\sigma_{i}\sigma_{j}E\left( \varepsilon_{it}\left( \lambda_{T}\right) \varepsilon_{jt}\left( \lambda_{T}\right) \right) \\ & \leq\frac{1}{n}\sum_{i=1}^{n}\sum_{j=1}^{n}\left\Vert \boldsymbol{\gamma }_{i}\right\Vert \left\Vert \boldsymbol{\gamma}_{j}\right\Vert \left\vert \sigma_{i}\sigma_{j}\right\vert \left\vert E\left( \varepsilon_{it}\left( \lambda_{T}\right) \varepsilon_{jt}\left( \lambda_{T}\right) \right) \right\vert \\ & \leq\left( \sup_{i}\left\Vert \boldsymbol{\gamma}_{i}\right\Vert ^{2}\right) \left( \sup_{i}\sigma_{i}^{2}\right) \left[ \frac{1}{n} \sum_{i=1}^{n}\sum_{j=1}^{n}\left\vert E\left( \varepsilon_{it}\left( \lambda_{T}\right) \varepsilon_{jt}\left( \lambda_{T}\right) \right) \right\vert \right] . \end{align*} Given ((ref)) and boundedness of $\boldsymbol{\gamma}_{i}$ and $\sigma_{i}$, it follows \begin{equation} E\left\Vert \boldsymbol{\kappa}_{t,n}\left( \lambda_{T}\right) \right\Vert ^{2}\leq\left( \sup_{i}\left\Vert \boldsymbol{\gamma}_{i}\right\Vert ^{2}\right) \left( \sup_{i}\sigma_{i}^{2}\right) \left[ \frac{1}{n} \sum_{i=1}^{n}\sum_{j=1}^{n}\left\vert E\left( \varepsilon_{it}\left( \lambda_{T}\right) \varepsilon_{jt}\left( \lambda_{T}\right) \right) \right\vert \right] =O\left( 1\right) \end{equation} and $\boldsymbol{\kappa}_{t,n}\left( \lambda_{T}\right) $ is therefore $O_{p}\left( 1\right) $. Also, under Assumption (ref) $\boldsymbol{\kappa}_{t,n}\left( \lambda_{T}\right) $ is serially independent and we have $T^{-1}\sum_{t=1}^{T}\boldsymbol{\kappa}_{t,n}\left( \lambda_{T}\right) \boldsymbol{\kappa}_{t,n}^{\prime}\left( \lambda _{T}\right) =O_{p}\left( 1\right) $. Using this result together with ((ref)) we then have \begin{equation} g_{3,nT}\left( \lambda_{T}\right) =o_{p}(1). \end{equation} Next consider $g_{4,nT}\left( \lambda_{T}\right) $ and note that by Cauchy-Schwarz inequality \[ \left\Vert \frac{1}{T}\sum_{t=1}^{T}\boldsymbol{\kappa}_{t,n}\left( \lambda_{T}\right) \upsilon_{t,nT}\left( \lambda_{T}\right) \right\Vert \leq\left( \frac{1}{T}\sum_{t=1}^{T}\left\Vert \boldsymbol{\kappa} _{t,n}\left( \lambda_{T}\right) \right\Vert ^{2}\right) ^{1/2}\left( \frac{1}{T}\sum_{t=1}^{T}\upsilon_{t,nT}^{2}\left( \lambda_{T}\right) \right) ^{1/2}=O_{p}\left( \frac{\ln\left( nT\right) }{\sqrt{T}}\right) \] where the equation holds by ((ref))\ and ((ref)). Then using the above results it also follows that \begin{equation} g_{4,nT}\left( \lambda_{T}\right) =\sqrt{T}\left( \boldsymbol{\varphi} _{nT}-\boldsymbol{\varphi}_{n}\right) ^{\prime}\left( \frac{1}{T}\sum _{t=1}^{T}\boldsymbol{\kappa}_{t,n}\left( \lambda_{T}\right) \upsilon _{t,nT}\left( \lambda_{T}\right) \right) =o_{p}(1). \end{equation} Similarly, note $\xi_{t,n}\left( \lambda_{T}\right) =n^{-1/2}\sum_{i=1} ^{n}a_{i,n}\varepsilon_{it}\left( \lambda_{T}\right) $ and by Cauchy-Schwarz inequality, \[ \frac{1}{T}\sum_{t=1}^{T}\boldsymbol{\kappa}_{t,n}\left( \lambda_{T}\right) \xi_{t,n}\left( \lambda_{T}\right) \leq\left( \frac{1}{T}\sum_{t=1} ^{T}\left\Vert \boldsymbol{\kappa}_{t,n}\left( \lambda_{T}\right) \right\Vert ^{2}\right) ^{1/2}\left( \frac{1}{T}\sum_{t=1}^{T}\xi_{t,n} ^{2}\left( \lambda_{T}\right) \right) ^{1/2} \] where given ((ref)), $\sup_{i}a_{i,n}^{2}<C$ and $\sup_{i}\sigma _{i}^{-2}<C$, it further follows that \begin{align} E\left( \xi_{t,n}^{2}\left( \lambda_{T}\right) \right) & =\frac{1} {n}\sum_{i=1}^{n}\sum_{j=1}^{n}a_{i,n}a_{j,n}E\left( \varepsilon_{it}\left( \lambda_{T}\right) \varepsilon_{jt}\left( \lambda_{T}\right) \right) \nonumber\\ & <\left( \sup_{i}a_{i,n}^{2}\right) \frac{1}{n}\sum_{i=1}^{n}\sum _{j=1}^{n}\left\vert E\left( \varepsilon_{it}\left( \lambda_{T}\right) \varepsilon_{jt}\left( \lambda_{T}\right) \right) \right\vert <C. \end{align} Hence, $T^{-1}\sum_{t=1}^{T}\boldsymbol{\kappa}_{t,n}\left( \lambda _{T}\right) \xi_{t,n}\left( \lambda_{T}\right) =O_{p}\left( 1\right) $, and again using ((ref)) it follows that \begin{equation} g_{5,nT}\left( \lambda_{T}\right) =\sqrt{T}\left( \boldsymbol{\varphi} _{nT}-\boldsymbol{\varphi}_{n}\right) ^{\prime}\left( \frac{1}{T}\sum _{t=1}^{T}\boldsymbol{\kappa}_{t,n}\left( \lambda_{T}\right) \xi _{t,n}\left( \lambda_{T}\right) \right) =o_{p}(1). \end{equation}
lemmaSuppose the data are generated by the latent factor model given by ((ref)) and ((ref)). Further denote \[ w_{nT}\left( \lambda_{T}\right) =\frac{T^{-1/2}\sum_{t=1}^{T}\xi _{t,n}\left( \lambda_{T}\right) \upsilon_{t,nT}\left( \lambda_{T}\right) }{T^{-1}\sum_{t=1}^{T}\xi_{t,n}^{2}\left( \lambda_{T}\right) }, \] where \begin{align*} \xi_{t,n}\left( \lambda_{T}\right) & =\frac{1}{\sqrt{n}}\sum_{i=1} ^{n}a_{i,n}\varepsilon_{it}\left( \lambda_{T}\right) ,a_{i,n}=1-\sigma _{i}\boldsymbol{\varphi}_{n}^{\prime}\boldsymbol{\gamma}_{i},\\ \upsilon_{t,nT}\left( \lambda_{T}\right) & =\frac{1}{\sqrt{n}}\sum _{i=1}^{n}\zeta_{it}\varepsilon_{it}\left( \lambda_{T}\right) ,\zeta _{it}=\frac{1}{\left( T^{-1}\boldsymbol{\varepsilon}_{i\circ}^{\prime }\mathbf{M}_{F}\boldsymbol{\varepsilon}_{i\circ}\right) ^{1/2}}-1, \end{align*} with $\boldsymbol{\varphi}_{n}=n^{-1}\sum_{i=1}^{n}\boldsymbol{\gamma} _{i}/\sigma_{i}$, $\boldsymbol{\varepsilon}_{i\circ}=\left( \varepsilon _{i1},\varepsilon_{i2},\ldots,\varepsilon_{iT}\right) ^{\prime}$, $\mathbf{M}_{F}=\mathbf{I}_{T}-\mathbf{F}\left( \mathbf{F}^{\prime} \mathbf{F}\right) ^{-1}\mathbf{F}^{\prime}$. Suppose that Assumptions (ref)-(ref) hold and $(n,T)\rightarrow\infty$, such that $n/T\rightarrow\kappa,$ and $0<\kappa<\infty$. Then $w_{nT}\left( \lambda _{T}\right) =o_{p}\left( 1\right) $.
proofConsider first the denominator of $w_{nT}\left( \lambda_{T}\right) $ and using ((ref)) note that \[ E\left( \frac{1}{T}\sum_{t=1}^{T}\xi_{t,n}^{2}\left( \lambda_{T}\right) \right) =\frac{1}{T}\sum_{t=1}^{T}E\left( \xi_{t,n}^{2}\left( \lambda _{T}\right) \right) <C. \] Hence, by Markov inequality it is obvious that the denominator of $w_{nT}\left( \lambda_{T}\right) $ is $O_{p}\left( 1\right) $. Consider now the numerator of $w_{nT}\left( \lambda_{T}\right) $, which is denoted as $r_{nT}\left( \lambda_{T}\right) $. For simplicity, let \[ \xi_{t,n}=\frac{1}{\sqrt{n}}\sum_{i=1}^{n}a_{i,n}\varepsilon_{it} ,\upsilon_{t,nT}=\frac{1}{\sqrt{n}}\sum_{i=1}^{n}\zeta_{it}\varepsilon_{it}, \] then \begin{align*} r_{nT}\left( \lambda_{T}\right) & =\frac{1}{\sqrt{T}}\sum_{t=1}^{T} \xi_{t,n}\upsilon_{t,nT}+\frac{1}{\sqrt{T}}\sum_{t=1}^{T}\left[ \xi _{t,n}\left( \lambda_{T}\right) -\xi_{t,n}\right] \upsilon_{t,nT}\left( \lambda_{T}\right) \\ & +\frac{1}{\sqrt{T}}\sum_{t=1}^{T}\xi_{t,n}\left( \lambda_{T}\right) \left[ \upsilon_{t,nT}\left( \lambda_{T}\right) -\upsilon_{t,nT}\right] -\frac{1}{\sqrt{T}}\sum_{t=1}^{T}\left[ \xi_{t,n}\left( \lambda_{T}\right) -\xi_{t,n}\right] \left[ \upsilon_{t,nT}\left( \lambda_{T}\right) -\upsilon_{t,nT}\right] \\ & =r_{1,nT}+\sum_{j=2}^{4}r_{j,nT}\left( \lambda_{T}\right) . \end{align*} To bound $r_{1,nT}$, we firstly note $\xi_{t,n}\upsilon_{t,nT}=\frac{1}{n} \sum_{i=1}^{n}\sum_{j=1}^{n}a_{in}\varepsilon_{it}\zeta_{jt},$ where for $i\neq j$ $\varepsilon_{it}$ and $\zeta_{jt}$ are distributed independently by parts (a) and (c) of Assumption (ref). Therefore, \[ E\left( \xi_{t,n}\upsilon_{t,nT}\right) =\frac{1}{n}\sum_{i=1}^{n} a_{in}E\left[ \varepsilon_{it}^{2}\left( \frac{1}{\left( T^{-1} \boldsymbol{\varepsilon}_{i\circ}^{\prime}\mathbf{M}_{F} \boldsymbol{\varepsilon}_{i\circ}\right) ^{1/2}}-1\right) \right] . \] Furthermore, since $a_{in}$ is bounded, and by result ((ref)) \[ E\left[ \varepsilon_{it}^{2}\left( \frac{1}{\left( T^{-1} \boldsymbol{\varepsilon}_{i\circ}^{\prime}\mathbf{M}_{F} \boldsymbol{\varepsilon}_{i\circ}\right) ^{1/2}}-1\right) \right] =O\left( \frac{1}{T}\right) . \] Then $E\left( \xi_{t,n}\upsilon_{t,nT}\right) =O\left( T^{-1}\right) $, and it follows that \begin{align} E\left( r_{1,nT}\right) & =\frac{1}{\sqrt{T}}\sum_{t=1}^{T}E\left( \xi_{t,n}\upsilon_{t,nT}\right) =O\left( \frac{1}{\sqrt{T}}\right) ,\\ Var\left( r_{1,nT}\right) & =E\left( r_{1,nT}^{2}\right) -\left[ E\left( r_{1,nT}\right) \right] ^{2}=\frac{1}{T}\sum_{t=1}^{T} \sum_{t^{\prime}=1}^{T}E\left( \xi_{t,n}\upsilon_{t,nT}\xi_{t^{\prime} ,n}v_{t^{\prime},nT}\right) +O\left( T^{-1}\right) . \end{align} Consider now the first term of $Var\left( r_{1,nT}\right) $, and using ((ref)) we have \begin{align*} & E\left( \xi_{t,n}\upsilon_{t,nT}\xi_{t^{\prime},n}v_{t^{\prime},nT}\right) \\ & \equiv E\left[ \left( \frac{1}{n}\sum_{i=1}^{n}\sum_{j=1}^{n}a_{in} a_{jn}\varepsilon_{it}\varepsilon_{jt^{\prime}}\right) \left( \frac{1} {n}\sum_{r=1}^{n}\sum_{s=1}^{n}\zeta_{rt}\zeta_{st^{\prime}}\right) \right] \\ & =\frac{1}{n^{2}}\sum_{i=1}^{n}\sum_{j=1}^{n}\sum_{r=1}^{n}\sum_{s=1} ^{n}a_{in}a_{jn}E\left\{ \left( \frac{1}{\left( T^{-1} \boldsymbol{\varepsilon}_{r\circ}^{\prime}\mathbf{M}_{F} \boldsymbol{\varepsilon}_{r\circ}\right) ^{1/2}}-1\right) \left( \frac {1}{\left( T^{-1}\boldsymbol{\varepsilon}_{s\circ}^{\prime}\mathbf{M} _{F}\boldsymbol{\varepsilon}_{s\circ}\right) ^{1/2}}-1\right) \varepsilon _{it}\varepsilon_{jt^{\prime}}\varepsilon_{rt}\varepsilon_{st^{\prime} }\right\} . \end{align*} Since by part (a) of Assumption (ref), $\varepsilon_{it}^{\prime}s$ are cross-sectionally independent and after some algebra we have \begin{equation} \frac{1}{T}\sum_{t=1}^{T}\sum_{t^{\prime}=1}^{T}E\left( \xi_{t,n} \upsilon_{t,nT}\xi_{t^{\prime},n}v_{t^{\prime},nT}\right) =\sum_{l=1} ^{6}D_{l,nT}, \end{equation} where \[ D_{1,nT}=\frac{1}{Tn^{2}}\sum_{t=1}^{T}\sum_{t^{\prime}=1}^{T}\sum_{i=1} ^{n}a_{in}^{2}E\left[ \varepsilon_{it}^{2}\varepsilon_{it^{\prime}} ^{2}\left( \frac{1}{\left( T^{-1}\boldsymbol{\varepsilon}_{i\circ}^{\prime }\mathbf{M}_{F}\boldsymbol{\varepsilon}_{i\circ}\right) ^{1/2}}-1\right) ^{2}\right] , \] \begin{align*} D_{2,nT} & =\frac{1}{Tn^{2}}\sum_{t=1}^{T}\sum_{t^{\prime}=1}^{T}\sum _{i=1}^{n}\sum_{r\neq i}^{n}a_{in}^{2}E\left[ \varepsilon_{it}\varepsilon _{it^{\prime}}^{2}\left( \frac{1}{\left( T^{-1}\boldsymbol{\varepsilon }_{i\circ}^{\prime}\mathbf{M}_{F}\boldsymbol{\varepsilon}_{i\circ}\right) ^{1/2}}-1\right) \right] \times\\ & E\left[ \left( \frac{1}{\left( T^{-1}\boldsymbol{\varepsilon}_{r\circ }^{\prime}\mathbf{M}_{F}\boldsymbol{\varepsilon}_{r\circ}\right) ^{1/2} }-1\right) \varepsilon_{rt}\right] , \end{align*} \begin{align*} D_{3,nT} & =\frac{1}{Tn^{2}}\sum_{t=1}^{T}\sum_{t^{\prime}=1}^{T}\sum _{i=1}^{n}\sum_{s\neq i}^{n}a_{in}^{2}E\left[ \varepsilon_{it}^{2} \varepsilon_{it^{\prime}}\left( \frac{1}{\left( T^{-1} \boldsymbol{\varepsilon}_{i\circ}^{\prime}\mathbf{M}_{F} \boldsymbol{\varepsilon}_{i\circ}\right) ^{1/2}}-1\right) \right] \times\\ & E\left[ \left( \frac{1}{\left( T^{-1}\boldsymbol{\varepsilon}_{s\circ }^{\prime}\mathbf{M}_{F}\boldsymbol{\varepsilon}_{s\circ}\right) ^{1/2} }-1\right) \varepsilon_{st^{\prime}}\right] , \end{align*} \[ D_{4,nT}=\frac{1}{Tn^{2}}\sum_{t=1}^{T}\sum_{t^{\prime}=1}^{T}\sum_{i=1} ^{n}\sum_{r\neq i}^{n}a_{in}^{2}E\left( \varepsilon_{it}\varepsilon _{it^{\prime}}\right) E\left[ \left( \frac{1}{\left( T^{-1} \boldsymbol{\varepsilon}_{r\circ}^{\prime}\mathbf{M}_{F} \boldsymbol{\varepsilon}_{r\circ}\right) ^{1/2}}-1\right) ^{2} \varepsilon_{rt}\varepsilon_{rt^{\prime}}\right] , \] \begin{align*} D_{5,nT} & =\frac{1}{Tn^{2}}\sum_{t=1}^{T}\sum_{t^{\prime}=1}^{T}\sum _{i=1}^{n}\sum_{j\neq i}^{n}a_{in}a_{jn}E\left[ \varepsilon_{it}^{2}\left( \frac{1}{\left( T^{-1}\boldsymbol{\varepsilon}_{i\circ}^{\prime} \mathbf{M}_{F}\boldsymbol{\varepsilon}_{i\circ}\right) ^{1/2}}-1\right) \right] \times\\ & E\left[ \varepsilon_{jt^{\prime}}^{2}\left( \frac{1}{\left( T^{-1}\boldsymbol{\varepsilon}_{j\circ}^{\prime}\mathbf{M}_{F} \boldsymbol{\varepsilon}_{j\circ}\right) ^{1/2}}-1\right) \right] , \end{align*} and \begin{align*} D_{6,nT} & =\frac{1}{Tn^{2}}\sum_{t=1}^{T}\sum_{t^{\prime}=1}^{T}\sum _{i=1}^{n}\sum_{j\neq i}^{n}a_{in}a_{jn}E\left[ \varepsilon_{it} \varepsilon_{it^{\prime}}\left( \frac{1}{\left( T^{-1} \boldsymbol{\varepsilon}_{i\circ}^{\prime}\mathbf{M}_{F} \boldsymbol{\varepsilon}_{i\circ}\right) ^{1/2}}-1\right) \right] \times\\ & E\left[ \varepsilon_{jt}\varepsilon_{jt^{\prime}}\left( \frac{1}{\left( T^{-1}\boldsymbol{\varepsilon}_{j\circ}^{\prime}\mathbf{M}_{F} \boldsymbol{\varepsilon}_{j\circ}\right) ^{1/2}}-1\right) \right] . \end{align*} For $D_{1,nT}$ we have \begin{align*} D_{1,nT} & =\frac{1}{n^{2}}\sum_{i=1}^{n}a_{in}^{2}\frac{1}{T}\sum_{t=1} ^{T}E\left[ \varepsilon_{it}^{4}\left( \frac{1}{\left( T^{-1} \boldsymbol{\varepsilon}_{i\circ}^{\prime}\mathbf{M}_{F} \boldsymbol{\varepsilon}_{i\circ}\right) ^{1/2}}-1\right) ^{2}\right] \\ + & \frac{1}{n^{2}}\sum_{i=1}^{n}a_{in}^{2}\frac{1}{T}\sum_{t^{\prime}\neq t}^{T}E\left[ \varepsilon_{it}^{2}\varepsilon_{it^{\prime}}^{2}\left( \frac{1}{\left( T^{-1}\boldsymbol{\varepsilon}_{i\circ}^{\prime} \mathbf{M}_{F}\boldsymbol{\varepsilon}_{i\circ}\right) ^{1/2}}-1\right) ^{2}\right] =D_{1,1,nT}+D_{1,2,nT}. \end{align*} It is clear that $D_{1,1,nT}=O(n^{-1})$. Also by Cauchy-Schwarz inequality \[ E\left[ \varepsilon_{it}^{2}\varepsilon_{it^{\prime}}^{2}\left( \frac {1}{\left( T^{-1}\boldsymbol{\varepsilon}_{i\circ}^{\prime}\mathbf{M} _{F}\boldsymbol{\varepsilon}_{i\circ}\right) ^{1/2}}-1\right) ^{2}\right] \leq\left[ E\left( \varepsilon_{it}^{4}\varepsilon_{it^{\prime}}^{4}\right) \right] ^{1/2}\left[ E\left( \frac{1}{\left( T^{-1}\boldsymbol{\varepsilon }_{i\circ}^{\prime}\mathbf{M}_{F}\boldsymbol{\varepsilon}_{i\circ}\right) ^{1/2}}-1\right) ^{4}\right] ^{1/2}, \] where by part (a) of Assumption (ref) we have $E\left( \varepsilon_{it}^{4}\varepsilon_{it^{\prime}}^{4}\right) =E\left( \varepsilon_{it}^{4}\right) E\left( \varepsilon_{it^{\prime}}^{4}\right) =O\left( 1\right) $. Further, in view of ((ref)) \begin{align*} E\left( \frac{1}{\left( T^{-1}\boldsymbol{\varepsilon}_{i\circ}^{\prime }\mathbf{M}_{F}\boldsymbol{\varepsilon}_{i\circ}\right) ^{1/2}}-1\right) ^{4} & =E\left[ \frac{1}{\left( T^{-1}\boldsymbol{\varepsilon}_{i\circ }^{\prime}\mathbf{M}_{F}\boldsymbol{\varepsilon}_{i\circ}\right) ^{2} }\right] -4E\left[ \frac{1}{\left( T^{-1}\boldsymbol{\varepsilon}_{i\circ }^{\prime}\mathbf{M}_{F}\boldsymbol{\varepsilon}_{i\circ}\right) ^{3/2} }\right] +6E\left[ \frac{1}{T^{-1}\boldsymbol{\varepsilon}_{i\circ}^{\prime }\mathbf{M}_{F}\boldsymbol{\varepsilon}_{i\circ}}\right] \\ & -4E\left[ \frac{1}{\left( T^{-1}\boldsymbol{\varepsilon}_{i\circ} ^{\prime}\mathbf{M}_{F}\boldsymbol{\varepsilon}_{i\circ}\right) ^{1/2} }\right] +1=O\left( \frac{1}{T}\right) . \end{align*} Using this result it follows that $D_{1,2,nT}=$ $O\left( \sqrt{T} n^{-1}\right) ,$ and overall we have \begin{equation} D_{1,nT}=O\left( \sqrt{T}n^{-1}\right) . \end{equation} Consider now $D_{2,nT}$ and note \begin{align*} \left\vert D_{2,nT}\right\vert & \leq\frac{1}{Tn^{2}}\sum_{t=1}^{T} \sum_{t^{\prime}=1}^{T}\sum_{i=1}^{n}\sum_{r\neq i}^{n}a_{in}^{2}\left\vert E\left[ \varepsilon_{it}\varepsilon_{it^{\prime}}^{2}\left( \frac{1}{\left( T^{-1}\boldsymbol{\varepsilon}_{i\circ}^{\prime}\mathbf{M}_{F} \boldsymbol{\varepsilon}_{i\circ}\right) ^{1/2}}-1\right) \right] \right\vert \times\\ & \left\vert E\left[ \left( \frac{1}{\left( T^{-1}\boldsymbol{\varepsilon }_{r\circ}^{\prime}\mathbf{M}_{F}\boldsymbol{\varepsilon}_{r\circ}\right) ^{1/2}}-1\right) \varepsilon_{rt}\right] \right\vert . \end{align*} By ((ref)) we have \begin{equation} E\left[ \left( \frac{1}{\left( T^{-1}\boldsymbol{\varepsilon}_{r\circ }^{\prime}\mathbf{M}_{F}\boldsymbol{\varepsilon}_{r\circ}\right) ^{1/2} }-1\right) \varepsilon_{rt}\right] =E\left[ \frac{\varepsilon_{rt}}{\left( T^{-1}\boldsymbol{\varepsilon}_{r\circ}^{\prime}\mathbf{M}_{F} \boldsymbol{\varepsilon}_{r\circ}\right) ^{1/2}}\right] =O\left( \frac {1}{T}\right) . \end{equation} Under Assumption (ref) and given results ((ref)) and ((ref)), it follows \begin{align} E\left[ \varepsilon_{it}^{2}\left( \frac{1}{\left( T^{-1} \boldsymbol{\varepsilon}_{i\circ}^{\prime}\mathbf{M}_{F} \boldsymbol{\varepsilon}_{i\circ}\right) ^{1/2}}-1\right) ^{2}\right] & =E\left[ \frac{\varepsilon_{it}^{2}}{\left( T^{-1}\boldsymbol{\varepsilon }_{i\circ}^{\prime}\mathbf{M}_{F}\boldsymbol{\varepsilon}_{i\circ}\right) }\right] -2E\left[ \frac{\varepsilon_{it}^{2}}{\left( T^{-1} \boldsymbol{\varepsilon}_{i\circ}^{\prime}\mathbf{M}_{F} \boldsymbol{\varepsilon}_{i\circ}\right) ^{1/2}}\right] +1\nonumber\\ & =\left[ E\left[ \frac{\varepsilon_{it}^{2}}{\left( T^{-1} \boldsymbol{\varepsilon}_{i\circ}^{\prime}\mathbf{M}_{F} \boldsymbol{\varepsilon}_{i\circ}\right) }\right] -1\right] -2\left[ E\left[ \frac{\varepsilon_{it}^{2}}{\left( T^{-1}\boldsymbol{\varepsilon }_{i\circ}^{\prime}\mathbf{M}_{F}\boldsymbol{\varepsilon}_{i\circ}\right) ^{1/2}}\right] -1\right] \nonumber\\ & =O\left( \frac{1}{T}\right) . \end{align} Further, by Cauchy-Schwarz inequality, \[ \left\vert E\left[ \varepsilon_{it}\varepsilon_{it^{\prime}}^{2}\left( \frac{1}{\left( T^{-1}\boldsymbol{\varepsilon}_{i\circ}^{\prime} \mathbf{M}_{F}\boldsymbol{\varepsilon}_{i\circ}\right) ^{1/2}}-1\right) \right] \right\vert \leq\left[ E\left( \varepsilon_{it^{\prime}} ^{4}\right) \right] ^{1/2}\left[ E\left[ \varepsilon_{it}^{2}\left( \frac{1}{\left( T^{-1}\boldsymbol{\varepsilon}_{i\circ}^{\prime} \mathbf{M}_{F}\boldsymbol{\varepsilon}_{i\circ}\right) ^{1/2}}-1\right) ^{2}\right] \right] ^{1/2}, \] whi