Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.
This text was truncated for display. The citation measures were computed over the complete text.
How to Detect Network Dependence in Latent Factor Models? A Bias-Corrected CD Test
abstractIn a recent paper juodis2022incidental (JR) show that the application
of the CD test proposed by pesaran2004general to residuals from panels
with latent factors results in over-rejection. They propose a randomized test
statistic to correct for over-rejection, and add a screening component to
achieve power. This paper considers the same problem but from a different
perspective, and shows that the standard CD test remains valid if the latent
factors are weak. A bias-corrected version, CD$^{\text{*}}$, is proposed which is shown to be asymptotically
standard normal under the null of error cross-sectional independence which has
power against network type alternatives. This result is shown to hold for pure
latent factor models as well as for panel regression models with latent
factors. The case where the errors are serially correlated is also considered.
Small sample properties of the CD$^{\text{*}}$ test are investigated by Monte Carlo experiments and are shown to have satisfactory small sample properties. In an empirical application, using the CD$^{\text{*}}$ test, it
is shown that there remains spatial error dependence in a panel data model for
real house price changes across 377 Metropolitan Statistical Areas in the
U.S., even after the effects of latent factors are filtered out.
JEL Classifications:\ { C18, C23, C55}
Key Words:{ \ Latent factor models, strong and weak
factors, error cross-sectional independence, spatial and network alternatives,
size and power.}
\pagenumbering{gobble}
\pagenumbering{arabic}
Introduction
It is now quite standard to use latent multi-factor models to characterize and
explain cross-sectional dependence in panels when the cross section dimension
$(n)$ and the time series dimension $(T)$ are both large. However, due to
uncertainty regarding the nature of error cross-sectional dependence, it is
arguable whether error cross-sectional dependence is fully accounted for by
latent factors. Some of the factors could be semi-strong, and the errors might
have spatial or network features that are not necessarily captured by common
factors alone. chudik2011weak provide an early discussion of the
different sources of cross-sectional dependence. Diagnostic tests of error
cross-sectional independence in panels are required to safeguard against
estimation bias and unreliable inference. See, for example,
bai2003inferential,bai2009panel,
phillips2003dynamic,phillips2007bias, bai2006confidence,
pesaran2006estimation, and pesaran2011large. Such tests are also
helpful to researchers interested in network or spatial dependence once the
influence of common factors are filtered out. See, for example,
bailey2016two, shi2017spatial, aquaro2021estimation, and
bai2021dynamic amongst others.
One standard test for error cross-sectional independence is the CD test
proposed by pesaran2004general,pesaran2021general, which has been
further developed. For example, hsiao2012diagnostic apply the CD
test to panel data models with limited dependent variables, while
pesaran2015testing uses it to test weak error cross-sectional
dependence in large panels. In a recent paper juodis2022incidental (JR)
show that the application of the CD test to the residuals from panels with
latent factors is invalid and can result in over-rejection of the null of
error cross-sectional independence.\footnote{In the empirical finance
literature gagliardini2019diagnostic propose a diagnostic criterion to
check if the errors from a (strong) factor model are weakly correlated or
contain missing strong factor(s). These authors do not propose a test of
cross-sectional error dependence but focus on the detection of potentially
omitted strong factors in asset pricing models.} They propose a randomized CD
test statistic as a solution. Their proposed test is constructed in two steps.
First, they multiply the residuals from panel regressions with independent
randomized weights to obtain their CD$_{W}$ statistic, which will have a zero
mean by construction. In this way they avoid the over-rejection problem of the
CD test, but by the very nature of the randomization process they recognize
that the CD$_{W}$ test will lack power. To overcome the problem of lack of
power, JR modify the CD$_{W}$ test statistic by adding to it a screening
component proposed by fan2015power which is expected to tend to zero
with probability approaching one under the null hypothesis, but to diverge at
a reasonably fast rate under the alternative. This further modification of the
CD$_{W}$ test is denoted by CD$_{W+}$. Accordingly, it is presumed that the
CD$_{W+}$ test can overcome both over-rejection and the low power problems.
However, JR do not provide a formal proof establishing conditions under which
the screening component tends to zero under the null and diverges sufficiently
fast under alternatives, including spatial or network dependence type
alternatives. Also, our Monte Carlo simulations show that the CD$_{W+}$ test
tends to over-reject when the errors are non-Gaussian and $n>>T$, and lacks power under spatial and network alternatives, which is likely to be
particularly important in empirical applications.
In this paper we show that the standard CD test is in fact valid for testing
error cross-sectional independence in panel data models with weak latent
factors. However, when the latent factors are semi-strong or strong the use of
the CD test will result in over-rejection and will no longer be valid,
extending JR's results to panels with semi-strong latent factors. In short,
whilst the CD$_{W+}$ test is a useful and welcome addition to testing for
error cross-sectional independence, it would be interesting to develop a
modified version of the test that simultaneously deals with the over-rejection
problem and does not compromise power for a general class of alternatives. To
that end, firstly we study testing for error cross-sectional independence in a
pure latent factor model, and derive an explicit expression for the bias of
the CD test statistic in terms of factor loadings and error variances. We then
propose a bias-corrected version of the CD test statistic, denoted by
$CD^{\ast}$, which is shown to have $\mathcal{N}(0,1)$ asymptotic distribution
under the null hypothesis irrespective of whether the latent factors are weak
or strong. When the latent factors are weak the correction tends to zero, $CD$
and $CD^{\ast}$ will be asymptotically equivalent. However, $CD-CD^{\ast}$
diverges if at least one of the underlying latent factors is (semi) strong. We
show that under the null of cross-sectional independence, $CD^{\ast}$
converges to a standard normal distribution when $n$ and $T$ tend to infinity
so long as $n/T\rightarrow\kappa$, where $0<\kappa<\infty$, and a test based
on $CD^{\ast}$ will have the correct size asymptotically. In addition, it is
shown that the CD$^{\text{*}}$ test has power against spatial and network type
alternatives. In particular, we are the first to give a formal derivation of
the power function for CD tests against general spatial and network
alternatives, which can be applied equally to panel data models without latent
factors and therefore supplement earlier research on CD tests\footnote{For
instance, pesaran2004general, which is the unpublished version of
pesaran2021general, also discusses the power of the CD test against
spatial dependence in Section 8.2 of his paper, for a specific connection matrix with $n$
fixed as $T\rightarrow\infty$.}.
We then consider the application of the $CD^{\ast}$ to test error
cross-sectional independence in the case of panel regression models with
latent factors, discussed in pesaran2006estimation. It is shown that
the asymptotic properties of the $CD^{\ast}$ in the case of pure latent factor
models also carry over to panel regression models with latent factors. We also
investigate the application of the $CD^{\ast}$ to panel data models with
serially correlated errors, and consider the method proposed by
baltagi2016testing as well as using an autoregressive distributed lag
(ARDL) representation which transforms the model with serially correlated
errors to one without error serial correlation.
The finite sample performance of the CD$^{\text{*}}$ test is investigated by
Monte Carlo simulations in the case of pure latent factor models, panel
regression models with latent factors with and without error serial
correlation. It is found that the CD$^{\text{*}}$ test avoids the
over-rejection problem under the null and has power against spatial and
network alternatives, and has desirable small sample properties regardless of
whether the errors are Gaussian or not, under different combinations of $n$
and $T$. We also find that both adjustments for dealing with error serial
correlation considered in the paper give desirable small sample properties.
Finally, as compared to JR's CD$_{W^{+}}$ test, the proposed bias-corrected CD
test is better in controlling the size of the test and has much better power
properties against spatial or network alternatives.
The use of the CD$^{\text{*}}$ test is illustrated by an empirical application
in modeling real house price changes in the U.S. Because it is evident that
real house price changes are driven by macroeconomic trends which can be
modeled by latent factors, it is necessary to filter out these factors before
testing for spillover effect. By applying the CD$^{\text{*}}$ test to real
house price changes in the U.S. we are able to show significant existence of
weak cross-sectional dependence in addition to latent factors.
The rest of the paper is organized as follows. Section (ref) sets
out the latent factor model and its assumptions. Section (ref)
introduces the estimation of latent factors and the CD test. The
bias-corrected test, CD$^{\ast}$, is introduced in Section (ref) and its
asymptotic distributions are derived under the null and the alternative
hypotheses. The extension to more general panel regression models with
observed covariates as well as latent factors are discussed in Section
(ref). Adjustments to the CD$^{\text{*}}$ test for panels with
serially correlated errors are discussed in Section (ref).
Using Monte Carlo techniques, the small sample properties of CD, CD$^{\ast}$,
and CD$_{W+}$ tests are discussed in Section (ref). An empirical
illustration is provided in Section (ref). Proofs of the propositions and
theorems are provided in an appendix. The auxiliary lemmas and the associated proofs are given in a supplement.
Notations: For the $n\times n$ matrix $\mathbf{A}=\left(
a_{ij}\right) $, we denote its largest eigenvalue by $\mu_{max}\left(
\mathbf{A}\right) $, its trace by $\mathrm{tr}\left( \mathbf{A}\right)
=\sum_{i=1}^{n}a_{ii}$, its spectral norm by $\left\Vert \mathbf{A}\right\Vert
=\mu_{\max}^{1/2}\left( \mathbf{A}^{\prime}\mathbf{A}\right) $, its maximum
absolute column sum norm by $\left\Vert \mathbf{A}\right\Vert _{1}=\max_{1\leq
j\leq n}\left( \sum_{i=1}^{n}\left\vert a_{ij}\right\vert \right) $, and its
maximum absolute row sum norm by $\left\Vert \mathbf{A}\right\Vert _{\infty
}=\max_{1\leq i\leq n}\left( \sum_{j=1}^{n}\left\vert a_{ij}\right\vert
\right) $. We write $\mathbf{A}>\mathbf{0}$ when $\mathbf{A}$ is positive
definite. For matrices $\mathbf{B}=\left( b_{ij}\right) $ and $\mathbf{C}
=\left( c_{ij}\right) $, $\mathbf{B\odot C=C\odot B}$ denote Hadamard
product with elements $b_{ij}c_{ij}$. $\rightarrow_{p}$ denotes convergence in
probability, $\rightarrow_{d}$ convergence in distribution, and
$\overset{a}{\thicksim}$ asymptotic equivalence in distribution. $O_{p}\left(
\cdot\right) $ and $o_{p}\left( \cdot\right) $ denote the stochastic order
relations. In particular, $o_{p}(1)$ indicates terms that tend to zero in
probability as $(n,T)\rightarrow\infty$, such that $n/T\rightarrow\kappa$,
where $0<\kappa<\infty$. $C$ and $c$ will be used to denote finite large and
non-zero small positive numbers, respectively, that are bounded in $n$ and
$T$. They can take different values at different instances. If $\left\{
f_{n}\right\} _{n=1}^{\infty}$ is any real sequence and $\left\{
g_{n}\right\} _{n=1}^{\infty}$ is a sequence of positive real numbers, then
$f_{n}=O(g_{n})$, if there exists $C$ such that $\left\vert f_{n}\right\vert
/g_{n}\leq C$ for all $n$. $f_{n}=o(g_{n})$ if $f_{n}/g_{n}\rightarrow0$ as
$n\rightarrow\infty$. If $\left\{ f_{n}\right\} _{n=1}^{\infty}$ and
$\left\{ g_{n}\right\} _{n=1}^{\infty}$ are both positive sequences of real
numbers, then $f_{n}=\ominus\left( g_{n}\right) $ if there exists $n_{0}
\geq1$ and positive finite constants $C_{0}$ and $C_{1}$, such that
$\inf_{n\geq n_{0}}\left( f_{n}/g_{n}\right) \geq C_{0},$ and $\sup_{n\geq
n_{0}}\left( f_{n}/g_{n}\right) \leq C_{1}$.
The latent factor model
To simplify the exposition and to highlight the main issue of concern, namely
the presence of latent (unobserved) factors in the panel regression
model, initially we focus on the approximate factor model, due to
chamberlain1983arbitrage, and assume that for each unit $i=1,2,\ldots
,n$,
equation[equation omitted — 117 chars of source]
where
equation[equation omitted — 167 chars of source]
$\sup_{i}\sigma_{i}^{2}<C<\infty$ and $\inf_{i}\sigma_{i}^{2}>c>0$, $\left\{
w_{ij}:j=1,2,\ldots,n\right\} $ represent the strengths of connections of
unit $i$ with the rest of units, $\mathbf{f}_{t}=\left( f_{1t},f_{2t}
,...,f_{m_{0}t}\right) ^{\prime}$ is an $m_{0}\times1$ vector of latent
factors with $m_{0}$ fixed, and $\boldsymbol{\gamma}_{i}=\left( \gamma
_{i1},\gamma_{i2},...,\gamma_{im_{0}}\right) ^{\prime}$ is the vector of
associated factor loadings.
We make the following assumptions that are mostly standard in the analysis of
latent factor models.
assumption(a) $\mathbf{f}_{t}$ is a covariance-stationary process
with zero means and the covariance matrix, $E\left( \mathbf{f}_{t}
\mathbf{f}_{t}^{\prime}\right) =\mathbf{\Sigma}_{ff}>\mathbf{0}$. (b)
$T^{-1}\sum_{t=1}^{T}\left[ \Vert\mathbf{f}_{t}\Vert^{j}-E\left(
\Vert\mathbf{f}_{t}\Vert^{j}\right) \right] {\rightarrow}_{p}0$, for
$j=3,4$, as $T\rightarrow\infty$. (c) There exists $T_{0}$ such that
for all $T>T_{0}$, $T^{-1}\sum_{t=1}^{T}\mathbf{f}_{t}\mathbf{f}_{t}^{\prime
}=T^{-1}\mathbf{F}^{\prime}\mathbf{F=\Sigma}_{T,ff}>\mathbf{0}$, and
$\mathbf{\Sigma}_{T,ff}\rightarrow_{p}E\left( T^{-1}\mathbf{F}^{\prime
}\mathbf{F}\right) =\mathbf{\Sigma}_{ff}>\mathbf{0}$, where $\mathbf{F=(f}
_{1},\mathbf{f}_{2},...,\mathbf{f}_{T})^{\prime}$. (d) There exist constants
$r_{1}$, $C_{0}$ and $C_{1}>0$ such that
\begin{equation}
\sup_{j,t}\Pr\left( \left\vert f_{jt}\right\vert >a\right) \leq C_{0}
\exp\left( -C_{1}a^{r_{1}}\right),
\end{equation}
all $a>0$.
assumption(a) $\varepsilon_{it}\sim IID\left( 0,1\right) $ for all
$i$ and $t$, and there exist constants $r_{2}$, $C_{2}$ and $C_{3}>0$ such
that
\begin{equation}
\sup_{i,t}\Pr\left( \left\vert \varepsilon_{it}\right\vert >a\right) \leq
C_{2}\exp\left( -C_{3}a^{r_{2}}\right),
\end{equation}
for all $a>0$. (b) $\mu_{\max}\left( \mathbf{V}_{\varepsilon T}\right)
=O_{p}(n/T)$, where $\mathbf{V}_{\varepsilon T}=T^{-1}\sum_{t=1}
^{T}\boldsymbol{\varepsilon}_{\circ t}\boldsymbol{\varepsilon}_{\circ
t}^{\prime}$, and $\boldsymbol{\varepsilon}_{\circ t}=(\varepsilon
_{1t},\varepsilon_{2t},...,\varepsilon_{nt})^{\prime}$. (c) $\varepsilon_{it}$
is distributed independently of $\mathbf{f}_{t^{\prime}}$, for all $i,t$ and
$t^{\prime}$, and there exists $v_{0}>0$ such that for all $v=T-m_{0}>v_{0}$
\begin{equation}
\inf_{i}\left( v^{-1}\boldsymbol{\varepsilon}_{i\circ}^{\prime}\mathbf{M}
_{F}\boldsymbol{\varepsilon}_{i\circ}\right) >c>0,
\end{equation}
where $\boldsymbol{\varepsilon}_{i\circ}=(\varepsilon_{i1},\varepsilon
_{i2},...,\varepsilon_{iT})^{\prime}$, and $\mathbf{M}_{F}=\mathbf{I}
_{T}-\mathbf{F(F}^{\prime}\mathbf{F)}^{-1}\mathbf{F}^{\prime}$.
assumptionThe $m_{0}\times1$ vector of factor loadings
$\boldsymbol{\gamma}_{i}$ is bounded such that $\sup_{i}\left\Vert
\boldsymbol{\gamma}_{i}\right\Vert <C$, $n^{-1}\sum_{i=1}^{n}
\boldsymbol{\gamma}_{i}\boldsymbol{\gamma}_{i}^{\prime}=\mathbf{\Sigma
}_{n,\gamma\gamma}\rightarrow\mathbf{\Sigma}_{\gamma\gamma}>\mathbf{0}$, and
\begin{equation}
1-\theta_{n}>0,
\end{equation}
for all $n>n_{0}$ and as $n\rightarrow\infty,$ where $\theta_{n}=1-n^{-1}
\sum_{i=1}^{n}a_{i,n}^{2}$, with $
\text{ }a_{i,n}=1-\sigma_{i}\boldsymbol{\varphi}_{n}^{\prime}
\boldsymbol{\gamma}_{i},
$ and $\boldsymbol{\varphi}_{n}=n^{-1}\sum_{i=1}^{n}\boldsymbol{\gamma}
_{i}/\sigma_{i}.$
remarkThe above assumptions relate closely to those made in the
literature on CD tests and high dimensional factor models. See, for example,
the assumptions in
pesaran2004general,pesaran2015testing,pesaran2021general, and
assumptions in bai2003inferential.\ Part (a) of Assumption
(ref) will be relaxed when we consider panel regression models
with observed regressors. The sub-exponential type conditions ((ref))
and ((ref)) are needed for bounding the probabilities across all $i$,
and are also adopted by fan2011high, fan2013large and
chudik2018one.
remarkSince $\mathbf{M}_{F}$ is an idempotent matrix with rank $v$,
then there exists the orthogonal transformation $\boldsymbol{\eta}_{i}
=(\eta_{i1},\eta_{i2},...,\eta_{iv})^{\prime}=\mathbf{H}
\boldsymbol{\varepsilon}_{i\circ}$ where $\mathbf{H}$ is a $v\times T$ matrix
such that $v^{-1}\boldsymbol{\varepsilon}_{i\circ}^{\prime}\mathbf{M}
_{F}\boldsymbol{\varepsilon}_{i\circ}=v^{-1}\boldsymbol{\eta}_{i}^{\prime
}\boldsymbol{\eta}_{i}>0,$ and $E\left( \eta_{it}^{2}\right) =1$. See, for
example, durbin1950testing. Therefore, there exists a finite
$v_{0}$ such that for all $v>v_{0}$ condition ((ref)) is met, and as
result we also have
\begin{equation}
E\left( \frac{\boldsymbol{\varepsilon}_{i\circ}^{\prime}\mathbf{M}
_{F}\boldsymbol{\varepsilon}_{i\circ}}{v}\right) ^{-s}<\frac{1}{c^{s}
}<C<\infty,
\end{equation}
for all $i$ and any fixed $s>0$.
The focus of this paper is on testing the null hypothesis of error
cross-sectional independence:
equation[equation omitted — 47 chars of source]
where $\lambda_{T}$ is defined by equation ((ref)). For the analysis of
power we consider local alternatives:
equation[equation omitted — 67 chars of source]
with $c_{\lambda}\neq0$. We also introduce the following assumption on
$\mathbf{W}=\left( w_{ij}\right) $.
assumptionThe connection matrix $\mathbf{W}$ has bounded maximum absolute
column and row sum norms:
\begin{equation}
\left\Vert \mathbf{W}\right\Vert _{1}=\sup_{j}\sum_{i=1}^{n}\left\vert
w_{ij}\right\vert <C, and \left\Vert \mathbf{W}\right\Vert _{\infty
}=\sup_{i}\sum_{j=1}^{n}\left\vert w_{ij}\right\vert <C,
\end{equation}
and $w_{ii}=0$ for all $i$.
remarkThe connection matrix does not need to be symmetric. To see this, we can
consider a more generalized setup of idiosyncratic errors,
\begin{equation}
\frac{u_{it}}{\sigma_{i}}=\varepsilon_{it}+\lambda_{T}{\sum\limits_{j=1}^{n}
}\frac{\rho_{i}\sigma_{j}\mathring{w}_{ij}}{\sigma_{i}}\varepsilon_{jt},
\end{equation}
where $\left\vert \lambda_{T}\right\vert <C$ and $\left\vert \rho
_{i}\right\vert <C$. It is clear that ((ref)) reduces to $u_{it}$ in
((ref)) by letting $w_{ij}=\rho_{i}\sigma_{i}^{-1}\sigma_{j}\mathring
{w}_{ij}$. In this way, $\mathbf{W}$ need not be symmetric even if the
connection matrix $\left( \mathring{w}_{ij}\right) $ is symmetric. Under
local alternatives and Assumption (ref), the specification of
((ref)) allows for a wide range of spatial and network dependence
characterized by the connection matrix. It is in accord with an early
discussion in chudik2011weak that spatial dependence can be captured by
a weak factor model, so long as the number of weak factors tends to infinity
with $n$, which is ruled out in standard factor models where the number of
latent factors is assumed to be fixed.
remarkIt is also easily seen that $\varepsilon_{it}\left( \lambda_{T}\right) $
defined by ((ref)) is sub-exponential for any $\left\vert \lambda
_{T}\right\vert <C$. This follows since by Assumption (ref)
$\left\{ \varepsilon_{it}\right\} $ are independently and identically
distributed sub-exponential processes, and by Assumption (ref)
$\sup_{i}\sum_{j=1}^{n}\left\vert w_{ij}\right\vert <C.$ For a proof see part
(b) of Theorem 5.5 in goldie1998subexponential or Theorem 2.8.2 in
vershynin2018high.
Estimation of latent factors and the CD test
Following the literature we use principal component (PC) analysis to estimate
the latent factors and their loadings. Let $\mathbf{Y=(y}_{1},\mathbf{y}
_{2},\ldots,\mathbf{y}_{n}\mathbf{)}$ be the $T\times n$ matrix of
observations on $y_{it}$, where $\mathbf{y}_{i}=\left( y_{i1},y_{i2}
,\ldots,y_{iT}\right) ^{^{\prime}}$ and denote the first $m_{0}$ largest
eigenvalues of $\mathbf{Y}^{^{\prime}}\mathbf{Y}$ by $(\hat{\rho}_{1}
,\hat{\rho}_{2},...,\hat{\rho}_{m_{0}})$, and its associated $n\times m_{0}$
matrix of orthonormal eigenvectors by $\mathbf{\hat{Q}}$. The PC estimators of
factors $\mathbf{F}=\left( \mathbf{f}_{1},\mathbf{f}_{2},\ldots
,\mathbf{f}_{T}\right) ^{^{\prime}}$ and their loadings $\boldsymbol{\Gamma
}=\left( \boldsymbol{\gamma}_{1},\boldsymbol{\gamma}_{2},\ldots
,\boldsymbol{\gamma}_{n}\right) ^{^{\prime}}$ are then given by
equation[equation omitted — 396 chars of source]
By construction $n^{-1}\boldsymbol{\hat{\Gamma}}^{\prime}\boldsymbol{\hat
{\Gamma}}=\mathbf{I}_{m_{0}},$ and $T^{-1}\mathbf{\hat{F}}^{^{\prime}
}\mathbf{\hat{F}=D}_{nT}$, where $\mathbf{D}_{nT}=(nT)^{-1}\mathrm{diag}
(\hat{\rho}_{1},\hat{\rho}_{2},...,\hat{\rho}_{m_{0}})$. Under Assumptions
(ref)-(ref) the asymptotic results derived by
bai2003inferential for PCs continue to apply here, and $u_{it}$ can be
consistently estimated by
equation[equation omitted — 111 chars of source]
The CD test is based on the standardized residuals,
equation[equation omitted — 98 chars of source]
where $\hat{\sigma}_{i,T}=\left( T^{-1}\sum_{t=1}^{T}\hat{u}_{it}^{2}\right)
^{1/2}=\left( T^{-1}\mathbf{y}_{i}^{\prime}\mathbf{M}_{\hat{F}}\mathbf{y}
_{i}\right) ^{1/2}$, and $\mathbf{M}_{\hat{F}}=\mathbf{I}_{T}-\mathbf{\hat
{F}(\hat{F}}^{\prime}\mathbf{\hat{F})}^{-1}\mathbf{\hat{F}}^{\prime}$. Only
units with non-zero $\hat{\sigma}_{i,T}^{2}$ are included in the construction
of the CD test, namely
equation[equation omitted — 70 chars of source]
The standard CD test statistic based on the residuals, ((ref)), is given
by
equation[equation omitted — 122 chars of source]
where $\hat{\rho}_{ij,T}=T^{-1}\sum_{t=1}^{T}\tilde{\varepsilon}_{it,T}
\tilde{\varepsilon}_{jt,T}$.
juodis2022incidental apply the CD test to a panel regression model with
latent factors, assuming that all the factors are strong. They show in that
case $CD=O_{p}\left( \sqrt{T}\right) $, and its use will lead to gross
over-rejection of the null of error cross-sectional independence. To deal with
the over-rejection problem, these authors propose a randomized CD test,
CD$_{W+}$. However, as shown in Section (ref) of the supplement,
the CD$_{W+}$ test is likely to over-reject and tends to lack power against
spatial and network alternatives. See also Section
(ref) for Monte Carlo evidence on the small sample
performance of the CD$_{W+}$ test.
The bias-corrected CD test
The main reason for the failure of the standard CD test in the case of latent
factor models lies in the fact that both the factors and their loadings are
unobserved and need to be estimated, and the differences between
$\boldsymbol{\hat{\gamma}}_{i}^{^{\prime}}\mathbf{\hat{f}}_{t}$ and
$\boldsymbol{\gamma}_{i}^{^{\prime}}\mathbf{f}_{t}$ do not tend to zero at a
sufficiently fast rate for the CD test to be valid. Since the errors from
estimation of $\boldsymbol{\gamma}_{i}^{^{\prime}}\mathbf{f}_{t}$ are included
in the residuals $\hat{u}_{it}$, the resultant CD statistic tends to
over-state the degree of underlying error cross-sectional dependence. This
problem also arises when latent factors are proxied by cross section averages,
as is the case when panel data models are estimated using correlated common
effect (CCE) estimators proposed by pesaran2006estimation, which we
shall address below in Section (ref).
We propose a bias-corrected CD test statistic, which we denote by $CD^{\ast}$,
that directly corrects the asymptotic bias of the CD test using the
estimates of the factor loadings and error variances. To obtain the expression
for the bias we first note under the null hypothesis of cross-sectional
independence, $CD=z_{nT}+o_{p}(1)$ and
equation[equation omitted — 142 chars of source]
where
equation[equation omitted — 183 chars of source]
$\boldsymbol{\varphi}_{n}=n^{-1}\sum_{i=1}^{n}\boldsymbol{\gamma}_{i}
/\sigma_{i}$,$\boldsymbol{\ }$which is established in the proof of Proposition
(ref) in the Appendix. Since $a_{i,n}$ are given
constants, then $E\left( \xi_{t,n}\right) =0$,
equation[equation omitted — 231 chars of source]
and
equation[equation omitted — 193 chars of source]
where $\kappa_{2}=E\left( \varepsilon_{it}^{4}\right) -3$. Clearly, when the
errors are Gaussian then $E\left( \varepsilon_{it}^{4}\right) =3$, and the
second term of $Var\left( \xi_{t,n}^{2}\right) $ defined by ((ref)) is
exactly zero. But even for non-Gaussian errors the second term of $Var\left(
\xi_{t,n}^{2}\right) $ is negligible when $n$ is sufficiently large. To see
this note that under Assumptions (ref) and (ref)
\[
\frac{1}{n^{2}}\sum_{i=1}^{n}a_{i,n}^{4}=\frac{1}{n^{2}}\sum_{i=1}
^{n}(1-\sigma_{i}\boldsymbol{\varphi}_{n}^{\prime}\boldsymbol{\gamma}_{i}
)^{4}\leq\frac{C}{n},
\]
where $C$ is a positive constant. Since $\varepsilon_{it}$ (and henceforth
$\xi_{t,n}$) are assumed to be serially independent, then we can also compute
the mean and the variance of $z_{nT}$ as
align*[align* omitted — 332 chars of source]
The above expressions for $E\left( z_{nT}\right) $ give the source of the
asymptotic bias of $CD$ as $E\left( z_{nT}\right) $ rises with $\sqrt{T}$,
unless
\[
\lim_{n\rightarrow\infty}\omega_{n}^{2}=\lim_{n\rightarrow\infty}n^{-1}
\sum_{i=1}^{n}\left( 1-\sigma_{i}\boldsymbol{\varphi}_{n}^{\prime
}\boldsymbol{\gamma}_{i}\right) ^{2}=1.
\]
A bias-corrected version of $CD$ can be defined by
equation[equation omitted — 105 chars of source]
where
equation[equation omitted — 169 chars of source]
and $1-\theta_{n}>0$ by condition ((ref)). Also upon using
((ref))
equation[equation omitted — 328 chars of source]
The main difference between $CD$ and $CD^{\ast}\left( \theta_{n}\right) $
depends on the magnitude of $\sqrt{T}\theta_{n}$, which in turn depends on the
strengths of the factor loadings. Following bailey2021measurement, we
measure the strength of factor $j$ by $\alpha_{j},$ defined by the rate at
which the sum of absolute values of factor loadings rises with $n$, namely
equation[equation omitted — 147 chars of source]
where $0\leq\alpha_{j}\leq1$. Using ((ref)) it is now easily
established that $\theta_{n}$ $=\ominus\left( n^{\alpha-1}\right) $, where
$\alpha=max_{j=1,2,...,m_{0}}(\alpha_{j})$, and $\theta_{n}$ does not tend to
zero when there is at least one strong factor in the panel data
model.\footnote{For a proof see Section (ref) of the supplement.}. Therefore, based on ((ref)), the relationship between $CD$
and $CD^{\ast}\left( \theta_{n}\right) $ is essentially controlled by the
maximum factor strength $\alpha\ $as $\sqrt{T}\theta_{n}=O\left(
T^{1/2}n^{\alpha-1}\right) $. Suppose now $T=\ominus\left( n^{d}\right) $
for some $d>0$, then $\sqrt{T}\theta_{n} =\ominus\left( n^{\alpha+d/2-1}\right) ,$ and the bias
correction becomes negligible if $\alpha<1-d/2$. Under the required relative
expansion rates of $n$ and $T$ entertained in this paper, we need to set
$d=1$, and for this choice the bias correction term, $\sqrt{T}\theta_{n}$,
becomes negligible if $\alpha<1/2$, and as a result $CD$ and
$CD^{\ast}\left( \theta_{n}\right) $ will be asymptotically equivalent. In fact,
the case of strong factors assumed in the PCA literature corresponds to
$\alpha_{j}=1$ for $j=1,2,\ldots,m_{0}$, which is also fulfilled by Assumption
(ref) and used in our mathematical derivations.
The theoretical results for $CD^{\ast}(\theta_{n})$ are summarized in the
following proposition.
propositionSuppose that observations on $y_{it}$, for
$i=1,2,\ldots,n,$ and $t=1,2,\ldots,T$ are generated from the pure latent
factor model given by ((ref)) and ((ref)), where the number of
factors, $m_{0}$, is known. Consider the statistic $CD^{\ast}\left(
\theta_{n}\right) $ defined by ((ref)) and assume $(n,T)\rightarrow
\infty$, such that $n/T\rightarrow\kappa,$ and $0<\kappa<\infty$.
(a) Under the null hypothesis $H_{0}$, defined by ((ref)), and supposing
that Assumptions (ref) to (ref) hold, then
\begin{equation}
CD^{\ast}\left( \theta_{n}\right) \rightarrow_{d}\mathcal{N}(0,1).
\end{equation}
(b) Under local alternatives $H_{1T}$, defined by ((ref)), and supposing
that Assumptions (ref) to (ref) hold, then
\begin{equation}
CD^{\ast}\left( \theta_{n}\right) \rightarrow_{d}\mathcal{N}(\phi,1),
\end{equation}
where $\phi=\lim_{n\rightarrow\infty}\phi_{n}$ and
\begin{equation}
\phi_{n}=\frac{\sqrt{2}c_{\lambda}}{1-\theta_{n}}n^{-1}\mathbf{a}_{n}^{\prime
}\mathbf{Wa}_{n},
\end{equation}
$\mathbf{W}=\left( w_{ij}\right) $ is the connection matrix, $\mathbf{a}
_{n}=\left( a_{1,n},a_{2,n},\ldots,a_{n,n}\right) ^{\prime}$, with $a_{i,n}$
and $\theta_{n}$ defined by ((ref)) and ((ref)), respectively.
For a proof see the Appendix.
The bias-corrected test statistic, $CD^{\ast}(\theta_{n}),$ depends on the
unknown parameter, $\theta_{n}$, which can be estimated by
equation[equation omitted — 94 chars of source]
where
equation[equation omitted — 274 chars of source]
The following proposition establishes the probability order of the difference
between $\hat{\theta}_{nT}$ and $\theta_{n}$.
propositionSuppose that observations on $y_{it}$, for
$i=1,2,\ldots,n,$ and $t=1,2,\ldots,T$ are generated from the pure latent
factor model given by ((ref)) and ((ref)), where the number of
factors, $m_{0}$, is known, and $\lambda_{T}=c_{\lambda}T^{-1/2}$ with
$\left\vert c_{\lambda}\right\vert <\infty$. Consider the term $\theta_{n}$ in
the $CD^{\ast}\left( \theta_{n}\right) $ statistic given by ((ref))
and its estimator $\hat{\theta}_{nT}$ given by ((ref)). Let Assumptions
(ref) to (ref) hold and $(n,T)\rightarrow\infty$, such that
$n/T\rightarrow\kappa$, where $0<\kappa<\infty$. Then
\begin{equation}
\sqrt{T}\left( \hat{\theta}_{nT}-\theta_{n}\right) =o_{p}(1).
\end{equation}
For a proof see the Appendix.
Consider now the following feasible version of $CD^{\ast}\left( \theta
_{n}\right) $,
equation[equation omitted — 141 chars of source]
and note that in view of ((ref)) and ((ref)) we have
align*[align* omitted — 366 chars of source]
Also,
\[
\frac{1-\theta_{n}}{1-\hat{\theta}_{nT}}=1+\frac{\sqrt{T}\left( \hat{\theta
}_{nT}-\theta_{n}\right) }{\sqrt{T}\left( 1-\theta_{n}\right) -\sqrt
{T}\left( \hat{\theta}_{nT}-\theta_{n}\right) }=1+o_{p}(1),
\]
and hence $CD^{\ast}\left( \hat{\theta}_{nT}\right) =CD^{\ast}(\theta
_{n})+o_{p}(1)$. We refer to $CD^{\ast}\left( \hat{\theta}_{nT}\right) $
simply as $CD^{\ast}$ and the test based on it as the CD$^{\text{*}}$ test.
The main result of the paper for pure latent factor models is summarized in
the following theorem.
theoremSuppose that observations on $y_{it}$, for $i=1,2,\ldots
,n,$ and $t=1,2,\ldots,T$ are generated from the pure latent factor model
given by ((ref)) and ((ref)), where the number of factors, $m_{0}$,
is known. Consider the statistic $CD^{\ast}$ defined by ((ref)), and
assume $(n,T)\rightarrow\infty$, such that $n/T\rightarrow\kappa,$ and
$0<\kappa<\infty$.
(a) Under the null hypothesis $H_{0}$, defined by ((ref)), and supposing
that Assumptions (ref) to (ref) hold, then
\[
CD^{\ast}\rightarrow_{d}\mathcal{N}(0,1).
\]
(b) Under local alternatives $H_{1T}$, defined by ((ref)), and supposing
that Assumptions (ref) to (ref) hold, then
\[
CD^{\ast}\rightarrow_{d}\mathcal{N}\left( \phi,1\right) ,
\]
where $\phi=\lim_{n\rightarrow\infty}\phi_{n}$, and $\phi_{n}$ is defined by
((ref)).
For a proof see the Appendix.
This theorem establishes the conditions under which the proposed CD$^{\ast}$
test has the correct size asymptotically. It also shows that the CD$^{\ast}$
test has power against network alternatives if the limit of $\phi_{n}$ defined
by ((ref)) is nonzero, namely so long as $\lim_{n\rightarrow\infty
}n^{-1}\mathbf{a}_{n}^{\prime}\mathbf{Wa}_{n}\neq0$. This condition is likely
to be satisfied if the connection matrix, $\mathbf{W}$, is not too sparse,
although it must be sufficiently sparse so that Assumption (ref) is
met. In the case where there are no latent factors, $\mathbf{a}_{n}
=(1,1,...,1)^{\prime}$, it is sufficient that $n^{-1}\sum_{i=1}^{n}
\sum_{j=1}^{n}w_{ij}\neq0$.
To our knowledge, this is the first paper to provide a formal derivation of the power
function of CD tests against spatial and network alternatives, which applies
equally to the CD test for panel data models without latent factors. Hence,
our derivation of the power function can be used to supplement earlier
research on CD tests.
As we shall see from the Monte Carlo results reported below, the CD$^{\ast}$
test performs well even if some of the latent factors happen to be weak with
$\alpha_{j}\in(0,1/2]$ or semi-strong with $\alpha_{j}\in\left( 1/2,1\right)
$. This is because when a factor is weak, it does not matter if its estimation
by PCA is not consistent at the standard rate of $\delta_{nT}=\min
(n^{1/2},T^{1/2})$, and its inclusion or exclusion from the analysis has no
material impact on the $CD^{\ast}$ statistics for $n$ and $T$ sufficiently
large. In view of this result, in the mathematical derivations it is
sufficient to consider the case of strong factors, and let the weak factors to
be absorbed in the error term.
However, it should be acknowledged that our derivations do not take account of
the case when one or more of the factors are semi-strong. Such an extension is
beyond the scope of the present paper, although recent studies by
bai2023approximate and jiang2023revisiting show that PCA
estimation is asymptotically valid for factor models so long as factor
strengths are all above $1/2$. It is therefore reasonable to conjecture that
the CD$^{\text{*}}$ test applied to PCA residuals will be asymptotically valid
even if some of the factors are semi-strong, namely if $1/2<\alpha_{j}<1$.
In practice, the true number of factors, $m_{0}$, is unknown. In cases where
the estimated number of factors, $\hat{m}$, is underestimated ($\hat{m}<m_{0}
$), the CD$^{\text{*}}$ test has power against the missing strong factors.
However, the rejection of the null hypothesis by the CD$^{\text{*}}$ test does
not necessarily mean there are missing factors, since the rejection could be
due to network error dependence. It is, therefore, important for the
investigator to decide on the number of strong latent factors before the
implementation of the proposed CD$^{\text{*}}$ test. To that end, we refer the
reader to the information criterion approach advanced by
bai2002determining and the eigenvalue ratio test of
ahn2013eigenvalue, for example.
The CD$^{\text{*}}$ test for panel regression models with interactive
effects
Consider now the factor model ((ref)) augmented with observed regressors
equation[equation omitted — 189 chars of source]
where $\mathbf{d}_{t}$ is a $k_{d}\times1$ vector of observed common factors,
$\mathbf{x}_{it}$ is a $k_{x}\times1$ vector of unit-specific observed
covariates, $\boldsymbol{\alpha}_{i}=(\alpha_{i1},\alpha_{i2},...,\alpha
_{ik_{d}})^{\prime}$, and $\boldsymbol{\beta}_{i}=(\beta_{i1},\beta
_{i2},...,\beta_{ik_{x}})^{\prime}$ are their associated unknown coefficients.
To highlight the relevance of the CD test for this set up, model ((ref))
can be written alternatively as
equation[equation omitted — 146 chars of source]
where the errors, $v_{it}$, follow the factor structure
equation[equation omitted — 89 chars of source]
The CD test is applicable, without any modifications, to test the null
hypothesis that the errors of the panel regression model, $v_{it}$, are
cross-sectionally independent, so long as the regressors, $\mathbf{d}_{t}$ and
$\mathbf{x}_{it}$, are strictly exogenous with respect to $v_{it}$. When the
regressors are correlated with the errors, the least squares estimates of
$v_{it}$ become inconsistent and the standard CD test will fail. One important
example of endogeneity arises when both $y_{it}$ and $\mathbf{x}_{it}$ are
driven by the same latent factor(s). pesaran2006estimation formalizes
this form of endogeneity by assuming that
equation[equation omitted — 168 chars of source]
where $\mathbf{A}_{i}$ and $\boldsymbol{\Gamma}_{i}$\ are $k_{d}\times k_{x}$
and $m_{0}\times k_{x}$ factor loading matrices and $\boldsymbol{\varepsilon
}_{xit}$ are distributed independently of $\mathbf{f}_{t}$. The system of
equations ((ref)), ((ref)) and ((ref)) fully specify the
dependence of $\mathbf{x}_{it}$ and $v_{it}$, and allows consistent estimation
of $v_{it}$ which can then be used to test the hypothesis that $u_{it}$ are
cross-sectionally independent in the pure latent factor model ((ref)). We
now show that the CD$^{\text{*}}$ test applied to these residuals will be
valid. To this end we make the following additional standard assumptions.
assumption(a) The $k_{d}\times1$\ vector $\mathbf{d}_{t}$ is a
covariance stationary process, with absolute summable autocovariances and
$\mathbf{d}_{t}$ is distributed independently of $\mathbf{f}_{t^{^{\prime}}},$
for all $t$ and $t^{^{\prime}}$, such that $T^{-1}\mathbf{D}^{^{\prime}
}\mathbf{F}=O_{p}\left( T^{-1/2}\right) $, where $\mathbf{D}=\left(
\mathbf{d}_{1},\mathbf{d}_{2},\ldots,\mathbf{d}_{T}\right) ^{^{\prime}}$ and
$\mathbf{F}=\left( \mathbf{f}_{1},\mathbf{f}_{2},\ldots,\mathbf{f}
_{T}\right) ^{^{\prime}}$ are matrices of observations on $\mathbf{d}_{t}$
and $\mathbf{f}_{t}$. (b) $\left( \mathbf{d}_{t},\mathbf{f}_{t}\right) $ is
distributed independently of $u_{is}$ and $\boldsymbol{\varepsilon}_{xis}$ for
all $i,t,s.$
assumptionThe unobserved factor loadings $\boldsymbol{\Gamma
}_{i}$ are bounded, i.e. $\left\Vert \boldsymbol{\Gamma}_{i}\,\right\Vert
_{2}<C$ for all $i.$
assumptionThe individual-specific errors $\varepsilon_{it}$ in ((ref)) and
$\boldsymbol{\varepsilon}_{xi,t^{\prime}}$ are distributed independently for
all $i,j,t$ and $t^{\prime},$ and $\boldsymbol{\varepsilon}_{xit}$ follows the
linear stationary process $\boldsymbol{\varepsilon}_{xit}=\sum_{l=0}^{\infty
}\mathbf{S}_{il}\boldsymbol{\eta}_{xi,t-l},$ where for each $i$,
$\boldsymbol{\eta}_{xit}$ is a $k_{x}\times1$ vector of serially uncorrelated
random variables with mean zero, the variance matrix $\mathbf{I}_{k_{x}}$, and
finite fourth-order cumulants. For each $i$, the coefficient matrices
$\mathbf{S}_{il}$ satisfy the condition
\[
Var\left( \boldsymbol{\varepsilon}_{xit}\right) =\sum_{l=0}^{\infty
}\mathbf{S}_{il}\mathbf{S}_{il}^{^{\prime}}=\boldsymbol{\Sigma}_{xi},
\]
where $\boldsymbol{\Sigma}_{xi}$ is a positive definite matrix, such that
$\sup_{i}\left\vert \left\vert \boldsymbol{\Sigma}_{xi}\right\vert \right\vert
_{2}<C$.
assumptionLet $\tilde{\boldsymbol{\Gamma}}=E\left(
\boldsymbol{\gamma}_{i},\boldsymbol{\Gamma}_{i}\right) .$ We assume that
$Rank\left( \tilde{\boldsymbol{\Gamma}}\right) =m_{0}.$
assumptionConsider\ the cross-sectional averages of the
individual-specific variables, $\mathbf{z}_{it}=\left( y_{it},\mathbf{x}
_{it}^{^{\prime}}\right) ^{^{\prime}}$ defined by $\bar{\mathbf{z}}
_{t}=n^{-1}\sum_{i=1}^{n}\mathbf{z}_{it}$, and let $\mathbf{\bar{M}
}=\mathbf{I}_{T}-\mathbf{\bar{H}}\left( \mathbf{\bar{H}}^{^{\prime}
}\mathbf{\bar{H}}\right) ^{-1}\mathbf{\bar{H}}^{^{\prime}}$, and
$\mathbf{M}_{g}=\mathbf{I}_{T}-\mathbf{G}\left( \mathbf{G}^{^{\prime}
}\mathbf{G}\right) ^{-1}\mathbf{G}^{^{\prime}}$, where $\mathbf{\bar{H}
=}\left( \mathbf{D},\mathbf{\bar{Z}}\right) ,$ $\mathbf{G}=\left(
\mathbf{D,F}\right) ,$ and $\mathbf{\bar{Z}=}\left( \mathbf{\bar{z}}
_{1},\mathbf{\bar{z}}_{2},\ldots,\mathbf{\bar{z}}_{T}\right) ^{\prime}$ is
the $T\times\left( k_{x}+1\right) $ matrix of observations on the
cross-sectional averages. Let $\mathbf{X}_{i}=\mathbf{(x}_{i1},\mathbf{x}
_{i2},...,\mathbf{x}_{iT})^{\prime}$, then the $k\times k$ matrices
$\mathbf{\hat{\Psi}}_{i,T}=T^{-1}\mathbf{X}_{i}^{^{\prime}}\mathbf{\bar{M}
X}_{i}$ and $\mathbf{\Psi}_{ig}=T^{-1}\mathbf{X}_{i}^{^{\prime}}\mathbf{M}
_{g}\mathbf{X}_{i}$ are non-singular, and $\mathbf{\hat{\Psi}}_{i,T}^{-1}$ and
$\mathbf{\Psi}_{ig}^{-1}$ have finite second-order moments for all $i.$
\begin{remark}
The above assumptions are standard in the panel data models with multi-factor
error structure. See, for example, pesaran2006estimation. But in our
setup under Assumption (ref) we require the error term,
$\varepsilon_{it}$, to be serially independent, since our focus is on testing
$\varepsilon_{it}$ for cross-sectional independence, and this assumption is
needed for asymptotic normality of the bias-corrected CD test. Later in
Section (ref), we will consider models with serially correlated
errors and show that the bias-corrected CD test remains valid. Nevertheless,
we allow $\varepsilon_{xit}$, the errors in the $\mathbf{x}_{it}$ equations to
be serially correlated. Assumption (ref) separates the observed
and the latent factors, as in Assumption 11 of pesaran2011large. This
assumption is required to obtain the probability order of estimated residuals
needed for computation of $CD^{\ast}$ statistic. A necessary condition for the
rank condition in Assumption (ref) to hold is $k_{x}\geq
m_{0}-1$.
\end{remark}
To estimate $v_{it}$ we first filter out the effects of observed covariates
using the CCE estimators proposed in pesaran2006estimation, namely for
each $i$ we estimate $\boldsymbol{\beta}_{i}$ by
equation[equation omitted — 210 chars of source]
and following pesaran2011large, estimate $\boldsymbol{\alpha}_{i}$ by
equation[equation omitted — 223 chars of source]
Then we have the following estimator of $v_{it}$
equation[equation omitted — 176 chars of source]
Using results in pesaran2011large (p. 189) it follows that under
Assumptions (ref)-(ref)
equation[equation omitted — 170 chars of source]
Note when $\boldsymbol{\alpha}_{i}=\mathbf{0}$ and $\boldsymbol{\beta}
_{i}=\mathbf{0}$, ((ref)) reduces to the pure latent factor model,
((ref)), where PCA can be applied to $v_{it}=y_{it}$ directly. In the case of
panel regressions $\hat{v}_{it}$ can be used instead of $v_{it}$ to compute
the bias-corrected CD statistic given by ((ref)). The errors involved
will become asymptotically negligible in view of the fast rate of convergence
of $\hat{v}_{it}$ to $v_{it}$, uniformly for each $i$ and $t$. Specifically,
as in the case of the pure latent factor model, we first compute $m_{0}$ PCs
of $\left\{ \hat{v}_{it};\text{ }i=1,\ldots,n;\text{ and }t=1,\ldots
,T\right\} $ and the associated factor loadings, $(\boldsymbol{\hat{\gamma}
}_{i},\mathbf{\hat{f}}_{t}),$ subject to the normalization $n^{-1}\sum
_{i=1}^{n}\boldsymbol{\hat{\gamma}}_{i}\boldsymbol{\hat{\gamma}}_{i}
^{^{\prime}}=\mathbf{I}_{m_{0}}$. The residuals
equation[equation omitted — 169 chars of source]
can then be used to compute the standard CD statistic, ((ref)), and its
bias-corrected version, $CD^{\ast}$, using ((ref)).
remarkIt is important to bear in mind that $\hat{u}_{it}$ is not the
same as the CCE residuals that result from running the panel regressions of
$y_{it}$ on $(\mathbf{d}_{t},\mathbf{x}_{it},\mathbf{\bar{z}}_{t})$. As shown
by juodis2022incidental, the standard CD test applied to the CCE
residuals will result in over-rejection and is not recommended. In our
approach, we filter out the latent factors from $\hat{v}_{it}$ and use the
filtered residuals, $\hat{u}_{it}$, to compute the CD statistic and correct
it, as in CD$^{\ast}$, to allow for errors associated with estimation of
factors and their loadings.
The following theorem extends Theorem (ref) to panel
regression models with observed regressors.
theoremSuppose that observations on $y_{it}$, for $i=1,2,\ldots
,n,$ and $t=1,2,\ldots,T$ are generated from the panel regression model
defined by ((ref)),\ ((ref)) and ((ref)), where the number of
latent factors in ((ref)), $m_{0}$, is known. Consider the statistic
$CD^{\ast}$ given\ by ((ref)) using the filtered residuals defined by
((ref)). Suppose that $(n,T)\rightarrow\infty$, such that
$n/T\rightarrow\kappa,$ and $0<\kappa<\infty$.
(a) Under the null hypothesis $H_{0}$, defined by ((ref)), and supposing
that Assumptions (ref) to (ref) and Assumptions
(ref) to (ref) hold, then
\[
CD^{\ast}\rightarrow_{d}\mathcal{N}(0,1).
\]
(b) Under local alternatives $H_{1T}$, defined by ((ref)), and supposing
that Assumptions (ref) to (ref) hold, then
\[
CD^{\ast}\rightarrow_{d}\mathcal{N}\left( \phi,1\right) ,
\]
where $\phi=\lim_{n\rightarrow\infty}\phi_{n}$ and $\phi_{n}$ defined by
((ref)).
For a proof see the Appendix.
CD$^{\text{*}}$ tests for models with serially correlated
errors
As shown by baltagi2016testing, when the errors $u_{it}$ in
((ref)) are serially correlated the variance of the standard CD test
statistic is not unity (even asymptotically) and the test is no longer valid.
The same also applies to the CD$^{\text{*}}$ test. To deal with this problem,
we propose two solutions which involve different ways of adjusting the
CD$^{\ast}$ test so that it will become applicable to panels with serially
correlated errors. The first method closely follows the variance adjustment
proposed by baltagi2016testing, in which $CD^{\ast}$ is scaled by
$\varpi$ where
equation[equation omitted — 429 chars of source]
with $\boldsymbol{\tilde{\varepsilon}}_{i,T}=(\tilde{\varepsilon}
_{i1,T},\tilde{\varepsilon}_{i2,T},\ldots,\tilde{\varepsilon}_{iT,T})^{\prime
}$, $\tilde{\varepsilon}_{it,T}$ defined in ((ref)) and
\[
\boldsymbol{\tilde{\varepsilon}}_{\left( ij\right) ,T}=\frac{1}{n-2}
\sum_{1\leq\tau\neq i,j\leq n}\boldsymbol{\tilde{\varepsilon}}_{\tau,T}.
\]
The expression in ((ref)) is the equivalent to that provided in
Theorem 3 of baltagi2016testing but the factor of $2$ in
((ref)) is missing in their paper. The same adjustment is
also applied to the CD$_{W+}$ test to allow for serially correlated errors.
Alternatively, following pesaran2004general, we first transform the
panel regression model to eliminate the error serial correlation and then
apply the CD$^{\text{*}}$ test to the residuals of the transformed model. This
is possible so long as the error serial correlation can be approximated by a
finite order stationary autoregressive process. As a simple illustration
consider the pure latent factor model $y_{it}=\gamma_{i}f_{t}+u_{it},$ in
which factor $f_{t}$ and loading $\gamma_{i}$ are both latent, and the errors
$u_{it}$ are generated as $AR(1)$ processes, $u_{it}=\rho_{i}u_{it-1}
+\epsilon_{it},$ where $\rho_{i}$ is the autoregression coefficient and
$\epsilon_{it}$ is serially independent, as well as being distributed
independently of $f_{t^{\prime}}$ for all $i$ and $t,t^{\prime}=1,2,\ldots,T$.
Testing the cross-sectional independence of $u_{it}$ is equivalent to testing
the cross-sectional independence of $\epsilon_{it}$ in the following
autoregressive distributed lag (ARDL) representation of $y_{it}$
\[
y_{it}=\rho_{i}y_{i,t-1}+\gamma_{i}f_{t}-\rho_{i}\gamma_{i}f_{t-1}
+\epsilon_{it},
\]
which can be written equivalently as a multi-factor AR panel regression model
equation[equation omitted — 145 chars of source]
where $\mathbf{\mathring{f}}_{t}=(f_{t},f_{t-1})^{\prime}$, and
$\boldsymbol{\mathring{\gamma}}_{i}=\left( \gamma_{i},-\rho_{i}\gamma
_{i}\right) ^{\prime}$. Since $y_{i,t-1}$ is weakly exogenous, the
transformed model satisfies the setup of panel regression model ((ref))
with $\mathbf{\mathring{f}}_{t}$ viewed as a vector of latent variables with
the associated factor loadings, $\boldsymbol{\mathring{\gamma}}_{i}$. It
therefore follows that the CD$^{\text{*}}$ test can now be applied to test the
cross-sectional independence of $\epsilon_{it}$ in ((ref)). We
refer to this test as the ARDL adjusted CD$^{\text{*}}$ test.
The same approach can also be used for panels with observed covariates. In
general, testing cross-sectional independence of $u_{it}$ in model
((ref)) is equivalent to testing the cross-sectional independence of
$\epsilon_{it}$ in
equation[equation omitted — 256 chars of source]
where $\mathbf{h}_{t}$\ is an extended set of latent factors (that encompass
$\mathbf{f}_{t}$), and $\mathbf{g}_{i}$ are the associated factor loadings.
The number of lags $S$ is determined by the order of the AR specification
assumed for $u_{it}$ in ((ref)).
The variance adjustment is simpler to implement but it requires theoretical
justification in the context of panel data models with latent factors. The
ARDL adjustment is theoretically justified so long as the underlying errors
follow finite order AR processes. As we shall see both approaches work well in
dealing with serially correlated errors, at least in the context of the
limited MC designs that we are considering. Clearly, further theoretical and
Monte Carlo investigations are needed for a better understanding of the
relative merits of the two approaches.
Small sample properties of CD$^{\ast}$ and CD$_{W^{+}}$
tests
Data generating process
We consider the following data generating process
equation[equation omitted — 243 chars of source]
where $\varepsilon_{it}\left( \lambda\right) $ follows the first order
spatial autoregressive process, $SAR\left( 1\right) $, such that
equation[equation omitted — 165 chars of source]
$\text{a}_{i}$ is a unit-specific effect$,$ $d_{t}$ is the observed common
factor, $x_{it}$ is the observed regressor that varies across $i$ and $t$,
$\mathbf{f}_{t}$ is the $m_{0}\times1$ vector of unobserved factors,
$\boldsymbol{\gamma}_{i}$ is the vector of associated factor loadings. The
scalar constants, $\sigma_{i}>0$, are generated as $\sigma_{i}^{2}
=0.5+\frac{1}{2}\left( s_{i}^{2}-1\right) $, with $s_{i}^{2}\sim IID\chi
^{2}(2)$, which ensures that $E(\sigma_{i}^{2})=1$.
DGP under the null hypothesis
Under the null hypothesis, we set $\lambda=0$ and $c=1$, and consider both
serially independent errors and serially correlated errors, which are
generated by both Gaussian and non-Gaussian distributions:
itemize• Serially independent errors: Gaussian errors, $\varepsilon_{it}\sim
IID\mathcal{N}(0,1)$; chi-squared distributed errors, $\varepsilon_{it}\sim
IID\left( \frac{\chi^{2}(2)-2}{2}\right) $.
• Serially correlated errors: $\varepsilon_{it}=\rho_{\varepsilon
}\varepsilon_{it-1}+\sqrt{1-\rho_{\varepsilon}^{2}}$ $e_{\varepsilon it}$, for
$i=1,2,\ldots,n$ and $t=1,2,\ldots,T$, where $\rho_{\varepsilon}=0.5$ and
$e_{\varepsilon it}$ are generated as Gaussian errors, $e_{\varepsilon it}\sim
IID\mathcal{N}\left( 0,1\right) $, or chi-squared distributed errors,
$e_{\varepsilon it}\thicksim IID\left( \frac{\chi^{2}(2)-2}{2}\right) $.
The focus of the experiments is on testing the null hypothesis that
$\varepsilon_{it}$ are cross-sectional independent, whilst allowing for the
presence of $m_{0}$ unobserved factors, $\mathbf{f}_{t}=(f_{1t},f_{2t}
,...,f_{m_{0}t})^{\prime}$. We consider $m_{0}=1$ and $m_{0}=2$, and generate
the factor loadings $\boldsymbol{\gamma}_{i}=(\gamma_{i1},\gamma_{i2}
)^{\prime}$ as:
align*[align* omitted — 372 chars of source]
In the one-factor case ($m_{0}=1$), we only include $f_{1t}$ as the latent
factor and denote its factor strength by $\alpha$. Three values of $\alpha$
are considered, namely $\alpha=1,2/3,1/2$, respectively representing strong,
semi-strong and weak factors. Similarly, in the two-factor case ($m_{0}=2$),
we include both $f_{1t}$ and $f_{2t}$ as the latent factors and consider the
following combinations of factor strengths: $(\alpha_{1},\alpha_{2})=\left[
(1,1),(1,2/3),(2/3,1/2)\right] $. The intercepts a$_{i}$ are generated as
$IID\mathcal{N}(1,2)$ and fixed thereafter. The observed common factor is
generated as $d_{t}=\rho_{d}d_{t-1}+\sqrt{1-\rho_{d}^{2}}$ $v_{dt}$, with
$\rho_{d}=0.8,$ and $v_{dt}\thicksim IID\mathcal{N}(0,1)$, thus ensuring that
$E(d_{t})=0$ and $Var(d_{t})=1$. The observed unit-specific regressors,
$x_{it}$, for $i=1,2,\ldots,n$ are generated to have non-zero correlations with
the unobserved factors:
equation[equation omitted — 82 chars of source]
where $f_{jt}=r_{j}f_{j,t-1}+\sqrt{1-r_{j}^{2}}$ $v_{jt}$, with $r_{j}=0.9$
and $v_{jt}\sim IID\left( \frac{\chi^{2}(2)-2}{2}\right) $, for $j=1,2$. The
factor loadings in ((ref)) are generated as $\gamma_{xi1}\sim IIDU\left(
0.25,0.75\right) $ and $\gamma_{xi2}\sim IIDU\left( 0.1,0.5\right) $. The
error term of ((ref)) is generated as $e_{xit}=\rho_{i}e_{xi,t-1}
+\sqrt{1-\rho_{i}^{2}}$ $v_{xit}\text{, }$where $\rho_{i}\sim IIDU(0,0.95)$
and $v_{xit}\thicksim IID\mathcal{N}(0,1)$.
We will examine the small sample properties of the CD and the bias-corrected
CD tests for both the pure latent factor model and for the panel regression
model which also includes observed covariates.
itemize• In the case of the pure latent factor model we set $\beta_{i1}
=\beta_{i2}=0$.
• In the case of the panel regression model with latent factors, we allow
for heterogeneous slopes and generate the slopes of observed covariates,
$d_{t}$ and $x_{it}$, as $\beta_{i1}\sim IID\mathcal{N}(\mu_{\beta1}
,\sigma_{\beta1}^{2}),$ and $\beta_{i2}\sim IID\mathcal{N}(\mu_{\beta2}
,\sigma_{\beta2}^{2})$ where $\mu_{\beta1}=\mu_{\beta2}=0.5$ and
$\sigma_{\beta1}^{2}=\sigma_{\beta2}^{2}=0.25,$ respectively.
As our theoretical results show the null distributions of the CD and the
bias-corrected CD tests do not depend on a$_{i}$, $\beta_{i1}$ and $\beta
_{i2}$, it is therefore innocuous what values are chosen for these parameters.
Moreover, the average fit of the panel is controlled in terms of the limiting
value of the pooled R-squared defined by
equation[equation omitted — 207 chars of source]
Since the underlying processes, ((ref)) and ((ref)), are stationary
and $E\left( \varepsilon_{it}^{2}\right) =1$, we have
\[
\lim_{T\rightarrow\infty}PR_{nT}^{2}=PR_{n}^{2}=\frac{n^{-1}\sum_{i=1}
^{n}\sigma_{i}^{2}\left[ \beta_{i1}^{2}+\beta_{i2}^{2}Var\left(
x_{it}\right) +m_{0}^{-1}\boldsymbol{\gamma}_{i}^{\prime}\boldsymbol{\gamma
}_{i}+2Cov\left( x_{it},\boldsymbol{\gamma}_{i}^{^{\prime}}\mathbf{f}
_{t}\right) \right] }{n^{-1}\sum_{i=1}^{n}Var\left( y_{it}\right) },
\]
where $\boldsymbol{\gamma}_{i}=\left( \gamma_{i1},\gamma_{i2}\right)
^{\prime},$ $Var\left( x_{it}\right) =\boldsymbol{\gamma}_{xi}^{^{\prime}
}\boldsymbol{\gamma}_{xi}+1,$ $Cov\left( x_{it},\boldsymbol{\gamma}
_{i}^{^{\prime}}\mathbf{f}_{t}\right) =\boldsymbol{\gamma}_{xi}^{^{\prime}
}\boldsymbol{\gamma}_{i},$ $\boldsymbol{\gamma}_{xi}=\left( \gamma
_{xi1},\gamma_{xi2}\right) ^{\prime}$, and
\[
Var\left( y_{it}\right) =\sigma_{i}^{2}\left[ \beta_{i1}^{2}+\beta_{i2}
^{2}Var\left( x_{it}\right) +m_{0}^{-1}\boldsymbol{\gamma}_{i}^{\prime
}\boldsymbol{\gamma}_{i}+2m_{0}^{-1/2}Cov\left( x_{it},\boldsymbol{\gamma
}_{i}^{^{\prime}}\mathbf{f}_{t}\right) +1\right] .
\]
Also since $\sigma_{i}^{2}$ and $\beta_{ij}$ are independently distributed and
$E(\sigma_{i}^{2})=1$, it then readily follows that $\lim_{n\rightarrow\infty
}PR_{n}^{2}=\eta^{2}/(1+\eta^{2})$, where
\[
\eta^{2}=\mu_{\beta1}^{2}+\sigma_{\beta1}^{2}+\left( \mu_{\beta2}^{2}
+\sigma_{\beta2}^{2}\right) \left[ 1+E\left( \boldsymbol{\gamma}
_{xi}^{^{\prime}}\boldsymbol{\gamma}_{xi}\right) \right] +\frac{2\mu
_{\beta2}E\left( \boldsymbol{\gamma}_{xi}^{^{\prime}}\boldsymbol{\gamma}
_{i}\right) }{\sqrt{m_{0}}}+\frac{E\left( \boldsymbol{\gamma}_{i}^{^{\prime
}}\boldsymbol{\gamma}_{i}\right) }{m_{0}}.
\]
By controlling the value of $\eta^{2}$ across the experiments we ensure that
the pooled R$^{2}$ in large samples is the same for all values of $\sigma
_{i}^{2}$. In particular, in the case of the pure latent model we have
$\eta^{2}=m_{0}^{-1}E\left( \boldsymbol{\gamma}_{i}^{^{\prime}}
\boldsymbol{\gamma}_{i}\right) =O\left( n^{\alpha-1}\right) ,$ where
$\alpha=max(\alpha_{1},\alpha_{2})$.
DGP under alternative hypotheses
Under alternative hypotheses, using ((ref)), we consider a spatial
alternative defined by
equation[equation omitted — 201 chars of source]
where $\boldsymbol{\varepsilon}_{\circ t}\left( \lambda\right) =\left(
\varepsilon_{1t}\left( \lambda\right) ,\varepsilon_{2t}\left(
\lambda\right) ,\ldots,\varepsilon_{nt}\left( \lambda\right) \right)
^{\prime},$ $\mathbf{W}=(w_{ij})$, and $\boldsymbol{\varepsilon}_{\circ
t}=(\varepsilon_{1t},\varepsilon_{2t},\ldots,\varepsilon_{nt})^{\prime}$. The
errors $\varepsilon_{it}$ are generated as described above. For the spatial
weights $w_{ij}$, we first set $w_{ij}^{0}=1$ if $j=i-2,i-1,i+1,i+2,$ and zero
otherwise. We then row normalize the weights such that $w_{ij}=\left(
\sum_{j=1}^{n}w_{ij}^{0}\right) ^{-1}w_{ij}^{0}$. We also set $c\left(
\lambda\right) ^{2}=n/$$tr$$\left[ \left( \mathbf{I}_{n}
-\lambda\mathbf{W}\right) ^{-1}\left( \mathbf{I}_{n}-\lambda\mathbf{W}
\right) ^{\prime-1}\right] $, which ensures that $n^{-1}\sum_{i=1}
^{n}Var(\varepsilon_{it}\left( \lambda\right) )=1$, for all values of
$\lambda$. In practice, only positive values of $\lambda$ are of interest, and
the power function need not be symmetric for all positive and negative values
of $\lambda$.
CD, CD$^{\ast}$ and CD$_{W^{+}}$ tests
All experiments are carried out for $n=100,200,500,1000$ and $T=100,200,500$,
and the number of replications is set to $2000$. Firstly we consider the DGPs
with serially independent errors. For the pure latent factor models, we
compute the filtered residuals as $\hat{v}_{it}=y_{it}-\hat{\text{a}}_{i}$,
where $\hat{\text{a}}_{i}=T^{-1}\sum_{t=1}^{T}y_{it}$. For the panel
regressions with latent factors, the filtered residuals are computed as
equation[equation omitted — 126 chars of source]
where $\left( \hat{\text{a}}_{CCE,i},\hat{\beta}_{CCE,i1},\hat{\beta
}_{CCE,i2}\right) $ is the CCE estimator of a$_{i}$, $\beta_{i1}$ and
$\beta_{i2}$, as set out in pesaran2006estimation. The residuals
$\{\hat{v}_{it};$ $i=1,2,\ldots,n;$ and $t=1,2,\ldots,T\}$, together with
their first $m$ PCs and the associated factor loadings, $(\boldsymbol{\hat
{\gamma}}_{i},\hat{\mathbf{f}}_{t}),$ are then used to compute the filtered
residuals, $\hat{u}_{it}=\hat{v}_{it}-\boldsymbol{\hat{\gamma}}_{i}^{\prime
}\hat{\mathbf{f}}_{t}$, to compute the CD test statistics, $CD$ and $CD^{\ast
}$, given by ((ref)) and ((ref)), respectively. For comparison, we
also consider the power enhanced version of the randomized CD test statistic
proposed by JR given by
equation[equation omitted — 56 chars of source]
where
equation[equation omitted — 228 chars of source]
The weights $w_{i}$, for $i=1,2,...,n$ are independently drawn from a
Rademacher distribution and
equation[equation omitted — 209 chars of source]
where $\hat{\rho}_{ij,T}=T^{-1}\sum_{t=1}^{T}\tilde{\varepsilon}_{it,T}
\tilde{\varepsilon}_{jt,T}$, and $\tilde{\varepsilon}_{it,T}$ is defined by
((ref)). As shown by JR, $CD_{W}$ has a zero mean by construction and avoids the over-rejection problem of the CD test, but it can also lack power by
the very nature of the randomization process. JR further suggest $CD_{W+}$ by
adding a screening component $\Delta_{nT}$ proposed by fan2015power,
which enhances the power of the test since $\Delta_{nT}$ converges to zero as
$n$ and $T\rightarrow\infty$ under the null hypotheses, but can diverge under
alternatives with a sufficient number of $(i,j)$ pairs with non-zero
correlations, $\rho_{ij}$.
As discussed in Section (ref), the CD$^{\text{*}}$ test is not
valid when the errors are serially correlated. In the simulations, we apply
the variance and ARDL adjustments to $CD$, $CD_{W+}$, and $CD^{\ast}$. The
variance adjusted versions are computed by scaling the original statistics by
the standard deviation of the CD statistics using the expression in
((ref)) with $\tilde{\varepsilon}_{it,T}=\hat{u}_{it}/\hat{\sigma}_{i,T}$, where $\hat{u}_{it}=\hat{v}_{it}-\boldsymbol{\hat{\gamma}
}_{i}^{\prime}\hat{\mathbf{f}}_{t}$. The ARDL adjusted versions of $CD$,
$CD^{\ast}$, and $CD_{W+}$, are computed using the residuals from the
following dynamic panel data model with latent factors,
equation[equation omitted — 204 chars of source]
In the simulations we set $S=1$, but higher order values can also be
considered. The number of latent factors in $\mathbf{h}_{t}$ depends on $S$
and is given by $m_{h}=(S+1)m_{0}$. Accordingly, the number of selected PCs,
$\hat{m}$, should satisfy $\hat{m}\geq(S+1)m_{0}$. In the simulations if
$S=0$, we consider $\hat{m}=1$ and $2$ if $m_{0}=1$, and $\hat{m}=2$ and $4$
if $m_{0}=2$. But if $S=1$ we consider $\hat{m}=2$ and $4$ if $m_{0}=1$, and
$\hat{m}=4$ and $6$ if $m_{0}=2$. Seen from this perspective, the variance
adjustment approach to dealing with error serial correlation seems preferable
since it does not require specifying the lag order $S$.
Simulation results
We first report the simulation results for the DGPs with normally distributed
errors, followed by the results based on DGPs with chi-squared distributed
errors. Next, we report simulation results for the DGPs with serially
correlated errors, using the variance and ARDL adjusted CD tests discussed in
Section (ref). Finally, to investigate the power of the
CD$^{\text{*}}$ test we consider the spatial $SAR(1)$ alternative with
$\lambda=0.25$. As to be expected the power rises very quickly as $\lambda$
deviates from $0$.\footnote{Simulated power functions are provided in the
supplement for $\lambda=\pm0.05$, $\pm0.1$, $\pm0.2$, $\pm0.3$,
$\pm0.4$, $\pm0.5$, $\pm0.6$, $\pm0.7$, $\pm0.8$, $\pm0.9$, $\pm0.95$.}
Serially independent errors: normally distributed errors
The simulation results for the DGPs with the errors following Gaussian
distribution are shown in Tables (ref) to (ref).
Tables (ref) and (ref) report the test results for the
latent factor model with one factor. Table (ref) gives the results
for the case where the number of selected PCs, denoted by $\hat{m}$, is the
same as the true number of factors ($m_{0}=1$), while Table (ref)
reports the results when $\hat{m}=2$. As to be expected the standard CD test
over-rejects when the factor is strong, namely when $\alpha=1$. By comparison,
the rejection frequencies of both CD$^{\ast}$ and CD$_{W+}$ tests under null
($\lambda=0)$ are generally around the nominal size of $5$ per cent. Under the
alternative (when $\lambda=0.25$), the CD$^{\ast}$ test has satisfactory power
properties with significantly high rejection frequencies even when the sample
size is small. But the CD$_{W+}$ test performs quite poorly under the spatial
alternative, especially when $T$ is small.
Tables (ref) and (ref) summarize the size and power
results for the latent factor model with $m_{0}=2,$ and reports the results
when $\hat{m}$ (the selected number of PCs) is set to $2$ (Table
(ref)) and $4$ (Table (ref)). The results are
qualitatively similar to the ones reported for the one factor model. The CD
test over-rejects if at least one of the factors is strong, and the empirical
sizes of CD$^{\ast}$ and CD$_{W+}$ tests are close to their nominal value of
$5$ per cent, although we now observe some mild over-rejection when $n=100$
and the selected number of PCs is 4. In terms of power, the CD$^{\ast}$ test
performs well, although there is some loss of power as the numbers of factors
and selected PCs rise. Similarly, the power of the CD$_{W^{+}}$ test is now
even lower and quite close to $5$ per cent when $T<500$ even if the number of
PCs is set to $m_{0}=2$.
Turning to panel regression models with latent factors estimated by CCE, the
associated simulation results are summarized in Tables (ref) to
(ref). As can be seen, the results are very close to the ones
reported in Tables (ref) to (ref) for the latent
factor model, and are in line with the asymptotic result in ((ref))
that underlies the use of CCE approach to filter out the effects of observed
covariates, as well as latent factors.
table[table omitted — 3,523 chars of source]
table[table omitted — 2,957 chars of source]
table[table omitted — 3,244 chars of source]
table[table omitted — 3,057 chars of source]
table[table omitted — 3,202 chars of source]
table[table omitted — 2,952 chars of source]
table[table omitted — 3,310 chars of source]
table[table omitted — 3,069 chars of source]
Serially independent errors: chi-squared distributed errors
To save space, the simulation results for the DGPs with chi-squared errors are
provided in Tables (ref) to (ref) in the
supplement. For the standard CD test and its biased-corrected
version, CD$^{\ast}$, as shown in Tables (ref) and
(ref), the results are very similar to the ones with Gaussian
errors, suggesting that the CD$^{\text{*}}$ test is likely to be robust to
departures from Gaussianity. As with the experiments with Gaussian errors, the
standard CD test continues to over-reject unless $\alpha<2/3$, and the
CD$^{\text{*}}$ test has the correct size for all $n$ and $T$ combinations,
except when the number of selected PCs is large relative to $m_{0}$, and
$T=100$. The main difference between the results with and without Gaussian
errors is the tendency for the CD$_{W^{+}}$ test to over-reject when $n>T$,
which seems to be a universal feature of this test and holds for all
choices of $m_{0}$ and the number of selected PCs, irrespective of whether the
factors are strong or weak. This could be due to the screening component of
the CD$_{W+}$ test not tending to zero sufficiently fast with $n$ and $T$.
Furthermore, the CD$^{\text{*}}$ test continues to have satisfactory power,
but the CD$_{W^{+}}$ test clearly lacks power against spatial or network
alternatives that are of primary interest.
Similar results are obtained for panel regression models with latent factors,
summarized in Tables (ref) and (ref) in the
supplement.
Serially correlated errors
To save space, the results for the DGPs with serially correlated errors are
summarized in Tables (ref) to (ref) in the
supplement. Tables (ref) to (ref) give
the simulation results for the variance adjusted CD tests, whilst Tables
(ref) to (ref) provide the results for the
ARDL adjusted tests. Overall, the results corroborate our earlier findings
obtained for DGPs with serially independent errors. Both adjustments for
serial error correlation work well, with size and power of the adjusted
CD$^{\text{*}}$ tests being quite close to the results already reported for
DGPs with serially independent errors. It is also clear that without
adjustments for latent factors and error serial correlation, the standard CD
test will lead to large size distortions when the latent factors are strong.
But in line with our theoretical results, the standard CD test, when adjusted
for error serial correlation if needed, tends to have the correct size when
the latent factors are weak.
Comparing the two types of adjustments for error serial correlations (for pure
latent factor models as well as for panel regression models with latent
factors), the variance adjusted CD$^{\text{*}}$ test works particularly well,
and only shows mild over-rejection in the case where $T=100$ and $n>T$. In
contrast, the CD$_{W+}$ test with variance adjustment over-rejects for all
combinations of $n$ and $T$.
The ARDL adjusted version of the CD$^{\text{*}}$ test also works well when the
number of PCs is not too large, and tends to have the correct size for all
$(n,T)$ combinations and only shows slight over-rejection when $n=100$. The
CD$_{W+}$ test using ARDL adjustment does better in controlling for the size
when the errors are Gaussian, but tends to over-reject when the errors are
chi-squared distributed and $n>T$. Both adjusted versions of the CD$_{W+}$
test continue to lack power against spatial or network alternatives.
Empirical application
It is well known that house price changes are spatially correlated, but it is
unclear if such correlations are mainly due to common factors (national or
regional) or arise from spatial spillover effects not related to the common
factors, a phenomenon also referred to as the ripple effect. See, for example,
holly2011spatial, tsai2015spillover, chiang2016ripple,
bailey2016two, and aquaro2021estimation. To test for the
presence of ripple effects the influence of common factors must first be
filtered out and this is often a challenging exercise due to the latent nature
of regional and national factors. Therefore, to find if there exist local
spillover effects, one needs to test for significant residual cross-sectional
dependence once the effects of common factors are filtered out.
We consider quarterly data on real house prices at the level of Metropolitan
Statistical Areas (MSAs) in the U.S. There are 381 MSAs, under the February
2013 definition provided by the U.S. Office of Management and Budget (OMB). We
use quarterly data on real house price changes compiled by
yang2021common which covers $n=377$ MSAs from the contiguous United
States over the period 1975Q1-2014Q4 ($T=160$ quarters). To allow for possible
regional factors, we also follow bailey2016two and start with the
Bureau of Economic Analysis eight regional classification, namely New England,
Mideast, Great Lakes, Plains, Southeast, Southwest, Rocky Mountain and Far
West. But due to the low number of MSAs in New England and Rocky Mountain
regions, we combine New England and Mideast, and Southwest and Rocky Mountain
as two regions. We end up with a six region classification ($R=6$), each
covering a reasonable number of MSAs.
We model house price changes and consider an extended factor model with
deterministic seasonal dummies to allow for seasonal movements in house
prices. bailey2016two find evidence of regional factors in U.S. house
price changes which might not be picked up when using PCA. Given this finding,
our model includes observed regional and national factors, as well as latent
factors. Specifically, we suppose
equation[equation omitted — 224 chars of source]
where $\pi_{irt}$ is the real house price change in MSA $i$ located in region
$r=1,2,\ldots,R$, l$\left\{ q_{t}=j\right\} $ is the index for quarter $j$, and
$\mathbf{f}_{t}$ is the $m_{0}\times1$ vector of latent factors. $\bar{\pi
}_{rt}=n_{r}^{-1}\sum_{i=1}^{n_{r}}\pi_{irt}$, where $n=\sum_{r=1}^Rn_r$, and $n_r$ is the number of MSAs in region $r$, and $\bar{\pi}_{t}=n^{-1}
\sum_{r=1}^{R}\sum_{i=1}^{n_{r}}\pi_{irt}$ are proxies for the regional and
national factors. To filter out the effects of seasonal dummies as well as
observed factors, we first run the least squares regression of $\pi_{irt}$ on
an intercept and $\left( \text{l}\left\{ q_{t}=j\right\} ,\bar{\pi}
_{rt},\bar{\pi}_{t}\right) $ for each $i$ to generate the residuals
equation[equation omitted — 201 chars of source]
and then apply PCA to $\left\{ \hat{v}_{irt}:i=1,2,\ldots,n_r,r=1,2,\ldots
,R,t=1,2,\ldots,T\right\} $ to obtain $\boldsymbol{\hat{\gamma}}_{ir}$ and
$\mathbf{\hat{f}}_{t}$, yielding the residuals
equation[equation omitted — 265 chars of source]
For the case without adjusting for error serial correlation, the above
residuals are used to compute $CD$, $CD^{\ast}$ and $CD_{W+}$, given by
((ref)), ((ref)) and ((ref)). For the case with serially
correlated errors, the variance adjusted versions of the three CD statistics
are generated by scaling original test statistics using ((ref)),
where $\tilde{\varepsilon}_{it,T}$ is replaced by the standardized residuals
generated from ((ref)), while the ARDL adjusted versions are computed
using the residuals from the following dynamic panel data model with latent
factors, $\mathbf{h}_{t}$:
align[align omitted — 383 chars of source]
To estimate the number of latent factors, $m$, we consider the information
criteria $IC_{P1}$ and $IC_{P2}$ proposed by bai2002determining, and
the $ER$ and $GR$ criteria proposed by ahn2013eigenvalue. Given the
spatial diversity of U.S. housing market, we set $m_{\max}=$ $10$, although
once we allow for national and regional factors we would expect $m_{0}$ and
its estimate, $\hat{m}$, to be relatively small. The estimated number
of factors and the associated CD test statistics are summarized in Table
(ref). The first four columns of the table report $\hat{m}$, $CD$,
$CD^{\ast}$ and $CD_{W+}$ statistics that are not adjusted for error serial
correlation, whilst the middle and the final four columns report the variance adjusted and ARDL adjusted versions of these statistics, respectively.
As can be seen in the case of no error serial correlations, there are large
differences in the number of factors selected by the different criteria, with
$IC_{p1}$ selecting the assumed maximum number of factors, $IC_{p2}$ selecting
$4$, and $ER$ and $GR$ both selecting $2$ factors. These estimates are not
affected when we allow for error serial correlations and consider variance
adjustment.
CD, CD$^{\ast}$ and CD$_{W+}$ tests all reject the null hypothesis of
cross-sectional independence, irrespective of the choices of $\hat{m}$ and
whether we allow for error serial correlation. In view of the theoretical and
finite sample results reported in this paper, it is advisable to focus on the
CD$^{\text{*}}$ test results and recognize that the relatively large
magnitudes obtained for CD and CD$_{W+}$ test statistics could be due to their
tendencies to over-rejection in the presence of strong latent factors and
non-Gaussian errors. Focusing on $CD^{\ast}$, we find that even with $\hat
{m}=10$ the CD$^{\text{*}}$ test strongly rejects the null of cross-sectional
independence with $CD^{\ast}$ statistic of $25.5$, compared to 95 per cent
critical value of $1.96$. There is clear evidence that in addition to latent
factors, spatial modeling of the type carried out in bailey2016two and
aquaro2021estimation is likely to be necessary to account for the
remaining error cross-sectional dependence.
table[table omitted — 2,070 chars of source]
Concluding remarks
This paper revisits the problem of testing error cross-sectional independence
in panel data models with latent factors. Starting with a pure latent
multi-factor model we show that the standard CD test proposed by
pesaran2004general remains valid if the latent factors are weak, but
over-reject when one or more of the latent factors are strong. The
over-rejection of the CD test in the case of strong factors is also
established by juodis2022incidental, who propose a randomized test
statistic to correct for over-rejection and add a screening component to
achieve power. However, as we show, JR's CD$_{W^{+}}$ test is not guaranteed
to have the correct size and need not be powerful against spatial or network
alternatives. Such alternatives are of particular interest in the analyses of
ripple effects in housing markets, and clustering of firms within industries
in capital or arbitrage asset pricing models. In fact, using Monte Carlo
experiments we show that under non-Gaussian errors the JR test continues to
over-reject when the cross section dimension ($n$) is larger than the time
dimension ($T$), and often has power close to size against spatial
alternatives. To overcome some of these shortcomings, we propose a simple
bias-corrected CD test statistic, labeled $CD^{\ast}$, which is shown to be
asymptotically $\mathcal{N}(0,1)$ under the null when $n$ and $T\rightarrow
\infty$ such that $n/T\rightarrow\kappa$, for a fixed constant $\kappa$. In
addition, the CD$^{\text{*}}$ test is shown to have power against network type
dependence. These results hold for pure latent factor models as well as for
panel regression models with latent factors. To deal with possible error
serial dependence, following baltagi2016testing, we also consider a
variance adjusted version of $CD^{\ast}$, as well as an alternative ARDL
adjusted version that eliminates the error serial dependence before the
application of the CD$^{\text{*}}$ test procedure. Both of these approaches
are shown to perform well within the Monte Carlo set up of the paper.
\setcounter{equation}{0}
\setcounter{section}{0}
center[center omitted — 40 chars of source]
In this appendix we provide proofs of the propositions and and theorems.
The auxiliary lemmas and the associated proofs are given in the supplement.
Proof of Proposition (ref)
Here we provide a proof for part (b) of Proposition (ref).
The proof for part (a) follows trivially by setting $\lambda_{T}=0$. To this
end we first note that the $CD$ statistic given by ((ref)) can be written
as (for a proof see Lemma (ref) of the supplement)
\[
CD=\left( \sqrt{\frac{n}{n-1}}\right) \frac{1}{\sqrt{2T}}\sum_{t=1}
^{T}\left[ \left( \frac{1}{\sqrt{n}}\sum_{i=1}^{n}\frac{\hat{u}_{it}}
{\hat{\sigma}_{i,T}}\right) ^{2}-1\right] ,
\]
where $\hat{u}_{it}$ is defined by ((ref)). Using Lemma (ref) of
the supplement we also note that
equation[equation omitted — 57 chars of source]
where
equation[equation omitted — 242 chars of source]
with $\omega_{i,T}=\left( T^{-1}\sigma_{i}^{2}\boldsymbol{\varepsilon
}_{i\circ}^{\prime}\mathbf{M}_{F}\boldsymbol{\varepsilon}_{i\circ}\right)
^{1/2}$, $\boldsymbol{\varepsilon}_{i\circ}=\left( \varepsilon_{i1}
,\varepsilon_{i2},\ldots,\varepsilon_{iT}\right) ^{\prime}$ and
$\mathbf{M}_{F}=\mathbf{I}_{T}-\mathbf{F}\left( \mathbf{F}^{\prime}
\mathbf{F}\right) ^{-1}\mathbf{F}^{\prime}$. Also letting
equation[equation omitted — 163 chars of source]
where $\lambda_{T}=c_{\lambda}T^{-1/2}$ with $c_{\lambda}\neq0$,
$\mathbf{w}_{i0}=\left( w_{i1},w_{i2},\ldots,w_{in}\right) ^{\prime}$ and
$\boldsymbol{\varepsilon}_{\circ t}=\left( \varepsilon_{1t},\varepsilon
_{2t},\ldots,\varepsilon_{nt}\right) ^{\prime}$, under ((ref)) $\hat
{u}_{it}$ can now be expressed as
equation[equation omitted — 410 chars of source]
Let $\boldsymbol{\delta}_{i,T}=\boldsymbol{\gamma}_{i}/\omega_{i,T}$, and
$\hat{\boldsymbol{\delta}}_{i,T}=\hat{\boldsymbol{\gamma}}_{i}/\omega_{i,T}$.
Then
equation[equation omitted — 445 chars of source]
Also, subject to the normalization $n^{-1}\sum_{j=1}^{n}\boldsymbol{\hat
{\gamma}}_{j}\boldsymbol{\hat{\gamma}}_{j}^{\prime}=\mathbf{I}_{m_{0}}$ and
$n^{-1}\sum_{j=1}^{n}\boldsymbol{\gamma}_{j}\boldsymbol{\gamma}_{j}^{\prime
}=\mathbf{I}_{m_{0}}$ we have
\[
\mathbf{\hat{f}}_{t}=n^{-1}\sum_{j=1}^{n}\boldsymbol{\hat{\gamma}}_{j}
y_{jt}=\left( n^{-1}\sum_{j=1}^{n}\boldsymbol{\hat{\gamma}}_{j}
\boldsymbol{\gamma}_{j}^{\prime}\right) \mathbf{f}_{t}+n^{-1}\sum_{j=1}
^{n}\boldsymbol{\hat{\gamma}}_{j}\sigma_{j}\varepsilon_{jt}\left( \lambda
_{T}\right) ,
\]
and hence
\[
\mathbf{\hat{f}}_{t}-\mathbf{f}_{t}=\left[ n^{-1}\sum_{j=1}^{n}\left(
\boldsymbol{\hat{\gamma}}_{j}-\boldsymbol{\gamma}_{j}\right)
\boldsymbol{\gamma}_{j}^{\prime}\right] \mathbf{f}_{t}+n^{-1}\sum_{j=1}
^{n}\left( \boldsymbol{\hat{\gamma}}_{j}-\boldsymbol{\gamma}_{j}\right)
\sigma_{j}\varepsilon_{jt}\left( \lambda_{T}\right) +n^{-1}\sum_{j=1}
^{n}\boldsymbol{\gamma}_{j}\sigma_{j}\varepsilon_{jt}\left( \lambda
_{T}\right) .
\]
Using this result in ((ref)) we obtain
align[align omitted — 919 chars of source]
and summing over $i$ yields
align*[align* omitted — 1,023 chars of source]
where $\boldsymbol{\varphi}_{nT}=n^{-1}\sum_{i=1}^{n}\boldsymbol{\delta}
_{i,T}$. Written more compactly
equation[equation omitted — 191 chars of source]
where
align[align omitted — 818 chars of source]
Further, let
equation[equation omitted — 242 chars of source]
where $\boldsymbol{\varphi}_{n}=n^{-1}\sum_{i=1}^{n}\boldsymbol{\delta}_{i},$
and $\boldsymbol{\delta}_{i}=\boldsymbol{\gamma}_{i}/\sigma_{i}$. Then
$\psi_{t,nT}\left( \lambda_{T}\right) ,$ given by ((ref)), can be
written as
align*[align* omitted — 1,266 chars of source]
where
equation[equation omitted — 263 chars of source]
Writing $\psi_{t,nT}\left( \lambda_{T}\right) $ more compactly we have
equation[equation omitted — 284 chars of source]
where
align[align omitted — 568 chars of source]
Using ((ref)) in ((ref)) and after some algebra we have (where
we have made the dependence of $\widetilde{CD}$ on $\lambda_{T}$ explicit)
\[
\widetilde{CD}(\lambda_{T})=\left( \sqrt{\frac{n}{n-1}}\right) \frac
{1}{\sqrt{T}}\sum_{t=1}^{T}\left( \frac{\psi_{t,nT}^{2}\left( \lambda
_{T}\right) -1}{\sqrt{2}}\right) +\left( \sqrt{\frac{n}{n-1}}\right)
\left( p_{nT}\left( \lambda_{T}\right) -q_{nT}\left( \lambda_{T}\right)
\right) ,
\]
where $\psi_{t,nT}\left( \lambda_{T}\right) $ is defined by ((ref)),
equation[equation omitted — 125 chars of source]
and
equation[equation omitted — 160 chars of source]
By Lemma (ref) of the supplement $p_{nT}\left( \lambda
_{T}\right) =o_{p}(1)$, and $q_{nT}\left( \lambda_{T}\right) =o_{p}(1)$.
Hence
equation[equation omitted — 202 chars of source]
Now consider $T^{-1/2}\sum_{t=1}^{T}\psi_{t,nT}^{2}\left( \lambda_{T}\right)
$ and using ((ref)) note that
align[align omitted — 1,246 chars of source]
By Lemma (ref) of the supplement, it follows that
equation[equation omitted — 388 chars of source]
Consider now the bias-corrected version of $\widetilde{CD}\left( \lambda
_{T}\right) $ defined by
equation[equation omitted — 172 chars of source]
where $\theta_{n}=1-\frac{1}{n}\sum_{i=1}^{n}a_{i,n}^{2},$ and $a_{i,n}
=1-\sigma_{i}\boldsymbol{\varphi}_{n}^{\prime}\boldsymbol{\gamma}_{i}$. Using
((ref)) in ((ref)), we have
align*[align* omitted — 461 chars of source]
Now using ((ref)) in the above and after some re-arrangement of the
terms we obtain
\[
\widetilde{CD}^{\ast}\left( \lambda_{T}\right) =\frac{\left[ \frac{1}
{\sqrt{T}}\sum_{t=1}^{T}\left( \frac{\xi_{t,n}^{2}\left( \lambda_{T}\right)
-\left( 1-\theta_{n}\right) }{\sqrt{2}}\right) \right] \left( 1+\frac
{2}{\sqrt{T}}w_{nT}\left( \lambda_{T}\right) \right) }{1-\theta_{n}}
+\sqrt{2}w_{nT}\left( \lambda_{T}\right) +o_{p}\left( 1\right) ,
\]
where
\[
w_{nT}\left( \lambda_{T}\right) =\frac{T^{-1/2}\sum_{t=1}^{T}\xi
_{t,n}\left( \lambda_{T}\right) \upsilon_{t,nT}\left( \lambda_{T}\right)
}{T^{-1}\sum_{t=1}^{T}\xi_{t,n}^{2}\left( \lambda_{T}\right) }.
\]
By Lemma (ref) of the supplement, $w_{nT}\left( \lambda
_{T}\right) =o_{p}\left( 1\right) $. Hence\
\[
\widetilde{CD}^{\ast}\left( \lambda_{T}\right) =\frac{\frac{1}{\sqrt{T}}
\sum_{t=1}^{T}\left( \frac{\xi_{t,n}^{2}\left( \lambda_{T}\right) -\left(
1-\theta_{n}\right) }{\sqrt{2}}\right) }{1-\theta_{n}}+o_{p}\left(
1\right) .
\]
In particular, since $\varepsilon_{it}(\lambda_{T})=\varepsilon_{it}
+\lambda_{T}\mathbf{w}_{i0}^{\prime}\boldsymbol{\varepsilon}_{\circ t}$, then
align*[align* omitted — 322 chars of source]
Using this result, $\widetilde{CD}^{\ast}(\lambda_{T})$ can be written as
align[align omitted — 912 chars of source]
The second term of ((ref)) can be written as
align*[align* omitted — 270 chars of source]
where $\phi_{n}=\frac{\sqrt{2}c_{\lambda}}{\left( 1-\theta_{n}\right)
}T^{-1}\sum_{t=1}^{T}E\left( A_{nt}B_{nt}\right) $. Further
align*[align* omitted — 500 chars of source]
with $\mathbf{a}_{n}=(a_{1,n},a_{2,n},...,a_{n,n})^{\prime}$. Hence
equation[equation omitted — 267 chars of source]
Under part (a) of Assumption (ref), $\varepsilon_{it}\sim
IID\left( 0,1\right) $ for all $i$ and $t$, with $E\left( \varepsilon
_{it}^{8}\right) <C$, it then follows that $A_{nt}B_{nt}-E\left(
A_{nt}B_{nt}\right) $ will be serially independent with zero means and finite
variances and by weak law of large numbers $T^{-1}\sum_{t=1}^{T}\left[
A_{nt}B_{nt}-E\left( A_{nt}B_{nt}\right) \right] =o_{p}(1)$. Therefore
equation[equation omitted — 60 chars of source]
Similarly, for the third term of ((ref)) we first note that
align*[align* omitted — 543 chars of source]
and hence
equation[equation omitted — 273 chars of source]
Using ((ref)) and ((ref)) in ((ref)) now yields
equation[equation omitted — 152 chars of source]
Consider the first term of ((ref)), $\widetilde{CD}^{\ast}\left(
0\right) $, and note that $A_{nt}$ can be written as $\,A_{nt}=n^{-1/2}
\mathbf{a}_{n}^{\prime}\boldsymbol{\varepsilon}_{\circ t}$, where
$\boldsymbol{a}_{n}^{\prime}=(a_{1,n},a_{2,n},...,a_{n,n})$, and we have
(recall that $a_{i}=1-\sigma_{i}\boldsymbol{\varphi}_{n}^{\prime
}\boldsymbol{\gamma}_{i}$)
\[
E\left( A_{nt}^{2}\right) =n^{-1}E\left( \boldsymbol{\varepsilon}_{\circ
t}^{\prime}\mathbf{a}_{n}\mathbf{a}_{n}^{\prime}\boldsymbol{\varepsilon
}_{\circ t}\right) =n^{-1}\mathbf{a}_{n}^{\prime}\mathbf{a}_{n}=n^{-1}
\sum_{i=1}^{n}a_{i,n}^{2}=\frac{1}{n}\sum_{i=1}^{n}\left( 1-\sigma
_{i}\boldsymbol{\varphi}_{n}^{\prime}\boldsymbol{\gamma}_{i}\right)
^{2}=1-\theta_{n}>0,
\]
and (using result (S.7) of Lemma 6 in pesaran2024testing)
align*[align* omitted — 474 chars of source]
where $\kappa_{2}=E\left( \varepsilon_{it}^{4}\right) -3$ and $\mathbf{A}
=\mathbf{a}_{n}\mathbf{a}_{n}^{\prime}$. Hence
\[
Var\left( A_{nt}^{2}\right) =2\left( \frac{1}{n}\sum_{i=1}^{n}a_{i,n}
^{2}\right) ^{2}-\frac{\kappa_{2}}{n}\left( \frac{1}{n}\sum_{i=1}^{n}
a_{i,n}^{4}\right) .
\]
Furthermore, since $\sup_{i}\left\vert a_{i,n}\right\vert =\sup_{i}\left\vert
1-\sigma_{i}\boldsymbol{\varphi}_{n}^{\prime}\boldsymbol{\gamma}
_{i}\right\vert <1+\left( \sup_{i}\sigma_{i}\right) \left( \sup
_{i}\left\Vert \boldsymbol{\gamma}_{i}\right\Vert \right) \left\Vert
\boldsymbol{\varphi}_{n}\right\Vert <C$, then $n^{-2}\sum_{i=1}^{n}a_{i,n}
^{4}=O(n^{-1})$, and $Var\left( A_{nt}^{2}\right) =2\left( 1-\theta
_{n}\right) ^{2}+O(n^{-1})$. Using the above results it now readily follows
that,
equation[equation omitted — 327 chars of source]
Since under part (a) of Assumption (ref), $\varepsilon_{it}\sim
IID\left( 0,1\right) $ for all $i$ and $t$, then $A_{nt}=\frac{1}{\sqrt{n}
}\sum_{i=1}^{n}a_{i,n}\varepsilon_{it}$, and $A_{nt}^{2}$ are also
independently distributed over $t$ with finite second order moments. Then by
Lindeberg-L\'{e}vy central limit theorem it follows that $\widetilde{CD}
^{\ast}(0)\rightarrow_{d}\mathcal{N}(0,1),$ as $n$ and $T\rightarrow\infty$.
Using this result in ((ref)) we further have $\widetilde{CD}^{\ast
}\left( \lambda_{T}\right) \rightarrow_{d}\mathcal{N}\left( \phi,1\right)
$, where $\phi=\lim_{n\rightarrow\infty}\phi_{n}$. By Lemma (ref) of
the supplement we have $CD=\widetilde{CD}+o_{p}(1)$, then it follows
align*[align* omitted — 235 chars of source]
where the final line holds by ((ref)). Now result ((ref)) is established as required.
Proof of Proposition (ref)
Note that $\theta_{n}$ define by ((ref)) can be written as $\theta
_{n}=2g_{n}-\boldsymbol{\varphi}_{n}^{\prime}\mathbf{H}_{n}\boldsymbol{\varphi
}_{n}$, where $g_{n}=n^{-1}\sum_{i=1}^{n}\sigma_{i}\boldsymbol{\varphi}
_{n}^{\prime}\boldsymbol{\gamma}_{i}$, $\mathbf{H}_{n}=n^{-1}\sum_{i=1}
^{n}\sigma_{i}\left( \boldsymbol{\gamma}_{i}\boldsymbol{\gamma}_{i}^{\prime
}\right) $, $\boldsymbol{\varphi}_{n}=n^{-1}\sum_{i=1}^{n}\boldsymbol{\delta
}_{i},$ and $\boldsymbol{\delta}_{i}=\boldsymbol{\gamma}_{i}/\sigma_{i}$.
Similarly using ((ref)) we have $\hat{\theta}_{nT}=2\hat{g}
_{nT}-\boldsymbol{\hat{\varphi}}_{nT}^{\prime}\mathbf{\hat{H}}_{nT}
\boldsymbol{\hat{\varphi}}_{nT}$, where $\hat{g}_{nT}=n^{-1}\sum_{i=1}^{n}
\hat{\sigma}_{i,T}\boldsymbol{\hat{\varphi}}_{nT}^{\prime}\boldsymbol{\hat
{\gamma}}_{i},$ $\mathbf{\hat{H}}_{nT}=n^{-1}\sum_{i=1}^{n}\hat{\sigma}
_{i,T}^{2}\left( \boldsymbol{\hat{\gamma}}_{i}\boldsymbol{\hat{\gamma}}
_{i}^{\prime}\right) ,$ $\boldsymbol{\hat{\varphi}}_{nT}=n^{-1}\sum_{i=1}
^{n}\boldsymbol{\hat{\delta}}_{i,nT},$ and\ $\boldsymbol{\hat{\delta}}
_{i,nT}=\boldsymbol{\hat{\gamma}}_{i}/\hat{\sigma}_{i,T}$. Then
equation[equation omitted — 329 chars of source]
Consider the first term of the above
align[align omitted — 466 chars of source]
and since $\sigma_{i}$ and $\boldsymbol{\gamma}_{i}$ are bounded then
$n^{-1}\sum_{i=1}^{n}\sigma_{i}\boldsymbol{\gamma}_{i}=O(1)$. Also by
((ref)) of Lemma (ref) in the supplement we have
$\sqrt{T}\left( \boldsymbol{\hat{\varphi}}_{nT}-\boldsymbol{\varphi}
_{n}\right) =o_{p}(1)$, and hence the first term of the above is $o_{p}(1)$.
To establish the probability order of the second term of ((ref)),
we first note that
align[align omitted — 737 chars of source]
But by ((ref)) and ((ref)) of the supplement,
$\boldsymbol{\varphi}_{n}=O_{p}(1)$ and $n^{-1}\sum_{i=1}^{n}\left(
\hat{\sigma}_{i,T}\boldsymbol{\hat{\gamma}}_{i}-\sigma_{i}\boldsymbol{\gamma
}_{i}\right) =O_{p}\left( \ln\left( n\right) /T\right) ,$ which also
establishes that the second term of ((ref)) is $o_{p}(1)$.
Therefore overall we have
equation[equation omitted — 89 chars of source]
Consider now the second term of ((ref)) and note that
align[align omitted — 688 chars of source]
where $\mathbf{\hat{H}}_{nT}=\frac{1}{n}\sum_{i=1}^{n}\hat{\sigma}_{i,T}
^{2}\left( \boldsymbol{\hat{\gamma}}_{i}\boldsymbol{\hat{\gamma}}_{i}
^{\prime}\right) $, and
align[align omitted — 865 chars of source]
The first two terms of ((ref)) are $o_{p}(1)$, since $\left\Vert
\boldsymbol{\varphi}_{n}\right\Vert <C$, $\sqrt{T}\left( \boldsymbol{\hat
{\varphi}}_{nT}-\boldsymbol{\varphi}_{n}\right) =o_{p}(1)$, and $n^{-1}
\sum_{i=1}^{n}\hat{\sigma}_{i,T}^{2}\left( \boldsymbol{\hat{\gamma}}
_{i}\boldsymbol{\hat{\gamma}}_{i}^{\prime}\right) =O_{p}(1)$. To establish
the probability order of the third term of ((ref)), since
$\left\Vert \boldsymbol{\varphi}_{n}\right\Vert <C$ it is sufficient to
consider the four terms of $\sqrt{T}\left( \mathbf{\hat{H}}_{nT}
-\mathbf{H}_{n}\right) $. It is clear that $\mathbf{D}_{2,nT}$ is dominated
by $\mathbf{D}_{1,nT\ }$ and by ((ref)) of Lemma (ref) of the
supplement,
\[
\mathbf{D}_{1,nT\ }=\frac{\sqrt{T}}{n}\sum_{i=1}^{n}\sigma_{i}^{2}\left(
\boldsymbol{\hat{\gamma}}_{i}\boldsymbol{\hat{\gamma}}_{i}^{\prime
}-\boldsymbol{\gamma}_{i}\boldsymbol{\gamma}_{i}^{\prime}\right)
=O_{p}\left( \sqrt{\frac{\ln\left( n\right) }{n}}\right) =o_{p}(1).
\]
Using ((ref)) of Lemma (ref) of the supplement and replacing
$b_{ni}$ with $\gamma_{ij}\gamma_{ij^{^{\prime}}}$ for $j,j^{\prime
}=1,2,\ldots,m_{0}$, it then follows that
\[
\mathbf{D}_{3,nT}=\frac{\sqrt{T}}{n}\sum_{i=1}^{n}\left( \hat{\sigma}
_{i,T}^{2}-\omega_{i,T}^{2}\right) \boldsymbol{\gamma}_{i}\boldsymbol{\gamma
}_{i}^{\prime}=O_{p}\left( \frac{\ln\left( n\right) }{\sqrt{T}}\right)
=o_{p}(1).
\]
Finally, denote the $\left( j,j^{\prime}\right) $ element of $\mathbf{D}
_{4,nT}$ by $d_{4,nT}(j,j^{\prime})$ and note that
\[
d_{4,nT}(j,j^{\prime})=\frac{1}{n}\sum_{i=1}^{n}\left( \sigma_{i}^{2}
\gamma_{ij}\gamma_{ij^{\prime}}\right) \sqrt{T}\left( \frac
{\boldsymbol{\varepsilon}_{i\circ}^{^{\prime}}\mathbf{M}_{F}
\boldsymbol{\varepsilon}_{i\circ}}{T}-1\right) \text{, for }j,j^{\prime
}=1,2,...,m_{0}.
\]
But under Assumptions (ref) and (ref), $\left\vert
\sigma_{i}^{2}\gamma_{ij}\gamma_{ij^{\prime}}\right\vert <C$, and $\sqrt
{T}\left( T^{-1}\boldsymbol{\varepsilon}_{i\circ}^{^{\prime}}\mathbf{M}
_{F}\boldsymbol{\varepsilon}_{i\circ}-1\right) $, for $i=1,2,...,n$ are
identically and independently distributed across $i$, with mean $1/\sqrt{T}$
and a finite variance\footnote{The mean and variance of $\sqrt{T}\left(
T^{-1}\boldsymbol{\varepsilon}_{i\circ}^{^{\prime}}\mathbf{M}_{F}
\boldsymbol{\varepsilon}_{i\circ}-1\right) $ can be obtained using
((ref)) and ((ref)) in Lemma (ref) of the supplement.}. Then by standard law of large numbers, for each $(j,j^{\prime})$,
$d_{4,nT}(j,j^{\prime})\rightarrow_{p}0$, as $n$ and $T\rightarrow\infty$, and
hence we also have $\mathbf{D}_{4,nT}=o_{p}(1)$. Overall, $\mathbf{\hat{H}
}_{nT}-\mathbf{H}_{n}=o_{p}(1)$, and we have $\sqrt{T}\left( \boldsymbol{\hat
{\varphi}}_{nT}^{\prime}\mathbf{\hat{H}}_{nT}\boldsymbol{\hat{\varphi}}
_{nT}-\boldsymbol{\varphi}_{n}^{\prime}\mathbf{H}_{n}\boldsymbol{\varphi}
_{n}\right) =o_{p}(1)$. Using this result and ((ref)) in
((ref)) now yields $\sqrt{T}\left( \hat{\theta}_{nT}-\theta
_{n}\right) =o_{p}(1)$, as required.
Proof of Theorem (ref)
Recall from ((ref)) that $CD^{\ast}$ is given by
\[
CD^{\ast}=\frac{CD+\sqrt{\frac{T}{2}}\hat{\theta}_{nT}}{1-\hat{\theta}_{nT}},
\]
where $\hat{\theta}_{nT}=1-\frac{1}{n}\sum_{i=1}^{n}\hat{a}_{i,n}^{2}$,
$\hat{a}_{i,n}=1-\hat{\sigma}_{i,T}\left( \boldsymbol{\varphi}_{nT}^{\prime
}\boldsymbol{\hat{\gamma}}_{i}\right) ,$ and $\boldsymbol{\hat{\varphi}}
_{nT}=n^{-1}\sum_{i=1}^{n}\boldsymbol{\hat{\gamma}}_{i}/\hat{\sigma}_{i,T}$,
subject to the normalization $n^{-1}\sum_{i=1}^{n}\boldsymbol{\hat{\gamma}
}_{i}\boldsymbol{\hat{\gamma}}_{i}^{^{\prime}}=\mathbf{I}_{m_{0}}$. By result
((ref)) of Proposition (ref), $\sqrt{T}\left(
\hat{\theta}_{nT}-\theta_{n}\right) =o_{p}(1)$, and hence
align*[align* omitted — 329 chars of source]
Theorem (ref) is then established by following Proposition
(ref).
Proof of Theorem (ref)
Let $v_{it}=y_{it}-\boldsymbol{\alpha}_{i}^{\prime}\mathbf{d}_{t}
-\boldsymbol{\beta}_{i}^{^{\prime}}\mathbf{x}_{it}$, and $u_{it}
=y_{it}-\boldsymbol{\alpha}_{i}^{\prime}\mathbf{d}_{t}-\boldsymbol{\beta}
_{i}^{^{\prime}}\mathbf{x}_{it}-\boldsymbol{\gamma}_{i}^{^{\prime}}
\mathbf{f}_{t}=v_{it}-\boldsymbol{\gamma}_{i}^{^{\prime}}\mathbf{f}_{t}$, and
consider the following two optimization problems
align[align omitted — 350 chars of source]
where
align[align omitted — 710 chars of source]
We need to show that solving problem ((ref)) is asymptotically
equivalent to solving problem ((ref)). First, using the results in
pesaran2011large and the fact that $\mathbf{d}_{t}$\ and $\mathbf{x}
_{it}$ are (stochastically) bounded\footnote{See equation (31) in
pesaran2011large.},
align[align omitted — 491 chars of source]
then rewrite the criterion for ((ref)) with ((ref)),
align[align omitted — 1,830 chars of source]
Therefore, using ((ref)) and ((ref)), then
$A_{2,nT}=O_{p}\left( \frac{1}{\sqrt{T}}\right) +O_{p}\left( \frac{1}
{n}\right) +O_{p}\left( \frac{1}{\sqrt{nT}}\right) $ and $A_{3,nT}
=O_{p}\left( \frac{1}{\sqrt{T}}\right) +O_{p}\left( \frac{1}{n}\right)
+O_{p}\left( \frac{1}{\sqrt{nT}}\right) $. Also, consider the fourth term of
((ref)) and note that by Cauchy-Schwarz inequality,
align*[align* omitted — 490 chars of source]
where $u_{it}=v_{it}-\boldsymbol{\gamma}_{i}^{^{\prime}}\mathbf{f}_{t}
=y_{it}-\boldsymbol{\alpha}_{i}^{\prime}\mathbf{d}_{t}-\boldsymbol{\beta}
_{i}^{^{\prime}}\mathbf{x}_{it}-\boldsymbol{\gamma}_{i}^{^{\prime}}
\mathbf{f}_{t}$, so given ((ref)) we have $A_{4,nT}=O_{p}\left(
\frac{1}{\sqrt{T}}\right) +O_{p}\left( \frac{1}{n}\right) +O_{p}\left(
\frac{1}{\sqrt{nT}}\right) $. Similarly, we can show $A_{5,nT}$ and
$A_{6,nT}$ share the same probability order as $A_{4,nT}$. Since in both
optimization problems $\boldsymbol{\gamma}_{i}$ and $\mathbf{f}_{t}$ are only
identified up to $m_{0}\times m_{0}$ rotation matrices, it follows that
align*[align* omitted — 444 chars of source]
Hence, PCs based on $\hat{v}_{it}$ are asymptotically equivalent to those
based on $v_{it}$. The remaining proof of Theorem (ref)
follows from the proof of Theorem (ref).
{
thebibliography\bibitem[Ahn and Horenstein, 2013]{ahn2013eigenvalue}
Ahn, S. C. and Horenstein, A. R. (2013).
\newblock Eigenvalue ratio test for the number of factors.
\newblock {\em Econometrica}, 81(3):1203--1227.
\bibitem[Aquaro et al., 2021]{aquaro2021estimation}
Aquaro, M., Bailey, N., and Pesaran, M. H. (2021).
\newblock Estimation and inference for spatial models with heterogeneous
coefficients: an application to {U.S.} house prices.
\newblock {\em Journal of Applied Econometrics}, 36(1):18--44.
\bibitem[Bai, 2003]{bai2003inferential}
Bai, J. (2003).
\newblock Inferential theory for factor models of large dimensions.
\newblock {\em Econometrica}, 71(1):135--171.
\bibitem[Bai, 2009]{bai2009panel}
Bai, J. (2009).
\newblock Panel data models with interactive fixed effects.
\newblock {\em Econometrica}, 77(4):1229--1279.
\bibitem[Bai and Li, 2021]{bai2021dynamic}
Bai, J. and Li, K. (2021).
\newblock Dynamic spatial panel data models with common shocks.
\newblock {\em Journal of Econometrics}, 224(1):134--160.
\bibitem[Bai and Ng, 2002]{bai2002determining}
Bai, J. and Ng, S. (2002).
\newblock Determining the number of factors in approximate factor models.
\newblock {\em Econometrica}, 70(1):191--221.
\bibitem[Bai and Ng, 2006]{bai2006confidence}
Bai, J. and Ng, S. (2006).
\newblock Confidence intervals for diffusion index forecasts and inference for
factor-augmented regressions.
\newblock {\em Econometrica}, 74(4):1133--1150.
\bibitem[Bai and Ng, 2023]{bai2023approximate}
Bai, J. and Ng, S. (2023).
\newblock Approximate factor models with weaker loadings.
\newblock {\em Journal of Econometrics}, 235(2):1893--1916.
\bibitem[Bailey et al., 2016]{bailey2016two}
Bailey, N., Holly, S., and Pesaran, M. H. (2016).
\newblock A two-stage approach to spatio-temporal analysis with strong and weak
cross-sectional dependence.
\newblock {\em Journal of Applied Econometrics}, 31(1):249--280.
\bibitem[Bailey et al., 2021]{bailey2021measurement}
Bailey, N., Kapetanios, G., and Pesaran, M. H. (2021).
\newblock Measurement of factor strength: Theory and practice.
\newblock {\em Journal of Applied Econometrics}, 36(2):431--453.
\bibitem[Baltagi et al., 2016]{baltagi2016testing}
Baltagi, B. H., Kao, C., and Peng, B. (2016).
\newblock Testing cross-sectional correlation in large panel data models with
serial correlation.
\newblock {\em Econometrics}, 4(4):44.
\bibitem[Chamberlain and Rothschild, 1983]{chamberlain1983arbitrage}
Chamberlain, G. and Rothschild, M. (1983).
\newblock Arbitrage, factor structure, and mean-variance analysis on large
asset markets.
\newblock {\em Econometrica}, 51(1):1281--1304.
\bibitem[Chiang and Tsai, 2016]{chiang2016ripple}
Chiang, M.-C. and Tsai, I.-C. (2016).
\newblock Ripple effect and contagious effect in the {U.S.} regional housing
markets.
\newblock {\em The Annals of Regional Science}, 56(1):55--82.
\bibitem[Chudik et al., 2018]{chudik2018one}
Chudik, A., Kapetanios, G., and Pesaran, M. H. (2018).
\newblock A one covariate at a time, multiple testing approach to variable
selection in high-dimensional linear regression models.
\newblock {\em Econometrica}, 86(4):1479--1512.
\bibitem[Chudik et al., 2011]{chudik2011weak}
Chudik, A., Pesaran, M., and Tosetti, E. (2011).
\newblock Weak and strong cross-section dependence and estimation of large
panels.
\newblock {\em The Econometrics Journal}, 14(1):C45--C90.
\bibitem[Durbin and Watson, 1950]{durbin1950testing}
Durbin, J. and Watson, G. S. (1950).
\newblock Testing for serial correlation in least squares regression. {I}.
\newblock {\em Biometrika}, 37(3-4):409--428.
\bibitem[Fan et al., 2011]{fan2011high}
Fan, J., Liao, Y., and Mincheva, M. (2011).
\newblock High dimensional covariance matrix estimation in approximate factor
models.
\newblock {\em Annals of Statistics}, 39(6):3320--3356.
\bibitem[Fan et al., 2013]{fan2013large}
Fan, J., Liao, Y., and Mincheva, M. (2013).
\newblock Large covariance estimation by thresholding principal orthogonal
complements.
\newblock {\em Journal of the Royal Statistical Society Series B: Statistical
Methodology}, 75(4):603--680.
\bibitem[Fan et al., 2015]{fan2015power}
Fan, J., Liao, Y., and Yao, J. (2015).
\newblock Power enhancement in high-dimensional cross-sectional tests.
\newblock {\em Econometrica}, 83(4):1497--1541.
\bibitem[Gagliardini et al., 2019]{gagliardini2019diagnostic}
Gagliardini, P., Ossola, E., and Scaillet, O. (2019).
\newblock A diagnostic criterion for approximate factor structure.
\newblock {\em Journal of Econometrics}, 212(2):503--521.
\bibitem[Goldie and Kl{\"u}ppelberg, 1998]{goldie1998subexponential}
Goldie, C. M. and Kl{\"u}ppelberg, C. (1998).
\newblock Subexponential distributions.
\newblock In R. J. Adler, R. E. Feldman, and M. S. Taqqu (Eds.), {\em A Practical Guide to Heavy Tails: Statistical Techniques and
Applications}, pages 435--459. Birkh{\"a}user Boston Inc., U.S.
\bibitem[Holly et al., 2011]{holly2011spatial}
Holly, S., Pesaran, M. H., and Yamagata, T. (2011).
\newblock The spatial and temporal diffusion of house prices in the {U.K.}
\newblock {\em Journal of Urban Economics}, 69(1):2--23.
\bibitem[Hsiao et al., 2012]{hsiao2012diagnostic}
Hsiao, C., Pesaran, M. H., and Pick, A. (2012).
\newblock Diagnostic tests of cross-section independence for limited dependent
variable panel data models.
\newblock {\em Oxford Bulletin of Economics and Statistics}, 74(2):253--277.
\bibitem[Jiang et al., 2023]{jiang2023revisiting}
Jiang, P., Uematsu, Y., and Yamagata, T. (2023).
\newblock Revisiting asymptotic theory for principal component estimators of
approximate factor models.
\newblock {\em arXiv preprint arXiv:2311.00625}.
\bibitem[Juodis and Reese, 2022]{juodis2022incidental}
Juodis, A. and Reese, S. (2022).
\newblock The incidental parameters problem in testing for remaining
cross-section correlation.
\newblock {\em Journal of Business & Economic Statistics}, 40(3):1191--1203.
\bibitem[Pesaran, 2004]{pesaran2004general}
Pesaran, M. H. (2004).
\newblock General diagnostic tests for cross-sectional dependence in panels.
\newblock {\em University of Cambridge, Cambridge Working Papers in Economics},
435.
\bibitem[Pesaran, 2006]{pesaran2006estimation}
Pesaran, M. H. (2006).
\newblock Estimation and inference in large heterogeneous panels with a
multifactor error structure.
\newblock {\em Econometrica}, 74(4):967--1012.
\bibitem[Pesaran, 2015]{pesaran2015testing}
Pesaran, M. H. (2015).
\newblock Testing weak cross-sectional dependence in large panels.
\newblock {\em Econometric Reviews}, 34(6-10):1089--1117.
\bibitem[Pesaran, 2021]{pesaran2021general}
Pesaran, M. H. (2021).
\newblock General diagnostic tests for cross-sectional dependence in panels.
\newblock {\em Empirical Economics}, 60(1):13--50.
\bibitem[Pesaran and Tosetti, 2011]{pesaran2011large}
Pesaran, M. H. and Tosetti, E. (2011).
\newblock Large panels with common factors and spatial correlation.
\newblock {\em Journal of Econometrics}, 161(2):182--202.
\bibitem[Pesaran and Yamagata, 2024]{pesaran2024testing}
Pesaran, M. H. and Yamagata, T. (2024).
\newblock Testing for alpha in linear factor pricing models with a large number
of securities.
\newblock {\em Journal of Financial Econometrics}, 22(2):407--460.
\bibitem[Phillips and Sul, 2003]{phillips2003dynamic}
Phillips, P. C. B. and Sul, D. (2003).
\newblock Dynamic panel estimation and homogeneity testing under cross section
dependence.
\newblock {\em The Econometrics Journal}, 6(1):217--259.
\bibitem[Phillips and Sul, 2007]{phillips2007bias}
Phillips, P. C. B. and Sul, D. (2007).
\newblock Bias in dynamic panel estimation with fixed effects, incidental
trends and cross section dependence.
\newblock {\em Journal of Econometrics}, 137(1):162--188.
\bibitem[Shi and Lee, 2017]{shi2017spatial}
Shi, W. and Lee, L.-f. (2017).
\newblock Spatial dynamic panel data models with interactive fixed effects.
\newblock {\em Journal of Econometrics}, 197(2):323--347.
\bibitem[Tsai, 2015]{tsai2015spillover}
Tsai, I.-C. (2015).
\newblock Spillover effect between the regional and the national housing
markets in the {U.K.}
\newblock {\em Regional Studies}, 49(12):1957--1976.
\bibitem[Vershynin, 2018]{vershynin2018high}
Vershynin, R. (2018).
\newblock {\em High-dimensional probability: An introduction with applications
in data science}, volume 47.
\newblock Cambridge university press.
\bibitem[Yang, 2021]{yang2021common}
Yang, C. F. (2021).
\newblock Common factors and spatial dependence: An application to {U.S.} house
prices.
\newblock {\em Econometric Reviews}, 40(1):14--50.
}
center[center omitted — 302 chars of source]
\setcounter{equation}{0}
\setcounter{section}{0}
\setcounter{page}{1}
\setcounter{theorem}{1}
\setcounter{footnote}{0}
\setcounter{table}{0}
\setcounter{figure}{0}
\setcounter{lemma}{0}
\setcounter{remark}{0}
This supplement is in four sections. Section (ref) states
and establishes the auxiliary lemmas used in the proofs of propositions and
theorems in the paper. Section (ref) derives the order of
$\theta_{n}$, defined by ((ref)) in the paper, in terms of the factor
strengths. Section (ref) considers the CD$_{W+}$ test proposed by
Juodis and Reese (2022), and discusses some of its properties. Section
(ref) reports simulation results for the experiments discussed in
Section (ref) of the main paper.
Statement and proofs of the lemmas
This section provides auxiliary lemmas and the associated proofs, which are
required to establish the main results of the paper.
lemmaThe CD statistic defined by ((ref)) can be written equivalently
as,
\begin{equation}
CD=\left( \sqrt{\frac{n}{n-1}}\right) \frac{1}{\sqrt{2T}}\sum_{t=1}
^{T}\left[ \left( \frac{1}{\sqrt{n}}\sum_{i=1}^{n}\frac{\hat{u}_{it}}
{\hat{\sigma}_{i,T}}\right) ^{2}-1\right] .
\end{equation}
proofUsing $\hat{\rho}_{ij,T}=\left( \frac{1}{T}\sum_{t=1}^{T}\hat{u}_{it}\hat
{u}_{jt}\right) /\hat{\sigma}_{i,T}\hat{\sigma}_{j,T}$ in ((ref)) we
have:
\begin{equation}
CD=\sqrt{\frac{2T}{n(n-1)}}\sum_{i=1}^{n-1}\sum_{j=i+1}^{n}\frac{\frac{1}
{T}\sum_{t=1}^{T}\hat{u}_{it}\hat{u}_{jt}}{\hat{\sigma}_{i,T}\hat{\sigma
}_{j,T}}=\sqrt{\frac{2T}{n(n-1)}}\frac{1}{T}\sum_{t=1}^{T}\left( \sum
_{i=1}^{n-1}\sum_{j=i+1}^{n}\left( \frac{\hat{u}_{it}}{\hat{\sigma}_{i,T}
}\right) \left( \frac{\hat{u}_{jt}}{\hat{\sigma}_{j,T}}\right) \right) .
\end{equation}
Further, we note that
\[
\frac{1}{n}\sum_{i=1}^{n-1}\sum_{j=i+1}^{n}\left( \frac{\hat{u}_{it}}
{\hat{\sigma}_{i,T}}\right) \left( \frac{\hat{u}_{jt}}{\hat{\sigma}_{j,T}
}\right) =\frac{1}{2}\left[ \left( \frac{1}{\sqrt{n}}\sum_{i=1}^{n}
\frac{\hat{u}_{it}}{\hat{\sigma}_{i,T}}\right) ^{2}-\frac{1}{n}\sum_{i=1}
^{n}\left( \frac{\hat{u}_{it}}{\hat{\sigma}_{i,T}}\right) ^{2}\right] .
\]
Then using this result in ((ref)), and after some algebra, we have
\begin{align*}
CD & =\sqrt{\frac{2Tn^{2}}{n(n-1)}}\frac{1}{2T}\sum_{t=1}^{T}\left[ \left(
\frac{1}{\sqrt{n}}\sum_{i=1}^{n}\frac{\hat{u}_{it}}{\hat{\sigma}_{i,T}
}\right) ^{2}-\frac{1}{n}\sum_{i=1}^{n}\left( \frac{\hat{u}_{it}}
{\hat{\sigma}_{i,T}}\right) ^{2}\right] \\
& =\sqrt{\frac{2Tn^{2}}{n(n-1)}}\frac{1}{2}\left[ \frac{1}{T}\sum_{t=1}
^{T}\left( \frac{1}{\sqrt{n}}\sum_{i=1}^{n}\frac{\hat{u}_{it}}{\hat{\sigma
}_{i,T}}\right) ^{2}-\frac{1}{n}\sum_{i=1}^{n}\frac{1}{T}\sum_{t=1}
^{T}\left( \frac{\hat{u}_{it}}{\hat{\sigma}_{i,T}}\right) ^{2}\right] \\
& =\left( \sqrt{\frac{n}{n-1}}\right) \frac{1}{\sqrt{2T}}\sum_{t=1}
^{T}\left[ \left( \frac{1}{\sqrt{n}}\sum_{i=1}^{n}\frac{\hat{u}_{it}}
{\hat{\sigma}_{i,T}}\right) ^{2}-1\right] ,
\end{align*}
as required.
lemmaConsider the latent factor model given by ((ref)) and
((ref)). The latent factors, $\mathbf{f}_{t}$, and their loadings,
$\boldsymbol{\gamma}_{i}$, are estimated by principal components,
$\mathbf{\hat{f}}_{t}$ and $\boldsymbol{\hat{\gamma}}_{i}$, given by
((ref)). Suppose that Assumptions (ref)-(ref) hold and
$\left( n,T\right) \rightarrow\infty$, such that $n/T\rightarrow\kappa\,$
for $0<\kappa<\infty$. Then
\begin{align}
\left\Vert \mathbf{\hat{F}}-\mathbf{F}\right\Vert _{F} & =O_{p}\left(
\frac{\sqrt{T}}{\delta_{nT}}\right) ,\\
\left\Vert \boldsymbol{\hat{\Gamma}}-\boldsymbol{\Gamma}\right\Vert _{F} &
=O_{p}\left( \frac{\sqrt{n}}{\delta_{nT}}\right) ,\\
\left\Vert \mathbf{U}\left( \lambda_{T}\right) ^{^{\prime}}\mathbf{(\hat
{F}-F)}\right\Vert _{F} & =O_{p}\left( \frac{\sqrt{nT}}{\delta_{nT}
}\right) ,\\
\left\Vert \boldsymbol{\Gamma}^{\prime}(\boldsymbol{\hat{\Gamma}
}-\boldsymbol{\Gamma})\right\Vert _{F} & =O_{p}\left( \frac{n}{\delta_{nT}
}\right) ,\\
\left\Vert \mathbf{F}^{\prime}(\mathbf{\hat{F}}-\mathbf{F})\right\Vert _{F}
& =O_{p}\left( \frac{T}{\delta_{nT}}\right) ,\\
\left( \mathbf{\hat{F}-F}\right) ^{^{\prime}}\mathbf{F} & =O_{p}\left(
\frac{T}{\delta_{nT}^{2}}\right) ,\\
\left( \mathbf{\hat{F}-F}\right) ^{^{\prime}}\mathbf{\hat{F}} &
=O_{p}\left( \frac{T}{\delta_{nT}^{2}}\right) ,\\
(\boldsymbol{\hat{\Gamma}}-\boldsymbol{\Gamma})^{\prime}\mathbf{u}_{\circ
t}\left( \lambda_{T}\right) & =O_{p}\left( \frac{n}{\delta_{nT}^{2}
}\right) ,
\end{align}
where $\mathbf{F}=\left( \mathbf{f}_{1},\mathbf{f}_{2},\ldots,\mathbf{f}
_{T}\right) ^{\prime}$, $\mathbf{\hat{F}}=\left( \mathbf{\hat{f}}
_{1},\mathbf{\hat{f}}_{2},\ldots,\mathbf{\hat{f}}_{T}\right) ^{\prime}$,
$\boldsymbol{\Gamma}=\left( \boldsymbol{\gamma}_{1},\boldsymbol{\gamma}
_{2},...,\boldsymbol{\gamma}_{n}\right) ^{\prime}$, $\boldsymbol{\hat{\Gamma
}}=\left( \boldsymbol{\hat{\gamma}}_{1},\boldsymbol{\hat{\gamma}}
_{2},...,\boldsymbol{\hat{\gamma}}_{n}\right) ^{\prime}$, $\mathbf{U}\left(
\lambda_{T}\right) =$ $(\mathbf{u}_{\circ1}\left( \lambda_{T}\right)
,\mathbf{u}_{\circ2}\left( \lambda_{T}\right) ,...,\mathbf{u}_{\circ
T}\left( \lambda_{T}\right) )^{\prime}$, $\mathbf{u}_{\circ t}\left(
\lambda_{T}\right) =(\sigma_{1}\varepsilon_{1t}\left( \lambda_{T}\right)
,\sigma_{2}\varepsilon_{2t}\left( \lambda_{T}\right) ,...,\sigma
_{n}\varepsilon_{nt}\left( \lambda_{T}\right) )^{\prime}$, and
$\varepsilon_{it}\left( \lambda_{T}\right) =\varepsilon_{it}+\lambda_{T}
\sum_{j=1}^{n}w_{ij}\varepsilon_{jt}$.
proofSince Assumptions (ref)-(ref) are a sub-set of assumptions
made by Bai (2003), so results ((ref)) to ((ref)), ((ref))
and ((ref)) follow directly from Lemmas B.1, B.2 and B.3, and Theorems 1
and 2 of Bai (2003). Results ((ref)) and ((ref)) can be
established analogously.
lemmaSuppose that Assumptions (ref)-(ref)
hold and $\left( n,T\right) \rightarrow\infty$, such that $n/T\rightarrow
\kappa\,,$ for $0<\kappa<\infty$. Then
\begin{align}
\sup_{i}\left( T^{-1}\left\Vert \boldsymbol{\varepsilon}_{i\circ}\right\Vert
^{2}\right) & =O_{p}\left( 1\right) ,\\
\sup_{i}\left\Vert \frac{\mathbf{F}^{\prime}\boldsymbol{\varepsilon}_{i\circ}
}{T}\right\Vert & =O_{p}\left( \sqrt{\frac{\ln(n)}{T}}\right)
,\\
\sup_{t}\left\Vert \frac{\boldsymbol{\Gamma}^{\prime}\boldsymbol{\varepsilon
}_{\circ t}}{n}\right\Vert & =O_{p}\left( \sqrt{\frac{\ln(T)}{n}}\right)
,\\
\sup_{i}\left\Vert \frac{1}{nT}\sum_{t=1}^{T}\sum_{j=1}^{n}\sigma
_{j}\boldsymbol{\gamma}_{j}\varepsilon_{it}\varepsilon_{jt}\right\Vert &
=O_{p}\left( \sqrt{\frac{\ln\left( n\right) }{nT}}\right) ,
\end{align}
where $\boldsymbol{\varepsilon}_{i\circ}=\left( \varepsilon_{i1}
,\varepsilon_{i2},\ldots,\varepsilon_{iT}\right) ^{\prime}$ and
$\boldsymbol{\varepsilon}_{\circ t}=\left( \varepsilon_{1t},\varepsilon
_{2t},\ldots,\varepsilon_{nt}\right) ^{\prime}$.
proofConsider ((ref)) and note
\[
T^{-1}\left\Vert \boldsymbol{\varepsilon}_{i\circ}\right\Vert ^{2}=\frac{1}
{T}\sum_{t=1}^{T}\left[ \varepsilon_{it}^{2}-E\left( \varepsilon_{it}
^{2}\right) \right] +\frac{1}{T}\sum_{t=1}^{T}E\left( \varepsilon_{it}
^{2}\right) =\frac{1}{T}\sum_{t=1}^{T}z_{it}+1,
\]
where $z_{it}=\varepsilon_{it}^{2}-E\left( \varepsilon_{it}^{2}\right) $.
Then
\begin{equation}
\sup_{i}\left( T^{-1}\left\Vert \boldsymbol{\varepsilon}_{i\circ}\right\Vert
^{2}\right) \leq\sup_{i}\left\vert \frac{1}{T}\sum_{t=1}^{T}z_{it}\right\vert
+1.
\end{equation}
To establish the probability of the first term, consider the filtration
$\mathcal{I}_{it}^{\left( 1\right) }=\{\varepsilon_{i\tau}:\tau
=t-1,t-2,\ldots\}$ and, given the serial independence of $\varepsilon_{it}$,
note that $E\left( z_{it}|\mathcal{I}_{i,t-1}^{\left( 1\right) }\right)
=0$, so $z_{it}$ is a martingale difference process with respect to
$\mathcal{I}_{i,t-1}^{\left( 1\right) }$. In addition, $Var\left(
z_{it}\right) =Var\left( \varepsilon_{it}^{2}\right) =E\left(
\varepsilon_{it}^{4}\right) -\left( E\varepsilon_{it}^{2}\right) ^{2}$
which is bounded by assumption. Also, as $\varepsilon_{it}$ is sub-exponential
by part (a) of Assumption (ref), $\varepsilon_{it}^{2}$ (and hence
$z_{it}$) is sub-exponential, and there exist positive constants $C_{4}$,
$C_{5}$ and $r_{3}$ such that
\[
\sup_{i}\Pr\left( \left\vert z_{it}\right\vert >a\right) \leq C_{4}
\exp\left( -C_{5}a^{r_{3}}\right) ,\text{ for all }a>0.\text{ }
\]
Then by Lemma A3 in the online theory supplement of Chudik et al. (2018), for
$\varsigma_{T}=\ominus\left( T^{\mu}\right) $ and $0<\mu<\left(
r_{3}+1\right) /\left( r_{3}+2\right) $, there exists a positive constant
$C_{6}$ such that$\Pr\left( \left\vert \sum_{t=1}^{T}z_{it}\right\vert
>\varsigma_{T}\right) \leq\exp\left( -C_{6}T^{-1}\varsigma_{T}^{2}\right)
$, and if $\mu>\left( r_{3}+1\right) /\left( r_{3}+2\right) $ there exists
a positive constant $C_{7}$ such that $\Pr\left( \left\vert \sum_{t=1}
^{T}z_{it}\right\vert >\varsigma_{T}\right) \leq\exp\left( -C_{7}\left(
\varsigma_{T}\right) ^{\frac{r_{3}}{r_{3}+1}}\right) . $ By Boole's
inequality, we have
\[
\Pr\left( \sup_{i}\left\vert \sum_{t=1}^{T}z_{it}\right\vert >\varsigma
_{T}\right) \leq\exp\left( \ln\left( n\right) -C_{6}T^{-1}\varsigma
_{T}^{2}\right) ,\text{ if }0<\mu<\left( r_{3}+1\right) /\left(
r_{3}+2\right) ,
\]
\[
\Pr\left( \sup_{i}\left\vert \sum_{t=1}^{T}z_{it}\right\vert >\varsigma
_{T}\right) \leq\exp\left( \ln\left( n\right) -C_{7}\left( \varsigma
_{T}\right) ^{\frac{r_{3}}{r_{3}+1}}\right) ,\text{ if }\mu>\left(
r_{3}+1\right) /\left( r_{3}+2\right) .
\]
Let $\varsigma_{T}=C_{8}\sqrt{T\ln\left( n\right) }$ where $C_{8}$ is a
finite but sufficiently large constant. Then for $0<\mu<\left( r_{3}
+1\right) /\left( r_{3}+2\right) $, we have
\begin{align*}
\Pr\left( \sup_{i}\left\vert \frac{1}{T}\sum_{t=1}^{T}z_{it}\right\vert
>C_{8}\sqrt{\frac{\ln\left( n\right) }{T}}\right) & =\Pr\left( \sup
_{i}\left\vert \sum_{t=1}^{T}\varepsilon_{it}-E\left( \varepsilon_{it}
^{2}\right) \right\vert >C_{8}\sqrt{T\ln\left( n\right) }\right) \\
& \leq\exp\left[ \ln\left( n\right) -C_{6}T^{-1}C_{8}^{2}T\ln\left(
n\right) \right] \\
& =\exp\left[ \ln\left( n\right) -C_{6}C_{8}^{2}\ln\left( n\right)
\right] ,
\end{align*}
which is $o\left( 1\right) $ given $C_{8}$ is sufficiently large. Also for
$\mu\geq\left( r_{3}+1\right) /\left( r_{3}+2\right) $, we have
\begin{align*}
\Pr\left( \sup_{i}\left\vert \frac{1}{T}\sum_{t=1}^{T}\varepsilon
_{it}\right\vert >C_{8}\sqrt{\frac{\ln\left( n\right) }{T}}\right) &
=\Pr\left( \sup_{i}\left\vert \sum_{t=1}^{T}\varepsilon_{it}\right\vert
>C_{8}\sqrt{T\ln\left( n\right) }\right) \\
& \leq\exp\left[ \ln\left( n\right) -C_{7}\left( C_{8}\sqrt{T\ln\left(
n\right) }\right) ^{\frac{r_{3}}{r_{3}+1}}\right] ,
\end{align*}
which is also $o\left( 1\right) $ as$\ n$ and $T$ are of the same order of
magnitude and sufficiently we have
\[
\frac{\ln\left( n\right) }{\left[ \sqrt{n\ln\left( n\right) }\right]
^{\frac{r_{3}}{r_{3}+1}}}=\frac{\left[ \ln\left( n\right) \right]
^{\frac{r_{3}+2}{2\left( r_{3}+1\right) }}}{n^{\frac{r_{3}}{2\left(
r_{3}+1\right) }}}=\left[ \frac{\left( \ln\left( n\right) \right)
^{\frac{r_{3}+2}{r_{3}}}}{n}\right] ^{\frac{r_{3}}{2\left( r_{3}+1\right)
}}\rightarrow0.
\]
Therefore, $\sup_{i}\left\vert T^{-1}\sum_{t=1}^{T}z_{it}\right\vert
=O_{p}\left( \sqrt{\frac{\ln\left( n\right) }{T}}\right) $ and
((ref)) follows from ((ref)). Next, consider ((ref))
and note by Assumptions (ref) and (ref), $\mathbf{f}
_{t}$ is independent from $\varepsilon_{it^{\prime}}$ for all $t,t^{\prime
}=1,2,\ldots,T$, also $\varepsilon_{it}$ is serially independent, then for a
suitable choice of $\mathcal{I}_{i,t-1}^{\left( 2\right) }=\{\mathbf{f}
_{\tau}\varepsilon_{i\tau}:$ $\tau=t-1,t-2,\ldots\}$ and $i=1,2,\ldots,n$,
$E\left( \mathbf{f}_{t}\varepsilon_{it}|\mathcal{I}_{i,t-1}^{\left(
2\right) }\right) =E\left( \mathbf{f}_{t}|\mathcal{I}_{i,t-1}^{\left(
2\right) }\right) E\left( \varepsilon_{it}\right) =\mathbf{0}$ so
$\mathbf{f}_{t}\varepsilon_{it}$ is a martingale difference sequence with
respect to the filtration $\mathcal{I}_{i,t-1}^{\left( 2\right) }$. In
addition, $E\left( \mathbf{f}_{t}\varepsilon_{it}\right) =E\left(
\mathbf{f}_{t}\right) E\left( \varepsilon_{it}\right) =\mathbf{0}$ and
$Var\left( \mathbf{f}_{t}\varepsilon_{it}\right) =E\left( \mathbf{f}
_{t}\mathbf{f}_{t}^{\prime}\right) E\left( \varepsilon_{it}^{2}\right) $,
which is bounded by Assumptions (ref) and (ref). Also
by assumptions both $\mathbf{f}_{t}$ and $\varepsilon_{it}$ are
sub-exponential, then it also follows that $\mathbf{f}_{t}\varepsilon_{it}$ is
sub-exponential. Hence, the method of proof used above can also be applied to
$\left\Vert \sum_{t=1}^{T}\mathbf{f}_{t}\varepsilon_{it}\right\Vert $, and
result ((ref)) follows. Similarly, ((ref)) can be
established by the symmetry of the standard factor models in
$\boldsymbol{\gamma}_{i}$ and $\mathbf{f}_{t}$. Now consider
((ref)), and note that we have the following decomposition,
\begin{align*}
\mathbf{q}_{i,nT} & =\frac{1}{nT}\sum_{t=1}^{T}\sum_{j=1}^{n}\sigma
_{j}\boldsymbol{\gamma}_{j}\varepsilon_{it}\varepsilon_{jt}\\
& =\frac{1}{nT}\sum_{t=1}^{T}\sum_{j=1}^{n}\left[ \sigma_{j}
\boldsymbol{\gamma}_{j}\varepsilon_{it}\varepsilon_{jt}-\sigma_{j}
\boldsymbol{\gamma}_{j}E\left( \varepsilon_{it}\varepsilon_{jt}\right)
\right] +\frac{1}{nT}\sum_{t=1}^{T}\sum_{j=1}^{n}\sigma_{j}\boldsymbol{\gamma
}_{j}E\left( \varepsilon_{it}\varepsilon_{jt}\right) =\mathbf{q}
_{i,nT}^{(a)}+\mathbf{q}_{i,nT}^{(b)}.
\end{align*}
Since $E(\varepsilon_{it}\varepsilon_{jt})=0$ if $i\neq j$, and $E(\varepsilon
_{it}\varepsilon_{jt})=1$, if $i=j$, then $\mathbf{q}_{i,nT}^{(b)}=\frac
{1}{nT}\sum_{t=1}^{T}\sum_{j=1}^{n}\sigma_{j}\boldsymbol{\gamma}_{j}E\left(
\varepsilon_{it}\varepsilon_{jt}\right) =n^{-1}\sigma_{j}\boldsymbol{\gamma
}_{j}, $ and $\sup_{i}\left\Vert \mathbf{q}_{i,nT}^{(b)}\right\Vert
=O(n^{-1})$. Consider now the first term and note that $\mathbf{q}
_{i,nT}^{(a)}=\frac{1}{nT}\sum_{t=1}^{T}\sum_{j=1}^{n}\mathbf{s}_{i,jt}, $
where $\mathbf{s}_{i,jt}=\sigma_{j}\boldsymbol{\gamma}_{j}\left[
\varepsilon_{it}\varepsilon_{jt}-E\left( \varepsilon_{it}\varepsilon
_{jt}\right) \right] $. Since by assumption $\varepsilon_{it}$ are
independently distributed over all $i$ and $t$, then $E\left( \mathbf{s}
_{i,jt}\left\vert \mathcal{I}_{i,t-1}^{\left( 3\right) }\right. \right)
=\mathbf{0}$, where $\mathcal{I}_{i,t-1}^{\left( 3\right) }=\{\varepsilon
_{i\tau}\varepsilon_{j\tau}:$ $j=1,2,...,n$ and $\tau=t-1,t-2,...\}$. Hence
$\mathbf{s}_{i,jt}$ is a martingale difference process with respect to the
filtration, $\mathcal{I}_{i,t-1}^{\left( 3\right) }$. The variance of
$\mathbf{s}_{i,jt}$ is $\left( \sigma_{j}^{2}\boldsymbol{\gamma}
_{j}\boldsymbol{\gamma}_{j}^{\prime}\right) Var(\varepsilon_{it}
\varepsilon_{jt})$ where $Var(\varepsilon_{it}\varepsilon_{jt})=1$ if $i\neq
j$, and $Var(\varepsilon_{it}\varepsilon_{jt})=Var(\varepsilon_{it}
^{2})=E(\varepsilon_{it}^{4})-1$ if $i=j$, so that by Assumption
(ref), $\left\Vert Var(\mathbf{s}_{i,jt})\right\Vert <C$. Also,
since by assumption $\varepsilon_{it}$ is sub-exponential, then it follows
that $\mathbf{s}_{i,jt}$ is also sub-exponential, and the above method of
proof can be applied to all elements of $\mathbf{q}_{i,nT}^{(a)}$.
Specifically $\sup_{i}\left\Vert \mathbf{q}_{i,nT}^{(a)}\right\Vert
=O_{p}\left( \sqrt{\frac{\ln(n)}{nT}}\right) $, $\sup_{i}\left\Vert
\mathbf{q}_{i,nT}\right\Vert =O_{p}\left( \sqrt{\frac{\ln(n)}{nT}}\right)
+O(n^{-1})=O_{p}\left( \sqrt{\frac{\ln(n)}{nT}}\right) $, and result
((ref)) follows, as required.
lemmaConsider the latent factor model given by ((ref)) and
((ref)). The latent factors, $\mathbf{f}_{t}$, and their loadings,
$\boldsymbol{\gamma}_{i}$, are estimated by principal components,
$\mathbf{\hat{f}}_{t}$ and $\boldsymbol{\hat{\gamma}}_{i}$, given by
((ref)). Suppose that Assumptions (ref)-(ref) hold and
$\left( n,T\right) \rightarrow\infty$, such that $n/T\rightarrow\kappa\,,$
for $0<\kappa<\infty$. Then
\begin{align}
\sup_{i}\left( T^{-1}\left\Vert \boldsymbol{\varepsilon}_{i\circ}\left(
\lambda_{T}\right) \right\Vert ^{2}\right) & =O_{p}\left( 1\right)
,\\
\sup_{i}\left\Vert \frac{\mathbf{F}^{\prime}\boldsymbol{\varepsilon}_{i\circ
}\left( \lambda_{T}\right) }{T}\right\Vert & =O_{p}\left( \sqrt{\frac
{\ln(n)}{T}}\right) ,\\
\sup_{t}\left\Vert \frac{\boldsymbol{\Gamma}^{\prime}\boldsymbol{\varepsilon
}_{\circ t}\left( \lambda_{T}\right) }{n}\right\Vert & =O_{p}\left(
\sqrt{\frac{\ln(T)}{n}}\right) ,\\
\sup_{i}\left\Vert \frac{1}{nT}\sum_{t=1}^{T}\sum_{j=1}^{n}\sigma
_{j}\boldsymbol{\gamma}_{j}\varepsilon_{it}\left( \lambda_{T}\right)
\varepsilon_{jt}\left( \lambda_{T}\right) \right\Vert & =O_{p}\left(
\sqrt{\frac{\ln\left( n\right) }{nT}}\right) ,\\
\sup_{i}\left\Vert \boldsymbol{\hat{\gamma}}_{i}-\boldsymbol{\gamma}
_{i}\right\Vert & =O_{p}\left( \sqrt{\frac{\ln(n)}{T}}\right)
,\\
\sup_{t}\left\Vert \mathbf{\hat{f}}_{t}-\mathbf{f}_{t}\right\Vert &
=O_{p}\left( \sqrt{\frac{\ln(T)}{n}}\right) ,
\end{align}
where $\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) =\left(
\varepsilon_{i1}\left( \lambda_{T}\right) ,\varepsilon_{i2}\left(
\lambda_{T}\right) ,\ldots,\varepsilon_{iT}\left( \lambda_{T}\right)
\right) ^{\prime}$ and $\boldsymbol{\varepsilon}_{\circ t}\left( \lambda
_{T}\right) =\left( \varepsilon_{1t}\left( \lambda_{T}\right)
,\varepsilon_{2t}\left( \lambda_{T}\right) ,\ldots,\varepsilon_{nt}\left(
\lambda_{T}\right) \right) ^{\prime}$.
proofConsider ((ref)) and note by definition
\begin{equation}
\varepsilon_{it}\left( \lambda_{T}\right) =\varepsilon_{it}+\lambda
_{T}\mathbf{w}_{i0}^{\prime}\boldsymbol{\varepsilon}_{\circ t},
\end{equation}
where $\mathbf{w}_{i0}=\left( w_{i1},w_{i2},\ldots,w_{in}\right) ^{\prime}$
and $\boldsymbol{\varepsilon}_{\circ t}=\left( \varepsilon_{1t}
,\varepsilon_{2t},\ldots,\varepsilon_{nt}\right) ^{\prime}$. Then
\[
\varepsilon_{it}^{2}\left( \lambda_{T}\right) =\varepsilon_{it}^{2}
+2\lambda_{T}\mathbf{w}_{i0}^{\prime}\boldsymbol{\varepsilon}_{\circ
t}\varepsilon_{it}+\lambda_{T}^{2}\mathbf{w}_{i0}^{\prime}
\boldsymbol{\varepsilon}_{\circ t}\boldsymbol{\varepsilon}_{\circ t}^{\prime
}\mathbf{w}_{i0},
\]
and hence
\[
\frac{1}{T}\sum_{t=1}^{T}\varepsilon_{it}^{2}\left( \lambda_{T}\right)
=\frac{1}{T}\sum_{t=1}^{T}\varepsilon_{it}^{2}+2\lambda_{T}\mathbf{w}
_{i0}^{\prime}\left( \frac{1}{T}\sum_{t=1}^{T}\boldsymbol{\varepsilon}_{\circ
t}\varepsilon_{it}\right) +\lambda_{T}^{2}\mathbf{w}_{i0}^{\prime}
\mathbf{V}_{\varepsilon T}\mathbf{w}_{i0},
\]
where $\mathbf{V}_{\varepsilon T}=T^{-1}\sum_{t=1}^{T}\boldsymbol{\varepsilon
}_{\circ t}\boldsymbol{\varepsilon}_{\circ t}^{\prime}$ and $\left\Vert
\mathbf{V}_{\varepsilon T}\right\Vert =O_{p}\left( 1\right) $ by part (b) of
Assumption (ref). It follows
\[
\left\vert \frac{1}{T}\sum_{t=1}^{T}\varepsilon_{it}^{2}\left( \lambda
_{T}\right) \right\vert \leq\left\vert \frac{1}{T}\sum_{t=1}^{T}
\varepsilon_{it}^{2}\right\vert +2\left\vert \lambda_{T}\right\vert \left\Vert
\mathbf{w}_{i0}\right\Vert \left\Vert \frac{1}{T}\sum_{t=1}^{T}
\boldsymbol{\varepsilon}_{\circ t}\varepsilon_{it}\right\Vert +\lambda_{T}
^{2}\left\Vert \mathbf{w}_{i0}\right\Vert ^{2}\left\Vert \mathbf{V}
_{\varepsilon T}\right\Vert .
\]
Denote $\mathbf{e}_{i}$ as $n\times1$ selection vector with $1$ on its
$i^{th}$ element and zeros elsewhere, and note that
\[
\sup_{i}\left\Vert \frac{1}{T}\sum_{t=1}^{T}\boldsymbol{\varepsilon}_{\circ
t}\varepsilon_{it}\right\Vert =\sup_{i}\left\Vert \frac{1}{T}\sum_{t=1}
^{T}\boldsymbol{\varepsilon}_{\circ t}\boldsymbol{\varepsilon}_{\circ
t}^{\prime}\mathbf{e}_{i}\right\Vert \leq\left\Vert \frac{1}{T}\sum_{t=1}
^{T}\boldsymbol{\varepsilon}_{\circ t}\boldsymbol{\varepsilon}_{\circ
t}^{\prime}\right\Vert \left( \sup_{i}\left\Vert \mathbf{e}_{i}\right\Vert
\right) =\left\Vert \mathbf{V}_{\varepsilon T}\right\Vert .
\]
Using this result we now have (recalling that $\lambda_{T}=c_{\lambda}
T^{-1/2})$
\begin{align*}
\sup_{i}\left\vert \frac{1}{T}\sum_{t=1}^{T}\varepsilon_{it}^{2}\left(
\lambda_{T}\right) \right\vert & \leq\sup_{i}\left\vert \frac{1}{T}
\sum_{t=1}^{T}\varepsilon_{it}^{2}\right\vert +2c_{\lambda}T^{-1/2}\left\Vert
\mathbf{V}_{\varepsilon T}\right\Vert \left( \sup_{i}\left\Vert
\mathbf{w}_{i0}\right\Vert \right) +\\
& c_{\lambda}^{2}T^{-1}\left\Vert \mathbf{V}_{\varepsilon T}\right\Vert
\left( \sup_{i}\left\Vert \mathbf{w}_{i0}\right\Vert ^{2}\right) .
\end{align*}
Therefore, since $\sup_{i}\left\Vert \mathbf{w}_{i0}\right\Vert <C$ and by
assumption $\left\Vert \mathbf{V}_{\varepsilon T}\right\Vert =O_{p}(1)$, then
\begin{equation}
\sup_{i}\left\vert \frac{1}{T}\sum_{t=1}^{T}\varepsilon_{it}^{2}\left(
\lambda_{T}\right) \right\vert =\sup_{i}\left\vert \frac{1}{T}\sum_{t=1}
^{T}\varepsilon_{it}^{2}\right\vert +O_{p}\left( \frac{1}{\sqrt{T}}\right) .
\end{equation}
Result ((ref)) now follows from ((ref)). Similarly, to
establish ((ref)) note that
\[
\frac{\mathbf{F}^{\prime}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda
_{T}\right) }{T}=\frac{1}{T}\sum_{t=1}^{T}\mathbf{f}_{t}\varepsilon
_{it}\left( \lambda_{T}\right) =\frac{1}{T}\sum_{t=1}^{T}\mathbf{f}
_{t}\varepsilon_{it}+\frac{\lambda_{T}}{T}\sum_{t=1}^{T}\mathbf{f}
_{t}\boldsymbol{\varepsilon}_{\circ t}^{\prime}\mathbf{w}_{i0},
\]
and
\[
\left\Vert \frac{\mathbf{F}^{\prime}\boldsymbol{\varepsilon}_{i\circ}\left(
\lambda_{T}\right) }{T}\right\Vert \leq\left\Vert \frac{1}{T}\sum_{t=1}
^{T}\mathbf{f}_{t}\varepsilon_{it}\right\Vert +\left\vert \lambda
_{T}\right\vert \left\Vert \mathbf{w}_{i0}\right\Vert \left\Vert \frac{1}
{T}\sum_{t=1}^{T}\mathbf{f}_{t}\boldsymbol{\varepsilon}_{\circ t}^{\prime
}\right\Vert .
\]
Applying the supremum operator to both sides yields
\[
\sup_{i}\left\Vert \frac{\mathbf{F}^{\prime}\boldsymbol{\varepsilon}_{i\circ
}\left( \lambda_{T}\right) }{T}\right\Vert \leq\sup_{i}\left\Vert \frac
{1}{T}\sum_{t=1}^{T}\mathbf{f}_{t}\varepsilon_{it}\right\Vert +\left\vert
\lambda_{T}\right\vert \left( \sup_{i}\left\Vert \mathbf{w}_{i0}\right\Vert
\right) \left\Vert \frac{1}{T}\sum_{t=1}^{T}\mathbf{f}_{t}
\boldsymbol{\varepsilon}_{\circ t}^{\prime}\right\Vert .
\]
Also
\begin{align*}
E\left\Vert \frac{1}{T}\sum_{t=1}^{T}\mathbf{f}_{t}\boldsymbol{\varepsilon
}_{\circ t}^{\prime}\right\Vert ^{2} & \leq E\left\Vert \frac{1}{T}
\sum_{t=1}^{T}\mathbf{f}_{t}\boldsymbol{\varepsilon}_{\circ t}^{\prime
}\right\Vert _{F}^{2}=\mathrm{tr}\left[ E\left[ \left( \frac{1}{T}
\sum_{t=1}^{T}\mathbf{f}_{t}\boldsymbol{\varepsilon}_{\circ t}^{\prime
}\right) \left( \frac{1}{T}\sum_{t=1}^{T}\mathbf{f}_{t}
\boldsymbol{\varepsilon}_{\circ t}^{\prime}\right) ^{\prime}\right] \right]
\\
& =\mathrm{tr}\left( \frac{1}{T^{2}}\sum_{t=1}^{T}\sum_{t^{\prime}=1}
^{T}E\left( \mathbf{f}_{t}\boldsymbol{\varepsilon}_{\circ t}^{\prime
}\boldsymbol{\varepsilon}_{\circ t^{\prime}}\mathbf{f}_{t^{\prime}}^{\prime
}\right) \right) =\frac{1}{T^{2}}\mathrm{tr}\left( \sum_{t=1}^{T}E\left(
\mathbf{f}_{t}\mathbf{f}_{t}^{\prime}\right) E\left( \boldsymbol{\varepsilon
}_{\circ t}^{\prime}\boldsymbol{\varepsilon}_{\circ t}\right) \right) \\
& =\frac{n}{T}\mathrm{tr}\left( \mathbf{\Sigma}_{ff}\right) =O\left(
\frac{nm_{0}}{T}\right) ,
\end{align*}
which establishes that $T^{-1}\sum_{t=1}^{T}\mathbf{f}_{t}
\boldsymbol{\varepsilon}_{\circ t}^{\prime}=O_{p}\left( 1\right) $ since by
assumption $n$ and $T$ have the same orders of magnitudes. Given $\lambda
_{T}=c_{\lambda}T^{-1/2}$ and $\sup_{i}\left\Vert \mathbf{w}_{i0}\right\Vert
<C$, we now have
\[
\sup_{i}\left\Vert \frac{\mathbf{F}^{\prime}\boldsymbol{\varepsilon}_{i\circ
}\left( \lambda_{T}\right) }{T}\right\Vert =\sup_{i}\left\Vert \frac{1}
{T}\sum_{t=1}^{T}\mathbf{f}_{t}\varepsilon_{it}\right\Vert +O_{p}\left(
\frac{1}{\sqrt{T}}\right) ,
\]
and ((ref)) follows using ((ref)). Similarly, ((ref))
can be established using result ((ref)). Next, consider
((ref)) and using the definition of $\varepsilon_{it}\left(
\lambda_{T}\right) $ in ((ref)) yields
\begin{align*}
\frac{1}{nT}\sum_{t=1}^{T}\sum_{j=1}^{n}\sigma_{j}\boldsymbol{\gamma}
_{j}\varepsilon_{it}\left( \lambda_{T}\right) \varepsilon_{jt}\left(
\lambda_{T}\right) & =\frac{1}{nT}\sum_{t=1}^{T}\sum_{j=1}^{n}\sigma
_{j}\boldsymbol{\gamma}_{j}\varepsilon_{it}\varepsilon_{jt}+\frac{\lambda_{T}
}{nT}\sum_{t=1}^{T}\sum_{j=1}^{n}\sigma_{j}\boldsymbol{\gamma}_{j}
\mathbf{w}_{j0}^{\prime}\boldsymbol{\varepsilon}_{\circ t}\varepsilon_{it}+\\
& \frac{\lambda_{T}}{nT}\sum_{t=1}^{T}\sum_{j=1}^{n}\sigma_{j}
\boldsymbol{\gamma}_{j}\mathbf{w}_{i0}^{\prime}\boldsymbol{\varepsilon}_{\circ
t}\varepsilon_{jt}+\frac{\lambda_{T}^{2}}{n}\sum_{j=1}^{n}\sigma
_{j}\boldsymbol{\gamma}_{j}\mathbf{w}_{i0}^{\prime}\left( \frac{1}{T}
\sum_{t=1}^{T}\boldsymbol{\varepsilon}_{\circ t}\boldsymbol{\varepsilon
}_{\circ t}^{\prime}\right) \mathbf{w}_{j0},
\end{align*}
which implies
\begin{align*}
\left\Vert \frac{1}{nT}\sum_{t=1}^{T}\sum_{j=1}^{n}\sigma_{j}
\boldsymbol{\gamma}_{j}\varepsilon_{it}\left( \lambda_{T}\right)
\varepsilon_{jt}\left( \lambda_{T}\right) \right\Vert & \leq\left\Vert
\frac{1}{nT}\sum_{t=1}^{T}\sum_{j=1}^{n}\sigma_{j}\boldsymbol{\gamma}
_{j}\varepsilon_{it}\varepsilon_{jt}\right\Vert +\left\Vert \frac{\lambda_{T}
}{nT}\sum_{t=1}^{T}\sum_{j=1}^{n}\sigma_{j}\boldsymbol{\gamma}_{j}
\mathbf{w}_{j0}^{\prime}\boldsymbol{\varepsilon}_{\circ t}\varepsilon
_{it}\right\Vert +\\
& \left\Vert \frac{\lambda_{T}}{nT}\sum_{t=1}^{T}\sum_{j=1}^{n}\sigma
_{j}\boldsymbol{\gamma}_{j}\mathbf{w}_{i0}^{\prime}\boldsymbol{\varepsilon
}_{\circ t}\varepsilon_{jt}\right\Vert +\left\Vert \frac{\lambda_{T}^{2}}
{n}\sum_{j=1}^{n}\sigma_{j}\boldsymbol{\gamma}_{j}\mathbf{w}_{i0}^{\prime
}\left( \frac{1}{T}\sum_{t=1}^{T}\boldsymbol{\varepsilon}_{\circ
t}\boldsymbol{\varepsilon}_{\circ t}^{\prime}\right) \mathbf{w}
_{j0}\right\Vert .
\end{align*}
Taking the supremum on both sides of this inequality yields
\begin{align}
& \sup_{i}\left\Vert \frac{1}{nT}\sum_{t=1}^{T}\sum_{j=1}^{n}\sigma
_{j}\boldsymbol{\gamma}_{j}\varepsilon_{it}\left( \lambda_{T}\right)
\varepsilon_{jt}\left( \lambda_{T}\right) \right\Vert \nonumber\\
& \leq\sup_{i}\left\Vert \frac{1}{nT}\sum_{t=1}^{T}\sum_{j=1}^{n}\sigma
_{j}\boldsymbol{\gamma}_{j}\varepsilon_{it}\varepsilon_{jt}\right\Vert
+\sup_{i}\left\Vert \frac{\lambda_{T}}{nT}\sum_{t=1}^{T}\sum_{j=1}^{n}
\sigma_{j}\boldsymbol{\gamma}_{j}\mathbf{w}_{j0}^{\prime}
\boldsymbol{\varepsilon}_{\circ t}\varepsilon_{it}\right\Vert +\nonumber\\
& \sup_{i}\left\Vert \frac{\lambda_{T}}{nT}\sum_{t=1}^{T}\sum_{j=1}^{n}
\sigma_{j}\boldsymbol{\gamma}_{j}\mathbf{w}_{i0}^{\prime}
\boldsymbol{\varepsilon}_{\circ t}\varepsilon_{jt}\right\Vert +\sup
_{i}\left\Vert \frac{\lambda_{T}^{2}}{n}\sum_{j=1}^{n}\sigma_{j}
\boldsymbol{\gamma}_{j}\mathbf{w}_{i0}^{\prime}\left( \frac{1}{T}\sum
_{t=1}^{T}\boldsymbol{\varepsilon}_{\circ t}\boldsymbol{\varepsilon}_{\circ
t}^{\prime}\right) \mathbf{w}_{j0}\right\Vert .
\end{align}
Note that $\varepsilon_{it}=\boldsymbol{\varepsilon}_{\circ t}^{\prime
}\mathbf{e}_{i}$ where $\mathbf{e}_{i}$ is an $n\times1$ selection vector with
$1$ on its $i^{th}$ element and zero elsewhere. Then the second term of the
above can be bounded as
\begin{align*}
\sup_{i}\left\Vert \frac{\lambda_{T}}{nT}\sum_{t=1}^{T}\sum_{j=1}^{n}
\sigma_{j}\boldsymbol{\gamma}_{j}\mathbf{w}_{j0}^{\prime}
\boldsymbol{\varepsilon}_{\circ t}\varepsilon_{it}\right\Vert & =\sup
_{i}\left\Vert \frac{\lambda_{T}}{n}\sum_{j=1}^{n}\sigma_{j}\boldsymbol{\gamma
}_{j}\mathbf{w}_{j0}^{\prime}\left( \frac{1}{T}\sum_{t=1}^{T}
\boldsymbol{\varepsilon}_{\circ t}\boldsymbol{\varepsilon}_{\circ t}^{\prime
}\right) \mathbf{e}_{i}\right\Vert \\
& \leq\left\vert \lambda_{T}\right\vert \left\Vert \frac{1}{n}\sum_{j=1}
^{n}\sigma_{j}\boldsymbol{\gamma}_{j}\mathbf{w}_{j0}^{\prime}\left( \frac
{1}{T}\sum_{t=1}^{T}\boldsymbol{\varepsilon}_{\circ t}\boldsymbol{\varepsilon
}_{\circ t}^{\prime}\right) \right\Vert \left( \sup_{i}\left\Vert
\mathbf{e}_{i}\right\Vert \right) \\
& \leq\left\vert \lambda_{T}\right\vert \left\Vert \frac{1}{n}\sum_{j=1}
^{n}\sigma_{j}\boldsymbol{\gamma}_{j}\mathbf{w}_{j0}^{\prime}\right\Vert
\left\Vert \mathbf{V}_{\varepsilon T}\right\Vert .
\end{align*}
Since $\sup_{i}\sigma_{i}^{2}<C$, and by Assumptions (ref)
-(ref), $\sup_{s,i}\gamma_{si}^{2}<C$ and $\sup_{i}\sum_{j=1}
^{n}\left\vert w_{ji}\right\vert <C$, then it follows
\begin{align*}
\left\Vert \frac{1}{n}\sum_{j=1}^{n}\sigma_{j}\boldsymbol{\gamma}
_{j}\mathbf{w}_{j0}^{\prime}\right\Vert ^{2} & \leq\left\Vert \frac{1}
{n}\sum_{j=1}^{n}\sigma_{j}\boldsymbol{\gamma}_{j}\mathbf{w}_{j0}^{\prime
}\right\Vert _{F}^{2}=\frac{1}{n^{2}}\sum_{s=1}^{m_{0}}\sum_{i=1}^{n}\left(
\sum_{j=1}^{n}\sigma_{j}\gamma_{sj}w_{ji}\right) ^{2}\\
& \leq\frac{1}{n^{2}}\sum_{s=1}^{m_{0}}\sum_{i=1}^{n}\left( \sum_{j=1}
^{n}\left\vert \sigma_{j}\gamma_{sj}\right\vert \left\vert w_{ji}\right\vert
\right) ^{2}\\
& \leq\left( \sup_{i}\sigma_{i}^{2}\right) \left( \sup_{s,i}\gamma
_{si}^{2}\right) \left[ \frac{1}{n}\sum_{s=1}^{m_{0}}\left( \frac{1}{n}
\sum_{i=1}^{n}\left( \sum_{j=1}^{n}\left\vert w_{ji}\right\vert \right)
^{2}\right) \right] \\
& =O\left( \frac{1}{n}\right) .
\end{align*}
In addition, $\lambda_{T}=c_{\lambda}T^{-1/2}$ and $\left\Vert \mathbf{V}
_{\varepsilon T}\right\Vert =O_{p}\left( n/T\right) $ with $0<n/T<C$. Hence,
we have $\sup_{i}\left\Vert \frac{\lambda_{T}}{nT}\sum_{t=1}^{T}\sum_{j=1}
^{n}\sigma_{j}\boldsymbol{\gamma}_{j}\mathbf{w}_{j0}^{\prime}
\boldsymbol{\varepsilon}_{\circ t}\varepsilon_{it}\right\Vert =O_{p}\left(
\left( nT\right) ^{-1/2}\right) $. Similarly the third term of
((ref)) is also $O_{p}\left( \left( nT\right) ^{-1/2}\right) $. For
the fourth term of ((ref)),
\begin{align*}
\sup_{i}\left\Vert \frac{\lambda_{T}^{2}}{n}\sum_{j=1}^{n}\sigma
_{j}\boldsymbol{\gamma}_{j}\mathbf{w}_{i0}^{\prime}\left( \frac{1}{T}
\sum_{t=1}^{T}\boldsymbol{\varepsilon}_{\circ t}\boldsymbol{\varepsilon
}_{\circ t}^{\prime}\right) \mathbf{w}_{j0}\right\Vert & \leq\lambda
_{T}^{2}\left( \sup_{i}\left\Vert \mathbf{w}_{i0}\right\Vert \right) \left(
\frac{1}{n}\sum_{j=1}^{n}\sigma_{j}\left\Vert \boldsymbol{\gamma}
_{j}\right\Vert \left\Vert \mathbf{w}_{j0}\right\Vert \right) \left\Vert
\mathbf{V}_{\varepsilon T}\right\Vert \\
& =O_{p}\left( \frac{1}{T}\right) .
\end{align*}
Using the above results in ((ref)) we now have
\[
\sup_{i}\left\Vert \frac{1}{nT}\sum_{t=1}^{T}\sum_{j=1}^{n}\sigma
_{j}\boldsymbol{\gamma}_{j}\varepsilon_{it}\left( \lambda_{T}\right)
\varepsilon_{jt}\left( \lambda_{T}\right) \right\Vert \leq\sup_{i}\left\Vert
\frac{1}{nT}\sum_{t=1}^{T}\sum_{j=1}^{n}\sigma_{j}\boldsymbol{\gamma}
_{j}\varepsilon_{it}\varepsilon_{jt}\right\Vert +O_{p}\left( \sqrt{\frac
{\ln\left( n\right) }{nT}}\right) ,
\]
and ((ref)) follows using ((ref)) to establish the
order of the first term of the above. To establish (ref)) note
that by definition of $\boldsymbol{\hat{\gamma}}_{i}$,
\begin{align*}
\boldsymbol{\hat{\gamma}}_{i}-\boldsymbol{\gamma}_{i} & =\left(
\mathbf{\hat{F}}^{\prime}\mathbf{\hat{F}}\right) ^{-1}\mathbf{\hat{F}
}^{\prime}\left( \mathbf{F}\boldsymbol{\gamma}_{i}+\sigma_{i}
\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) \right)
-\boldsymbol{\gamma}_{i}=\left( \mathbf{\hat{F}}^{\prime}\mathbf{\hat{F}
}\right) ^{-1}\mathbf{\hat{F}}^{\prime}\left[ \left( \mathbf{F-\hat{F}
}\right) \boldsymbol{\gamma}_{i}+\mathbf{\hat{F}}\boldsymbol{\gamma}
_{i}+\sigma_{i}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right)
\right] -\boldsymbol{\gamma}_{i}\\
& =\left( \frac{\mathbf{\hat{F}}^{\prime}\mathbf{\hat{F}}}{T}\right)
^{-1}\left( \frac{\mathbf{\hat{F}}-\mathbf{F+F}}{T}\right) ^{\prime}\left[
\left( \mathbf{F-\hat{F}}\right) \boldsymbol{\gamma}_{i}+\sigma
_{i}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) \right] \\
& =\left( \frac{\mathbf{\hat{F}}^{\prime}\mathbf{\hat{F}}}{T}\right)
^{-1}\left[ \frac{\left( \mathbf{\hat{F}}-\mathbf{F}\right) ^{\prime
}\left( \mathbf{F-\hat{F}}\right) }{T}\right] \boldsymbol{\gamma}
_{i}+\left( \frac{\mathbf{\hat{F}}^{\prime}\mathbf{\hat{F}}}{T}\right)
^{-1}\left[ \frac{\mathbf{F}^{\prime}\left( \mathbf{F-\hat{F}}\right) }
{T}\right] \boldsymbol{\gamma}_{i}\\
& +\left( \frac{\mathbf{\hat{F}}^{\prime}\mathbf{\hat{F}}}{T}\right)
^{-1}\left[ \frac{\sigma_{i}\left( \mathbf{\hat{F}}-\mathbf{F}\right)
^{\prime}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) }
{T}\right] +\left( \frac{\mathbf{\hat{F}}^{\prime}\mathbf{\hat{F}}}
{T}\right) ^{-1}\left( \frac{\sigma_{i}\mathbf{F}^{\prime}
\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) }{T}\right)
=\sum_{j=1}^{4}\mathbf{a}_{j,iT},
\end{align*}
and
\begin{equation}
\left\Vert \boldsymbol{\hat{\gamma}}_{i}-\boldsymbol{\gamma}_{i}\right\Vert
\leq\sum_{j=1}^{4}\left\Vert \mathbf{a}_{j,iT}\right\Vert .
\end{equation}
Firstly we have
\begin{align*}
\left\Vert \mathbf{a}_{1,iT}\right\Vert & \leq\left\Vert \left(
\frac{\mathbf{\hat{F}}^{\prime}\mathbf{\hat{F}}}{T}\right) ^{-1}\right\Vert
\left\Vert \left[ \frac{\left( \mathbf{\hat{F}}-\mathbf{F}\right) ^{\prime
}\left( \mathbf{F-\hat{F}}\right) }{T}\right] \right\Vert \left\Vert
\boldsymbol{\gamma}_{i}\right\Vert \\
& \leq\left\Vert \left( \frac{\mathbf{\hat{F}}^{\prime}\mathbf{\hat{F}}}
{T}\right) ^{-1}\right\Vert \left( \frac{\left\Vert \mathbf{\hat{F}
}-\mathbf{F}\right\Vert ^{2}}{T}\right) \left\Vert \boldsymbol{\gamma}
_{i}\right\Vert ,
\end{align*}
which implies
\[
\sup_{i}\left\Vert \mathbf{a}_{1,iT}\right\Vert \leq\left\Vert \left(
\frac{\mathbf{\hat{F}}^{\prime}\mathbf{\hat{F}}}{T}\right) ^{-1}\right\Vert
\left( \frac{\left\Vert \mathbf{\hat{F}}-\mathbf{F}\right\Vert ^{2}}
{T}\right) \sup_{i}\left\Vert \boldsymbol{\gamma}_{i}\right\Vert .
\]
Using ((ref)) in Lemma (ref) and ((ref)) in Lemma (ref), we
note that
\begin{equation}
\frac{\mathbf{\hat{F}}^{^{\prime}}\mathbf{\hat{F}}}{T}=O_{p}(1),
\left( \frac{\mathbf{\hat{F}}^{^{\prime}}\mathbf{\hat{F}}}{T}\right)
^{-1}=O_{p}(1), and T^{-1}\left\Vert \mathbf{\hat{F}}-\mathbf{F}
\right\Vert _{F}^{2}=O_{p}\left( \frac{1}{\delta_{nT}^{2}}\right) .
\end{equation}
Using this result and $\sup_{i}\left\Vert \boldsymbol{\gamma}_{i}\right\Vert
<C$, we obtain
\begin{equation}
\sup_{i}\left\Vert \mathbf{a}_{1,iT}\right\Vert =O_{p}\left( \frac{1}
{\delta_{nT}^{2}}\right) .
\end{equation}
Similarly, $\left\Vert \mathbf{a}_{2,iT}\right\Vert \leq\left\Vert \left(
\frac{\mathbf{\hat{F}}^{\prime}\mathbf{\hat{F}}}{T}\right) ^{-1}\right\Vert
\left\Vert \frac{\mathbf{F}^{\prime}\left( \mathbf{F-\hat{F}}\right) }
{T}\right\Vert \left\Vert \boldsymbol{\gamma}_{i}\right\Vert ,$ so using
((ref)) and ((ref)) it yields
\begin{equation}
\sup_{i}\left\Vert \mathbf{a}_{2,iT}\right\Vert \leq\left\Vert \left(
\frac{\mathbf{\hat{F}}^{\prime}\mathbf{\hat{F}}}{T}\right) ^{-1}\right\Vert
\left( \frac{\left\Vert \mathbf{F}^{\prime}\left( \mathbf{F-\hat{F}}\right)
\right\Vert }{T}\right) \sup_{i}\left\Vert \boldsymbol{\gamma}_{i}\right\Vert
=O_{p}\left( \frac{1}{\delta_{nT}^{2}}\right) .
\end{equation}
Regarding $\mathbf{a}_{3,iT}$, by Cauchy-Schwarz inequality we have
\[
\left\Vert \mathbf{a}_{3,iT}\right\Vert \leq\left\Vert \left( \frac
{\mathbf{\hat{F}}^{\prime}\mathbf{\hat{F}}}{T}\right) ^{-1}\right\Vert
\left\Vert \frac{\sigma_{i}\left( \mathbf{\hat{F}}-\mathbf{F}\right)
^{\prime}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) }
{T}\right\Vert \leq\left\Vert \left( \frac{\mathbf{\hat{F}}^{\prime
}\mathbf{\hat{F}}}{T}\right) ^{-1}\right\Vert \left( \frac{\left\Vert
\mathbf{\hat{F}}-\mathbf{F}\right\Vert ^{2}}{T}\right) ^{1/2}\left(
\frac{\sigma_{i}^{2}\left\Vert \boldsymbol{\varepsilon}_{i\circ}\left(
\lambda_{T}\right) \right\Vert ^{2}}{T}\right) ^{1/2},
\]
and therefore
\[
\sup_{i}\left\Vert \mathbf{a}_{3,iT}\right\Vert \leq\left( \sup_{i}\sigma
_{i}\right) \left\Vert \left( \frac{\mathbf{\hat{F}}^{\prime}\mathbf{\hat
{F}}}{T}\right) ^{-1}\right\Vert \left( \frac{\left\Vert \mathbf{\hat{F}
}-\mathbf{F}\right\Vert ^{2}}{T}\right) ^{1/2}\left[ \sup_{i}\left(
\frac{\left\Vert \boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right)
\right\Vert ^{2}}{T}\right) \right] ^{1/2}.
\]
Now using ((ref)) and ((ref)) it follows that
\begin{equation}
\sup_{i}\left\Vert \mathbf{a}_{3,iT}\right\Vert =O_{p}\left( \frac{1}
{\delta_{nT}}\right) .
\end{equation}
Next, note
\[
\left\Vert \mathbf{a}_{4,iT}\right\Vert \leq\left\Vert \left( \frac
{\mathbf{\hat{F}}^{\prime}\mathbf{\hat{F}}}{T}\right) ^{-1}\right\Vert
\left\Vert \frac{\sigma_{i}\mathbf{F}^{\prime}\boldsymbol{\varepsilon}
_{i\circ}\left( \lambda_{T}\right) }{T}\right\Vert
\]
then by ((ref)) we also have
\begin{equation}
\sup_{i}\left\Vert \mathbf{a}_{4,iT}\right\Vert \leq\left( \sup_{i}\sigma
_{i}\right) \left\Vert \left( \frac{\mathbf{\hat{F}}^{\prime}\mathbf{\hat
{F}}}{T}\right) ^{-1}\right\Vert \left( \sup_{i}\left\Vert \frac
{\mathbf{F}^{\prime}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda
_{T}\right) }{T}\right\Vert \right) =O_{p}\left( \sqrt{\frac{\ln(n)}{T}
}\right) .
\end{equation}
Hence using ((ref))-((ref)) in ((ref)) we have
\[
\sup_{i}\left\Vert \boldsymbol{\hat{\gamma}}_{i}-\boldsymbol{\gamma}
_{i}\right\Vert \leq\sum_{j=1}^{4}\sup_{i}\left\Vert \mathbf{a}_{j,iT}
\right\Vert =O_{p}\left( \sqrt{\frac{\ln(n)}{T}}\right) ,
\]
as required. Result ((ref)) follows by symmetry.
lemmaConsider $\varepsilon_{it}\left( \lambda_{T}\right)
=\varepsilon_{it}+\lambda_{T}\mathbf{w}_{i0}^{\prime}\boldsymbol{\varepsilon
}_{\circ t}$, where $\boldsymbol{\varepsilon}_{\circ t}=(\varepsilon
_{1t},\varepsilon_{2t},...,\varepsilon_{nt})^{\prime}$, $\varepsilon_{it}\sim
IID\left( 0,1\right) $ for all $i$ and $t$, $\mathbf{w}_{i0}=\left(
w_{i1},w_{i2},\ldots,w_{in}\right) ^{\prime}$, and $\mathbf{W}=\left(
w_{ij}\right) $ satisfy the bounded conditions $\left\Vert \mathbf{W}
\right\Vert _{1}=\sup_{j}\sum_{i=1}^{n}\left\vert w_{ij}\right\vert <C,$ and
$\left\Vert \mathbf{W}\right\Vert _{\infty}=\sup_{i}\sum_{j=1}^{n}\left\vert
w_{ij}\right\vert <C$. Then for all $\left\vert \lambda_{T}\right\vert <C$ we
have
\[
\sup_{j}\sum_{i=1}^{n}\left\vert E\left[ \varepsilon_{it}\left( \lambda
_{T}\right) \varepsilon_{jt}\left( \lambda_{T}\right) \right] \right\vert
<C,\text{ }\sup_{i}\sum_{j=1}^{n}\left\vert E\left[ \varepsilon_{it}\left(
\lambda_{T}\right) \varepsilon_{jt}\left( \lambda_{T}\right) \right]
\right\vert <C,
\]
and
\begin{equation}
n^{-1}\sum_{i=1}^{n}\sum_{j=1}^{n}\left\vert E\left[ \varepsilon_{it}\left(
\lambda_{T}\right) \varepsilon_{jt}\left( \lambda_{T}\right) \right]
\right\vert <C.
\end{equation}
proofLet $\boldsymbol{\varepsilon}_{\circ t}(\lambda_{T})=$
$\boldsymbol{\varepsilon}_{\circ t}+\lambda_{T}\mathbf{W}
\boldsymbol{\varepsilon}_{\circ t}$, where $\mathbf{W}^{\prime}=(\mathbf{w}
_{10},\mathbf{w}_{20},...,\mathbf{w}_{n0})$. Then
\[
\boldsymbol{\varepsilon}_{\circ t}(\lambda_{T})\boldsymbol{\varepsilon}_{\circ
t}^{\prime}(\lambda_{T})=\boldsymbol{\varepsilon}_{\circ t}
\boldsymbol{\varepsilon}_{\circ t}^{\prime}+\lambda_{T}^{2}\mathbf{W}
\boldsymbol{\varepsilon}_{\circ t}\boldsymbol{\varepsilon}_{\circ t}^{\prime
}\mathbf{W}^{\prime}+\lambda_{T}\mathbf{W}\boldsymbol{\varepsilon}_{\circ
t}\boldsymbol{\varepsilon}_{\circ t}^{\prime}+\lambda_{T}
\boldsymbol{\varepsilon}_{\circ t}\boldsymbol{\varepsilon}_{\circ t}^{\prime
}\mathbf{W}^{\prime},
\]
and
\[
\mathbf{V}_{\varepsilon}(\lambda_{T})=E\left[ \boldsymbol{\varepsilon}_{\circ
t}(\lambda_{T})\boldsymbol{\varepsilon}_{\circ t}^{\prime}(\lambda
_{T})\right] =\mathbf{I}_{n}+\lambda_{T}\left( \mathbf{W+W}^{\prime}\right)
+\lambda_{T}^{2}\mathbf{WW}^{\prime}.
\]
Consider the maximum absolute column sum norm of $\mathbf{V}_{\varepsilon
}(\lambda_{T})$ and note that
\[
\left\Vert \mathbf{V}_{\varepsilon}(\lambda_{T})\right\Vert _{1}=\sup_{j}
\sum_{i=1}^{n}\left\vert E\left[ \varepsilon_{it}\left( \lambda_{T}\right)
\varepsilon_{jt}\left( \lambda_{T}\right) \right] \right\vert <1+\left\vert
\lambda_{T}\right\vert \left( \left\Vert \mathbf{W}\right\Vert _{1}
+\left\Vert \mathbf{W}\right\Vert _{\infty}\right) +\lambda_{T}^{2}\left\Vert
\mathbf{W}\right\Vert _{1}\left\Vert \mathbf{W}\right\Vert _{\infty}<C\text{.
}
\]
Similarly for the maximum absolute row sum norm of $\mathbf{V}_{\varepsilon
}(\lambda_{T})$
\[
\left\Vert \mathbf{V}_{\varepsilon}(\lambda_{T})\right\Vert _{\infty}=\sup
_{i}\sum_{j=1}^{n}\left\vert E\left[ \varepsilon_{it}\left( \lambda
_{T}\right) \varepsilon_{jt}\left( \lambda_{T}\right) \right] \right\vert
<1+\left\vert \lambda_{T}\right\vert \left( \left\Vert \mathbf{W}\right\Vert
_{\infty}+\left\Vert \mathbf{W}\right\Vert _{1}\right) +\lambda_{T}
^{2}\left\Vert \mathbf{W}\right\Vert _{\infty}\left\Vert \mathbf{W}\right\Vert
_{1}<C\text{,}
\]
and result ((ref)) follows.
lemmaConsider the latent factor model given by ((ref)) and
((ref)). Suppose that Assumptions (ref)-(ref) hold
and $\left( n,T\right) \rightarrow\infty$, such that $n/T\rightarrow
\kappa\,,$ for $0<\kappa<\infty$. Then for the estimator of factors, we have
\begin{equation}
\sup_{i}\left\Vert \frac{\left( \mathbf{\hat{F}-F}\right) ^{\prime
}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) }{T}\right\Vert
=O_{p}\left( \sqrt{\frac{\ln\left( n\right) }{nT}}\right) .
\end{equation}
proofBy (A.1) of Bai (2003) we note that
\[
\mathbf{\hat{f}}_{t}-\mathbf{f}_{t}=\frac{1}{T}\sum_{t^{\prime}=1}
^{T}\mathbf{\hat{f}}_{t^{\prime}}\eta_{tt^{\prime}}+\frac{1}{T}\sum
_{t^{\prime}=1}^{T}\mathbf{\hat{f}}_{t^{\prime}}\zeta_{tt^{\prime}}+\frac
{1}{T}\sum_{t^{\prime}=1}^{T}\mathbf{\hat{f}}_{t^{\prime}}\varkappa
_{tt^{\prime}}+\frac{1}{T}\sum_{t^{\prime}=1}^{T}\mathbf{\hat{f}}_{t^{\prime}
}\xi_{tt^{\prime}}
\]
where $\eta_{tt^{\prime}}=n^{-1}\sum_{i=1}^{n}\sigma_{i}^{2}E\left(
\varepsilon_{it}\left( \lambda_{T}\right) \varepsilon_{it^{\prime}}\left(
\lambda_{T}\right) \right) $, $\zeta_{tt^{\prime}}=n^{-1}\sum_{i=1}
^{n}\sigma_{i}^{2}\varepsilon_{it}^{2}\left( \lambda_{T}\right)
-\eta_{tt^{\prime}}$, $\varkappa_{tt^{\prime}}=n^{-1}\sum_{i=1}^{n}\sigma
_{i}\mathbf{f}_{t^{\prime}}^{\prime}\boldsymbol{\gamma}_{i}^{\prime
}\varepsilon_{it}\left( \lambda_{T}\right) $, and $\xi_{tt^{\prime}}
=n^{-1}\sum_{i=1}^{n}\sigma_{i}\mathbf{f}_{t}^{\prime}\boldsymbol{\gamma}
_{i}^{\prime}\varepsilon_{it^{\prime}}\left( \lambda_{T}\right) $. Hence
\begin{align*}
\frac{1}{T}\left( \mathbf{\hat{F}-F}\right) ^{\prime}\boldsymbol{\varepsilon
}_{i\circ}\left( \lambda_{T}\right) & =\frac{1}{T}\sum_{t=1}^{T}\left(
\mathbf{\hat{f}}_{t}-\mathbf{f}_{t}\right) \varepsilon_{it}\left(
\lambda_{T}\right) \\
& =\frac{1}{T^{2}}\sum_{t=1}^{T}\sum_{t^{\prime}=1}^{T}\mathbf{\hat{f}
}_{t^{\prime}}\eta_{tt^{\prime}}\varepsilon_{it}\left( \lambda_{T}\right)
+\frac{1}{T^{2}}\sum_{t=1}^{T}\sum_{t^{\prime}=1}^{T}\mathbf{\hat{f}
}_{t^{\prime}}\zeta_{tt^{\prime}}\varepsilon_{it}\left( \lambda_{T}\right) \\
& +\frac{1}{T^{2}}\sum_{t=1}^{T}\sum_{t^{\prime}=1}^{T}\mathbf{\hat{f}
}_{t^{\prime}}\varkappa_{tt^{\prime}}\varepsilon_{it}\left( \lambda
_{T}\right) +\frac{1}{T^{2}}\sum_{t=1}^{T}\sum_{t^{\prime}=1}^{T}
\mathbf{\hat{f}}_{t^{\prime}}\xi_{tt^{\prime}}\varepsilon_{it}\left(
\lambda_{T}\right) \\
& =\sum_{j=1}^{4}\mathbf{b}_{j,iT},
\end{align*}
and
\begin{equation}
\sup_{i}\left\Vert \frac{\left( \mathbf{\hat{F}-F}\right) ^{\prime
}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) }{T}\right\Vert
\leq\sum_{j=1}^{4}\sup_{i}\left\Vert \mathbf{b}_{j,iT}\right\Vert .
\end{equation}
Firstly, consider $\mathbf{b}_{1,iT}$ and note that
\begin{align*}
\mathbf{b}_{1,iT} & =\frac{1}{T^{2}}\sum_{t=1}^{T}\sum_{t^{\prime}=1}
^{T}\left( \mathbf{\hat{f}}_{t^{\prime}}-\mathbf{f}_{t^{\prime}}\right)
\eta_{tt^{\prime}}\varepsilon_{it}\left( \lambda_{T}\right) +\frac{1}{T^{2}
}\sum_{t=1}^{T}\sum_{t^{\prime}=1}^{T}\mathbf{f}_{t^{\prime}}\eta_{tt^{\prime
}}\varepsilon_{it}\left( \lambda_{T}\right) =\mathbf{b}_{1,1,iT}
+\mathbf{b}_{1,2,iT},
\end{align*}
so $\left\Vert \mathbf{b}_{1,iT}\right\Vert \leq\left\Vert \mathbf{b}
_{1,1,iT}\right\Vert +\left\Vert \mathbf{b}_{1,2,iT}\right\Vert . $ Note for
the first term on the right hand side,
\begin{align*}
\left\Vert \mathbf{b}_{1,1,iT}\right\Vert & =\left\Vert \frac{1}{T}
\sum_{t^{\prime}=1}^{T}\left( \mathbf{\hat{f}}_{t^{\prime}}-\mathbf{f}
_{t^{\prime}}\right) \left( \frac{1}{T}\sum_{t=1}^{T}\eta_{tt^{\prime}
}\varepsilon_{it}\left( \lambda_{T}\right) \right) \right\Vert \\
& \leq\left( \frac{1}{T}\sum_{t^{\prime}=1}^{T}\left\Vert \mathbf{\hat{f}
}_{t^{\prime}}-\mathbf{f}_{t^{\prime}}\right\Vert ^{2}\right) ^{1/2}\left[
\frac{1}{T}\sum_{t^{\prime}=1}^{T}\left( \frac{1}{T}\sum_{t=1}^{T}
\eta_{tt^{\prime}}\varepsilon_{it}\left( \lambda_{T}\right) \right)
^{2}\right] ^{1/2}\\
& \leq\frac{1}{\sqrt{T}}\left( \frac{1}{T}\sum_{t^{\prime}=1}^{T}\left\Vert
\mathbf{\hat{f}}_{t^{\prime}}-\mathbf{f}_{t^{\prime}}\right\Vert ^{2}\right)
^{1/2}\left( \frac{1}{T}\sum_{t^{\prime}=1}^{T}\sum_{t=1}^{T}\eta
_{tt^{\prime}}^{2}\right) ^{1/2}\left( \frac{1}{T}\sum_{t=1}^{T}
\varepsilon_{it}^{2}\left( \lambda_{T}\right) \right) ^{1/2},
\end{align*}
where the second and third lines hold by Cauchy-Schwarz inequality. By
definition of $\eta_{tt^{\prime}}$ and serial independence of $\varepsilon
_{it}\left( \lambda_{T}\right) ,$ $\eta_{tt^{\prime}}=\bar{\sigma}_{n}^{2}$
for $t=t^{\prime}$ but $0$ otherwise, where $\bar{\sigma}_{n}^{2}=n^{-1}
\sum_{i=1}^{n}\sigma_{i}^{2}E\left( \varepsilon_{it}^{2}\left( \lambda
_{T}\right) \right) $. Under assumption on weight $\left\{ w_{ij}\right\}
$, $E\left( \varepsilon_{it}^{2}\left( \lambda_{T}\right) \right)
=1+\lambda_{T}^{2}\left( \sum_{j=1}^{n}w_{ij}^{2}\right) <C $, so it follows that
\begin{equation}
\frac{1}{T}\sum_{t^{\prime}=1}^{T}\sum_{t=1}^{T}\eta_{tt^{\prime}}^{2}=\left(
\bar{\sigma}_{n}^{2}\right) ^{2}<C.
\end{equation}
Given results ((ref)), ((ref)) and ((ref)), we
further obtain
\begin{align*}
\sup_{i}\left\Vert \mathbf{b}_{1,1,iT}\right\Vert & \leq\frac{1}{\sqrt{T}
}\left( \frac{1}{T}\sum_{t^{\prime}=1}^{T}\left\Vert \mathbf{\hat{f}
}_{t^{\prime}}-\mathbf{f}_{t^{\prime}}\right\Vert ^{2}\right) ^{1/2}\left(
\frac{1}{T}\sum_{t^{\prime}=1}^{T}\sum_{t=1}^{T}\eta_{tt^{\prime}}^{2}\right)
^{1/2}\left( \sup_{i}\frac{1}{T}\sum_{t=1}^{T}\varepsilon_{it}^{2}\left(
\lambda_{T}\right) \right) ^{1/2}\\
& =O_{p}\left( T^{-1/2}\delta_{nT}^{-1}\right) .
\end{align*}
Now consider $\mathbf{b}_{1,2,iT}$. Using properties of $\eta_{tt^{\prime}}$
we have
\[
\frac{1}{T^{2}}\sum_{t=1}^{T}\sum_{t^{\prime}=1}^{T}\mathbf{f}_{t^{\prime}
}\eta_{tt^{\prime}}\varepsilon_{it}\left( \lambda_{T}\right) =\frac{1}
{T^{2}}\sum_{t=1}^{T}\mathbf{f}_{t}\eta_{tt}\varepsilon_{it}\left(
\lambda_{T}\right) +\frac{1}{T^{2}}\sum_{t=1}^{T}\sum_{t^{\prime}\neq
t^{\prime}}^{T}\mathbf{f}_{t^{\prime}}\eta_{tt^{\prime}}\varepsilon
_{it}\left( \lambda_{T}\right) =\bar{\sigma}_{n}^{2}\left( \frac{1}{T^{2}
}\sum_{t=1}^{T}\mathbf{f}_{t}\varepsilon_{it}\left( \lambda_{T}\right)
\right) ,
\]
and therefore
\[
\left\Vert \mathbf{b}_{1,2,iT}\right\Vert =\left\Vert \frac{1}{T^{2}}
\sum_{t=1}^{T}\sum_{t^{\prime}=1}^{T}\mathbf{f}_{t^{\prime}}\eta_{tt^{\prime}
}\varepsilon_{it}\left( \lambda_{T}\right) \right\Vert \leq\frac{1}{T}
\bar{\sigma}_{n}^{2}\left( \left\Vert \frac{1}{T}\sum_{t=1}^{T}\mathbf{f}
_{t}\varepsilon_{it}\left( \lambda_{T}\right) \right\Vert \right) .
\]
Then using ((ref)) we have
\[
\sup_{i}\left\Vert \mathbf{b}_{1,2,iT}\right\Vert \leq\frac{1}{T}\bar{\sigma
}_{n}^{2}\sup_{i}\left\Vert \frac{1}{T}\sum_{t=1}^{T}\mathbf{f}_{t}
\varepsilon_{it}\left( \lambda_{T}\right) \right\Vert =O_{p}\left( \frac
{1}{T}\sqrt{\frac{\ln\left( n\right) }{T}}\right) .
\]
Hence, combining the probability orders of $\sup_{i}\left\Vert \mathbf{b}
_{1,1,iT}\right\Vert $ and $\sup_{i}\left\Vert \mathbf{b}_{1,2,iT}\right\Vert
$, we have
\begin{equation}
\sup_{i}\left\Vert \mathbf{b}_{1,nT}\right\Vert \leq\sup_{i}\left\Vert
\mathbf{b}_{1,1,iT}\right\Vert +\sup_{i}\left\Vert \mathbf{b}_{1,2,iT}
\right\Vert =O_{p}\left( \frac{1}{T^{1/2}\delta_{nT}}\right) .
\end{equation}
Next, consider $\mathbf{b}_{2,iT}$ in ((ref)), which can be written
as
\begin{align*}
\mathbf{b}_{2,iT} & =\frac{1}{T^{2}}\sum_{t=1}^{T}\sum_{t^{\prime}=1}
^{T}\left( \mathbf{\hat{f}}_{t^{\prime}}-\mathbf{f}_{t^{\prime}}\right)
\zeta_{tt^{\prime}}\varepsilon_{it}\left( \lambda_{T}\right) +\frac{1}
{T^{2}}\sum_{t=1}^{T}\sum_{t^{\prime}=1}^{T}\mathbf{f}_{t^{\prime}}
\zeta_{tt^{\prime}}\varepsilon_{it}\left( \lambda_{T}\right) =\mathbf{b}
_{2,1,iT}+\mathbf{b}_{2,2,iT}.
\end{align*}
For the first term, we can apply Cauchy-Schwarz inequality to obtain
\begin{align*}
\left\Vert \mathbf{b}_{2,1,iT}\right\Vert & =\left\Vert \frac{1}{T^{2}}
\sum_{t^{\prime}=1}^{T}\left( \mathbf{\hat{f}}_{t^{\prime}}-\mathbf{f}
_{t^{\prime}}\right) \left( \sum_{t=1}^{T}\zeta_{tt^{\prime}}\varepsilon
_{it}\left( \lambda_{T}\right) \right) \right\Vert \\
& \leq\left( \frac{1}{T}\sum_{t^{\prime}=1}^{T}\left\Vert \mathbf{\hat{f}
}_{t^{\prime}}-\mathbf{f}_{t^{\prime}}\right\Vert ^{2}\right) ^{1/2}\left[
\frac{1}{T^{3}}\sum_{t^{\prime}=1}^{T}\left( \sum_{t=1}^{T}\zeta_{tt^{\prime
}}\varepsilon_{it}\left( \lambda_{T}\right) \right) ^{2}\right] ^{1/2}\\
& \leq\left( \frac{1}{T}\sum_{t^{\prime}=1}^{T}\left\Vert \mathbf{\hat{f}
}_{t^{\prime}}-\mathbf{f}_{t^{\prime}}\right\Vert ^{2}\right) ^{1/2}\left(
\frac{1}{T^{2}}\sum_{t^{\prime}=1}^{T}\sum_{t=1}^{T}\zeta_{tt^{\prime}}
^{2}\right) ^{1/2}\left( \frac{1}{T}\sum_{t=1}^{T}\varepsilon_{it}
^{2}\left( \lambda_{T}\right) \right) ^{1/2}.
\end{align*}
Since
\[
\zeta_{tt^{\prime}}=n^{-1}\sum_{i=1}^{n}\sigma_{i}^{2}\varepsilon_{it}
^{2}\left( \lambda_{T}\right) -\eta_{tt^{\prime}}=\frac{1}{n}\sum_{j=1}
^{n}\sigma_{j}^{2}\left[ \varepsilon_{jt}\left( \lambda_{T}\right)
\varepsilon_{jt^{\prime}}\left( \lambda_{T}\right) -E\left( \varepsilon
_{jt}\left( \lambda_{T}\right) \varepsilon_{jt^{\prime}}\left( \lambda
_{T}\right) \right) \right] ,
\]
then it follows that
\begin{align}
E\left( \frac{1}{T^{2}}\sum_{t^{\prime}=1}^{T}\sum_{t=1}^{T}\zeta
_{tt^{\prime}}^{2}\right) & =\frac{1}{T^{2}}\sum_{t^{\prime}=1}^{T}
\sum_{t=1}^{T}E\left( \zeta_{tt^{\prime}}^{2}\right) \nonumber\\
& =\frac{1}{T^{2}n^{2}}\sum_{t^{\prime}=1}^{T}\sum_{t=1}^{T}E\left(
\sum_{j=1}^{n}\sigma_{j}^{2}\left[ \varepsilon_{jt}\left( \lambda
_{T}\right) \varepsilon_{jt^{\prime}}\left( \lambda_{T}\right) -E\left(
\varepsilon_{jt}\left( \lambda_{T}\right) \varepsilon_{jt^{\prime}}\left(
\lambda_{T}\right) \right) \right] \right) ^{2}\nonumber\\
& =\frac{1}{T^{2}n^{2}}\sum_{t=1}^{T}\sum_{t^{\prime}=1}^{T}\sum_{j=1}
^{n}\sum_{j^{\prime}=1}^{n}\sigma_{j}^{2}\sigma_{j^{\prime}}^{2}E\left(
\varepsilon_{jt}\left( \lambda_{T}\right) \varepsilon_{jt^{\prime}}\left(
\lambda_{T}\right) \varepsilon_{j^{\prime}t}\left( \lambda_{T}\right)
\varepsilon_{j^{\prime}t^{\prime}}\left( \lambda_{T}\right) \right)
\nonumber\\
& -\frac{1}{T^{2}n^{2}}\sum_{t=1}^{T}\sum_{t^{\prime}=1}^{T}\sum_{j=1}
^{n}\sum_{j^{\prime}=1}^{n}\sigma_{j}^{2}\sigma_{j^{\prime}}^{2}E\left(
\varepsilon_{jt}\left( \lambda_{T}\right) \varepsilon_{jt^{\prime}}\left(
\lambda_{T}\right) \right) E\left( \varepsilon_{j^{\prime}t}\left(
\lambda_{T}\right) \varepsilon_{j^{\prime}t^{\prime}}\left( \lambda
_{T}\right) \right) .
\end{align}
For the first term of ((ref)), given the serial independence of
$\varepsilon_{it}\left( \lambda_{T}\right) $ and note $E\left(
\varepsilon_{it}\left( \lambda_{T}\right) \right) =E\left( \varepsilon
_{it}+\lambda_{T}\mathbf{w}_{i0}^{\prime}\boldsymbol{\varepsilon}_{\circ
t}\right) =0$, some algebra yields
\begin{align}
& \frac{1}{T^{2}n^{2}}\sum_{t=1}^{T}\sum_{t^{\prime}=1}^{T}\sum_{j=1}^{n}
\sum_{j^{\prime}=1}^{n}\sigma_{j}^{2}\sigma_{j^{\prime}}^{2}E\left(
\varepsilon_{jt}\left( \lambda_{T}\right) \varepsilon_{jt^{\prime}}\left(
\lambda_{T}\right) \varepsilon_{j^{\prime}t}\left( \lambda_{T}\right)
\varepsilon_{j^{\prime}t^{\prime}}\left( \lambda_{T}\right) \right)
\nonumber\\
& =\frac{1}{T^{2}n^{2}}\sum_{t=1}^{T}\sum_{j=1}^{n}\sigma_{j}^{4}E\left(
\varepsilon_{jt}^{4}\left( \lambda_{T}\right) \right) +\frac{1}{T^{2}n^{2}
}\sum_{t=1}^{T}\sum_{t^{\prime}\neq t}^{T}\sum_{j=1}^{n}\sigma_{j}^{4}E\left(
\varepsilon_{jt}^{2}\left( \lambda_{T}\right) \right) E\left(
\varepsilon_{jt^{\prime}}^{2}\left( \lambda_{T}\right) \right) +\nonumber\\
& \frac{1}{T^{2}n^{2}}\sum_{t=1}^{T}\sum_{j=1}^{n}\sum_{j^{\prime}\neq j}
^{n}\sigma_{j}^{2}\sigma_{j^{\prime}}^{2}E\left( \varepsilon_{jt}^{2}\left(
\lambda_{T}\right) \varepsilon_{j^{\prime}t}^{2}\left( \lambda_{T}\right)
\right) +\nonumber\\
& \frac{1}{T^{2}n^{2}}\sum_{t=1}^{T}\sum_{t^{\prime}\neq t}^{T}\sum_{j=1}
^{n}\sum_{j^{\prime}\neq j}^{n}\sigma_{j}^{2}\sigma_{j^{\prime}}^{2}E\left(
\varepsilon_{jt}\left( \lambda_{T}\right) \varepsilon_{jt^{\prime}}\left(
\lambda_{T}\right) \varepsilon_{j^{\prime}t}\left( \lambda_{T}\right)
\varepsilon_{j^{\prime}t^{\prime}}\left( \lambda_{T}\right) \right) .
\end{align}
To show the order of ((ref)), we note $\varepsilon_{it}=\mathbf{e}
_{i}^{\prime}\boldsymbol{\varepsilon}_{\circ t}$ where $\mathbf{e}_{i}$ is an
$n\times1$ selection vector with $1$ on its $i^{th}$ element and zero
elsewhere, then
\begin{align*}
\varepsilon_{jt}^{4}\left( \lambda_{T}\right) & =\left( \varepsilon
_{it}+\lambda_{T}\mathbf{w}_{i0}^{\prime}\boldsymbol{\varepsilon}_{\circ
t}\right) ^{4}=\left[ \left( \mathbf{e}_{i}+\lambda_{T}\mathbf{w}
_{i0}\right) ^{\prime}\boldsymbol{\varepsilon}_{\circ t}\right] ^{4}\\
& =\left[ \boldsymbol{\varepsilon}_{\circ t}^{\prime}\left( \mathbf{e}
_{i}+\lambda_{T}\mathbf{w}_{i0}\right) \left( \mathbf{e}_{i}+\lambda
_{T}\mathbf{w}_{i0}\right) ^{\prime}\boldsymbol{\varepsilon}_{\circ
t}\right] ^{2}=\left( \boldsymbol{\varepsilon}_{\circ t}^{\prime}
\mathbf{A}_{i}\boldsymbol{\varepsilon}_{\circ t}\right) ^{2},
\end{align*}
where $\mathbf{A}_{i}=\left( \mathbf{e}_{i}+\lambda_{T}\mathbf{w}
_{i0}\right) \left( \mathbf{e}_{i}+\lambda_{T}\mathbf{w}_{i0}\right)
^{\prime}.$ Using result (S.7) of Lemma 6 in Pesaran and Yamagata (2024) and
noting $w_{ii}=0$ for all $i$, we have
\begin{align}
E\left( \varepsilon_{jt}^{4}\left( \lambda_{T}\right) \right) &
=E\left( \boldsymbol{\varepsilon}_{\circ t}^{\prime}\mathbf{A}_{i}
\boldsymbol{\varepsilon}_{\circ t}\right) ^{2} =\kappa_{2}\mathrm{tr}\left[
\left( \mathbf{A}_{i}\odot\mathbf{A}_{i}\right) \right] +\left[
\mathrm{tr}\left( \mathbf{A}_{i}\right) \right] ^{2}+2\mathrm{tr}\left(
\mathbf{A}_{i}^{2}\right) ,\nonumber
\end{align}
where $\kappa_{2}=E(\varepsilon_{it}^{4})-3$. Also, by condition
((ref)) we have
\[
\sum_{j=1}^{n}w_{ij}^{2}\leq\left( \sum_{j=1}^{n}\left\vert w_{ij}\right\vert
\right) ^{2}<C,\sum_{j=1}^{n}w_{ij}^{4}\leq\left( \sum_{j=1}^{n}\left\vert
w_{ij}\right\vert \right) ^{4}<C,
\]
using which yields
\begin{equation}
E\left( \varepsilon_{jt}^{4}\left( \lambda_{T}\right) \right) =\kappa
_{2}\left( 1+\lambda_{T}^{4}\sum_{j=1}^{n}w_{ij}^{4}\right) +3\left(
1+\lambda_{T}^{2}\sum_{j=1}^{n}w_{ij}^{2}\right) =O\left( 1\right) .
\end{equation}
It therefore follows that $E\left( \varepsilon_{jt}^{2}\left( \lambda
_{T}\right) \right) $ and $E\left( \varepsilon_{jt}^{4}\left( \lambda
_{T}\right) \right) $ are bounded so that the first two terms of
((ref)) satisfy
\begin{align*}
\frac{1}{T^{2}n^{2}}\sum_{t=1}^{T}\sum_{j=1}^{n}\sigma_{j}^{4}E\left(
\varepsilon_{jt}^{4}\left( \lambda_{T}\right) \right) & =O\left(
\frac{1}{nT}\right) ,\\
\frac{1}{T^{2}n^{2}}\sum_{t=1}^{T}\sum_{t^{\prime}\neq t}^{T}\sum_{j=1}
^{n}\sigma_{j}^{4}E\left( \varepsilon_{jt}^{2}\left( \lambda_{T}\right)
\right) E\left( \varepsilon_{jt^{\prime}}^{2}\left( \lambda_{T}\right)
\right) & =O\left( \frac{1}{n}\right) .
\end{align*}
In addition, the third term of ((ref)) can be bounded using
Cauchy-Schwarz inequality,
\begin{align*}
& \frac{1}{T^{2}n^{2}}\sum_{t=1}^{T}\sum_{j=1}^{n}\sum_{j^{\prime}\neq j}
^{n}\sigma_{j}^{2}\sigma_{j^{\prime}}^{2}E\left( \varepsilon_{jt}^{2}\left(
\lambda_{T}\right) \varepsilon_{j^{\prime}t}^{2}\left( \lambda_{T}\right)
\right) \\
& =\frac{1}{T^{2}n^{2}}\sum_{t=1}^{T}\sum_{j=1}^{n}\sum_{j^{\prime}\neq
j}^{n}\sigma_{j}^{2}\sigma_{j^{\prime}}^{2}\left[ E\left( \varepsilon
_{jt}^{4}\left( \lambda_{T}\right) \right) \right] ^{1/2}\times\left[
E\left( \varepsilon_{j^{\prime}t}^{4}\left( \lambda_{T}\right) \right)
\right] ^{1/2}=O\left( \frac{1}{T}\right) .
\end{align*}
The fourth term of ((ref)) can be expanded based on the serial
independence of $\varepsilon_{it}\left( \lambda_{T}\right) $, so that
\begin{align*}
& \frac{1}{T^{2}n^{2}}\sum_{t=1}^{T}\sum_{t^{\prime}\neq t}^{T}\sum_{j=1}
^{n}\sum_{j^{\prime}\neq j}^{n}\sigma_{j}^{2}\sigma_{j^{\prime}}^{2}E\left(
\varepsilon_{jt}\left( \lambda_{T}\right) \varepsilon_{jt^{\prime}}\left(
\lambda_{T}\right) \varepsilon_{j^{\prime}t}\left( \lambda_{T}\right)
\varepsilon_{j^{\prime}t^{\prime}}\left( \lambda_{T}\right) \right) \\
& =\frac{1}{T^{2}n^{2}}\sum_{j=1}^{n}\sum_{j^{\prime}\neq j}^{n}\sigma
_{j}^{2}\sigma_{j^{\prime}}^{2}\left[ \sum_{t=1}^{T}E\left( \varepsilon
_{jt}\left( \lambda_{T}\right) \varepsilon_{j^{\prime}t}\left( \lambda
_{T}\right) \right) \right] \left[ \sum_{t^{\prime}\neq t}^{T}E\left(
\varepsilon_{jt^{\prime}}\left( \lambda_{T}\right) \varepsilon_{j^{\prime
}t^{\prime}}\left( \lambda_{T}\right) \right) \right] .
\end{align*}
Also note by definition of $\varepsilon_{jt}\left( \lambda_{T}\right) $, for
$j\neq j^{\prime}$,
\[
E\left( \varepsilon_{jt}\left( \lambda_{T}\right) \varepsilon_{j^{\prime}
t}\left( \lambda_{T}\right) \right) =\lambda_{T}g_{jj^{\prime}}+\lambda
_{T}^{2}\mathbf{w}_{j0}^{\prime}\mathbf{w}_{j^{\prime}0},
\]
where $g_{jj^{\prime}}=w_{jj^{\prime}}+w_{j^{\prime}j}$ and $\mathbf{w}
_{j0}=\left( w_{j1},w_{j2},\ldots,w_{jn}\right) ^{\prime}$, so that
\[
\left[ \sum_{t=1}^{T}E\left( \varepsilon_{jt}\left( \lambda_{T}\right)
\varepsilon_{j^{\prime}t}\left( \lambda_{T}\right) \right) \right] \left[
\sum_{t^{\prime}\neq t}^{T}E\left( \varepsilon_{jt^{\prime}}\left(
\lambda_{T}\right) \varepsilon_{j^{\prime}t^{\prime}}\left( \lambda
_{T}\right) \right) \right] =T\left( T-1\right) \left( \lambda
_{T}g_{jj^{\prime}}+\lambda_{T}^{2}\mathbf{w}_{j0}^{\prime}\mathbf{w}
_{j^{\prime}0}\right) ^{2}.
\]
Then it follows
\begin{align*}
& \frac{1}{T^{2}n^{2}}\sum_{t=1}^{T}\sum_{t^{\prime}\neq t}^{T}\sum_{j=1}
^{n}\sum_{j^{\prime}\neq j}^{n}\sigma_{j}^{2}\sigma_{j^{\prime}}^{2}E\left(
\varepsilon_{jt}\left( \lambda_{T}\right) \varepsilon_{jt^{\prime}}\left(
\lambda_{T}\right) \varepsilon_{j^{\prime}t}\left( \lambda_{T}\right)
\varepsilon_{j^{\prime}t^{\prime}}\left( \lambda_{T}\right) \right) \\
& =\frac{T\left( T-1\right) }{T^{2}n^{2}}\sum_{j=1}^{n}\sum_{j^{\prime}\neq
j}^{n}\sigma_{j}^{2}\sigma_{j^{\prime}}^{2}\left( \lambda_{T}g_{jj^{\prime}
}+\lambda_{T}^{2}\mathbf{w}_{j0}^{\prime}\mathbf{w}_{j^{\prime}0}\right)
^{2}\\
& \leq\frac{2T\left( T-1\right) }{T^{2}n^{2}}\sum_{j=1}^{n}\sum_{j^{\prime
}\neq j}^{n}\sigma_{j}^{2}\sigma_{j^{\prime}}^{2}\left( \lambda
_{T}g_{jj^{\prime}}\right) ^{2}+\frac{2T\left( T-1\right) }{T^{2}n^{2}}
\sum_{j=1}^{n}\sum_{j^{\prime}\neq j}^{n}\sigma_{j}^{2}\sigma_{j^{\prime}}
^{2}\left( \lambda_{T}^{2}\mathbf{w}_{j0}^{\prime}\mathbf{w}_{j^{\prime}
0}\right) ^{2}.
\end{align*}
Since $\sum_{j^{\prime}=1}^{n}\left\vert w_{jj^{\prime}}\right\vert ^{2}
\leq\left( \sum_{j^{\prime}=1}^{n}\left\vert w_{jj^{\prime}}\right\vert
\right) ^{2}\leq\left\Vert \mathbf{w}_{j0}\right\Vert ^{2}<C$ \ by
((ref)), $\lambda_{T}=c_{\lambda}T^{-1/2}$ and $\left\vert
\mathbf{w}_{j0}^{\prime}\mathbf{w}_{j^{\prime}0}\right\vert \leq\left\Vert
\mathbf{w}_{j0}\right\Vert \left\Vert \mathbf{w}_{j^{\prime}0}\right\Vert <C$,
we further have
\begin{align*}
\frac{2T\left( T-1\right) }{T^{2}n^{2}}\sum_{j=1}^{n}\sum_{j^{\prime}\neq
j}^{n}\sigma_{j}^{2}\sigma_{j^{\prime}}^{2}\left( \lambda_{T}g_{jj^{\prime}
}\right) ^{2} & =\frac{2T\left( T-1\right) \lambda_{T}^{2}}{T^{2}n^{2}
}\sum_{j=1}^{n}\sum_{j^{\prime}\neq j}^{n}\sigma_{j}^{2}\sigma_{j^{\prime}
}^{2}\left( w_{jj^{\prime}}+w_{j^{\prime}j}\right) ^{2}\\
& \leq\frac{4T\left( T-1\right) \lambda_{T}^{2}}{T^{2}n^{2}}\sum_{j=1}
^{n}\sum_{j^{\prime}\neq j}^{n}\sigma_{j}^{2}\sigma_{j^{\prime}}^{2}\left(
\left\vert w_{jj^{\prime}}\right\vert ^{2}+\left\vert w_{j^{\prime}
j}\right\vert ^{2}\right) \\
& \leq\frac{4T\left( T-1\right) \lambda_{T}^{2}}{T^{2}n^{2}}\left(
\sup_{i}\sigma_{i}^{4}\right) \sum_{j=1}^{n}\sum_{j^{\prime}\neq j}
^{n}\left( \left\vert w_{jj^{\prime}}\right\vert ^{2}+\left\vert
w_{j^{\prime}j}\right\vert ^{2}\right) \\
& =O\left( \frac{1}{nT}\right) ,
\end{align*}
and
\begin{align*}
\frac{2T\left( T-1\right) }{T^{2}n^{2}}\sum_{j=1}^{n}\sum_{j^{\prime}\neq
j}^{n}\sigma_{j}^{2}\sigma_{j^{\prime}}^{2}\left( \lambda_{T}^{2}
\mathbf{w}_{j0}^{\prime}\mathbf{w}_{j^{\prime}0}\right) ^{2} &
=\frac{2T\left( T-1\right) \lambda_{T}^{4}}{T^{2}n^{2}}\sum_{j=1}^{n}
\sum_{j^{\prime}\neq j}^{n}\sigma_{j}^{2}\sigma_{j^{\prime}}^{2}\left(
\mathbf{w}_{j0}^{\prime}\mathbf{w}_{j^{\prime}0}\right) ^{2}\\
& \leq\frac{2T\left( T-1\right) \lambda_{T}^{4}}{T^{2}}\left( \sup
_{i}\sigma_{i}^{4}\right) \left( \sup_{j,j^{\prime}}\left\vert
\mathbf{w}_{j0}^{\prime}\mathbf{w}_{j^{\prime}0}\right\vert \right) ^{2}\\
& =O_{p}\left( \frac{1}{T^{2}}\right) .
\end{align*}
Overall, using ((ref)) we are able to show the first term of
((ref)) is $O\left( T^{-1}\right) $ as $n,T\rightarrow\infty$ such
that $n/T=\kappa$ where $0<\kappa<\infty$. Compared to that, the second term
of ((ref)) satisfies
\begin{align*}
& \frac{1}{T^{2}n^{2}}\sum_{t=1}^{T}\sum_{t^{\prime}=1}^{T}\sum_{j=1}^{n}
\sum_{j^{\prime}=1}^{n}\sigma_{j}^{2}\sigma_{j^{\prime}}^{2}E\left(
\varepsilon_{jt}\left( \lambda_{T}\right) \varepsilon_{jt^{\prime}}\left(
\lambda_{T}\right) \right) E\left( \varepsilon_{j^{\prime}t}\left(
\lambda_{T}\right) \varepsilon_{j^{\prime}t^{\prime}}\left( \lambda
_{T}\right) \right) \\
& =\frac{1}{T^{2}n^{2}}\sum_{t=1}^{T}\sum_{j=1}^{n}\sum_{j^{\prime}=1}
^{n}\sigma_{j}^{2}\sigma_{j^{\prime}}^{2}E\left( \varepsilon_{jt}^{2}\left(
\lambda_{T}\right) \right) E\left( \varepsilon_{j^{\prime}t}^{2}\left(
\lambda_{T}\right) \right) =O\left( \frac{1}{T}\right) .
\end{align*}
As we have shown the orders of the two terms in ((ref)), it follows
that $E\left( \frac{1}{T^{2}}\sum_{t^{\prime}=1}^{T}\sum_{t=1}^{T}
\zeta_{tt^{\prime}}^{2}\right) =O\left( \frac{1}{T}\right) $, so by Markov
inequality $T^{-2}\sum_{t^{\prime}=1}^{T}\sum_{t=1}^{T}\zeta_{tt^{\prime}}
^{2}=O_{p}\left( \frac{1}{T}\right) $. Using this result, ((ref)), and
((ref)), it follows that
\begin{align*}
\sup_{i}\left\Vert \mathbf{b}_{2,1,iT}\right\Vert & \leq\left( \frac{1}
{T}\sum_{t^{\prime}=1}^{T}\left\Vert \mathbf{\hat{f}}_{t^{\prime}}
-\mathbf{f}_{t^{\prime}}\right\Vert ^{2}\right) ^{1/2}\left( \frac{1}{T^{2}
}\sum_{t^{\prime}=1}^{T}\sum_{t=1}^{T}\zeta_{tt^{\prime}}^{2}\right)
^{1/2}\left( \sup_{i}\frac{1}{T}\sum_{t=1}^{T}\varepsilon_{it}^{2}\left(
\lambda_{T}\right) \right) ^{1/2}\\
& =O_{p}\left( \frac{1}{\delta_{nT}}\right) \times O_{p}\left( \frac
{1}{\sqrt{n}}\right) =O_{p}\left( \frac{1}{\sqrt{n}\delta_{nT}}\right) .
\end{align*}
Note that $\mathbf{b}_{2,2,iT}$ can be written as
\[
\mathbf{b}_{2,2,iT}=\frac{1}{\sqrt{nT}}\frac{1}{T}\sum_{t=1}^{T}\mathbf{z}
_{t}\varepsilon_{it}\left( \lambda_{T}\right) ,
\]
where
\[
\mathbf{z}_{t}=\frac{1}{\sqrt{nT}}\sum_{t^{\prime}=1}^{T}\sum_{k=1}^{n}
\sigma_{k}\mathbf{f}_{t^{\prime}}\left[ \varepsilon_{kt^{\prime}}\left(
\lambda_{T}\right) \varepsilon_{kt}\left( \lambda_{T}\right) -E\left(
\varepsilon_{kt^{\prime}}\left( \lambda_{T}\right) \varepsilon_{kt}\left(
\lambda_{T}\right) \right) \right] ,
\]
and $\left\Vert \mathbf{z}_{t}\right\Vert ^{2}=O_{p}\left( 1\right) $. Then
by Cauchy-Schwarz inequality,
\[
\left\Vert \mathbf{b}_{2,2,iT}\right\Vert =\frac{1}{\sqrt{nT}}\left\Vert
\frac{1}{T}\sum_{t=1}^{T}\mathbf{z}_{t}\varepsilon_{it}\left( \lambda
_{T}\right) \right\Vert \leq\frac{1}{\sqrt{nT}}\left( \frac{1}{T}\sum
_{t=1}^{T}\left\Vert \mathbf{z}_{t}\right\Vert ^{2}\right) ^{1/2}\left(
\frac{1}{T}\sum_{t=1}^{T}\varepsilon_{it}^{2}\left( \lambda_{T}\right)
\right) ^{1/2},
\]
and in view of ((ref)) we have
\[
\sup_{i}\left\Vert \mathbf{b}_{2,2,iT}\right\Vert \leq\frac{1}{\sqrt{nT}
}\left( \frac{1}{T}\sum_{t=1}^{T}\left\Vert \mathbf{z}_{t}\right\Vert
^{2}\right) ^{1/2}\sup_{i}\left( \frac{1}{T}\sum_{t=1}^{T}\varepsilon
_{it}^{2}\left( \lambda_{T}\right) \right) ^{1/2}=O_{p}\left( \frac
{1}{\sqrt{nT}}\right) .
\]
Hence,
\[
\sup_{i}\left\Vert \mathbf{b}_{2,iT}\right\Vert \leq\sup_{i}\left\Vert
\mathbf{b}_{2,1,iT}\right\Vert +\sup_{i}\left\Vert \mathbf{b}_{2,2,iT}
\right\Vert =O_{p}\left( \frac{1}{\sqrt{n}\delta_{nT}}\right) +O_{p}\left(
\frac{1}{\sqrt{nT}}\right) .
\]
Now consider $\mathbf{b}_{3,iT}$ in ((ref)) and note that
\begin{align*}
\mathbf{b}_{3,iT} & =\frac{1}{T^{2}}\sum_{t=1}^{T}\sum_{t^{\prime}=1}
^{T}\left( \mathbf{\hat{f}}_{t^{\prime}}-\mathbf{f}_{t^{\prime}}\right)
\varkappa_{tt^{\prime}}\varepsilon_{it}\left( \lambda_{T}\right) +\frac
{1}{T^{2}}\sum_{t=1}^{T}\sum_{t^{\prime}=1}^{T}\mathbf{f}_{t^{\prime}
}\varkappa_{tt^{\prime}}\varepsilon_{it}\left( \lambda_{T}\right)
=\mathbf{b}_{3,1,iT}+\mathbf{b}_{3,2,iT}.
\end{align*}
To bound the first term, note that
\begin{align*}
\left\Vert \mathbf{b}_{3,1,iT}\right\Vert & =\left\Vert \frac{1}{T^{2}}
\sum_{t^{\prime}=1}^{T}\left( \mathbf{\hat{f}}_{t^{\prime}}-\mathbf{f}
_{t^{\prime}}\right) \left( \sum_{t=1}^{T}\varkappa_{tt^{\prime}}
\varepsilon_{it}\left( \lambda_{T}\right) \right) \right\Vert \\
& \leq\left( \frac{1}{T}\sum_{t^{\prime}=1}^{T}\left\Vert \mathbf{\hat{f}
}_{t^{\prime}}-\mathbf{f}_{t^{\prime}}\right\Vert ^{2}\right) ^{1/2}\left[
\frac{1}{T^{3}}\sum_{t^{\prime}=1}^{T}\left( \sum_{t=1}^{T}\varkappa
_{tt^{\prime}}\varepsilon_{it}\left( \lambda_{T}\right) \right)
^{2}\right] ^{1/2}\\
& \leq\left( \frac{1}{T}\sum_{t^{\prime}=1}^{T}\left\Vert \mathbf{\hat{f}
}_{t^{\prime}}-\mathbf{f}_{t^{\prime}}\right\Vert ^{2}\right) ^{1/2}\left(
\frac{1}{T^{2}}\sum_{t^{\prime}=1}^{T}\sum_{t=1}^{T}\varkappa_{tt^{\prime}
}^{2}\right) ^{1/2}\left( \frac{1}{T}\sum_{t=1}^{T}\varepsilon_{it}
^{2}\left( \lambda_{T}\right) \right) ^{1/2}.
\end{align*}
By definition of $\varkappa_{tt^{\prime}}$ and given ((ref)), it
follows that
\begin{align*}
E\left( \frac{1}{T^{2}}\sum_{t=1}^{T}\sum_{t^{\prime}=1}^{T}\varkappa
_{tt^{\prime}}^{2}\right) & =\frac{1}{T^{2}}\sum_{t^{\prime}=1}^{T}
\sum_{t=1}^{T}E\left( \frac{1}{n}\sum_{j=1}^{n}\sigma_{j}\mathbf{f}
_{t^{\prime}}^{\prime}\boldsymbol{\gamma}_{j}\varepsilon_{jt}\left(
\lambda_{T}\right) \right) ^{2}\\
& =\frac{1}{T^{2}n^{2}}\sum_{t^{\prime}=1}^{T}\sum_{t=1}^{T}\sum_{j=1}
^{n}\sum_{j^{\prime}=1}^{n}\sigma_{j}\sigma_{j^{\prime}}E\left(
\mathbf{f}_{t^{\prime}}^{\prime}\boldsymbol{\gamma}_{j}\mathbf{f}_{t^{\prime}
}\boldsymbol{\gamma}_{j^{\prime}}\right) E\left( \varepsilon_{jt}\left(
\lambda_{T}\right) \varepsilon_{j^{\prime}t}\left( \lambda_{T}\right)
\right) \\
& \leq\left( \sup_{i}\left\Vert \boldsymbol{\gamma}_{i}\right\Vert
^{2}\right) \left( \sup_{t}E\left\Vert \mathbf{f}_{t}\right\Vert
^{2}\right) \frac{1}{T^{2}n^{2}}\sum_{t^{\prime}=1}^{T}\sum_{t=1}^{T}
\sum_{j=1}^{n}\sum_{j^{\prime}=1}^{n}\sigma_{j}\sigma_{j^{\prime}}E\left(
\varepsilon_{jt}\left( \lambda_{T}\right) \varepsilon_{j^{\prime}t}\left(
\lambda_{T}\right) \right) \\
& \leq\left( \sup_{i}\left\Vert \boldsymbol{\gamma}_{i}\right\Vert
^{2}\right) \left( \sup_{t}E\left\Vert \mathbf{f}_{t}\right\Vert
^{2}\right) \left( \sup_{i}\sigma_{i}^{2}\right) \frac{1}{nT}\sum_{t=1}
^{T}\left( \frac{1}{n}\sum_{j=1}^{n}\sum_{j^{\prime}=1}^{n}\left\vert
E\left( \varepsilon_{jt}\left( \lambda_{T}\right) \varepsilon_{j^{\prime}
t}\left( \lambda_{T}\right) \right) \right\vert \right) \\
& =O\left( \frac{1}{n}\right) ,
\end{align*}
using which and Markov inequality yields $T^{-2}\sum_{t^{\prime}=1}^{T}
\sum_{t=1}^{T}\varkappa_{tt^{\prime}}^{2}=O_{p}\left( n^{-1}\right) $. Then
given this result and ((ref)), ((ref)), it follows
\begin{align*}
\sup_{i}\left\Vert \mathbf{b}_{3,1,iT}\right\Vert & =\left( \frac{1}{T}
\sum_{t^{\prime}=1}^{T}\left\Vert \mathbf{\hat{f}}_{t^{\prime}}-\mathbf{f}
_{t^{\prime}}\right\Vert ^{2}\right) ^{1/2}\left( \frac{1}{T^{2}}
\sum_{t^{\prime}=1}^{T}\sum_{t=1}^{T}\varkappa_{tt^{\prime}}^{2}\right)
^{1/2}\left( \sup_{i}\frac{1}{T}\sum_{t=1}^{T}\varepsilon_{it}^{2}\left(
\lambda_{T}\right) \right) ^{1/2}\\
& =O_{p}\left( \frac{1}{\delta_{nT}}\right) \times O_{p}\left( \frac
{1}{\sqrt{n}}\right) =O_{p}\left( \frac{1}{\sqrt{n}\delta_{nT}}\right) .
\end{align*}
Next, we consider $\left\Vert \mathbf{b}_{3,2,iT}\right\Vert $ and observe
that
\[
\frac{1}{T^{2}}\sum_{t=1}^{T}\sum_{t^{\prime}=1}^{T}\mathbf{f}_{t^{\prime}
}\varkappa_{tt^{\prime}}\varepsilon_{it}\left( \lambda_{T}\right) =\left(
\frac{1}{T}\sum_{t^{\prime}=1}^{T}\mathbf{f}_{t^{\prime}}\mathbf{f}
_{t^{\prime}}^{\prime}\right) \left( \frac{1}{Tn}\sum_{t=1}^{T}\sum
_{j=1}^{n}\sigma_{j}\boldsymbol{\gamma}_{j}\varepsilon_{jt}\left( \lambda
_{T}\right) \varepsilon_{it}\left( \lambda_{T}\right) \right) ,
\]
using which yields
\[
\left\Vert \mathbf{b}_{3,2,iT}\right\Vert =\left\Vert \frac{1}{T^{2}}
\sum_{t=1}^{T}\sum_{t^{\prime}=1}^{T}\mathbf{f}_{t^{\prime}}\varkappa
_{tt^{\prime}}\varepsilon_{it}\left( \lambda_{T}\right) \right\Vert
\leq\left( \frac{1}{T}\sum_{t^{\prime}=1}^{T}\left\Vert \mathbf{f}
_{t^{\prime}}\mathbf{f}_{t^{\prime}}^{\prime}\right\Vert \right) \left(
\left\Vert \frac{1}{nT}\sum_{t=1}^{T}\sum_{j=1}^{n}\sigma_{j}
\boldsymbol{\gamma}_{j}\varepsilon_{jt}\left( \lambda_{T}\right)
\varepsilon_{it}\left( \lambda_{T}\right) \right\Vert \right) .
\]
Then given ((ref)),
\[
\sup_{i}\left\Vert \mathbf{b}_{3,2,iT}\right\Vert \leq\left( \frac{1}{T}
\sum_{t^{\prime}=1}^{T}\left\Vert \mathbf{f}_{t^{\prime}}\mathbf{f}
_{t^{\prime}}^{\prime}\right\Vert \right) \left( \sup_{i}\left\Vert \frac
{1}{nT}\sum_{t=1}^{T}\sum_{j=1}^{n}\sigma_{j}\boldsymbol{\gamma}
_{j}\varepsilon_{jt}\left( \lambda_{T}\right) \varepsilon_{it}\left(
\lambda_{T}\right) \right\Vert \right) =O_{p}\left( \sqrt{\frac{\ln\left(
n\right) }{nT}}\right) .
\]
Hence, $\sup_{i}\left\Vert \mathbf{b}_{3,iT}\right\Vert \leq\sup_{i}\left\Vert
\mathbf{b}_{3,1,iT}\right\Vert +\sup_{i}\left\Vert \mathbf{b}_{3,2,iT}
\right\Vert =O_{p}\left( \sqrt{\frac{\ln\left( n\right) }{nT}}\right) . $
Similarly, the probability order of $\sup_{i}\left\Vert \mathbf{b}
_{4,iT}\right\Vert $ can also be shown to be $O_{p}\left( \sqrt{\ln\left(
n\right) /\left( nT\right) }\right) $. Overall, result ((ref))
follows as we can use ((ref)) to show
\begin{align*}
\sup_{i}\left\Vert \frac{\left( \mathbf{\hat{F}-F}\right) ^{\prime
}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) }{T}\right\Vert
& \leq\sup_{i}\left( \sum_{j=1}^{4}\left\Vert \mathbf{b}_{j,iT}\right\Vert
\right) \leq\sum_{j=1}^{4}\sup_{i}\left\Vert \mathbf{b}_{j,iT}\right\Vert \\
& =O_{p}\left( \frac{1}{\sqrt{T}\delta_{nT}}\right) +O_{p}\left( \frac
{1}{\sqrt{n}\delta_{nT}}\right) +O_{p}\left( \frac{1}{\sqrt{nT}}\right)
+O_{p}\left( \sqrt{\frac{\ln\left( n\right) }{nT}}\right) \\
& =O_{p}\left( \sqrt{\frac{\ln\left( n\right) }{nT}}\right) .
\end{align*}
lemmaDenote $\boldsymbol{\varepsilon}_{i\circ}=\left(
\varepsilon_{i1},\varepsilon_{i2},\ldots,\varepsilon_{iT}\right) ^{\prime}$
and $\mathbf{b}_{i}=\left( b_{i1},b_{i2},\ldots,b_{iT}\right) ^{\prime}$
with $b_{it}=\mathbf{w}_{i0}^{\prime}\boldsymbol{\varepsilon}_{\circ t}$,
$\mathbf{w}_{i0}=\left( w_{i1},w_{is},\ldots,w_{in}\right) ^{\prime}$ and
$\boldsymbol{\varepsilon}_{\circ t}=\left( \varepsilon_{1t},\varepsilon
_{2t},\ldots,\varepsilon_{nt}\right) ^{\prime}$. Suppose that Assumptions
(ref) and (ref) hold. Then as $\left( n,T\right)
\rightarrow\infty$, such that $n/T\rightarrow\kappa\,,$ for $0<\kappa<\infty
$,
\begin{align}
\sup_{i}\left\vert \frac{\boldsymbol{\varepsilon}_{i\circ}^{\prime}
\mathbf{b}_{i}}{T}\right\vert & =O_{p}\left( \sqrt{\frac{\ln\left(
n\right) }{T}}\right) ,\\
\sup_{i}\left\vert \frac{\mathbf{b}_{i}^{\prime}\mathbf{b}_{i}}{T}\right\vert
& =O_{p}\left( 1\right) .
\end{align}
proofConsider ((ref)) and denote $\mathcal{I}_{i,t-1}^{\left(
\varepsilon b\right) }=\left\{ \varepsilon_{i\tau}b_{i\tau}:\tau
=t-1,t-2,\ldots\right\} $ and $i=1,2,\ldots,n$. By assumption, $\varepsilon
_{it}$ is cross-sectionally independent, and is independent from $b_{it}$ as
$w_{ii}=0$ for all $i$. Then $E\left( \varepsilon_{it}b_{it}|\mathcal{I}
_{i,t-1}^{\left( \varepsilon b\right) }\right) =$ $E\left( \varepsilon
_{it}b_{it}\right) =0$. Also $Var\left( \varepsilon_{it}b_{it}\right)
=Var\left( \varepsilon_{it}\right) Var\left( b_{it}\right) =E\left(
b_{it}^{2}\right) , $ which is bounded as by condition ((ref)),
\begin{equation}
E\left( b_{it}^{2}\right) =E\left( \mathbf{w}_{i0}^{\prime}
\boldsymbol{\varepsilon}_{\circ t}\right) ^{2}=\mathbf{w}_{i0}^{\prime
}E\left( \boldsymbol{\varepsilon}_{\circ t}\boldsymbol{\varepsilon}_{\circ
t}^{\prime}\right) \mathbf{w}_{i0}=\sum_{s=1}^{n}w_{is}^{2}<C.
\end{equation}
In addition, since $\varepsilon_{it}\left( \lambda_{T}\right) =$
$\varepsilon_{it}+\lambda_{T}b_{it}$ is sub-exponential for any $\left\vert
\lambda_{T}\right\vert <C$, then it follows that $\varepsilon_{it}$ and
$b_{it}$ are both sub-exponential, and hence $\varepsilon_{it}b_{it}$ is also
sub-exponential. Hence, result ((ref)) can be established by
applying the method of proof used for result ((ref)). Now consider
((ref)) and note
\[
\frac{\mathbf{b}_{i}^{\prime}\mathbf{b}_{i}}{T}=\frac{1}{T}\sum_{t=1}
^{T}\left[ b_{it}^{2}-E\left( b_{it}^{2}\right) \right] +\frac{1}{T}
\sum_{t=1}^{T}E\left( b_{it}^{2}\right) ,
\]
which further implies
\[
\sup_{i}\left\vert \frac{\mathbf{b}_{i}^{\prime}\mathbf{b}_{i}}{T}\right\vert
\leq\sup_{i}\left\vert \frac{1}{T}\sum_{t=1}^{T}\left[ b_{it}^{2}-E\left(
b_{it}^{2}\right) \right] \right\vert +\sup_{i}\left( \frac{1}{T}\sum
_{t=1}^{T}E\left( b_{it}^{2}\right) \right) .
\]
Denote $\mathcal{I}_{i,t-1}^{\left( b\right) }=\left\{ b_{i\tau}
:\tau=t-1,t-2,\ldots\right\} $ and $i=1,2,\ldots,n$, then $E\left(
b_{it}^{2}-E\left( b_{it}^{2}\right) |\mathcal{I}_{i,t-1}^{\left( b\right)
}\right) =$ $E\left( b_{it}^{2}-E\left( b_{it}^{2}\right) \right) =0$.
Besides, $Var\left( b_{it}^{2}\right) =E\left( b_{it}^{4}\right) -\left[
E\left( b_{it}^{2}\right) \right] ^{2}$ is bounded as ((ref)) shows
$E\left( b_{it}^{4}\right) =E\left( \mathbf{w}_{i0}^{\prime}
\boldsymbol{\varepsilon}_{\circ t}\right) ^{4}=O\left( 1\right) $. Also
$b_{it}^{2}-E\left( b_{it}^{2}\right) $ is sub-exponential given that it is
already established that $b_{it}$ is sub-exponential. Therefore, it follows
that $\sup_{i}\left\vert \frac{1}{T}\sum_{t=1}^{T}\left[ b_{it}^{2}-E\left(
b_{it}^{2}\right) \right] \right\vert =O_{p}\left( \sqrt{\frac{\ln\left(
n\right) }{T}}\right) $ by applying the method of proof used for result
((ref)). Also by ((ref)), we have $\sup_{i}\left( \frac{1}
{T}\sum_{t=1}^{T}E\left( b_{it}^{2}\right) \right) =O\left( 1\right) $.
Result ((ref)) now follows straightforwardly.
lemmaConsider the latent factor model given by ((ref)) and
((ref)). Let $\hat{\sigma}_{i,T}=\left( T^{-1}\mathbf{e}_{i}^{\prime
}\mathbf{e}_{i}\right) ^{1/2}$, where $\mathbf{e}_{i}=\mathbf{M}_{\hat{F}
}\mathbf{y}_{i}$, $\mathbf{M}_{\hat{F}}=\mathbf{I}_{T}-\mathbf{\hat{F}(\hat
{F}}^{\prime}\mathbf{\hat{F})}^{-1}\mathbf{\hat{F}}^{\prime}$, $\mathbf{y}
_{i}=(y_{i1},y_{i2},...,y_{iT})^{\prime}$, and $\mathbf{\hat{F}}$ is given by
((ref)). Also let $\omega_{i,T}=\left( T^{-1}\sigma_{i}^{2}
\boldsymbol{\varepsilon}_{i\circ}^{^{\prime}}\mathbf{M}_{F}
\boldsymbol{\varepsilon}_{i\circ}\right) ^{1/2},$ where $\mathbf{M}
_{F}=\mathbf{I}_{T}-\mathbf{F(F}^{\prime}\mathbf{F)}^{-1}\mathbf{F}^{\prime}$.
Suppose that Assumptions (ref)-(ref) hold and $\left(
n,T\right) \rightarrow\infty$, such that $n/T\rightarrow\kappa\,,$ for
$0<\kappa<\infty$. Then
\begin{align}
\sup_{i}\left\vert \hat{\sigma}_{i,T}^{2}-\omega_{i,T}^{2}\right\vert &
=O_{p}\left( \frac{\ln\left( n\right) }{T}\right) ,
\\
\sup_{i}\left\vert \hat{\sigma}_{i,T}-\omega_{i,T}\right\vert &
=O_{p}\left( \frac{\ln\left( n\right) }{T}\right) ,\\
\sup_{i}\left\vert \frac{1}{\hat{\sigma}_{i,T}}-\frac{1}{\omega_{i,T}
}\right\vert & =O_{p}\left( \frac{\ln\left( n\right) }{T}\right) ,
\end{align}
and
\begin{align}
\frac{1}{n}\sum_{i=1}^{n}\left\vert \hat{\sigma}_{i,T}^{2}-\omega_{i,T}
^{2}\right\vert & =O_{p}\left( \frac{\ln\left( n\right) }{T}\right)
,\\
\frac{1}{n}\sum_{i=1}^{n}\left\vert \hat{\sigma}_{i,T}-\omega_{i,T}
\right\vert & =O_{p}\left( \frac{\ln\left( n\right) }{T}\right)
,\\
\frac{1}{n}\sum_{i=1}^{n}\left\vert \frac{1}{\hat{\sigma}_{i,T}}-\frac
{1}{\omega_{i,T}}\right\vert & =O_{p}\left( \frac{\ln\left( n\right) }
{T}\right) .
\end{align}
proofNote that by ((ref)), $\mathbf{y}_{i}=\mathbf{F}\boldsymbol{\gamma}
_{i}+\sigma_{i}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) $,
where $\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) =\left(
\varepsilon_{i1}\left( \lambda_{T}\right) ,\varepsilon_{i2}\left(
\lambda_{T}\right) ,\ldots,\varepsilon_{iT}\left( \lambda_{T}\right)
\right) ^{\prime}$, which in turn implies
\begin{align*}
\mathbf{e}_{i} & =\mathbf{M}_{\hat{F}}\mathbf{y}_{i}=\mathbf{M}_{\hat{F}
}\left( \mathbf{F}\boldsymbol{\gamma}_{i}+\sigma_{i}\boldsymbol{\varepsilon
}_{i\circ}\left( \lambda_{T}\right) \right) \\
& =\sigma_{i}\mathbf{M}_{F}\boldsymbol{\varepsilon}_{i\circ}\left(
\lambda_{T}\right) +\sigma_{i}\left( \mathbf{M}_{\hat{F}}-\mathbf{M}
_{F}\right) \boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right)
+\mathbf{M}_{\hat{F}}\mathbf{F}\boldsymbol{\gamma}_{i}.
\end{align*}
Then $\hat{\sigma}_{i,T}^{2}$ can be decomposed as
\begin{align}
\hat{\sigma}_{i,T}^{2}-\omega_{i,T}^{2} & =\frac{\boldsymbol{\gamma}
_{i}^{^{\prime}}\mathbf{F}^{^{\prime}}\mathbf{M}_{\hat{F}}\mathbf{F}
\boldsymbol{\gamma}_{i}}{T}+\frac{\sigma_{i}^{2}\boldsymbol{\varepsilon
}_{i\circ}\left( \lambda_{T}\right) ^{\prime}\left( \mathbf{M}_{\hat{F}
}-\mathbf{M}_{F}\right) \left( \mathbf{M}_{\hat{F}}-\mathbf{M}_{F}\right)
\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) }{T}\nonumber\\
& +\frac{2\sigma_{i}^{2}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda
_{T}\right) ^{\prime}\mathbf{M}_{F}\left( \mathbf{M}_{\hat{F}}
-\mathbf{M}_{F}\right) \boldsymbol{\varepsilon}_{i\circ}\left( \lambda
_{T}\right) }{T}+\frac{2\sigma_{i}^{2}\boldsymbol{\varepsilon}_{i\circ
}\left( \lambda_{T}\right) ^{\prime}\mathbf{M}_{F}\mathbf{M}_{\hat{F}
}\mathbf{F}\boldsymbol{\gamma}_{i}}{T}\nonumber\\
& +\frac{2\sigma_{i}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda
_{T}\right) ^{\prime}\left( \mathbf{M}_{\hat{F}}-\mathbf{M}_{F}\right)
\mathbf{M}_{\hat{F}}\mathbf{F}\boldsymbol{\gamma}_{i}}{T}+\frac{\sigma_{i}
^{2}\left( \boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right)
+\boldsymbol{\varepsilon}_{i\circ}\right) ^{\prime}\mathbf{M}_{F}\left(
\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right)
-\boldsymbol{\varepsilon}_{i\circ}\right) }{T}\nonumber\\
& =\sum_{j=1}^{6}B_{j,iT},
\end{align}
and $\left\vert \hat{\sigma}_{i,T}^{2}-\omega_{i,T}^{2}\right\vert \leq
\sum_{j=1}^{6}\left\vert B_{j,iT}\right\vert $. Starting with $B_{1,iT}$, note
that
\[
\left\vert B_{1,iT}\right\vert =\left\vert \frac{\boldsymbol{\gamma}
_{i}^{^{\prime}}\mathbf{F}^{^{\prime}}\mathbf{M}_{\hat{F}}\mathbf{F}
\boldsymbol{\gamma}_{i}}{T}\right\vert =\left\vert \frac{\boldsymbol{\gamma
}_{i}^{^{\prime}}\left( \mathbf{F-\hat{F}}\right) ^{^{\prime}}
\mathbf{M}_{\hat{F}}\left( \mathbf{F-\hat{F}}\right) \boldsymbol{\gamma}
_{i}}{T}\right\vert \leq\left\Vert \boldsymbol{\gamma}_{i}\right\Vert
^{2}\left\Vert \mathbf{M}_{\hat{F}}\right\Vert \left( \frac{\left\Vert
\mathbf{\hat{F}-F}\right\Vert ^{2}}{T}\right) .
\]
Also $\sup_{i}\left\Vert \boldsymbol{\gamma}_{i}\right\Vert <C$ and
$\left\Vert \mathbf{M}_{\hat{F}}\right\Vert =1$, and using ((ref)) of
Lemma (ref) we have
\[
\sup_{i}\left\vert B_{1,iT}\right\vert \leq\left( \sup_{i}\left\Vert
\boldsymbol{\gamma}_{i}\right\Vert ^{2}\right) \left\Vert \mathbf{M}_{\hat
{F}}\right\Vert \left( \frac{\left\Vert \mathbf{\hat{F}-F}\right\Vert ^{2}
}{T}\right) =O_{p}\left( \frac{1}{\delta_{nT}^{2}}\right) .
\]
To establish the probability orders of the remaining terms of ((ref)), we
first observe that
\[
\frac{\mathbf{\hat{F}}^{^{\prime}}\mathbf{\hat{F}}}{T}-\frac{\mathbf{F}
^{^{\prime}}\mathbf{F}}{T}=\frac{\left( \mathbf{\hat{F}-F}\right)
^{^{\prime}}\left( \mathbf{\hat{F}-F}\right) }{T}+\frac{\left(
\mathbf{\hat{F}-F}\right) ^{^{\prime}}\mathbf{F}}{T}+\frac{\mathbf{F}
^{^{\prime}}\left( \mathbf{\hat{F}-F}\right) }{T}.
\]
Using results ((ref)) and ((ref)) it follows that
\begin{equation}
\left\Vert \frac{\mathbf{\hat{F}}^{^{\prime}}\mathbf{\hat{F}}}{T}
-\frac{\mathbf{F}^{^{\prime}}\mathbf{F}}{T}\right\Vert \leq\frac{\left\Vert
\mathbf{\hat{F}-F}\right\Vert ^{2}}{T}+\frac{\left\Vert \left( \mathbf{\hat
{F}-F}\right) ^{\prime}\mathbf{F}\right\Vert }{T}+\frac{\left\Vert
\mathbf{F}^{\prime}\left( \mathbf{\hat{F}-F}\right) \right\Vert }{T}
=O_{p}\left( \frac{1}{\delta_{nT}^{2}}\right) .
\end{equation}
By assumption $T^{-1}\mathbf{F}^{^{\prime}}\mathbf{F}$ is a positive definite
matrix, then
\[
\left\Vert \left( \frac{\mathbf{\hat{F}}^{^{\prime}}\mathbf{\hat{F}}}
{T}\right) ^{-1}-\left( \frac{\mathbf{F}^{^{\prime}}\mathbf{F}}{T}\right)
^{-1}\right\Vert \leq\left\Vert \left( \frac{\mathbf{\hat{F}}^{^{\prime}
}\mathbf{\hat{F}}}{T}\right) ^{-1}\right\Vert \left\Vert \frac{\mathbf{\hat
{F}}^{^{\prime}}\mathbf{\hat{F}}}{T}-\frac{\mathbf{F}^{^{\prime}}\mathbf{F}
}{T}\right\Vert \left\Vert \left( \frac{\mathbf{F}^{^{\prime}}\mathbf{F}}
{T}\right) ^{-1}\right\Vert =O_{p}\left( \frac{1}{\delta_{nT}^{2}}\right)
,
\]
so it follows that
\begin{equation}
\frac{\mathbf{\hat{F}}^{^{\prime}}\mathbf{\hat{F}}}{T}=\frac{\mathbf{F}
^{^{\prime}}\mathbf{F}}{T}+O_{p}\left( \frac{1}{\delta_{nT}^{2}}\right)
, and \left( \frac{\mathbf{\hat{F}}^{^{\prime}}\mathbf{\hat{F}}}
{T}\right) ^{-1}=\left( \frac{\mathbf{F}^{^{\prime}}\mathbf{F}}{T}\right)
^{-1}+O_{p}\left( \frac{1}{\delta_{nT}^{2}}\right) .
\end{equation}
Now we consider $B_{2,iT},$ and note that
\begin{align}
B_{2,iT} & =T^{-1}\sigma_{i}^{2}\boldsymbol{\varepsilon}_{i\circ}\left(
\lambda_{T}\right) ^{\prime}\left[ \mathbf{\hat{F}}\left( \mathbf{\hat{F}
}^{^{\prime}}\mathbf{\hat{F}}\right) ^{-1}\mathbf{\hat{F}}^{^{\prime}
}-\mathbf{F}\left( \mathbf{F}^{^{\prime}}\mathbf{F}\right) ^{-1}
\mathbf{F}^{^{\prime}}\right] \boldsymbol{\varepsilon}_{i\circ}\left(
\lambda_{T}\right) +\nonumber\\
& T^{-1}\sigma_{i}^{2}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda
_{T}\right) ^{\prime}\left( \mathbf{I}_{m_{0}}-\mathbf{\hat{F}}\left(
\mathbf{\hat{F}}^{^{\prime}}\mathbf{\hat{F}}\right) ^{-1}\mathbf{\hat{F}
}^{^{\prime}}\right) \mathbf{F}\left( \mathbf{F}^{^{\prime}}\mathbf{F}
\right) ^{-1}\mathbf{F}^{^{\prime}}\boldsymbol{\varepsilon}_{i\circ}\left(
\lambda_{T}\right) +\nonumber\\
& T^{-1}\sigma_{i}^{2}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda
_{T}\right) ^{\prime}\mathbf{F}\left( \mathbf{F}^{^{\prime}}\mathbf{F}
\right) ^{-1}\mathbf{F}^{^{\prime}}\left( \mathbf{I}_{m_{0}}-\mathbf{\hat
{F}}\left( \mathbf{\hat{F}}^{^{\prime}}\mathbf{\hat{F}}\right)
^{-1}\mathbf{\hat{F}}^{^{\prime}}\right) \boldsymbol{\varepsilon}_{i\circ
}\left( \lambda_{T}\right) .
\end{align}
We further note
\begin{align*}
\mathbf{\hat{F}}\left( \mathbf{\hat{F}}^{^{\prime}}\mathbf{\hat{F}}\right)
^{-1}\mathbf{\hat{F}}^{^{\prime}}-\mathbf{F}\left( \mathbf{F}^{^{\prime}
}\mathbf{F}\right) ^{-1}\mathbf{F}^{^{\prime}} & =\left( \mathbf{\hat
{F}-F}\right) \left( \mathbf{\hat{F}}^{^{\prime}}\mathbf{\hat{F}}\right)
^{-1}\left( \mathbf{\hat{F}-F}\right) ^{^{\prime}}+\left[ \mathbf{F}\left(
\mathbf{\hat{F}}^{^{\prime}}\mathbf{\hat{F}}\right) ^{-1}\mathbf{F}
^{^{\prime}}-\mathbf{F}\left( \mathbf{F}^{^{\prime}}\mathbf{F}\right)
^{-1}\mathbf{F}^{^{\prime}}\right] \\
& +\mathbf{F}\left( \mathbf{\hat{F}}^{^{\prime}}\mathbf{\hat{F}}\right)
^{-1}\left( \mathbf{\hat{F}-F}\right) ^{^{\prime}}+\left( \mathbf{\hat
{F}-F}\right) \left( \mathbf{\hat{F}}^{^{\prime}}\mathbf{\hat{F}}\right)
^{-1}\mathbf{F}^{\prime}.
\end{align*}
Using ((ref)), we have $B_{2,iT}=\sum_{j=1}^{6}B_{2,j,iT}$, where
\begin{align*}
B_{2,1,iT} & =T^{-1}\sigma_{i}^{2}\boldsymbol{\varepsilon}_{i\circ}\left(
\lambda_{T}\right) ^{\prime}\left( \mathbf{\hat{F}-F}\right) \left(
\mathbf{\hat{F}}^{^{\prime}}\mathbf{\hat{F}}\right) ^{-1}\left(
\mathbf{\hat{F}-F}\right) ^{\prime}\boldsymbol{\varepsilon}_{i\circ}\left(
\lambda_{T}\right) ,\\
B_{2,2,iT} & =T^{-1}\sigma_{i}^{2}\boldsymbol{\varepsilon}_{i\circ}\left(
\lambda_{T}\right) ^{\prime}\left[ \mathbf{F}\left( \mathbf{\hat{F}
}^{^{\prime}}\mathbf{\hat{F}}\right) ^{-1}\mathbf{F}^{^{\prime}}
-\mathbf{F}\left( \mathbf{F}^{^{\prime}}\mathbf{F}\right) ^{-1}
\mathbf{F}^{^{\prime}}\right] \boldsymbol{\varepsilon}_{i\circ}\left(
\lambda_{T}\right) ,\\
B_{2,3,iT} & =B_{2,4,iT}=T^{-1}\sigma_{i}^{2}\boldsymbol{\varepsilon
}_{i\circ}\left( \lambda_{T}\right) ^{\prime}\mathbf{F}\left(
\mathbf{\hat{F}}^{^{\prime}}\mathbf{\hat{F}}\right) ^{-1}\left(
\mathbf{\hat{F}-F}\right) ^{\prime}\boldsymbol{\varepsilon}_{i\circ}\left(
\lambda_{T}\right) ,\\
B_{2,5,iT} & =B_{2,6,iT}=T^{-1}\sigma_{i}^{2}\boldsymbol{\varepsilon
}_{i\circ}\left( \lambda_{T}\right) ^{\prime}\left( \mathbf{I}_{m_{0}
}-\mathbf{\hat{F}}\left( \mathbf{\hat{F}}^{^{\prime}}\mathbf{\hat{F}}\right)
^{-1}\mathbf{\hat{F}}^{^{\prime}}\right) \mathbf{F}\left( \mathbf{F}
^{^{\prime}}\mathbf{F}\right) ^{-1}\mathbf{F}^{\prime}\boldsymbol{\varepsilon
}_{i\circ}\left( \lambda_{T}\right) ,
\end{align*}
and $B_{2,iT}\leq\sum_{j=1}^{6}\left\vert B_{2,j,iT}\right\vert $. Starting
with the first term we note that
\begin{align*}
\left\vert B_{2,1,iT}\right\vert & =\left\vert T^{-1}\sigma_{i}
^{2}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) ^{\prime
}\left( \mathbf{\hat{F}-F}\right) \left( \mathbf{\hat{F}}^{^{\prime}
}\mathbf{\hat{F}}\right) ^{-1}\left( \mathbf{\hat{F}-F}\right) ^{\prime
}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) \right\vert \\
& \leq\sigma_{i}^{2}\left( \frac{1}{T}\left\Vert \boldsymbol{\varepsilon
}_{i\circ}\left( \lambda_{T}\right) \right\Vert ^{2}\right) \left(
\frac{\left\Vert \mathbf{\hat{F}-F}\right\Vert ^{2}}{T}\right) \left\Vert
\left( \frac{\mathbf{\hat{F}}^{^{\prime}}\mathbf{\hat{F}}}{T}\right)
^{-1}\right\Vert .
\end{align*}
By ((ref)) and ((ref)) we have
\[
\sup_{i}\left\vert B_{2,1,iT}\right\vert \leq\left( \sup_{i}\sigma_{i}
^{2}\right) \left( \sup_{i}\frac{1}{T}\left\Vert \boldsymbol{\varepsilon
}_{i\circ}\left( \lambda_{T}\right) \right\Vert ^{2}\right) \left(
\frac{\left\Vert \mathbf{\hat{F}-F}\right\Vert ^{2}}{T}\right) \left\Vert
\left( \frac{\mathbf{\hat{F}}^{^{\prime}}\mathbf{\hat{F}}}{T}\right)
^{-1}\right\Vert =O_{p}\left( \frac{1}{\delta_{nT}^{2}}\right) .
\]
Next, consider $B_{2,2,iT}$ and note that
\begin{align*}
\left\vert B_{2,2,iT}\right\vert & =\left\vert \frac{1}{T}\sigma_{i}
^{2}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) ^{\prime
}\left[ \mathbf{F}\left( \mathbf{\hat{F}}^{^{\prime}}\mathbf{\hat{F}
}\right) ^{-1}\mathbf{F}^{^{\prime}}-\mathbf{F}\left( \mathbf{F}^{^{\prime}
}\mathbf{F}\right) ^{-1}\mathbf{F}^{^{\prime}}\right]
\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) \right\vert \\
& =\sigma_{i}^{2}\left\vert \left( \frac{\boldsymbol{\varepsilon}_{i\circ
}\left( \lambda_{T}\right) ^{\prime}\mathbf{F}}{T}\right) \left[ \left(
\frac{\mathbf{\hat{F}}^{^{\prime}}\mathbf{\hat{F}}}{T}\right) ^{-1}-\left(
\frac{\mathbf{F}^{^{\prime}}\mathbf{F}}{T}\right) ^{-1}\right] \left(
\frac{\mathbf{F}^{^{\prime}}\boldsymbol{\varepsilon}_{i\circ}\left(
\lambda_{T}\right) }{T}\right) \right\vert \\
& \leq\sigma_{i}^{2}\left( \left\Vert \frac{\boldsymbol{\varepsilon}
_{i\circ}\left( \lambda_{T}\right) ^{\prime}\mathbf{F}}{T}\right\Vert
^{2}\right) \left\Vert \left( \frac{\mathbf{\hat{F}}^{^{\prime}}
\mathbf{\hat{F}}}{T}\right) ^{-1}-\left( \frac{\mathbf{F}^{^{\prime}
}\mathbf{F}}{T}\right) ^{-1}\right\Vert .
\end{align*}
Using ((ref)) (from Lemma (ref)) and ((ref))), we further
have
\begin{align*}
\sup_{i}\left\vert B_{2,2,iT}\right\vert & \leq\left( \sup_{i}\sigma
_{i}^{2}\right) \left( \sup_{i}\left\Vert \frac{\boldsymbol{\varepsilon
}_{i\circ}\left( \lambda_{T}\right) ^{\prime}\mathbf{F}}{T}\right\Vert
^{2}\right) \left\Vert \left( \frac{\mathbf{\hat{F}}^{^{\prime}}
\mathbf{\hat{F}}}{T}\right) ^{-1}-\left( \frac{\mathbf{F}^{^{\prime}
}\mathbf{F}}{T}\right) ^{-1}\right\Vert \\
& =O_{p}\left( \frac{\ln\left( n\right) }{T}\right) \times O_{p}\left(
\frac{1}{\delta_{nT}^{2}}\right) =O_{p}\left( \frac{\ln\left( n\right)
}{T\delta_{nT}^{2}}\right) .
\end{align*}
Next, consider
\begin{align*}
\left\vert B_{2,3,iT}\right\vert & =\sigma_{i}^{2}\left\vert \left(
\frac{\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) ^{\prime
}\mathbf{F}}{T}\right) \left( \frac{\mathbf{\hat{F}}^{^{\prime}}
\mathbf{\hat{F}}}{T}\right) ^{-1}\left[ \frac{\left( \mathbf{\hat{F}
-F}\right) ^{\prime}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda
_{T}\right) }{T}\right] \right\vert \\
& \leq\sigma_{i}^{2}\left\Vert \frac{\boldsymbol{\varepsilon}_{i\circ}\left(
\lambda_{T}\right) ^{\prime}\mathbf{F}}{T}\right\Vert \left\Vert \left(
\frac{\mathbf{\hat{F}}^{^{\prime}}\mathbf{\hat{F}}}{T}\right) ^{-1}
\right\Vert \left\Vert \frac{\left( \mathbf{\hat{F}-F}\right) ^{\prime
}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) }{T}\right\Vert
,
\end{align*}
and using ((ref)), ((ref)) and ((ref)) now yields
\begin{align*}
\sup_{i}\left\vert B_{2,3,iT}\right\vert & \leq\left\Vert \left(
\frac{\mathbf{\hat{F}}^{^{\prime}}\mathbf{\hat{F}}}{T}\right) ^{-1}
\right\Vert \left( \sup_{i}\left\Vert \frac{\boldsymbol{\varepsilon}_{i\circ
}\left( \lambda_{T}\right) ^{\prime}\mathbf{F}}{T}\right\Vert \right)
\left( \sup_{i}\left\Vert \frac{\left( \mathbf{\hat{F}-F}\right) ^{\prime
}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) }{T}\right\Vert
\right) \\
& =O_{p}\left( \sqrt{\frac{\ln\left( n\right) }{T}}\right) \times
O_{p}\left( \sqrt{\frac{\ln\left( n\right) }{nT}}\right) .
\end{align*}
Similarly,
\begin{align}
\left\vert B_{2,5,iT}\right\vert & =\left\vert T^{-1}\sigma_{i}
^{2}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) ^{\prime
}\left( \mathbf{I}_{m_{0}}-\mathbf{\hat{F}}\left( \mathbf{\hat{F}}
^{^{\prime}}\mathbf{\hat{F}}\right) ^{-1}\mathbf{\hat{F}}^{^{\prime}}\right)
\mathbf{F}\left( \mathbf{F}^{^{\prime}}\mathbf{F}\right) ^{-1}
\mathbf{F}^{\prime}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda
_{T}\right) \right\vert \nonumber\\
& \leq\sigma_{i}^{2}\left\vert \frac{\boldsymbol{\varepsilon}_{i\circ}\left(
\lambda_{T}\right) ^{\prime}\mathbf{P}_{F}\boldsymbol{\varepsilon}_{i\circ
}\left( \lambda_{T}\right) }{T}\right\vert +\sigma_{i}^{2}\left\vert
\frac{\boldsymbol{\varepsilon}_{i\circ}^{\prime}\left( \lambda_{T}\right)
\mathbf{P}_{F}\left( \mathbf{P}_{\hat{F}}-\mathbf{P}_{F}\right)
\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) }{T}\right\vert ,
\end{align}
where $\mathbf{P}_{F}=\mathbf{F}\left( \mathbf{F}^{^{\prime}}\mathbf{F}
\right) ^{-1}\mathbf{F}^{^{\prime}}$ and $\mathbf{P}_{\hat{F}}=\mathbf{\hat
{F}}\left( \mathbf{\hat{F}}^{^{\prime}}\mathbf{\hat{F}}\right)
^{-1}\mathbf{\hat{F}}^{^{\prime}}$. Further
\[
\left\vert \frac{\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right)
^{\prime}\mathbf{P}_{F}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda
\right) }{T}\right\vert =\left( \frac{\boldsymbol{\varepsilon}_{i\circ
}\left( \lambda_{T}\right) ^{\prime}\mathbf{F}}{T}\right) \left(
\frac{\mathbf{F}^{^{\prime}}\mathbf{F}}{T}\right) ^{-1}\left( \frac
{\mathbf{F}^{^{\prime}}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda
_{T}\right) }{T}\right) \leq\left\Vert \left( \frac{\mathbf{F}^{^{\prime}
}\mathbf{F}}{T}\right) ^{-1}\right\Vert \left\Vert \frac
{\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) ^{\prime
}\mathbf{F}}{T}\right\Vert ^{2},
\]
and using ((ref)) it follows that
\[
\sup_{i}\left\vert \frac{\boldsymbol{\varepsilon}_{i\circ}\left( \lambda
_{T}\right) ^{\prime}\mathbf{P}_{F}\boldsymbol{\varepsilon}_{i\circ}\left(
\lambda_{T}\right) }{T}\right\vert \leq\left\Vert \left( \frac
{\mathbf{F}^{^{\prime}}\mathbf{F}}{T}\right) ^{-1}\right\Vert \left(
\sup_{i}\left\Vert \frac{\boldsymbol{\varepsilon}_{i\circ}\left( \lambda
_{T}\right) ^{\prime}\mathbf{F}}{T}\right\Vert ^{2}\right) =O_{p}\left(
\frac{\ln\left( n\right) }{T}\right) .
\]
Also, the probability order of the first term in ((ref)) dominates the
second term, and we have
\[
\sup_{i}\left\vert B_{2,5,iT}\right\vert \leq\left( \sup_{i}\sigma_{i}
^{2}\right) \left( \sup_{i}\left\vert \frac{\boldsymbol{\varepsilon}
_{i\circ}\left( \lambda_{T}\right) ^{\prime}\mathbf{P}_{F}
\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) }{T}\right\vert
\right) =O_{p}\left( \frac{\ln\left( n\right) }{T}\right) .
\]
Overall, using the above results we obtain
\[
\sup_{i}\left\vert B_{2,iT}\right\vert \leq\sup_{i}\left( \sum_{j=1}
^{6}\left\vert B_{2,j,iT}\right\vert \right) \leq\sum_{j=1}^{6}\sup
_{i}\left\vert B_{2,j,iT}\right\vert =O_{p}\left( \frac{\ln\left( n\right)
}{T}\right) .
\]
Next, consider $B_{3,iT}$ which can be rewritten as
\begin{align*}
B_{3,iT} & =\frac{2\sigma_{i}^{2}\boldsymbol{\varepsilon}_{i\circ}\left(
\lambda_{T}\right) ^{\prime}\left( \mathbf{I}_{m_{0}}-\mathbf{P}_{F}\right)
\left( \mathbf{P}_{F}-\mathbf{P}_{\hat{F}}\right) \boldsymbol{\varepsilon
}_{i\circ}\left( \lambda_{T}\right) }{T}=\frac{2\sigma_{i}^{2}
\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) ^{\prime}\left(
\mathbf{P}_{F}-\mathbf{P}_{\hat{F}}-\mathbf{P}_{F}+\mathbf{P}_{F}
\mathbf{P}_{\hat{F}}\right) \boldsymbol{\varepsilon}_{i\circ}\left(
\lambda_{T}\right) }{T}\\
& =\frac{2\sigma_{i}^{2}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda
_{T}\right) ^{\prime}\mathbf{P}_{\hat{F}}\boldsymbol{\varepsilon}_{i\circ
}\left( \lambda_{T}\right) }{T}+\frac{2\sigma_{i}^{2}\boldsymbol{\varepsilon
}_{i\circ}\left( \lambda_{T}\right) ^{\prime}\mathbf{P}_{F}\mathbf{P}
_{\hat{F}}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) }
{T}=B_{3,1,nT}+B_{3,2,nT}.
\end{align*}
Note that
\begin{align*}
\left\vert B_{3,1,iT}\right\vert & =2\left\vert \sigma_{i}^{2}\left(
\frac{\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) ^{\prime
}\mathbf{\hat{F}}}{T}\right) \left( \frac{\mathbf{\hat{F}}^{\prime
}\mathbf{\hat{F}}}{T}\right) ^{-1}\left( \frac{\mathbf{\hat{F}}^{\prime
}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) }{T}\right)
\right\vert \leq2\sigma_{i}^{2}\left\Vert \frac{\boldsymbol{\varepsilon
}_{i\circ}\left( \lambda_{T}\right) ^{\prime}\mathbf{\hat{F}}}{T}\right\Vert
^{2}\left\Vert \left( \frac{\mathbf{\hat{F}}^{\prime}\mathbf{\hat{F}}}
{T}\right) ^{-1}\right\Vert ^{2}\\
& \leq4\sigma_{i}^{2}\left( \left\Vert \frac{\boldsymbol{\varepsilon
}_{i\circ}\left( \lambda_{T}\right) ^{\prime}\mathbf{F}}{T}\right\Vert
^{2}+\left\Vert \frac{\boldsymbol{\varepsilon}_{i\circ}\left( \lambda
_{T}\right) ^{\prime}\left( \mathbf{\hat{F}}-\mathbf{F}\right) }
{T}\right\Vert ^{2}\right) \left\Vert \left( \frac{\mathbf{\hat{F}}^{\prime
}\mathbf{\hat{F}}}{T}\right) ^{-1}\right\Vert ^{2}.
\end{align*}
Then using ((ref)) and ((ref)), we have
\begin{align*}
\sup_{i}\left\vert B_{3,1,iT}\right\vert & \leq4\left( \sup_{i}\sigma
_{i}^{2}\right) \sup_{i}\left( \left\Vert \frac{\boldsymbol{\varepsilon
}_{i\circ}\left( \lambda_{T}\right) ^{\prime}\mathbf{F}}{T}\right\Vert
^{2}+\left\Vert \frac{\boldsymbol{\varepsilon}_{i\circ}\left( \lambda
_{T}\right) ^{\prime}\left( \mathbf{\hat{F}}-\mathbf{F}\right) }
{T}\right\Vert ^{2}\right) \left\Vert \left( \frac{\mathbf{\hat{F}}^{\prime
}\mathbf{\hat{F}}}{T}\right) ^{-1}\right\Vert ^{2}\\
& \leq4\left( \sup_{i}\sigma_{i}^{2}\right) \left( \sup_{i}\left\Vert
\frac{\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) ^{\prime
}\mathbf{F}}{T}\right\Vert ^{2}+\sup_{i}\left\Vert \frac
{\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) ^{\prime}\left(
\mathbf{\hat{F}}-\mathbf{F}\right) }{T}\right\Vert ^{2}\right)
\times\left\Vert \left( \frac{\mathbf{\hat{F}}^{\prime}\mathbf{\hat{F}}}
{T}\right) ^{-1}\right\Vert ^{2}\\
& =O_{p}\left( \frac{\ln\left( n\right) }{T}\right) .
\end{align*}
Also,
\begin{align*}
\left\vert B_{3,2,iT}\right\vert & =2\sigma_{i}^{2}\left\vert \left(
\frac{\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) ^{\prime
}\mathbf{F}}{T}\right) \left( \frac{\mathbf{F}^{\prime}\mathbf{F}}
{T}\right) ^{-1}\left( \frac{\mathbf{F}^{\prime}\mathbf{\hat{F}}}{T}\right)
\left( \frac{\mathbf{\hat{F}}^{\prime}\mathbf{\hat{F}}}{T}\right)
^{-1}\left( \frac{\mathbf{\hat{F}}^{\prime}\boldsymbol{\varepsilon}_{i\circ
}\left( \lambda_{T}\right) }{T}\right) \right\vert \\
& \leq2\sigma_{i}^{2}\left\vert \left( \frac{\boldsymbol{\varepsilon
}_{i\circ}\left( \lambda_{T}\right) ^{\prime}\mathbf{F}}{T}\right) \left(
\frac{\mathbf{F}^{\prime}\mathbf{F}}{T}\right) ^{-1}\left( \frac
{\mathbf{F}^{\prime}\mathbf{\hat{F}}}{T}\right) \left( \frac{\mathbf{\hat
{F}}^{\prime}\mathbf{\hat{F}}}{T}\right) ^{-1}\left[ \frac{\mathbf{F}
^{\prime}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) }
{T}+\frac{\left( \mathbf{\hat{F}-F}\right) ^{\prime}\boldsymbol{\varepsilon
}_{i\circ}\left( \lambda_{T}\right) }{T}\right] \right\vert \\
& \leq2\sigma_{i}^{2}\left\Vert \frac{\boldsymbol{\varepsilon}_{i\circ
}\left( \lambda_{T}\right) ^{\prime}\mathbf{F}}{T}\right\Vert \left(
\left\Vert \frac{\mathbf{F}^{\prime}\boldsymbol{\varepsilon}_{i\circ}\left(
\lambda_{T}\right) }{T}\right\Vert +\left\Vert \frac{\left( \mathbf{\hat
{F}-F}\right) ^{\prime}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda
_{T}\right) }{T}\right\Vert \right) \left\Vert \left( \frac{\mathbf{F}
^{\prime}\mathbf{F}}{T}\right) ^{-1}\right\Vert \times\\
& \left\Vert \left( \frac{\mathbf{\hat{F}}^{\prime}\mathbf{\hat{F}}}
{T}\right) ^{-1}\right\Vert \left\Vert \frac{\mathbf{F}^{\prime}
\mathbf{\hat{F}}}{T}\right\Vert .
\end{align*}
By ((ref)) and ((ref)) we obtain
\begin{align}
& \sup_{i}\sigma_{i}^{2}\left\Vert \frac{\boldsymbol{\varepsilon}_{i\circ
}\left( \lambda_{T}\right) ^{\prime}\mathbf{F}}{T}\right\Vert \left(
\left\Vert \frac{\mathbf{F}^{\prime}\boldsymbol{\varepsilon}_{i\circ}\left(
\lambda_{T}\right) }{T}\right\Vert +\left\Vert \frac{\left( \mathbf{\hat
{F}-F}\right) ^{\prime}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda
_{T}\right) }{T}\right\Vert \right) \nonumber\\
& \leq\left( \sup_{i}\sigma_{i}^{2}\right) \left[ \sup_{i}\left\Vert
\frac{\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) ^{\prime
}\mathbf{F}}{T}\right\Vert ^{2}+\sup_{i}\left\Vert \frac
{\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) ^{\prime
}\mathbf{F}}{T}\right\Vert \left( \sup_{i}\left\Vert \frac{\left(
\mathbf{\hat{F}-F}\right) ^{\prime}\boldsymbol{\varepsilon}_{i\circ}\left(
\lambda_{T}\right) }{T}\right\Vert \right) \right] \nonumber\\
& =O_{p}\left( \frac{\ln\left( n\right) }{T}\right) +O_{p}\left(
\sqrt{\frac{\ln\left( n\right) }{T}}\right) \times O_{p}\left( \sqrt
{\frac{\ln\left( n\right) }{nT}}\right) =O_{p}\left( \frac{\ln\left(
n\right) }{T}\right) .
\end{align}
Further, by ((ref)),
\begin{equation}
\left\Vert \frac{\mathbf{F}^{\prime}\mathbf{\hat{F}}}{T}\right\Vert
\leq\left\Vert \frac{\mathbf{F}^{\prime}\mathbf{F}}{T}\right\Vert +\left\Vert
\frac{\mathbf{F}^{\prime}\left( \mathbf{\hat{F}-F}\right) }{T}\right\Vert
=O_{p}\left( 1\right) .
\end{equation}
Using ((ref)), ((ref)), and ((ref)) now yields
\begin{align*}
\sup_{i}\left\vert B_{3,2,iT}\right\vert & \leq2\left[ \sup_{i}\sigma
_{i}^{2}\left\Vert \frac{\boldsymbol{\varepsilon}_{i\circ}\left( \lambda
_{T}\right) ^{\prime}\mathbf{F}}{T}\right\Vert \left( \left\Vert
\frac{\mathbf{F}^{\prime}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda
_{T}\right) }{T}\right\Vert +\left\Vert \frac{\left( \mathbf{\hat{F}
-F}\right) ^{\prime}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda
_{T}\right) }{T}\right\Vert \right) \right] \times\\
& \left\Vert \left( \frac{\mathbf{F}^{\prime}\mathbf{F}}{T}\right)
^{-1}\right\Vert \left\Vert \left( \frac{\mathbf{\hat{F}}^{\prime
}\mathbf{\hat{F}}}{T}\right) ^{-1}\right\Vert \left\Vert \frac{\mathbf{F}
^{\prime}\mathbf{\hat{F}}}{T}\right\Vert =O_{p}\left( \frac{\ln\left(
n\right) }{T}\right) .
\end{align*}
Hence,
\[
\sup_{i}\left\vert B_{3,iT}\right\vert \leq\sup_{i}\left( \left\vert
B_{3,1,iT}\right\vert +\left\vert B_{3,2,iT}\right\vert \right) \leq\sup
_{i}\left\vert B_{3,1,iT}\right\vert +\sup_{i}\left\vert B_{3,2,iT}\right\vert
=O_{p}\left( \frac{\ln\left( n\right) }{T}\right) .
\]
Now consider $B_{4,iT}$, and note that
\begin{align*}
B_{4,iT} & =\frac{2\sigma_{i}\boldsymbol{\gamma}_{i}^{\prime}\left(
\mathbf{F-\hat{F}}\right) ^{\prime}\mathbf{M}_{\hat{F}}\mathbf{M}
_{F}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) }{T}\\
& =-\frac{2\sigma_{i}\boldsymbol{\gamma}_{i}^{\prime}\left( \mathbf{\hat
{F}-F}\right) ^{\prime}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda
_{T}\right) }{T}+\frac{2\sigma_{i}\boldsymbol{\gamma}_{i}^{\prime}\left(
\mathbf{\hat{F}-F}\right) ^{\prime}\mathbf{\hat{F}}\left( \mathbf{\hat{F}
}^{\prime}\mathbf{\hat{F}}\right) ^{-1}\mathbf{\hat{F}}^{\prime
}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) }{T}\\
& -\frac{2\sigma_{i}\boldsymbol{\gamma}_{i}^{\prime}\left( \mathbf{\hat
{F}-F}\right) ^{\prime}\mathbf{M}_{\hat{F}}\left( \mathbf{\hat{F}-F}\right)
\left( \mathbf{F}^{\prime}\mathbf{F}\right) ^{-1}\mathbf{F}^{\prime
}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) }{T}=\sum
_{j=1}^{3}B_{4,j,1T}.
\end{align*}
For the first term of the above equation, we have
\[
\left\vert B_{4,1,iT}\right\vert \leq2\left\vert \frac{\sigma_{i}
\boldsymbol{\gamma}_{i}^{\prime}\left( \mathbf{\hat{F}-F}\right) ^{\prime
}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) }{T}\right\vert
\leq2\sigma_{i}\left\Vert \boldsymbol{\gamma}_{i}\right\Vert \left\Vert
\frac{\left( \mathbf{\hat{F}-F}\right) ^{\prime}\boldsymbol{\varepsilon
}_{i\circ}\left( \lambda_{T}\right) }{T}\right\Vert ,
\]
where $\sigma_{i}$\ and $\boldsymbol{\gamma}_{i}$ are bounded. Then using
((ref)) it follows that
\[
\sup_{i}\left\vert B_{4,1,iT}\right\vert \leq2\left( \sup_{i}\sigma
_{i}\right) \left( \sup_{i}\left\Vert \boldsymbol{\gamma}_{i}\right\Vert
\right) \left( \sup_{i}\left\Vert \frac{\left( \mathbf{\hat{F}-F}\right)
^{\prime}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) }
{T}\right\Vert \right) =O_{p}\left( \sqrt{\frac{\ln\left( n\right) }{nT}
}\right) .
\]
Similarly,
\begin{align*}
\left\vert B_{4,2,iT}\right\vert & =\left\vert \frac{2\sigma_{i}
\boldsymbol{\gamma}_{i}^{\prime}\left( \mathbf{\hat{F}-F}\right) ^{\prime
}\mathbf{\hat{F}}\left( \mathbf{\hat{F}}^{\prime}\mathbf{\hat{F}}\right)
^{-1}\mathbf{\hat{F}}^{\prime}\boldsymbol{\varepsilon}_{i\circ}\left(
\lambda_{T}\right) }{T}\right\vert \\
& \leq2\sigma_{i}\left\Vert \boldsymbol{\gamma}_{i}\right\Vert \left\Vert
\frac{\left( \mathbf{\hat{F}-F}\right) ^{\prime}\mathbf{\hat{F}}}
{T}\right\Vert \left\Vert \left( \frac{\mathbf{\hat{F}}^{^{\prime}
}\mathbf{\hat{F}}}{T}\right) ^{-1}\right\Vert \left( \frac{\left\Vert
\mathbf{F}^{\prime}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda
_{T}\right) \right\Vert }{T}+\frac{\left\Vert \left( \mathbf{\hat{F}
-F}\right) ^{\prime}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda
_{T}\right) \right\Vert }{T}\right) ,
\end{align*}
and using ((ref)) and ((ref)) we have
\begin{align*}
\sup_{i}\left\vert B_{4,2,iT}\right\vert & \leq2\left( \sup_{i}\sigma
_{i}\right) \left( \sup_{i}\left\Vert \boldsymbol{\gamma}_{i}\right\Vert
\right) \left( \sup_{i}\frac{\left\Vert \mathbf{F}^{\prime}
\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) \right\Vert }
{T}+\sup_{i}\frac{\left\Vert \left( \mathbf{\hat{F}-F}\right) ^{\prime
}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) \right\Vert }
{T}\right) \\
& \times\left\Vert \frac{\left( \mathbf{\hat{F}-F}\right) ^{\prime
}\mathbf{\hat{F}}}{T}\right\Vert \left\Vert \left( \frac{\mathbf{\hat{F}
}^{^{\prime}}\mathbf{\hat{F}}}{T}\right) ^{-1}\right\Vert =O_{p}\left(
\sqrt{\frac{\ln\left( n\right) }{T}}\right) \times O_{p}\left( \frac
{1}{\delta_{nT}^{2}}\right) .
\end{align*}
Moreover,
\begin{align*}
\left\vert B_{4,3,iT}\right\vert & =\left\vert \frac{2\sigma_{i}
\boldsymbol{\gamma}_{i}^{\prime}\left( \mathbf{\hat{F}-F}\right) ^{\prime
}\mathbf{M}_{\hat{F}}\left( \mathbf{\hat{F}-F}\right) \left( \mathbf{F}
^{\prime}\mathbf{F}\right) ^{-1}\mathbf{F}^{\prime}\boldsymbol{\varepsilon
}_{i\circ}\left( \lambda_{T}\right) }{T}\right\vert \\
& \leq2\sigma_{i}\left\Vert \boldsymbol{\gamma}_{i}\right\Vert \left(
\frac{\left\Vert \mathbf{\hat{F}-F}\right\Vert ^{2}}{T}\right) \left\Vert
\left( \frac{\mathbf{F}^{\prime}\mathbf{F}}{T}\right) ^{-1}\right\Vert
\left\Vert \frac{\mathbf{F}^{\prime}\boldsymbol{\varepsilon}_{i\circ}\left(
\lambda_{T}\right) }{T}\right\Vert ,
\end{align*}
then taking the supremum,
\begin{align*}
\sup_{i}\left\vert B_{4,3,iT}\right\vert & \leq2\left( \sup_{i}\sigma
_{i}\right) \left( \sup_{i}\left\Vert \boldsymbol{\gamma}_{i}\right\Vert
\right) \left( \sup_{i}\left\Vert \frac{\mathbf{F}^{\prime}
\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) }{T}\right\Vert
\right) \left( \frac{\left\Vert \mathbf{\hat{F}-F}\right\Vert ^{2}}
{T}\right) \left\Vert \left( \frac{\mathbf{F}^{\prime}\mathbf{F}}{T}\right)
^{-1}\right\Vert \\
& =O_{p}\left( \sqrt{\frac{\ln\left( n\right) }{T}}\right) \times
O_{p}\left( \frac{1}{\delta_{nT}^{2}}\right) .
\end{align*}
Hence, it follows that
\[
\sup_{i}\left\vert B_{4,iT}\right\vert \leq\sup_{i}\sum_{j=1}^{3}\left\vert
B_{4,j,iT}\right\vert \leq\sum_{j=1}^{3}\sup_{i}\left\vert B_{4,j,iT}
\right\vert =O_{p}\left( \sqrt{\frac{\ln\left( n\right) }{nT}}\right) .
\]
Similarly, $\sup_{i}\left\vert B_{5,iT}\right\vert =O_{p}\left( \sqrt
{\ln\left( n\right) /\left( nT\right) }\right) $. For the final term
$B_{6,iT}$, note that
\begin{align*}
B_{6,iT} & =\frac{\sigma_{i}^{2}\left( \boldsymbol{\varepsilon}_{i\circ
}\left( \lambda_{T}\right) +\boldsymbol{\varepsilon}_{i\circ}\right)
^{\prime}\left( \boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right)
-\boldsymbol{\varepsilon}_{i\circ}\right) }{T}-\frac{\sigma_{i}^{2}\left(
\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right)
+\boldsymbol{\varepsilon}_{i\circ}\right) ^{\prime}\mathbf{F}}{T}\left(
\frac{\mathbf{F}^{\prime}\mathbf{F}}{T}\right) ^{-1}\frac{\mathbf{F}^{\prime
}\left( \boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right)
-\boldsymbol{\varepsilon}_{i\circ}\right) }{T}\\
& =B_{6,1,iT}+B_{6,2,iT}.
\end{align*}
Since $\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right)
=\boldsymbol{\varepsilon}_{i\circ}+\lambda_{T}\mathbf{b}_{i}$ where
$\boldsymbol{\varepsilon}_{i\circ}=\left( \varepsilon_{i1},\varepsilon
_{i2},\ldots,\varepsilon_{iT}\right) ^{\prime}$, $\mathbf{b}_{i}=\left(
b_{i1},b_{i2},\ldots,b_{iT}\right) ^{\prime}$, $b_{it}=\mathbf{w}
_{i0}^{\prime}\boldsymbol{\varepsilon}_{\circ t}$, $\mathbf{w}_{i0}=\left(
w_{i1},w_{is},\ldots,w_{in}\right) ^{\prime}$ and $\boldsymbol{\varepsilon
}_{\circ t}=\left( \varepsilon_{1t},\varepsilon_{2t},\ldots,\varepsilon
_{nt}\right) ^{\prime}$, then
\[
B_{6,1,iT}=\sigma_{i}^{2}\left( \frac{2\lambda_{T}\boldsymbol{\varepsilon
}_{i\circ}^{\prime}\mathbf{b}_{i}}{T}+\frac{\lambda_{T}^{2}\mathbf{b}
_{i}^{\prime}\mathbf{b}_{i}}{T}\right) .
\]
Further, using results ((ref)) and ((ref)) we have
\[
\sup_{i}\left\vert B_{6,1,iT}\right\vert \leq\left( \sup_{i}\sigma_{i}
^{2}\right) \left( 2\left\vert \lambda_{T}\right\vert \sup_{i}\left\vert
\frac{\boldsymbol{\varepsilon}_{i\circ}^{\prime}\mathbf{b}_{i}}{T}\right\vert
+\lambda_{T}^{2}\sup_{i}\left\vert \frac{\mathbf{b}_{i}^{\prime}\mathbf{b}
_{i}}{T}\right\vert \right) =O_{p}\left( \frac{\sqrt{\ln\left( n\right) }
}{T}\right) +O_{p}\left( \frac{1}{T}\right) .
\]
Now consider $B_{6,2,iT}$ and note that
\[
\left\vert B_{6,2,iT}\right\vert \leq\sigma_{i}^{2}\left\Vert \frac{\left(
\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right)
+\boldsymbol{\varepsilon}_{i\circ}\right) ^{\prime}\mathbf{F}}{T}\right\Vert
\left\Vert \left( \frac{\mathbf{F}^{\prime}\mathbf{F}}{T}\right)
^{-1}\right\Vert \left\Vert \frac{\mathbf{F}^{\prime}\left(
\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right)
-\boldsymbol{\varepsilon}_{i\circ}\right) }{T}\right\Vert .
\]
Further, using ((ref)) we have
\begin{align*}
\sup_{i}\left\Vert \frac{\left( \boldsymbol{\varepsilon}_{i\circ}\left(
\lambda_{T}\right) +\boldsymbol{\varepsilon}_{i\circ}\right) ^{\prime
}\mathbf{F}}{T}\right\Vert & \leq\sup_{i}\left\Vert \frac
{\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) ^{\prime
}\mathbf{F}}{T}\right\Vert +\sup_{i}\left\Vert \frac{\boldsymbol{\varepsilon
}_{i\circ}^{\prime}\mathbf{F}}{T}\right\Vert =O_{p}\left( \sqrt{\ln\left(
n\right) /T}\right) ,\\
\sup_{i}\left\Vert \frac{\mathbf{F}^{\prime}\left( \boldsymbol{\varepsilon
}_{i\circ}\left( \lambda_{T}\right) -\boldsymbol{\varepsilon}_{i\circ
}\right) }{T}\right\Vert & \leq\sup_{i}\left\Vert \frac
{\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) ^{\prime
}\mathbf{F}}{T}\right\Vert +\sup_{i}\left\Vert \frac{\boldsymbol{\varepsilon
}_{i\circ}^{\prime}\mathbf{F}}{T}\right\Vert =O_{p}\left( \sqrt{\ln\left(
n\right) /T}\right) .
\end{align*}
Hence, it follows that
\begin{align*}
\sup_{i}\left\vert B_{6,2,iT}\right\vert & \leq\left\Vert \left(
\frac{\mathbf{F}^{\prime}\mathbf{F}}{T}\right) ^{-1}\right\Vert \left(
\sup_{i}\sigma_{i}^{2}\right) \left( \sup_{i}\left\Vert \frac{\left(
\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right)
+\boldsymbol{\varepsilon}_{i\circ}\right) ^{\prime}\mathbf{F}}{T}\right\Vert
\right) \\
& \times\left( \sup_{i}\left\Vert \frac{\mathbf{F}^{\prime}\left(
\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right)
-\boldsymbol{\varepsilon}_{i\circ}\right) }{T}\right\Vert \right)
=O_{p}\left( \frac{\ln\left( n\right) }{T}\right) .
\end{align*}
Overall,
\[
\sup_{i}\left\vert B_{6,iT}\right\vert \leq\sup_{i}\left\vert B_{6,1,iT}
\right\vert +\sup_{i}\left\vert B_{6,2,iT}\right\vert =O_{p}\left( \frac
{\ln\left( n\right) }{T}\right) .
\]
Using the above results of $B_{1,iT}$ to $B_{6,iT}$, and noting that $n$ and
$T$ are assumed to be of the same order of magnitude, we obtain
\[
\sup_{i}\left\vert \hat{\sigma}_{i,T}^{2}-\omega_{i,T}^{2}\right\vert \leq
\sup_{i}\sum_{j=1}^{6}\left\vert B_{j,iT}\right\vert \leq\sum_{j=1}^{6}
\sup_{i}\left\vert B_{j,iT}\right\vert =O_{p}\left( \frac{\ln\left(
n\right) }{T}\right) ,
\]
so ((ref)) is established. To prove ((ref)),
note that
\[
\sup_{i}\left\vert \hat{\sigma}_{i,T}-\omega_{i,T}\right\vert =\sup_{i}
\frac{\left\vert \hat{\sigma}_{i,T}^{2}-\omega_{i,T}^{2}\right\vert }
{\hat{\sigma}_{i,T}+\omega_{i,T}}=\sup_{i}\left( \frac{1}{\hat{\sigma}
_{i,T}+\omega_{i,T}}\right) \sup_{i}\left\vert \hat{\sigma}_{i,T}^{2}
-\omega_{i,T}^{2}\right\vert .
\]
Since $\hat{\sigma}_{i,T}^{2}>0$ (by construction)
\[
\sup_{i}\left\vert \hat{\sigma}_{i,T}-\omega_{i,T}\right\vert \leq\sup
_{i}\left( \frac{1}{\omega_{i,T}}\right) \sup_{i}\left\vert \hat{\sigma
}_{i,T}^{2}-\omega_{i,T}^{2}\right\vert .
\]
By definition of $\omega_{i,T}$, we have $\omega_{i,T}^{-2}=\sigma_{i}
^{-2}\left( T^{-1}\boldsymbol{\varepsilon}_{i\circ}^{\prime}\mathbf{M}
_{F}\boldsymbol{\varepsilon}_{i\circ}\right) ^{-1}$, so that
\begin{align*}
\sup_{i}\left( \frac{1}{\omega_{i,T}}\right) & =\sup_{i}\left( \frac
{1}{\omega_{i,T}^{2}}\right) ^{1/2}\leq\left( \sup_{i}\frac{1}{\sigma
_{i}^{2}}\right) ^{1/2}\left( \sup_{i}\frac{1}{T^{-1}\boldsymbol{\varepsilon
}_{i\circ}^{\prime}\mathbf{M}_{F}\boldsymbol{\varepsilon}_{i\circ}}\right)
^{1/2}\\
& =\left( \frac{1}{\inf_{i}\sigma_{i}^{2}}\right) ^{1/2}\left( \frac
{1}{\inf_{i}\left( T^{-1}\boldsymbol{\varepsilon}_{i\circ}^{\prime}
\mathbf{M}_{F}\boldsymbol{\varepsilon}_{i\circ}\right) }\right) ^{1/2}.
\end{align*}
Since $\inf_{i}\sigma_{i}^{2}>c>0$ and by condition ((ref)) in
Assumption (ref), $\inf_{i}T^{-1}\boldsymbol{\varepsilon}_{i\circ
}^{\prime}\mathbf{M}_{F}\boldsymbol{\varepsilon}_{i\circ}>c>0$ as
$T\rightarrow\infty$, then it readily follows $\sup_{i}\omega_{i,T}
^{-1}<C<\infty$, which together with ((ref)) now establishes
result ((ref)). Similarly, note that
\[
\sup_{i}\left\vert \frac{1}{\hat{\sigma}_{i,T}}-\frac{1}{\omega_{i,T}
}\right\vert \leq\sup_{i}\left( \frac{1}{\hat{\sigma}_{i,T}}\right) \sup
_{i}\left( \frac{1}{\omega_{i,T}}\right) \sup_{i}\left\vert \hat{\sigma
}_{i,T}-\omega_{i,T}\right\vert ,
\]
where $\sup_{i}\left( \frac{1}{\hat{\sigma}_{i,T}}\right) <C$, by
construction. Hence, result ((ref)) can be established
using ((ref)). Finally, results ((ref)),((ref)) and
((ref)) follow using ((ref)), ((ref)) and
((ref)), respectively.
lemmaConsider the latent factor model given by ((ref)) and
((ref)). The latent factors, $\mathbf{f}_{t}$, and their loadings,
$\boldsymbol{\gamma}_{i}$, are estimated by principal components,
$\mathbf{\hat{f}}_{t}$ and $\boldsymbol{\hat{\gamma}}_{i}$, given by
((ref)). Suppose that Assumptions (ref)-(ref) hold and
$\left( n,T\right) \rightarrow\infty$, such that $n/T\rightarrow\kappa\,,$
for $0<\kappa<\infty$. Then
\begin{align}
\mathbf{d}_{1,nT} & =\frac{1}{n}\sum_{i=1}^{n}b_{in}(\boldsymbol{\hat
{\gamma}}_{i}-\boldsymbol{\gamma}_{i})=O_{p}\left( \sqrt{\frac{\ln\left(
n\right) }{nT}}\right) ,\\
\mathbf{d}_{2,nT} & =\frac{1}{n}\sum_{i=1}^{n}(\boldsymbol{\hat{\delta}
}_{i,T}-\boldsymbol{\delta}_{i,T})=O_{p}\left( \sqrt{\frac{\ln\left(
n\right) }{nT}}\right) ,\\
\mathbf{d}_{3,nT} & =\frac{1}{n}\sum_{i=1}^{n}b_{in}\left( \boldsymbol{\hat
{\gamma}}_{i}-\boldsymbol{\gamma}_{i}\right) \boldsymbol{\gamma}_{i}^{\prime
}=O_{p}\left( \sqrt{\frac{\ln\left( n\right) }{nT}}\right) ,\\
\mathbf{d}_{4,nT} & =\frac{1}{n}\sum_{i=1}^{n}\left( \omega_{i,T}
-\sigma_{i}\right) (\boldsymbol{\hat{\gamma}}_{i}-\boldsymbol{\gamma}
_{i})=O_{p}\left( \frac{1}{\delta_{nT}^{2}}\right) ,\\
\mathbf{d}_{5,nT} & =\frac{1}{n}\sum_{i=1}^{n}\left( \frac{1}{\omega_{i,T}
}-\frac{1}{\sigma_{i}}\right) (\boldsymbol{\hat{\gamma}}_{i}
-\boldsymbol{\gamma}_{i})=O_{p}\left( \frac{1}{\delta_{nT}^{2}}\right)
,\\
\mathbf{d}_{6,nT} & =\frac{1}{n}\sum_{i=1}^{n}\left( \hat{\sigma}
_{i,T}-\omega_{i,T}\right) (\boldsymbol{\hat{\gamma}}_{i}-\boldsymbol{\gamma
}_{i})=O_{p}\left[ \left( \frac{\ln\left( n\right) }{T}\right)
^{3/2}\right] ,\\
\mathbf{d}_{7,nT} & =\frac{1}{n}\sum_{i=1}^{n}\left( \frac{1}{\hat{\sigma
}_{i,T}}-\frac{1}{\omega_{i,T}}\right) (\boldsymbol{\hat{\gamma}}
_{i}-\boldsymbol{\gamma}_{i})=O_{p}\left[ \left( \frac{\ln\left( n\right)
}{T}\right) ^{3/2}\right] ,
\end{align}
where $\left\{ b_{in}\right\} _{i=1}^{n}$ is a sequence of fixed values
bounded in $n$, such that $n^{-1}\sum_{i=1}^{n}b_{in}^{2}=O(1)$,
$\boldsymbol{\delta}_{i,T}=\boldsymbol{\gamma}_{i}/\omega_{i,T},$
$\boldsymbol{\hat{\delta}}_{i,T}=\hat{\boldsymbol{\gamma}}_{i}/\omega_{i,T},$
and $\omega_{i,T}=\left( T^{-1}\sigma_{i}^{2}\boldsymbol{\varepsilon}
_{i\circ}^{\prime}\mathbf{M}_{F}\boldsymbol{\varepsilon}_{i\circ}\right)
^{1/2}.$
proofNote that in general
\begin{align}
\boldsymbol{\hat{\gamma}}_{i}-\boldsymbol{\gamma}_{i} & =\left(
\frac{\mathbf{\hat{F}}^{\prime}\mathbf{\hat{F}}}{T}\right) ^{-1}\left(
\frac{\mathbf{\hat{F}}^{\prime}\mathbf{F}\boldsymbol{\gamma}_{i}}{T}
+\frac{\sigma_{i}\mathbf{\hat{F}}^{\prime}\boldsymbol{\varepsilon}_{i\circ
}\left( \lambda_{T}\right) }{T}\right) -\left( \frac{\mathbf{\hat{F}
}^{\prime}\mathbf{\hat{F}}}{T}\right) ^{-1}\left( \frac{\mathbf{\hat{F}
}^{\prime}\mathbf{\hat{F}}}{T}\right) \boldsymbol{\gamma}_{i}\nonumber\\
& =-\left( \frac{\mathbf{\hat{F}}^{\prime}\mathbf{\hat{F}}}{T}\right)
^{-1}\left[ \frac{\mathbf{\hat{F}}^{\prime}\left( \mathbf{\hat{F}-F}\right)
\boldsymbol{\gamma}_{i}}{T}\right] +\left( \frac{\mathbf{\hat{F}}^{\prime
}\mathbf{\hat{F}}}{T}\right) ^{-1}\left( \frac{\sigma_{i}\mathbf{\hat{F}
}^{\prime}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) }
{T}\right) ,
\end{align}
and we have
\begin{align}
\mathbf{d}_{1,nT} & =\frac{1}{n}\sum_{i=1}^{n}b_{in}\left( \boldsymbol{\hat
{\gamma}}_{i}-\boldsymbol{\gamma}_{i}\right) \nonumber\\
& =-\left( \frac{\mathbf{\hat{F}}^{\prime}\mathbf{\hat{F}}}{T}\right)
^{-1}\left[ \frac{\mathbf{\hat{F}}^{\prime}\left( \mathbf{\hat{F}-F}\right)
}{T}\right] \left( \frac{1}{n}\sum_{i=1}^{n}b_{in}\boldsymbol{\gamma}
_{i}\right) +\left( \frac{\mathbf{\hat{F}}^{\prime}\mathbf{\hat{F}}}
{T}\right) ^{-1}T^{-1}\left( \frac{1}{n}\sum_{i=1}^{n}b_{in}\sigma
_{i}\mathbf{\hat{F}}^{\prime}\boldsymbol{\varepsilon}_{i\circ}\left(
\lambda_{T}\right) \right) .
\end{align}
Since by assumption $\left\Vert \boldsymbol{\gamma}_{i}\right\Vert <C$, we
have
\[
\left\Vert \frac{1}{n}\sum_{i=1}^{n}b_{in}\boldsymbol{\gamma}_{i}\right\Vert
\leq\left( \frac{1}{n}\sum_{i=1}^{n}b_{in}^{2}\right) ^{1/2}\left\Vert
\frac{1}{n}\sum_{i=1}^{n}\boldsymbol{\gamma}_{i}\boldsymbol{\gamma}
_{i}^{^{\prime}}\right\Vert ^{1/2}\leq\left( \frac{1}{n}\sum_{i=1}^{n}
b_{in}^{2}\right) ^{1/2}\left( \frac{1}{n}\sum_{i=1}^{n}\left\Vert
\boldsymbol{\gamma}_{i}\right\Vert ^{2}\right) ^{1/2}<C.
\]
Also by results ((ref)) and ((ref)), the first term of
((ref)) is $O_{p}\left( \delta_{nT}^{-2}\right) $. For the second term
of ((ref)), since $\left( T^{-1}\mathbf{\hat{F}}^{\prime}\mathbf{\hat
{F}}\right) ^{-1}=O_{p}(1)$, we note that
\[
T^{-1}\left( \mathbf{\hat{F}}-\mathbf{F+F}\right) ^{\prime}\left( \frac
{1}{n}\sum_{i=1}^{n}b_{in}\sigma_{i}\boldsymbol{\varepsilon}_{i\circ}\left(
\lambda_{T}\right) \right) =\frac{1}{n}\sum_{i=1}^{n}b_{in}\left(
\frac{\mathbf{\hat{F}}-\mathbf{F}}{T}\right) ^{\prime}\sigma_{i}
\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) +\frac{1}{Tn}
\sum_{i=1}^{n}b_{in}\mathbf{F}^{^{\prime}}\sigma_{i}\boldsymbol{\varepsilon
}_{i\circ}\left( \lambda_{T}\right) .
\]
Using result ((ref)), we have
\begin{align}
& \left\Vert \frac{1}{n}\sum_{i=1}^{n}b_{in}\left( \frac{\mathbf{\hat{F}
}-\mathbf{F}}{T}\right) ^{\prime}\sigma_{i}\boldsymbol{\varepsilon}_{i\circ
}\left( \lambda_{T}\right) \right\Vert \nonumber\\
& \leq\frac{1}{n}\sum_{i=1}^{n}\left\vert b_{in}\sigma_{i}\right\vert
\left\Vert \left( \frac{\mathbf{\hat{F}}-\mathbf{F}}{T}\right) ^{\prime
}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) \right\Vert
\leq\left( \sup_{i}\left\vert b_{in}\right\vert \right) \left( \sup
_{i}\sigma_{i}\right) \left( \sup_{i}\left\Vert \frac{\left( \mathbf{\hat
{F}}-\mathbf{F}\right) ^{\prime}\boldsymbol{\varepsilon}_{i\circ}\left(
\lambda_{T}\right) }{T}\right\Vert \right) \nonumber\\
& =O_{p}\left( \sqrt{\frac{\ln\left( n\right) }{nT}}\right) .
\end{align}
Under part (a) of Assumption (ref) and by the serial independence
of $\varepsilon_{it}$,
\begin{align*}
E\left\Vert \frac{1}{\sqrt{nT}}\sum_{i=1}^{n}\sum_{t=1}^{T}b_{in}
\mathbf{f}_{t}\sigma_{i}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda
_{T}\right) \right\Vert ^{2} & =\frac{1}{nT}\sum_{i=1}^{n}\sum_{j=1}
^{n}\sum_{t=1}^{T}b_{in}b_{jn}\sigma_{i}\sigma_{j}E\left\Vert \mathbf{f}
_{t}\right\Vert ^{2}E\left( \varepsilon_{it}\left( \lambda_{T}\right)
\varepsilon_{jt}\left( \lambda_{T}\right) \right) \\
& \leq E\left\Vert \mathbf{f}_{t}\right\Vert ^{2}\left( \sup_{i}b_{in}
^{2}\right) \left( \sup_{i}\sigma_{i}^{2}\right) \left[ \frac{1}{nT}
\sum_{t=1}^{T}\sum_{i=1}^{n}\sum_{j=1}^{n}\left\vert E\left( \varepsilon
_{it}\left( \lambda_{T}\right) \varepsilon_{jt}\left( \lambda_{T}\right)
\right) \right\vert \right]
\end{align*}
which is $O\left( 1\right) $ based on ((ref)) and the boundedness of
$E\left\Vert \mathbf{f}_{t}\right\Vert ^{2}$, $b_{in}^{2}$ and $\sigma_{i}
^{2}$ required by assumptions. So it follows
\begin{equation}
\frac{1}{nT}\sum_{i=1}^{n}b_{in}\mathbf{F}^{^{\prime}}\sigma_{i}
\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right) =\frac{1}
{\sqrt{nT}}\left( \frac{1}{\sqrt{nT}}\sum_{i=1}^{n}\sum_{t=1}^{T}
b_{in}\mathbf{f}_{t}\sigma_{i}\boldsymbol{\varepsilon}_{i\circ}\left(
\lambda_{T}\right) \right) =O_{p}\left( \frac{1}{\sqrt{nT}}\right) .
\end{equation}
Result ((ref)) now follows using ((ref)) and ((ref)) in
((ref)), and noting that by assumption $n$ and $T$ are of the same order.
Consider now ((ref)), which can be written as
\begin{align*}
\mathbf{d}_{2,nT} & =\frac{1}{n}\sum_{i=1}^{n}\left( \frac{\boldsymbol{\hat
{\gamma}}_{i}}{\omega_{i,T}}-\frac{\boldsymbol{\gamma}_{i}}{\omega_{i,T}
}\right) \\
& =\frac{1}{n}\sum_{i=1}^{n}\left( \frac{\boldsymbol{\hat{\gamma}}
_{i}-\boldsymbol{\gamma}_{i}}{\sigma_{i}}\right) \left( 1-\frac{\omega
_{i,T}-\sigma_{i}}{\omega_{i,T}}\right) \\
& =\frac{1}{n}\sum_{i=1}^{n}\left( \frac{\boldsymbol{\hat{\gamma}}
_{i}-\boldsymbol{\gamma}_{i}}{\sigma_{i}}\right) -\frac{1}{n}\sum_{i=1}
^{n}\left( \frac{\boldsymbol{\hat{\gamma}}_{i}-\boldsymbol{\gamma}_{i}
}{\sigma_{i}}\right) \left( 1-\frac{T}{\boldsymbol{\varepsilon}_{i\circ
}^{\prime}\mathbf{M}_{F}\boldsymbol{\varepsilon}_{i\circ}}\right) .
\end{align*}
The first term of the above has the same form as ((ref)), and becomes
identical to it if we replace $a_{i}$ in ((ref)) with $1/\sigma_{i}$,
since by assumption $\inf_{i}(\sigma_{i})>c$. Hence, the order of the first
term is $O_{p}(\sqrt{\ln\left( n\right) /\left( nT\right) })$. Also the
second term is dominated by the first term, since $1-\left( T^{-1}
\boldsymbol{\varepsilon}_{i\circ}^{\prime}\mathbf{M}_{F}
\boldsymbol{\varepsilon}_{i\circ}\right) ^{-1}=O_{p}(T^{-1})$ based on result
((ref)). Therefore, ((ref)) is established as required. For
((ref)), note that
\begin{align}
\mathbf{d}_{3,nT} & =\frac{1}{n}\sum_{i=1}^{n}b_{in}\left[ -\left(
\mathbf{\hat{F}}^{\prime}\mathbf{\hat{F}}\right) ^{-1}\mathbf{\hat{F}
}^{\prime}\left( \mathbf{\hat{F}-F}\right) \boldsymbol{\gamma}_{i}
+\sigma_{i}\left( \mathbf{\hat{F}}^{\prime}\mathbf{\hat{F}}\right)
^{-1}\mathbf{\hat{F}}^{\prime}\boldsymbol{\varepsilon}_{i\circ}\left(
\lambda_{T}\right) \right] \boldsymbol{\gamma}_{i}^{\prime}\nonumber\\
& =-\left( \frac{\mathbf{\hat{F}}^{\prime}\mathbf{\hat{F}}}{T}\right)
^{-1}\frac{\mathbf{\hat{F}}^{\prime}\left( \mathbf{\hat{F}-F}\right) }
{T}\left( \frac{1}{n}\sum_{i=1}^{n}b_{in}\boldsymbol{\gamma}_{i}
\boldsymbol{\gamma}_{i}^{\prime}\right) +\left( \frac{\mathbf{\hat{F}
}^{\prime}\mathbf{\hat{F}}}{T}\right) ^{-1}T^{-1}\mathbf{\hat{F}}^{\prime
}\left( \frac{1}{n}\sum_{i=1}^{n}b_{in}\sigma_{i}\boldsymbol{\varepsilon
}_{i\circ}\left( \lambda_{T}\right) \boldsymbol{\gamma}_{i}^{\prime}\right)
\nonumber\\
& =-\left( \frac{\mathbf{\hat{F}}^{\prime}\mathbf{\hat{F}}}{T}\right)
^{-1}\frac{\mathbf{\hat{F}}^{\prime}\left( \mathbf{\hat{F}-F}\right) }
{T}\left( \frac{1}{n}\sum_{i=1}^{n}b_{in}\boldsymbol{\gamma}_{i}
\boldsymbol{\gamma}_{i}^{\prime}\right) +\left( \frac{\mathbf{\hat{F}
}^{\prime}\mathbf{\hat{F}}}{T}\right) ^{-1}\left( \frac{1}{n}\sum_{i=1}
^{n}T^{-1}b_{in}\sigma_{i}\left( \mathbf{\hat{F}-F}\right) ^{\prime
}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right)
\boldsymbol{\gamma}_{i}^{\prime}\right) \nonumber\\
& +\left( \frac{\mathbf{\hat{F}}^{\prime}\mathbf{\hat{F}}}{T}\right)
^{-1}\left( \frac{1}{n}\sum_{i=1}^{n}T^{-1}b_{in}\sigma_{i}\mathbf{F}
^{\prime}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right)
\boldsymbol{\gamma}_{i}^{\prime}\right) .
\end{align}
Recall that $\left( T^{-1}\mathbf{\hat{F}}^{\prime}\mathbf{\hat{F}}\right)
^{-1}=O_{p}(1)$, and $n^{-1}\sum_{i=1}^{n}\boldsymbol{\gamma}_{i}
\boldsymbol{\gamma}_{i}^{\prime}=O_{p}(1)$. Also note that $b_{in}$ is bounded
in $n$. Then using ((ref)) it follows that ($n$ and $T$ being of the
same order)
\[
\left( \frac{\mathbf{\hat{F}}^{\prime}\mathbf{\hat{F}}}{T}\right) ^{-1}
\frac{\mathbf{\hat{F}}^{\prime}\left( \mathbf{\hat{F}-F}\right) }{T}\left(
n^{-1}\sum_{i=1}^{n}b_{in}\boldsymbol{\gamma}_{i}\boldsymbol{\gamma}
_{i}^{\prime}\right) =O_{p}\left( \frac{1}{\min(n,T)}\right) =O_{p}
(\delta_{nT\ }^{-2}).
\]
Similarly, using ((ref))
\[
\left( \frac{\mathbf{\hat{F}}^{\prime}\mathbf{\hat{F}}}{T}\right)
^{-1}\left( n^{-1}\sum_{i=1}^{n}T^{-1}\left( \mathbf{\hat{F}-F}\right)
^{\prime}b_{in}\sigma_{i}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda
_{T}\right) \boldsymbol{\gamma}_{i}^{\prime}\right) =O_{p}\left(
\sqrt{\frac{\ln\left( n\right) }{nT}}\right) .
\]
The last term of ( (ref))\ can be written as $\left( nT\right)
^{-1/2}\left( T^{-1}\mathbf{\hat{F}}^{\prime}\mathbf{\hat{F}}\right)
^{-1}\left( n^{-1/2}T^{-1/2}\sum_{i=1}^{n}b_{in}\sigma_{i}\mathbf{F}^{\prime
}\boldsymbol{\varepsilon}_{i\circ}\left( \lambda_{T}\right)
\boldsymbol{\gamma}_{i}^{\prime}\right) ,$ where $n^{-1/2}T^{-1/2}\sum
_{i=1}^{n}b_{in}\sigma_{i}\mathbf{F}^{\prime}\boldsymbol{\varepsilon}_{i\circ
}\left( \lambda_{T}\right) \boldsymbol{\gamma}_{i}^{\prime}$ is an
$m_{0}\times m_{0}$ matrix with its $(j,j^{\prime})$ element given by
$n^{-1/2}T^{-1/2}\sum_{i=1}^{n}\sum_{t=1}^{T}$ $b_{in}\sigma_{i}
f_{jt}\varepsilon_{it}\left( \lambda_{T}\right) \gamma_{ij^{\prime}}$ for
$j,j^{\prime}=1,2,\ldots,m_{0}$. It can be further shown
\begin{align*}
& E\left( n^{-1/2}T^{-1/2}\sum_{i=1}^{n}\sum_{t=1}^{T}b_{in}\sigma_{i}
f_{jt}\gamma_{ij^{\prime}}\varepsilon_{it}\left( \lambda_{T}\right) \right)
^{2}\\
& =\frac{1}{nT}\sum_{i=1}^{n}\sum_{i^{\prime}=1}^{n}\sum_{t=1}^{T}
b_{in}b_{i^{\prime}n}\sigma_{i}\sigma_{i^{\prime}}\gamma_{ij^{\prime}}
\gamma_{i^{\prime}j^{\prime}}E\left( f_{jt}^{2}\right) E\left(
\varepsilon_{it}\left( \lambda_{T}\right) \varepsilon_{i^{\prime}t}\left(
\lambda_{T}\right) \right) \\
& \leq\frac{1}{nT}\sum_{i=1}^{n}\sum_{i^{\prime}=1}^{n}\sum_{t=1}
^{T}\left\vert b_{in}b_{i^{\prime}n}\sigma_{i}\sigma_{i^{\prime}}
\gamma_{ij^{\prime}}\gamma_{i^{\prime}j^{\prime}}\right\vert E\left(
f_{jt}^{2}\right) \left\vert E\left( \varepsilon_{it}\left( \lambda
_{T}\right) \varepsilon_{i^{\prime}t}\left( \lambda_{T}\right) \right)
\right\vert \\
& \leq\left( \sup_{i}b_{in}^{2}\right) \left( \sup_{i}\gamma_{ij^{\prime}
}^{2}\right) \left( \sup_{i}\sigma_{i}^{2}\right) E\left( f_{jt}
^{2}\right) \left[ \frac{1}{nT}\sum_{i=1}^{n}\sum_{i^{\prime}=1}^{n}
\sum_{t=1}^{T}\left\vert E\left( \varepsilon_{it}\left( \lambda_{T}\right)
\varepsilon_{i^{\prime}t}\left( \lambda_{T}\right) \right) \right\vert
\right]
\end{align*}
which is $O\left( 1\right) $ based on ((ref)) and the boundedness of
$b_{in}^{2}$, $\gamma_{ij^{\prime}}^{2}$, $\sigma_{i}^{2}$ and $E\left(
f_{jt}^{2}\right) $. Consequentially, $n^{-1/2}T^{-1/2}\sum_{i=1}^{n}
b_{in}\sigma_{i}\mathbf{F}^{\prime}\boldsymbol{\varepsilon}_{i\circ}\left(
\lambda_{T}\right) \boldsymbol{\gamma}_{i}^{\prime}=O_{p}\left( 1\right) $
and the last term of ((ref)) are also $O_{p}(\delta_{nT\ }^{-2})$. Thus
result ((ref)) is established, as required. To prove ((ref)) we
first write it as
\[
\mathbf{d}_{4,nT}=\left( \sqrt{\frac{1}{T}}\right) \frac{1}{n}\sum_{i=1}
^{n}q_{iT}\left( \boldsymbol{\hat{\gamma}}_{i}-\boldsymbol{\gamma}
_{i}\right) ,
\]
where
\[
q_{iT}=\sqrt{T}\left( \omega_{i,T}-\sigma_{i}\right) =\sigma_{i}\sqrt
{T}\left[ \left( \frac{\boldsymbol{\varepsilon}_{i\circ}^{\prime}
\mathbf{M}_{F}\boldsymbol{\varepsilon}_{i\circ}}{T}\right) ^{1/2}-1\right]
,
\]
and conditional on $\mathbf{F}$ and $\sigma_{i}$, $q_{iT}$ are independently
distributed across $i$. Using results in Lemma (ref) it is easily seen
that $E\left( q_{iT}\right) =O(T^{-1/2})$ and $Var\left( q_{iT}\right)
=O(1)$, and hence $n^{-1}\sum_{i=1}^{n}q_{iT}^{2}=O_{p}(1)$. Also by
Cauchy-Schwarz inequality we have
\[
\left\Vert \mathbf{d}_{4,nT}\right\Vert \leq\left( \sqrt{\frac{1}{T}}\right)
\left( n^{-1}\sum_{i=1}^{n}q_{iT}^{2}\right) ^{1/2}\left( n^{-1/2}
\left\Vert \mathbf{\hat{\Gamma}}-\mathbf{\Gamma}\right\Vert \right) ,
\]
where $T^{-1}n=\ominus(1)$, and by ((ref)) $n^{-1/2}\left\Vert
\mathbf{\hat{\Gamma}}-\mathbf{\Gamma}\right\Vert =O_{p}(\delta_{nT}^{-1})$,
and ((ref)) is established. Result ((ref)) follows similarly, with
$q_{iT}$ re-defined as $q_{iT}=\sigma_{i}^{-1}\sqrt{T}\left[ \left(
\frac{\boldsymbol{\varepsilon}_{i\circ}^{\prime}\mathbf{M}_{F}
\boldsymbol{\varepsilon}_{i\circ}}{T}\right) ^{-1/2}-1\right] $, and noting
that $\sup_{i}(1/\sigma_{i}^{2})<C$, and using results in Lemma (ref).
Result ((ref)) is established as
\begin{align*}
\left\Vert \mathbf{d}_{6,nT}\right\Vert & =\left\Vert \frac{1}{n}\sum
_{i=1}^{n}\left( \hat{\sigma}_{i,T}-\omega_{i,T}\right) \left(
\boldsymbol{\hat{\gamma}}_{i}-\boldsymbol{\gamma}_{i}\right) \right\Vert
\leq\frac{1}{n}\sum_{i=1}^{n}\left\Vert \left( \hat{\sigma}_{i,T}
-\omega_{i,T}\right) \left( \boldsymbol{\hat{\gamma}}_{i}-\boldsymbol{\gamma
}_{i}\right) \right\Vert \\
& \leq\left( \frac{1}{n}\sum_{i=1}^{n}\left\vert \hat{\sigma}_{i,T}
-\omega_{i,T}\right\vert \right) \left( \sup_{i}\left\Vert \boldsymbol{\hat
{\gamma}}_{i}-\boldsymbol{\gamma}_{i}\right\Vert \right) =O_{p}\left[
\left( \frac{\ln\left( n\right) }{T}\right) ^{3/2}\right]
\end{align*}
where by ((ref)) $\sup_{i}\left\Vert \boldsymbol{\hat{\gamma}
}_{i}-\boldsymbol{\gamma}_{i}\right\Vert =O_{p}\left( \sqrt{\ln\left(
n\right) /T}\right) $, and by ((ref)) $n^{-1}\sum_{i=1}^{n}\left\vert
\hat{\sigma}_{i,T}-\omega_{i,T}\right\vert =O_{p}(\ln\left( n\right) /T)$.
Similarly by ((ref)) and ((ref)) we have
\begin{align*}
\left\Vert \mathbf{d}_{7,nT}\right\Vert & =\left\Vert \frac{1}{n}\sum
_{i=1}^{n}\left( \frac{1}{\hat{\sigma}_{i,T}}-\frac{1}{\omega_{i,T}}\right)
\left( \boldsymbol{\hat{\gamma}}_{i}-\boldsymbol{\gamma}_{i}\right)
\right\Vert \leq\frac{1}{n}\sum_{i=1}^{n}\left\Vert \left( \frac{1}
{\hat{\sigma}_{i,T}}-\frac{1}{\omega_{i,T}}\right) \left( \boldsymbol{\hat
{\gamma}}_{i}-\boldsymbol{\gamma}_{i}\right) \right\Vert \\
& \leq\left( \frac{1}{n}\sum_{i=1}^{n}\left\vert \frac{1}{\hat{\sigma}
_{i,T}}-\frac{1}{\omega_{i,T}}\right\vert \right) \left( \sup_{i}\left\Vert
\boldsymbol{\hat{\gamma}}_{i}-\boldsymbol{\gamma}_{i}\right\Vert \right)
=O_{p}\left[ \left( \frac{\ln\left( n\right) }{T}\right) ^{3/2}\right] .
\end{align*}
lemmaConsider the latent factor model given by ((ref)) and
((ref)). Suppose that Assumptions (ref)-(ref) hold
and $\left( n,T\right) \rightarrow\infty$, such that $n/T\rightarrow
\kappa\,,$ for $0<\kappa<\infty$. Then
\begin{align}
p_{nT}\left( \lambda_{T}\right) & =\frac{1}{\sqrt{T}}\sum_{t=1}
^{T}s_{t,nT}^{2}\left( \lambda_{T}\right) =o_{p}(1),\\
q_{nT}\left( \lambda_{T}\right) & =\frac{1}{\sqrt{T}}\sum_{t=1}^{T}
\psi_{t,nT}\left( \lambda_{T}\right) s_{t,nT}\left( \lambda_{T}\right)
=o_{p}(1),
\end{align}
where $\psi_{t,nT}\left( \lambda_{T}\right) $ and $s_{t,nT}\left(
\lambda_{T}\right) $ are defined by ((ref)) and ((ref)), respectively.
proofUsing ((ref)), recall that
\begin{align}
s_{t,nT}\left( \lambda_{T}\right) & =\boldsymbol{\varphi}_{nT}^{\prime
}\left[ n^{-1/2}\sum_{i=1}^{n}\left( \boldsymbol{\hat{\gamma}}
_{i}-\boldsymbol{\gamma}_{i}\right) \sigma_{i}\varepsilon_{it}\left(
\lambda_{T}\right) \right] +\boldsymbol{\varphi}_{nT}^{\prime}\left[
n^{-1/2}\sum_{i=1}^{n}\left( \boldsymbol{\hat{\gamma}}_{i}-\boldsymbol{\gamma
}_{i}\right) \boldsymbol{\gamma}_{i}^{\prime}\right] \mathbf{f}
_{t}\nonumber\\
& +\left[ n^{-1/2}\sum_{i=1}^{n}\left( \boldsymbol{\hat{\delta}}
_{i,T}-\boldsymbol{\delta}_{i,T}\right) ^{\prime}\right] \mathbf{f}
_{t}+\left[ n^{-1/2}\sum_{i=1}^{n}\left( \boldsymbol{\hat{\delta}}
_{i,T}-\boldsymbol{\delta}_{i,T}\right) ^{^{\prime}}\right] \left(
\mathbf{\hat{f}}_{t}-\mathbf{f}_{t}\right) .\nonumber\\
&
\end{align}
We also note that using ((ref)), $\psi_{t,nT}\left( \lambda
_{T}\right) $ can be written as
\begin{equation}
\psi_{t,nT}\left( \lambda_{T}\right) =\xi_{t,n}\left( \lambda_{T}\right)
-\left( \boldsymbol{\varphi}_{nT}-\boldsymbol{\varphi}_{n}\right) ^{\prime
}\boldsymbol{\kappa}_{t,n}\left( \lambda_{T}\right) +\upsilon_{t,nT}\left(
\lambda_{T}\right)
\end{equation}
where
\begin{align}
\xi_{t,n}\left( \lambda_{T}\right) & =\frac{1}{\sqrt{n}}\sum_{i=1}
^{n}a_{i,n}\varepsilon_{it}\left( \lambda_{T}\right) ,
a_{i,n}=1-\sigma_{i}\boldsymbol{\varphi}_{n}^{\prime}\boldsymbol{\gamma}
_{i},\\
\boldsymbol{\kappa}_{t,n}\left( \lambda_{T}\right) & =\frac{1}{\sqrt{n}
}\sum_{i=1}^{n}\boldsymbol{\gamma}_{i}\sigma_{i}\varepsilon_{it}\left(
\lambda_{T}\right) ,\\
\upsilon_{t,nT}\left( \lambda_{T}\right) & =\frac{1}{\sqrt{n}}\sum
_{i=1}^{n}\left[ \frac{1}{\left( \boldsymbol{\varepsilon}_{i\circ}^{\prime
}\mathbf{M}_{F}\boldsymbol{\varepsilon}_{i\circ}/T\right) ^{1/2}}-1\right]
\varepsilon_{it}\left( \lambda_{T}\right) .
\end{align}
After squaring $s_{t,nT}\left( \lambda_{T}\right) $, we end up with
$p_{nT}\left( \lambda_{T}\right) =\sum_{j=1}^{10}A_{j,nT}\left( \lambda
_{T}\right) $, composed of four squared terms and six cross product terms.
For the first square term we have
\[
A_{1,nT}\left( \lambda_{T}\right) =\sqrt{T}\boldsymbol{\varphi}_{nT}
^{\prime}\left( \frac{1}{T}\sum_{t=1}^{T}\mathbf{b}_{t,n}\left( \lambda
_{T}\right) \mathbf{b}_{t,n}^{^{\prime}}\left( \lambda_{T}\right) \right)
\boldsymbol{\varphi}_{nT},
\]
where $\mathbf{b}_{t,n}\left( \lambda_{T}\right) =n^{-1/2}\sum_{i=1}
^{n}\left( \boldsymbol{\hat{\gamma}}_{i}-\boldsymbol{\gamma}_{i}\right)
\sigma_{i}\varepsilon_{it}\left( \lambda_{T}\right) $. Let $\mathbf{u}
_{\circ t}\left( \lambda_{T}\right) =\left( \sigma_{1}\varepsilon
_{1t}\left( \lambda_{T}\right) ,\sigma_{2}\varepsilon_{2t}\left(
\lambda_{T}\right) ,\ldots,\sigma_{n}\varepsilon_{nt}\left( \lambda
_{T}\right) \right) ^{\prime}$ so that $\mathbf{b}_{t,n}\left( \lambda
_{T}\right) =n^{-1/2}\left( \boldsymbol{\hat{\Gamma}}-\boldsymbol{\Gamma
}\right) ^{^{\prime}}\mathbf{u}_{\circ t}\left( \lambda_{T}\right) $. Then
\begin{equation}
\left\vert A_{1,nT}\left( \lambda_{T}\right) \right\vert \leq\frac{\sqrt{T}
}{n}\left\Vert \boldsymbol{\varphi}_{nT}\right\Vert ^{2}\left\Vert
\boldsymbol{\hat{\Gamma}}-\boldsymbol{\Gamma}\right\Vert ^{2}\left\Vert
\mathbf{V}_{T}\left( \lambda_{T}\right) \right\Vert ,
\end{equation}
where $\mathbf{V}_{T}\left( \lambda_{T}\right) =T^{-1}\sum_{t=1}
^{T}\mathbf{u}_{\circ t}\left( \lambda_{T}\right) \mathbf{u}_{\circ
t}^{^{\prime}}\left( \lambda_{T}\right) .$ Since $\boldsymbol{\varphi}
_{n}=n^{-1}\sum_{i=1}^{n}\boldsymbol{\gamma}_{i}/\sigma_{i}=O(1)$ by
Assumption (ref), and $\boldsymbol{\varphi}_{nT}
=\boldsymbol{\varphi}_{n}+o_{p}\left( 1\right) $ by result ((ref))
in Lemma (ref), then
\begin{equation}
\boldsymbol{\varphi}_{nT}=O_{p}\left( 1\right) .
\end{equation}
Further, using ((ref)) in the paper, $\mathbf{u}_{\circ t}\left(
\lambda_{T}\right) =\mathbf{D}_{0}\left( \boldsymbol{\varepsilon}_{\circ
t}+\lambda_{T}\mathbf{W}\boldsymbol{\varepsilon}_{\circ t}\right) $, where
$\mathbf{D}_{0}=diag\left( \sigma_{1},\sigma_{2},\ldots,\sigma_{n}\right) $,
$\boldsymbol{\varepsilon}_{\circ t}=\left( \varepsilon_{1t},\varepsilon
_{2t},\ldots,\varepsilon_{nT}\right) ^{\prime}$, and $\mathbf{W}=\left(
w_{ij}\right) $. Therefore
\[
\mathbf{u}_{\circ t}\left( \lambda_{T}\right) \mathbf{u}_{\circ t}^{\prime
}\left( \lambda_{T}\right) =\mathbf{D}_{0}\left( \boldsymbol{\varepsilon
}_{\circ t}\boldsymbol{\varepsilon}_{\circ t}^{\prime}+\lambda_{T}
^{2}\mathbf{W}\boldsymbol{\varepsilon}_{\circ t}\boldsymbol{\varepsilon
}_{\circ t}^{\prime}\mathbf{W}^{\prime}+\lambda_{T}\boldsymbol{\varepsilon
}_{\circ t}\boldsymbol{\varepsilon}_{\circ t}^{\prime}\mathbf{W}^{\prime
}+\lambda_{T}\mathbf{W}\boldsymbol{\varepsilon}_{\circ t}
\boldsymbol{\varepsilon}_{\circ t}^{\prime}\right) \mathbf{D}_{0},
\]
and
\[
\mathbf{V}_{T}\left( \lambda_{T}\right) =\mathbf{D}_{0}\left(
\mathbf{V}_{\varepsilon T}+\lambda_{T}^{2}\mathbf{W\mathbf{V}}_{\varepsilon
T}\mathbf{W}^{\prime}+\lambda_{T}\mathbf{V}_{\varepsilon T}\mathbf{W}^{\prime
}+\lambda_{T}\mathbf{W\mathbf{V}}_{\varepsilon T}\right) \mathbf{D}_{0},
\]
where $\mathbf{V}_{\varepsilon T}=T^{-1}\sum_{t=1}^{T}\boldsymbol{\varepsilon
}_{\circ t}\boldsymbol{\varepsilon}_{\circ t}^{\prime}$. It follows that
\[
\left\Vert \mathbf{V}_{T}\left( \lambda_{T}\right) \right\Vert
\leq\left\Vert \mathbf{\mathbf{V}}_{\varepsilon T}\right\Vert \left\Vert
\mathbf{D}_{0}\right\Vert ^{2}\left( \left\Vert \mathbf{I}_{n}\right\Vert
+\lambda_{T}^{2}\left\Vert \mathbf{W}\right\Vert \left\Vert \mathbf{W}
^{\prime}\right\Vert +\left\vert \lambda_{T}\right\vert \left\Vert
\mathbf{W}^{\prime}\right\Vert +\left\vert \lambda_{T}\right\vert \left\Vert
\mathbf{W}\right\Vert \right) .
\]
Note that $\left\Vert \mathbf{D}_{0}\right\Vert $ and $\left\Vert
\mathbf{W}\right\Vert $ are both bounded, and by part (b) of Assumption
(ref) we have $\left\Vert \mathbf{V}_{\varepsilon T}\right\Vert
=\mu_{max}\left( \mathbf{V}_{\varepsilon T}\right) =O_{p}\left( \frac{n}
{T}\right) $. Hence,
\begin{equation}
\left\Vert \mathbf{V}_{T}\left( \lambda_{T}\right) \right\Vert =O_{p}\left(
\frac{n}{T}\right) .
\end{equation}
Since $n$ and $T$ are of the same order of magnitude, then using results
((ref)), ((ref)), and ((ref)) in ((ref)) yields
\[
\left\vert A_{1,nT}\left( \lambda_{T}\right) \right\vert =\frac{\sqrt{T}}
{n}O_{p}\left( \frac{n}{\delta_{nT}^{2}}\right) =O_{p}\left( \frac
{1}{\delta_{nT}}\right) .
\]
For the second squared term we have
\[
A_{2,nT}\left( \lambda_{T}\right) =\sqrt{T}\left[ n^{-1/2}\sum_{i=1}
^{n}\left( \boldsymbol{\hat{\delta}}_{i,T}-\boldsymbol{\delta}_{i,T}\right)
^{\prime}\right] \left( T^{-1}\mathbf{F}^{\prime}\mathbf{F}\right) \left[
n^{-1/2}\sum_{i=1}^{n}\left( \boldsymbol{\hat{\delta}}_{i,T}
-\boldsymbol{\delta}_{i,T}\right) \right] .
\]
where $T^{-1}\mathbf{F}^{\prime}\mathbf{F}=O_{p}(1)$ and using ((ref))
$n^{-1/2}\sum_{i=1}^{n}\left( \boldsymbol{\hat{\delta}}_{i,T}
-\boldsymbol{\delta}_{i,T}\right) =O_{p}\left( \sqrt{\ln\left( n\right)
/T}\right) $. Hence, $A_{2,nT}\left( \lambda_{T}\right) =$ $O_{p}\left(
\ln\left( n\right) /\sqrt{T}\right) =o_{p}(1)$. Similarly,
\begin{align*}
A_{3,nT}\left( \lambda_{T}\right) & =\sqrt{T}\boldsymbol{\varphi}
_{nT}^{\prime}\left[ n^{-1/2}\sum_{i=1}^{n}\left( \boldsymbol{\hat{\gamma}
}_{i}-\boldsymbol{\gamma}_{i}\right) \boldsymbol{\gamma}_{i}^{\prime}\right]
\left( T^{-1}\sum_{t=1}^{T}\mathbf{f}_{t}\mathbf{f}_{t}^{\prime}\right)
\left[ n^{-1/2}\sum_{i=1}^{n}\boldsymbol{\gamma}_{i}\left( \boldsymbol{\hat
{\gamma}}-\boldsymbol{\gamma}_{i}\right) ^{\prime}\right]
\boldsymbol{\varphi}_{nT}\\
& =\sqrt{T}\boldsymbol{\varphi}_{nT}^{\prime}\left[ n^{-1/2}\sum_{i=1}
^{n}\left( \boldsymbol{\hat{\gamma}}_{i}-\boldsymbol{\gamma}_{i}\right)
\boldsymbol{\gamma}_{i}^{\prime}\right] \left( T^{-1}\mathbf{F}^{\prime
}\mathbf{F}\right) \left[ n^{-1/2}\sum_{i=1}^{n}\boldsymbol{\gamma}
_{i}\left( \boldsymbol{\hat{\gamma}}-\boldsymbol{\gamma}_{i}\right)
^{\prime}\right] \boldsymbol{\varphi}_{nT},
\end{align*}
where $\Vert\boldsymbol{\varphi}_{nT}\Vert$ is bounded by ((ref)),
and by ((ref)) $n^{-1/2}\sum_{i=1}^{n}\left( \boldsymbol{\hat{\gamma}
}_{i}-\boldsymbol{\gamma}_{i}\right) \boldsymbol{\gamma}_{i}^{\prime}
=O_{p}\left( \sqrt{\ln\left( n\right) /T}\right) $. Hence, $A_{3,nT}
\left( \lambda_{T}\right) =$ $O_{p}\left( \ln\left( n\right) /\sqrt
{T}\right) =o_{p}(1)$. For the final squared term,
\[
A_{4,nT}\left( \lambda_{T}\right) =\sqrt{T}\left[ \frac{1}{\sqrt{n}}
\sum_{i=1}^{n}\left( \boldsymbol{\hat{\delta}}_{i,T}-\boldsymbol{\delta
}_{i,T}\right) \right] ^{\prime}\left( \frac{\left\Vert \hat{\mathbf{F}
}-\mathbf{F}\right\Vert ^{2}}{T}\right) \left[ \frac{1}{\sqrt{n}}\sum
_{i=1}^{n}\left( \boldsymbol{\hat{\delta}}_{i,T}-\boldsymbol{\delta}
_{i,T}\right) \right] ,
\]
By results ((ref)) and ((ref)), it follows $A_{4,nT}\left(
\lambda_{T}\right) =\sqrt{T}O_{p}\left( \delta_{nT\ }^{-2}\right) \times
O_{p}\left( \ln\left( n\right) /T\right) =o_{p}(1)$. The probability
orders of the cross product terms of $p_{nT}\left( \lambda_{T}\right) $,
namely $A_{5,NT}\left( \lambda_{T}\right) ,\ldots,A_{10,NT}\left(
\lambda_{T}\right) $, are also easily seen to be $o_{p}(1)$, by application
of the Cauchy-Schwarz inequality to the product pairs of the terms
$A_{1,NT}\left( \lambda_{T}\right) ,A_{2,NT}\left( \lambda_{T}\right)
,A_{3,NT}\left( \lambda_{T}\right) ,$ and $A_{4,NT}\left( \lambda
_{T}\right) $. Thus, overall $p_{nT}\left( \lambda_{T}\right) =o_{p}(1)$,
as required. Consider now $q_{nT}\left( \lambda_{T}\right) $ and note that
it can be written as (using ((ref)) in ((ref)))
\begin{align*}
q_{nT}\left( \lambda_{T}\right) & =\frac{1}{\sqrt{T}}\sum_{t=1}
^{T}s_{t,nT}\left( \lambda_{T}\right) \xi_{t,n}\left( \lambda_{T}\right)
-\left( \boldsymbol{\varphi}_{nT}-\boldsymbol{\varphi}_{n}\right) ^{\prime
}\frac{1}{\sqrt{T}}\sum_{t=1}^{T}s_{t,nT}\left( \lambda_{T}\right)
\boldsymbol{\kappa}_{t,n}\left( \lambda_{T}\right) \\
& +\frac{1}{\sqrt{T}}\sum_{t=1}^{T}s_{t,nT}\left( \lambda_{T}\right)
\upsilon_{t,nT}\left( \lambda_{T}\right) ,
\end{align*}
where $\xi_{t,n}\left( \lambda_{T}\right) $, $\upsilon_{t,nT}\left(
\lambda_{T}\right) $ and $\boldsymbol{\kappa}_{t,n}\left( \lambda
_{T}\right) \mathbf{\ }$are given by ((ref)), ((ref)) and
((ref)), respectively. The first term of the above can be written as
\begin{align*}
\frac{1}{\sqrt{T}}\sum_{t=1}^{T}s_{t,nT}\left( \lambda_{T}\right) \xi
_{t,n}\left( \lambda_{T}\right) & =\boldsymbol{\varphi}_{nT}^{\prime
}\left[ \frac{1}{\sqrt{n}}\sum_{i=1}^{n}\left( \boldsymbol{\hat{\gamma}}
_{i}-\boldsymbol{\gamma}_{i}\right) \frac{1}{\sqrt{T}}\sum_{t=1}^{T}\xi
_{t,n}\left( \lambda_{T}\right) \sigma_{i}\varepsilon_{it}\left(
\lambda_{T}\right) \right] \\
& +\boldsymbol{\varphi}_{nT}^{\prime}\left[ n^{-1/2}\sum_{i=1}^{n}\left(
\boldsymbol{\hat{\gamma}}_{i}-\boldsymbol{\gamma}_{i}\right)
\boldsymbol{\gamma}_{i}^{\prime}\right] \left( \frac{1}{\sqrt{T}}\sum
_{t=1}^{T}\xi_{t,n}\left( \lambda_{T}\right) \mathbf{f}_{t}\right) \\
& +\left[ n^{-1/2}\sum_{i=1}^{n}\left( \boldsymbol{\hat{\delta}}
_{i,T}-\boldsymbol{\delta}_{i,T}\right) ^{\prime}\right] \left( \frac
{1}{\sqrt{T}}\sum_{t=1}^{T}\xi_{t,n}\left( \lambda_{T}\right) \mathbf{f}
_{t}\right) \\
& +\left[ n^{-1/2}\sum_{i=1}^{n}\left( \boldsymbol{\hat{\delta}}
_{i,T}-\boldsymbol{\delta}_{i,T}\right) ^{^{\prime}}\right] \left[ \frac
{1}{\sqrt{T}}\sum_{t=1}^{T}\xi_{t,n}\left( \lambda_{T}\right) \left(
\mathbf{\hat{f}}_{t}-\mathbf{f}_{t}\right) \right] \\
& =\sum_{j=1}^{4}B_{j,nT}\left( \lambda_{T}\right) .
\end{align*}
Using ((ref)), $B_{1,nT}\left( \lambda_{T}\right) $ can be written as
\[
B_{1,nT}\left( \lambda_{T}\right) =\boldsymbol{\varphi}_{nT}^{\prime}\left[
\frac{\sqrt{T}}{\sqrt{n}}\sum_{i=1}^{n}\left( \boldsymbol{\hat{\gamma}}
_{i}-\boldsymbol{\gamma}_{i}\right) \frac{1}{\sqrt{n}}\sum_{j=1}^{n}
a_{j,n}\left( \frac{1}{T}\sum_{t=1}^{T}\sigma_{i}\varepsilon_{jt}\left(
\lambda_{T}\right) \varepsilon_{it}\left( \lambda_{T}\right) \right) ,
\right]
\]
where $a_{i,n}=1-\sigma_{i}\boldsymbol{\varphi}_{n}^{\prime}\boldsymbol{\gamma
}_{i}$ and $\boldsymbol{\varphi}_{nT}=O_{p}(1)$. Also since $\varepsilon
_{it}\left( \lambda_{T}\right) $ are independently distributed over $t$ and
weakly cross-sectionally dependent, and $n$ and $T$ are of the same order,
then
\[
B_{1,nT}\left( \lambda_{T}\right) =O_{p}\left( \frac{1}{\sqrt{n}}\sum
_{i=1}^{n}a_{i,n}\sigma_{i}\left( \boldsymbol{\hat{\gamma}}_{i}
-\boldsymbol{\gamma}_{i}\right) \right) .
\]
Further, letting $b_{in}=a_{i,n}\sigma_{i}$, it follows from ((ref)) that
$n^{-1/2}\sum_{i=1}^{n}a_{i,n}\sigma_{i}\left( \boldsymbol{\hat{\gamma}}
_{i}-\boldsymbol{\gamma}_{i}\right) =O_{p}(\sqrt{\ln\left( n\right) /T})$,
which in turn establishes that $B_{1,nT}\left( \lambda_{T}\right) =o_{p}
(1)$. Similarly, using ((ref)), $B_{2,nT}\left( \lambda_{T}\right) $
can be written as
\[
B_{2,nT}\left( \lambda_{T}\right) =\boldsymbol{\varphi}_{nT}^{\prime}\left[
n^{-1/2}\sum_{i=1}^{n}\left( \boldsymbol{\hat{\gamma}}_{i}-\boldsymbol{\gamma
}_{i}\right) \boldsymbol{\gamma}_{i}^{\prime}\right] \left( \frac{1}
{\sqrt{nT}}\sum_{j=1}^{n}\sum_{t=1}^{T}a_{j,n}\mathbf{f}_{t}\varepsilon
_{jt}\left( \lambda_{T}\right) \right) ,
\]
where $\boldsymbol{\varphi}_{nT}=O_{p}(1)$. Under parts (a) and (c) of
Assumption (ref)
$
\frac{1}{\sqrt{nT}}\sum_{j=1}^{n}\sum_{t=1}^{T}a_{j,n}\mathbf{f}
_{t}\varepsilon_{jt}\left( \lambda_{T}\right) =O_{p}(1).
$
Using this result together with ((ref)) it follows that $B_{2,nT}\left(
\lambda_{T}\right) =o_{p}(1)$. Similarly, using ((ref)) we can establish
that $B_{3,nT}\left( \lambda_{T}\right) =o_{p}(1)$. The final term,
$B_{4,nT}\left( \lambda_{T}\right) $, is dominated by the third term and is
also $o_{p}(1)$. Thus overall, $T^{-1/2}\sum_{t=1}^{T}s_{t,nT}\left(
\lambda_{T}\right) \xi_{t,n}\left( \lambda_{T}\right) =o_{p}(1)$. Using the
same line of reasoning, it is also readily established that $T^{-1/2}
\sum_{t=1}^{T}s_{t,nT}\left( \lambda_{T}\right) \boldsymbol{\kappa}
_{t,n}\left( \lambda_{T}\right) =o_{p}(1)$, considering that,
$\boldsymbol{\kappa}_{t,n}\left( \lambda_{T}\right) $ $=n^{-1/2}\sum
_{i=1}^{n}\boldsymbol{\gamma}_{i}\sigma_{i}\varepsilon_{it}\left( \lambda
_{T}\right) $ has the same format as $\xi_{t,n}\left( \lambda_{T}\right) $,
and in addition by ((ref)) $\boldsymbol{\varphi}_{nT}
-\boldsymbol{\varphi}_{n}=O_{p}(n^{-1/2}T^{-1/2})+O_{p}(T^{-1})$. Finally, the
last term of $q_{nT}$ is given by
\begin{align*}
\frac{1}{\sqrt{T}}\sum_{t=1}^{T}s_{t,nT}\left( \lambda_{T}\right)
\upsilon_{t,nT}\left( \lambda_{T}\right) & =\frac{1}{\sqrt{T}}\sum
_{t=1}^{T}\upsilon_{t,nT}\left( \lambda_{T}\right) \boldsymbol{\varphi}
_{nT}^{\prime}\left[ n^{-1/2}\sum_{i=1}^{n}\left( \boldsymbol{\hat{\gamma}
}_{i}-\boldsymbol{\gamma}_{i}\right) \sigma_{i}\varepsilon_{it}\left(
\lambda_{T}\right) \right] \\
& +\frac{1}{\sqrt{T}}\sum_{t=1}^{T}\upsilon_{t,nT}\left( \lambda_{T}\right)
\boldsymbol{\varphi}_{nT}^{\prime}\left[ n^{-1/2}\sum_{i=1}^{n}\left(
\boldsymbol{\hat{\gamma}}_{i}-\boldsymbol{\gamma}_{i}\right)
\boldsymbol{\gamma}_{i}^{\prime}\right] \mathbf{f}_{t}\\
& +\frac{1}{\sqrt{T}}\sum_{t=1}^{T}\upsilon_{t,nT}\left( \lambda_{T}\right)
\left[ n^{-1/2}\sum_{i=1}^{n}\left( \boldsymbol{\hat{\delta}}_{i,T}
-\boldsymbol{\delta}_{i,T}\right) ^{\prime}\right] \mathbf{f}_{t}\\
& +\frac{1}{\sqrt{T}}\sum_{t=1}^{T}\upsilon_{t,nT}\left( \lambda_{T}\right)
\left[ n^{-1/2}\sum_{i=1}^{n}\left( \boldsymbol{\hat{\delta}}_{i,T}
-\boldsymbol{\delta}_{i,T}\right) ^{^{\prime}}\right] \left( \mathbf{\hat
{f}}_{t}-\mathbf{f}_{t}\right) \\
& =\sum_{j=1}^{4}C_{j,nT}\left( \lambda_{T}\right) .
\end{align*}
Using ((ref)) we have
\begin{align*}
C_{1,nT}\left( \lambda_{T}\right) & =\frac{1}{\sqrt{T}}\sum_{t=1}^{T}
\frac{1}{\sqrt{n}}\sum_{j=1}^{n}\left( \frac{1}{\left(
\boldsymbol{\varepsilon}_{j\circ}^{\prime}\mathbf{M}_{F}
\boldsymbol{\varepsilon}_{j\circ}/T\right) ^{1/2}}-1\right) \varepsilon
_{jt}\left( \lambda_{T}\right) \boldsymbol{\varphi}_{nT}^{\prime}\left[
n^{-1/2}\sum_{i=1}^{n}\left( \boldsymbol{\hat{\gamma}}_{i}-\boldsymbol{\gamma
}_{i}\right) \sigma_{i}\varepsilon_{it}\left( \lambda_{T}\right) \right] \\
& =\sqrt{\frac{T}{n}}\boldsymbol{\varphi}_{nT}^{\prime}\sum_{j=1}^{n}\left[
n^{-1/2}\sum_{i=1}^{n}\left( \boldsymbol{\hat{\gamma}}_{i}-\boldsymbol{\gamma
}_{i}\right) \frac{1}{T}\sum_{t=1}^{T}\sigma_{i}\left( \frac{1}{\left(
\boldsymbol{\varepsilon}_{j\circ}^{\prime}\mathbf{M}_{F}
\boldsymbol{\varepsilon}_{j\circ}/T\right) ^{1/2}}-1\right) \varepsilon
_{it}\left( \lambda_{T}\right) \varepsilon_{jt}\left( \lambda_{T}\right)
\right] .
\end{align*}
Since $\varepsilon_{it}\left( \lambda_{T}\right) $ is distributed
independently over $t$ and weakly cross-sectionally dependent, then
\[
\frac{1}{T}\sum_{t=1}^{T}\sigma_{i}\left( \frac{1}{\left(
\boldsymbol{\varepsilon}_{j\circ}^{\prime}\mathbf{M}_{F}
\boldsymbol{\varepsilon}_{j\circ}/T\right) ^{1/2}}-1\right) \varepsilon
_{it}\left( \lambda_{T}\right) \varepsilon_{jt}\left( \lambda_{T}\right)
\rightarrow_{p}0\text{, if }i\neq j\text{, }
\]
and
\begin{align*}
& \frac{1}{T}\sum_{t=1}^{T}\sigma_{i}\left( \frac{1}{\left(
\boldsymbol{\varepsilon}_{j\circ}^{\prime}\mathbf{M}_{F}
\boldsymbol{\varepsilon}_{j\circ}/T\right) ^{1/2}}-1\right) \varepsilon
_{it}\left( \lambda_{T}\right) \varepsilon_{jt}\left( \lambda_{T}\right) \\
& \rightarrow_{p}\lim_{T\rightarrow\infty}\frac{1}{T}\sum_{t=1}^{T}\sigma
_{i}E\left\{ \left[ \left( \frac{\boldsymbol{\varepsilon}_{i\circ}^{\prime
}\mathbf{M}_{F}\boldsymbol{\varepsilon}_{i\circ}}{T}\right) ^{-1/2}-1\right]
\varepsilon_{it}^{2}\left( \lambda_{T}\right) \right\} , if i=j.
\end{align*}
Also, by definition $\varepsilon_{it}\left( \lambda_{T}\right)
=\varepsilon_{it}+\lambda_{T}b_{it}$, where $b_{it}=\mathbf{w}_{i0}^{\prime
}\boldsymbol{\varepsilon}_{\circ t}$, $w_{ii}=0$, and $\varepsilon_{it}$ is
independent from $b_{it}$, and therefore
\begin{align*}
& E\left\{ \left[ \left( \frac{\boldsymbol{\varepsilon}_{i\circ}^{\prime
}\mathbf{M}_{F}\boldsymbol{\varepsilon}_{i\circ}}{T}\right) ^{-1/2}-1\right]
\varepsilon_{it}^{2}\left( \lambda_{T}\right) \right\} \\
& =E\left\{ \left[ \left( \frac{\boldsymbol{\varepsilon}_{i\circ}^{\prime
}\mathbf{M}_{F}\boldsymbol{\varepsilon}_{i\circ}}{T}\right) ^{-1/2}-1\right]
\varepsilon_{it}^{2}\right\} +E\left[ \left( \frac{\boldsymbol{\varepsilon
}_{i\circ}^{\prime}\mathbf{M}_{F}\boldsymbol{\varepsilon}_{i\circ}}{T}\right)
^{-1/2}-1\right] \lambda_{T}^{2}E\left( b_{it}^{2}\right) \\
& =O\left( \frac{1}{T}\right) +O\left( \frac{1}{T^{2}}\right) ,
\end{align*}
where the last line holds by ((ref)), ((ref)) and ((ref)).
Moreover, by ((ref)) $n^{-1/2}\sum_{i=1}^{n}\left( \boldsymbol{\hat
{\gamma}}_{i}-\boldsymbol{\gamma}_{i}\right) =O_{p}(\sqrt{\ln\left(
n\right) /T})$. As $n$ and $T$ being of the same order, it then follows that
$C_{1,nT}\left( \lambda_{T}\right) =o_{p}(1)$. Similarly to $B_{2,nT}\left(
\lambda_{T}\right) $, we have
\begin{align*}
C_{2,nT}\left( \lambda_{T}\right) & =\boldsymbol{\varphi}_{nT}^{\prime
}\left[ n^{-1/2}\sum_{i=1}^{n}\left( \boldsymbol{\hat{\gamma}}
_{i}-\boldsymbol{\gamma}_{i}\right) \boldsymbol{\gamma}_{i}^{\prime}\right]
\left[ \frac{1}{\sqrt{nT}}\sum_{j=1}^{n}\sum_{t=1}^{T}\left( \frac
{1}{\left( \boldsymbol{\varepsilon}_{j\circ}^{\prime}\mathbf{M}
_{F}\boldsymbol{\varepsilon}_{j\circ}/T\right) ^{1/2}}-1\right)
\varepsilon_{jt}\left( \lambda_{T}\right) \mathbf{f}_{t}\right] \\
& =O_{p}\left( \sqrt{\frac{\ln\left( n\right) }{T}}\right) O_{p}
(1)=o_{p}(1).
\end{align*}
The same line of reasoning as used for $B_{3,nT}\left( \lambda_{T}\right) $
and $B_{4,nT}\left( \lambda_{T}\right) $ can be used to establish
$C_{j,nT}\left( \lambda_{T}\right) =o_{p}(1)$ for $j=3$ and $4.$ Hence,
$T^{-1/2}\sum_{t=1}^{T}s_{t,nT}\left( \lambda_{T}\right) \upsilon
_{t,nT}\left( \lambda_{T}\right) =o_{p}(1)$, and overall we have
$q_{nT}\left( \lambda_{T}\right) =o_{p}(1)$, as required.
lemmaConsider the latent factor model given by ((ref)) and
((ref)). Suppose that Assumptions (ref)-(ref) hold
and $\left( n,T\right) \rightarrow\infty$, such that $n/T\rightarrow
\kappa\,,$ for $0<\kappa<\infty$. Then
\begin{equation}
\sqrt{T}\left( \boldsymbol{\varphi}_{n}-\boldsymbol{\varphi}_{nT}\right)
=O_{p}\left( n^{-1/2}\right) +O_{p}\left( T^{-1/2}\right) ,
\end{equation}
\begin{equation}
\sqrt{T}\left( \boldsymbol{\hat{\varphi}}_{nT}-\boldsymbol{\varphi}
_{n}\right) =o_{p}(1),
\end{equation}
where $\boldsymbol{\varphi}_{n}=n^{-1}\sum_{i=1}^{n}\boldsymbol{\gamma}
_{i}/\sigma_{i},$ $\boldsymbol{\varphi}_{nT}=n^{-1}\sum_{i=1}^{n}
\boldsymbol{\gamma}_{i}/\omega_{i,T}$ with $\omega_{i,T}=\left( T^{-1}
\sigma_{i}^{2}\boldsymbol{\varepsilon}_{i\circ}^{\prime}\mathbf{M}
_{F}\boldsymbol{\varepsilon}_{i\circ}\right) ^{1/2}$, $\boldsymbol{\hat
{\varphi}}_{nT}=n^{-1}\sum_{i=1}^{n}\boldsymbol{\hat{\gamma}}_{i}/\hat{\sigma
}_{i,T}$, $\hat{\sigma}_{i,T}=\left( T^{-1}\mathbf{y}_{i}^{\prime}
\mathbf{M}_{\hat{F}}\mathbf{y}_{i}\right) ^{1/2}$ and $\boldsymbol{\hat
{\gamma}}_{i}$ and $\mathbf{\hat{F}}$ are the principal component estimators
of $\boldsymbol{\gamma}_{i}$ and $\mathbf{F}$.
proofFirst note that
\begin{align*}
\sqrt{T}\left( \boldsymbol{\varphi}_{n}-\boldsymbol{\varphi}_{nT}\right) &
=\frac{\sqrt{T}}{n}\sum_{i=1}^{n}\frac{\boldsymbol{\gamma}_{i}}{\sigma_{i}
}\left\{ \left( 1-\frac{\sigma_{i}}{\omega_{i,T}}\right) -\left[
1-E\left( \frac{\sigma_{i}}{\omega_{i,T}}\right) \right] \right\}
+\frac{\sqrt{T}}{n}\sum_{i=1}^{n}\frac{\boldsymbol{\gamma}_{i}}{\sigma_{i}
}\left[ 1-E\left( \frac{\sigma_{i}}{\omega_{i,T}}\right) \right] ,\\
& =\mathbf{g}_{1,nT}+\mathbf{g}_{2,nT},
\end{align*}
where
\begin{align*}
\mathbf{g}_{1,nT} & =-\frac{1}{n}\sum_{i=1}^{n}\sqrt{T}\left[ \frac
{\sigma_{i}}{\omega_{i,T}}-E\left( \frac{\sigma_{i}}{\omega_{i,T}}\right)
\right] \frac{\boldsymbol{\gamma}_{i}}{\sigma_{i}},\\
\mathbf{g}_{2,nT} & =\frac{\sqrt{T}}{n}\sum_{i=1}^{n}\left[ 1-E\left(
\frac{\sigma_{i}}{\omega_{i,T}}\right) \right] \frac{\boldsymbol{\gamma}
_{i}}{\sigma_{i}}.
\end{align*}
Since $\sigma_{i}/\omega_{i,T}=\left( T^{-1}\boldsymbol{\varepsilon}_{i\circ
}^{\prime}\mathbf{M}_{F}\boldsymbol{\varepsilon}_{i\circ}\right) ^{-1/2}$,
$\left\Vert \boldsymbol{\gamma}_{i}\right\Vert <C$, $\sigma_{i}<C$, then using
result ((ref)) we have $E\left( \frac{\sigma_{i}}{\omega_{i,T}}\right)
=1+O\left( T^{-1}\right) $, and $\mathbf{g}_{2,nT}=$ $O\left(
T^{-1/2}\right) $. The first term can be written as $\mathbf{g}_{1,nT}
=n^{-1}\sum_{i=1}^{n}\boldsymbol{\gamma}_{i}\sigma_{i}^{-1}\chi_{i,T}$, where
$\chi_{i,T}=-\sqrt{T}\left[ \sigma_{i}/\omega_{i,T}-E\left( \sigma
_{i}/\omega_{i,T}\right) \right] $. Conditional on $\mathbf{F}$ and
$\sigma_{i}$, $\chi_{i,T}$ are distributed independently over $i$ with mean
zero and bounded variances:\footnote{When $\varepsilon_{it}$ are normally
distributed we have the exact result $E\left( \frac{T}
{\boldsymbol{\varepsilon}_{i\circ}^{\prime}\mathbf{M}_{F}
\boldsymbol{\varepsilon}_{i\circ}}\right) =T/(T-m_{0}-2).$}
\[
Var\left( \chi_{i,T}\right) =T\left[ E\left( \frac{T}
{\boldsymbol{\varepsilon}_{i\circ}^{\prime}\mathbf{M}_{F}
\boldsymbol{\varepsilon}_{i\circ}}\right) -\left[ E\left( \frac{\sigma_{i}
}{\omega_{i,T}}\right) \right] ^{2}\right] =T\left[ 1+O\left( \frac{1}
{T}\right) -\left[ 1+O\left( \frac{1}{T}\right) \right] ^{2}\right]
=O(1).
\]
Hence, $\mathbf{g}_{1,nT}=O_{p}\left( n^{-1/2}\right) $, and the desired
result ((ref)) follows. Consider now ((ref)) and note
that it can be decomposed as
\begin{equation}
\sqrt{T}\left( \boldsymbol{\hat{\varphi}}_{nT}-\boldsymbol{\varphi}
_{n}\right) =\sqrt{T}\left( \boldsymbol{\varphi}_{nT}-\boldsymbol{\varphi
}_{n}\right) +\sqrt{T}\left( \boldsymbol{\hat{\varphi}}_{nT}
-\boldsymbol{\varphi}_{nT}\right) ,
\end{equation}
where it is already established that the first term is $o_{p}(1)$. Consider
now the second term of ((ref)) and note that it can be written as
\begin{align*}
\sqrt{T}\left( \boldsymbol{\hat{\varphi}}_{nT}-\boldsymbol{\varphi}
_{nT}\right) & =\frac{\sqrt{T}}{n}\sum_{i=1}^{n}\left( \frac
{\boldsymbol{\hat{\gamma}}_{i}}{\hat{\sigma}_{i,T}}\boldsymbol{-}
\frac{\boldsymbol{\gamma}_{i}}{\omega_{i,T}}\right) \\
& =\frac{\sqrt{T}}{n}\sum_{i=1}^{n}\boldsymbol{\gamma}_{i}\left( \frac
{1}{\hat{\sigma}_{i,T}}\boldsymbol{-}\frac{1}{\omega_{i,T}}\right)
+\frac{\sqrt{T}}{n}\sum_{i=1}^{n}\frac{1}{\sigma_{i}}\left( \boldsymbol{\hat
{\gamma}}_{i}-\boldsymbol{\gamma}_{i}\right) +\\
& \frac{\sqrt{T}}{n}\sum_{i=1}^{n}\left( \frac{1}{\omega_{i,T}}-\frac
{1}{\sigma_{i}}\right) \left( \boldsymbol{\hat{\gamma}}_{i}
-\boldsymbol{\gamma}_{i}\right) +\frac{\sqrt{T}}{n}\sum_{i=1}^{n}\left(
\frac{1}{\hat{\sigma}_{i,T}}-\frac{1}{\omega_{i,T}}\right) \left(
\boldsymbol{\hat{\gamma}}_{i}-\boldsymbol{\gamma}_{i}\right) .
\end{align*}
Now using ((ref)) of Lemma (ref)\ we have
\[
\frac{\sqrt{T}}{n}\sum_{i=1}^{n}\boldsymbol{\gamma}_{i}\left( \frac{1}
{\hat{\sigma}_{i,T}}-\frac{1}{\omega_{i,T}}\right) =O_{p}\left( \frac
{\ln\left( n\right) }{\sqrt{T}}\right) .
\]
Using this result as well as ((ref)), ((ref)) and ((ref)), we
have $\sqrt{T}\left( \boldsymbol{\hat{\varphi}}_{nT}-\boldsymbol{\varphi
}_{nT}\right) =o_{p}(1)$ as required.
lemmaSuppose $\mathbf{M}_{F}=\mathbf{I}_{T}-\mathbf{F(F}^{\prime
}\mathbf{F)}^{-1}\mathbf{F}^{\prime}$, where $\mathbf{F}$ is a $T\times m_{0}$
matrix, and $\boldsymbol{\tau}_{T}$ is a $T\times1$ vector of ones. Then
\begin{align}
\mathrm{tr}\left( \mathbf{M}_{F}\right) & =v, \mathrm{tr}\left(
\mathbf{M}_{F}\mathbf{\odot M}_{F}\right) =O\left( v\right) ,\\
\mathrm{tr}\left( \mathbf{M}_{F}\mathbf{\odot M}_{F}\mathbf{\odot M}
_{F}\right) & =O\left( v\right) ,\mathrm{tr}\left( \mathbf{M}
_{F}\mathbf{\odot M}_{F}\mathbf{\odot M}_{F}\mathbf{\odot M}_{F}\right)
=O\left( v\right) ,\\
\mathrm{tr}\left[ \left( \mathbf{I}_{T}\mathbf{\odot M}_{F}\right)
\mathbf{M}_{F}\right] & =O\left( v\right) ,\mathrm{tr}\left[ \left(
\mathbf{M}_{F}\mathbf{\odot M}_{F}\right) \mathbf{M}_{F}\right] =O\left(
v\right) ,
\end{align}
\begin{align}
\boldsymbol{\tau}_{T}^{\prime}\left( \mathbf{M}_{F}\mathbf{\odot M}
_{F}\right) \boldsymbol{\tau}_{T} & =O\left( v\right) ,\boldsymbol{\tau
}_{T}^{\prime}\left( \mathbf{M}_{F}\mathbf{\odot M}_{F}\mathbf{\odot M}
_{F}\right) \boldsymbol{\tau}_{T}=O\left( v\right) ,\\
\boldsymbol{\tau}_{T}^{\prime}\left( \mathbf{M}_{F}\mathbf{\odot M}
_{F}\mathbf{\odot M}_{F}\mathbf{\odot M}_{F}\right) \boldsymbol{\tau}_{T} &
=O\left( v\right) ,\boldsymbol{\tau}_{T}^{\prime}\left( \mathbf{I}
_{T}\mathbf{\odot M}_{F}\right) \left( \mathbf{M}_{F}\mathbf{\odot M}
_{F}\mathbf{\odot M}_{F}\right) \boldsymbol{\tau}_{T}=O\left( v\right)
,\\
\boldsymbol{\tau}_{T}^{\prime}\left( \mathbf{I}_{T}\mathbf{\odot M}
_{F}\right) \left( \mathbf{M}_{F}\mathbf{\odot M}_{F}\right) \left(
\mathbf{I}_{T}\mathbf{\odot M}_{F}\right) \boldsymbol{\tau}_{T} & =O\left(
v\right) ,\boldsymbol{\tau}_{T}^{\prime}\left( \mathbf{I}_{T}\mathbf{\odot
M}_{F}\right) \left( \mathbf{I}_{T}\mathbf{\odot M}_{F}\right)
\boldsymbol{\tau}_{T}=O\left( v\right) ,
\end{align}
\begin{align}
\boldsymbol{\tau}_{T}^{\prime}\left( \mathbf{I}_{T}\mathbf{\odot M}
_{F}\right) \mathbf{M}_{F}\left( \mathbf{I}_{T}\mathbf{\odot M}_{F}\right)
\boldsymbol{\tau}_{T} & =O\left( v^{3/2}\right) ,\boldsymbol{\tau}
_{T}^{\prime}\left( \mathbf{M}_{F}\mathbf{\odot M}_{F}\right) \mathbf{M}
_{F}\left( \mathbf{I}_{T}\mathbf{\odot M}_{F}\right) \boldsymbol{\tau}
_{T}=O\left( v^{3/2}\right) ,\\
\boldsymbol{\tau}_{T}^{\prime}\left( \mathbf{I}_{T}\mathbf{\odot M}
_{F}\right) \mathbf{M}_{F}\left( \mathbf{I}_{T}\mathbf{\odot M}
_{F}\mathbf{\odot M}_{F}\right) \boldsymbol{\tau}_{T} & =O\left(
v^{3/2}\right) ,
\end{align}
where $v=T-m_{0}$.
proofSee Lemma 10 of Pesaran and Yamagata (2024).
lemmaSuppose the $T\times1$ vector $\boldsymbol{\varepsilon}
=\mathbf{(}\varepsilon_{1},\varepsilon_{2},...,\varepsilon_{T})^{\prime}$ is
$\boldsymbol{\varepsilon}\thicksim IID(\mathbf{0},\mathbf{I}_{T})$, $\sup
_{t}E\left( \left\vert \varepsilon_{t}\right\vert ^{8+s}\right) <C$ for some
small $s>0$, and $\mathbf{M}_{F}=\mathbf{I}_{T}-\mathbf{F(F}^{\prime
}\mathbf{F)}^{-1}\mathbf{F}^{\prime}$, where the $T\times m_{0}$ matrix
$\mathbf{F}$ is distributed independently of $\boldsymbol{\varepsilon}$,
then
\begin{equation}
E\left( \frac{\boldsymbol{\varepsilon}^{\prime}\mathbf{M}_{F}
\boldsymbol{\varepsilon}}{v}\right) =1,
\end{equation}
\begin{equation}
E\left[ \left( \frac{\boldsymbol{\varepsilon}^{\prime}\mathbf{M}
_{F}\boldsymbol{\varepsilon}}{v}\right) ^{2}\right] =1+O\left( \frac{1}
{v}\right) ,
\end{equation}
\begin{equation}
E\left( q_{v}\right) =0\mathbf{,} E\left( q_{v}^{4}\right)
=O(1)\mathbf{,}
\end{equation}
where $v=T-m_{0}$, and $q_{v}=\sqrt{v}\left( \frac{\boldsymbol{\varepsilon
}^{\prime}\mathbf{M}_{F}\boldsymbol{\varepsilon}}{v}-1\right) $.
proofDenote $\boldsymbol{\tau}_{T}$ as a $T\times1$ vector of ones and suppose
$\gamma_{1}=E\left( \varepsilon_{t}^{3}\right) $, $\gamma_{2}=E\left(
\varepsilon_{t}^{4}\right) -3$, $\gamma_{3}=E\left( \varepsilon_{t}
^{5}\right) -10\gamma_{1}$, $\gamma_{4}=E\left( \varepsilon_{t}^{6}\right)
-15\gamma_{2}-10\gamma_{1}^{2}-15$, $\gamma_{6}=E\left( \varepsilon_{t}
^{8}\right) -28\gamma_{4}-56\gamma_{3}\gamma_{1}-35\gamma_{2}^{2}
-210\gamma_{2}-280\gamma_{1}^{2}-105$, which are all bounded as it is assumed
$\sup_{t}E\left( \left\vert \varepsilon_{t}\right\vert ^{8+\epsilon}\right)
<C$. Since $\boldsymbol{\varepsilon}\thicksim IID(\mathbf{0},\mathbf{I}_{T})$
and $\mathbf{M}_{F}=\left( m_{tt^{\prime}}\right) $ is an idempotent matrix
then results (S.6) to (S.9) of Lemma 6 in Pesaran and Yamagata (2024) apply
and we have (since $\mathrm{tr}\left( \mathbf{M}_{F}^{s}\right)
=\mathrm{tr}\left( \mathbf{M}_{F}\right) =v$, for $s=1,2,\ldots$)
\begin{equation}
E\left( \boldsymbol{\varepsilon}^{\prime}\mathbf{M}_{F}
\boldsymbol{\varepsilon}\right) =v,
\end{equation}
\begin{equation}
E\left[ \left( \boldsymbol{\varepsilon}^{\prime}\mathbf{M}_{F}
\boldsymbol{\varepsilon}\right) ^{2}\right] =v^{2}+2v+\gamma_{2}
\mathrm{tr}\left( \mathbf{M}_{F}\mathbf{\odot M}_{F}\right) ,
\end{equation}
\begin{align}
E\left[ \left( \boldsymbol{\varepsilon}^{\prime}\mathbf{M}_{F}
\boldsymbol{\varepsilon}\right) ^{3}\right] & =v^{3}+6v^{2}+8v+\gamma
_{4}\mathrm{tr}\left( \mathbf{M}_{F}\mathbf{\odot M}_{F}\mathbf{\odot M}
_{F}\right) +3\gamma_{2}\left( v+4\right) \mathrm{tr}\left( \mathbf{M}
_{F}\mathbf{\odot M}_{F}\right) \nonumber\\
& +6\gamma_{1}^{2}\left[ \boldsymbol{\tau}_{T}^{\prime}\left(
\mathbf{I}_{T}\mathbf{\odot M}_{F}\right) \mathbf{M}_{F}\left(
\mathbf{I}_{T}\mathbf{\odot M}_{F}\right) \boldsymbol{\tau}_{T}\right]
+4\gamma_{1}^{2}\left[ \boldsymbol{\tau}_{T}^{\prime}\left( \mathbf{M}
_{F}\mathbf{\odot M}_{F}\mathbf{\odot M}_{F}\right) \boldsymbol{\tau}
_{T}\right]
\end{align}
\begin{equation}
E\left[ \left( \boldsymbol{\varepsilon}^{\prime}\mathbf{M}_{F}
\boldsymbol{\varepsilon}\right) ^{4}\right] =v^{4}+12v^{3}+44v^{2}
+48v+\gamma_{2}f_{\gamma_{2}}+\gamma_{4}f_{\gamma_{4}}+\gamma_{6}f_{\gamma
_{6}}+\gamma_{1}^{2}f_{\gamma_{1}^{2}}+\gamma_{2}^{2}f_{\gamma_{2}^{2}}
+\gamma_{1}\gamma_{3}f_{\gamma_{1}\gamma_{3}},
\end{equation}
where
\begin{align*}
f_{\gamma_{2}} & =\left( 6v^{2}+48v\right) \mathrm{tr}\left(
\mathbf{M}_{F}\mathbf{\odot M}_{F}\right) +12\left[ \boldsymbol{\tau}
_{T}^{\prime}\left( \mathbf{M}_{F}\mathbf{\odot M}_{F}\right)
\boldsymbol{\tau}_{T}\mathrm{tr}\left( \mathbf{M}_{F}\mathbf{\odot M}
_{F}\right) \right] \\
& +96\mathrm{tr}\left[ \left( \mathbf{I}_{T}\mathbf{\odot M}_{F}\right)
\mathbf{M}_{F}\right] +48\boldsymbol{\tau}_{T}^{\prime}\left( \mathbf{I}
_{T}\mathbf{\odot M}_{F}\right) \left( \mathbf{I}_{T}\mathbf{\odot M}
_{F}\right) \boldsymbol{\tau}_{T},
\end{align*}
\[
f_{\gamma_{4}}=\left( 4v+24\right) \mathrm{tr}\left( \mathbf{M}
_{F}\mathbf{\odot M}_{F}\mathbf{\odot M}_{F}\right) ,
\]
\[
f_{\gamma_{6}}=\mathrm{tr}\left( \mathbf{M}_{F}\mathbf{\odot M}
_{F}\mathbf{\odot M}_{F}\mathbf{\odot M}_{F}\right) ,
\]
\begin{align*}
f_{\gamma_{1}^{2}} & =24v\boldsymbol{\tau}_{T}^{\prime}\left(
\mathbf{I}_{T}\mathbf{\odot M}_{F}\right) \mathbf{M}_{F}\left(
\mathbf{I}_{T}\mathbf{\odot M}_{F}\right) \boldsymbol{\tau}_{T}
+48\boldsymbol{\tau}_{T}^{\prime}\left( \mathbf{I}_{T}\mathbf{\odot M}
_{F}\right) \mathbf{M}_{F}\left( \mathbf{I}_{T}\mathbf{\odot M}_{F}\right)
\boldsymbol{\tau}_{T}\\
& +16v\boldsymbol{\tau}_{T}^{\prime}\left( \mathbf{M}_{F}\mathbf{\odot
M}_{F}\mathbf{\odot M}_{F}\right) \boldsymbol{\tau}_{T}+96\boldsymbol{\tau
}_{T}^{\prime}\left( \mathbf{M}_{F}\mathbf{\odot M}_{F}\right)
\mathbf{M}_{F}\left( \mathbf{I}_{T}\mathbf{\odot M}_{F}\right)
\boldsymbol{\tau}_{T}\\
& +96\mathrm{tr}\left[ \mathbf{M}_{F}\left( \mathbf{M}_{F}\mathbf{\odot
M}_{F}\right) \mathbf{M}_{F}\right] ,
\end{align*}
\begin{align*}
f_{\gamma_{2}^{2}} & =3\left[ \mathrm{tr}\left( \mathbf{M}_{F}
\mathbf{\odot M}_{F}\right) \right] ^{2}+24\boldsymbol{\tau}_{T}^{\prime
}\left( \mathbf{I}_{T}\mathbf{\odot M}_{F}\right) \left( \mathbf{M}
_{F}\mathbf{\odot M}_{F}\right) \left( \mathbf{I}_{T}\mathbf{\odot M}
_{F}\right) \boldsymbol{\tau}_{T}\\
& +8\boldsymbol{\tau}_{T}^{\prime}\left( \mathbf{M}_{F}\mathbf{\odot M}
_{F}\mathbf{\odot M}_{F}\mathbf{\odot M}_{F}\right) \boldsymbol{\tau}_{T},
\end{align*}
\[
f_{\gamma_{1}\gamma_{3}}=24\boldsymbol{\tau}_{T}^{\prime}\left(
\mathbf{I}_{T}\mathbf{\odot M}_{F}\right) \mathbf{M}_{F}\left(
\mathbf{I}_{T}\mathbf{\odot M}_{F}\mathbf{\odot M}_{F}\right)
\boldsymbol{\tau}_{T}+32\boldsymbol{\tau}_{T}^{\prime}\left( \mathbf{I}
_{T}\mathbf{\odot M}_{F}\right) \left( \mathbf{M}_{F}\mathbf{\odot M}
_{F}\mathbf{\odot M}_{F}\right) \boldsymbol{\tau}_{T}.
\]
Result ((ref)) follows from ((ref)). To establish ((ref)),
using ((ref)), we first note that
\[
E\left( \frac{\boldsymbol{\varepsilon}^{\prime}\mathbf{M}_{F}
\boldsymbol{\varepsilon}}{v}\right) ^{2}=1+\frac{2}{v}+\gamma_{2}
\frac{\mathrm{tr}\left( \mathbf{M}_{F}\mathbf{\odot M}_{F}\right) }{v^{2}}.
\]
But by ((ref)) $\mathrm{tr}\left( \mathbf{M}_{F}\mathbf{\odot M}
_{F}\right) =O\left( v\right) $ and by assumption $\gamma_{2}$ is bounded.
Hence $E\left( \frac{\boldsymbol{\varepsilon}^{\prime}\mathbf{M}
_{F}\boldsymbol{\varepsilon}}{v}\right) ^{2}=1+O(\frac{1}{v})$, as required.
To prove ((ref)), noting that $E\left( \boldsymbol{\varepsilon}^{\prime
}\mathbf{M}_{F}\boldsymbol{\varepsilon}\right) =v$ then $E(q_{v})=0$. Also
\[
q_{v}^{4}=v^{2}\left( \frac{\boldsymbol{\varepsilon}^{\prime}\mathbf{M}
_{F}\boldsymbol{\varepsilon}}{v}-1\right) ^{4}=v^{2}\left[ \left(
\frac{\boldsymbol{\varepsilon}^{\prime}\mathbf{M}_{F}\boldsymbol{\varepsilon}
}{v}\right) ^{4}-4\left( \frac{\boldsymbol{\varepsilon}^{\prime}
\mathbf{M}_{F}\boldsymbol{\varepsilon}}{v}\right) ^{3}+6\left(
\frac{\boldsymbol{\varepsilon}^{\prime}\mathbf{M}_{F}\boldsymbol{\varepsilon}
}{v}\right) ^{2}-4\left( \frac{\boldsymbol{\varepsilon}^{\prime}
\mathbf{M}_{F}\boldsymbol{\varepsilon}}{v}\right) +1\right] ,
\]
and taking expectation yields
\begin{align*}
E\left( q_{v}^{4}\right) & =v^{2}\left[ E\left[ \left( \frac
{\boldsymbol{\varepsilon}^{\prime}\mathbf{M}_{F}\boldsymbol{\varepsilon}}
{v}\right) ^{4}\right] -4E\left[ \left( \frac{\boldsymbol{\varepsilon
}^{\prime}\mathbf{M}_{F}\boldsymbol{\varepsilon}}{v}\right) ^{3}\right]
+6E\left[ \left( \frac{\boldsymbol{\varepsilon}^{\prime}\mathbf{M}
_{F}\boldsymbol{\varepsilon}}{v}\right) ^{2}\right] -4E\left(
\frac{\boldsymbol{\varepsilon}^{\prime}\mathbf{M}_{F}\boldsymbol{\varepsilon}
}{v}\right) +1\right] \\
& =\frac{1}{v^{2}}E\left[ \left( \boldsymbol{\varepsilon}^{\prime
}\mathbf{M}_{F}\boldsymbol{\varepsilon}\right) ^{4}\right] -\frac{4}
{v}E\left[ \left( \boldsymbol{\varepsilon}^{\prime}\mathbf{M}_{F}
\boldsymbol{\varepsilon}\right) ^{3}\right] +6E\left[ \left(
\boldsymbol{\varepsilon}^{\prime}\mathbf{M}_{F}\boldsymbol{\varepsilon
}\right) ^{2}\right] -4vE\left( \boldsymbol{\varepsilon}^{\prime}
\mathbf{M}_{F}\boldsymbol{\varepsilon}\right) +v^{2}.
\end{align*}
Now using the results in ((ref))-((ref)), and after some
algebra, we obtain
\begin{align}
E\left( q_{v}^{4}\right) & =\frac{1}{v^{2}}\left[
\begin{array}
[c]{c}
12v^{2}+48v+12\gamma_{2}\left[ \boldsymbol{\tau}_{T}^{\prime}\left(
\mathbf{M}_{F}\mathbf{\odot M}_{F}\right) \boldsymbol{\tau}_{T}\right]
\mathrm{tr}\left( \mathbf{M}_{F}\mathbf{\odot M}_{F}\right) \\
+96\gamma_{2}\mathrm{tr}\left[ \left( \mathbf{I}_{T}\mathbf{\odot M}
_{F}\right) \mathbf{M}_{F}\right] +48\gamma_{2}\left[ \boldsymbol{\tau}
_{T}^{\prime}\left( \mathbf{I}_{T}\mathbf{\odot M}_{F}\right) \left(
\mathbf{I}_{T}\mathbf{\odot M}_{F}\right) \boldsymbol{\tau}_{T}\right] \\
+\left( 4\gamma_{4}v+24\gamma_{4}\right) \mathrm{tr}\left( \mathbf{M}
_{F}\mathbf{\odot M}_{F}\mathbf{\odot M}_{F}\right) +\gamma_{6}
\mathrm{tr}\left( \mathbf{M}_{F}\mathbf{\odot M}_{F}\mathbf{\odot M}
_{F}\mathbf{\odot M}_{F}\right) \\
+48\gamma_{1}^{2}\left[ \boldsymbol{\tau}_{T}^{\prime}\left( \mathbf{I}
_{T}\mathbf{\odot M}_{F}\right) \mathbf{M}_{F}\left( \mathbf{I}
_{T}\mathbf{\odot M}_{F}\right) \boldsymbol{\tau}_{T}\right] +96\gamma
_{1}^{2}\left[ \boldsymbol{\tau}_{T}^{\prime}\left( \mathbf{M}
_{F}\mathbf{\odot M}_{F}\right) \mathbf{M}_{F}\left( \mathbf{I}
_{T}\mathbf{\odot M}_{F}\right) \boldsymbol{\tau}_{T}\right] \\
+96\gamma_{1}^{2}\mathrm{tr}\left[ \left( \mathbf{M}_{F}\mathbf{\odot M}
_{F}\right) \mathbf{M}_{F}\right] +3\gamma_{2}^{2}\left[ \mathrm{tr}\left(
\mathbf{M}_{F}\mathbf{\odot M}_{F}\right) \right] ^{2}\\
+24\gamma_{2}^{2}\left[ \boldsymbol{\tau}_{T}^{\prime}\left( \mathbf{I}
_{T}\mathbf{\odot M}_{F}\right) \left( \mathbf{M}_{F}\mathbf{\odot M}
_{F}\right) \left( \mathbf{I}_{T}\mathbf{\odot M}_{F}\right)
\boldsymbol{\tau}_{T}\right] +8\gamma_{2}^{2}\left[ \boldsymbol{\tau}
_{T}^{\prime}\left( \mathbf{M}_{F}\mathbf{\odot M}_{F}\mathbf{\odot M}
_{F}\mathbf{\odot M}_{F}\right) \boldsymbol{\tau}_{T}\right] \\
+24\gamma_{1}\gamma_{3}\left[ \boldsymbol{\tau}_{T}^{\prime}\left(
\mathbf{I}_{T}\mathbf{\odot M}_{F}\right) \mathbf{M}_{F}\left(
\mathbf{I}_{T}\mathbf{\odot M}_{F}\mathbf{\odot M}_{F}\right)
\boldsymbol{\tau}_{T}\right] \\
+32\gamma_{1}\gamma_{3}\left[ \boldsymbol{\tau}_{T}^{\prime}\left(
\mathbf{I}_{T}\mathbf{\odot M}_{F}\right) \left( \mathbf{M}_{F}\mathbf{\odot
M}_{F}\mathbf{\odot M}_{F}\right) \boldsymbol{\tau}_{T}\right]
\end{array}
\right] \nonumber\\
& =\sum_{s=1}^{15}a_{s,v}.
\end{align}
Further noting that $\gamma_{1},\gamma_{2}$,$\gamma_{3}$,$\gamma_{4}$
,$\gamma_{5}$, and $\gamma_{6}$ are all bounded, then using the results
((ref))-((ref)) we have
\[
a_{1,v}=12,a_{2,v}=\frac{48}{v},
\]
\[
a_{3,v}=\frac{12\gamma_{2}\left[ \boldsymbol{\tau}_{T}^{\prime}\left(
\mathbf{M}_{F}\mathbf{\odot M}_{F}\right) \boldsymbol{\tau}_{T}\right]
\mathrm{tr}\left( \mathbf{M}_{F}\mathbf{\odot M}_{F}\right) }{v^{2}
}=O\left( 1\right) ,
\]
\[
a_{4,v}=\frac{96\gamma_{2}\mathrm{tr}\left[ \left( \mathbf{I}_{T}
\mathbf{\odot M}_{F}\right) \mathbf{M}_{F}\right] }{v^{2}}=O\left( \frac
{1}{v}\right) ,
\]
\[
a_{5,v}=\frac{48\gamma_{2}\left[ \boldsymbol{\tau}_{T}^{\prime}\left(
\mathbf{I}_{T}\mathbf{\odot M}_{F}\right) \left( \mathbf{I}_{T}\mathbf{\odot
M}_{F}\right) \boldsymbol{\tau}_{T}\right] }{v^{2}}=O\left( \frac{1}
{v}\right) ,
\]
\[
a_{6,v}=\frac{\left( 4\gamma_{4}v+24\gamma_{4}\right) \mathrm{tr}\left(
\mathbf{M}_{F}\mathbf{\odot M}_{F}\mathbf{\odot M}_{F}\right) }{v^{2}
}=O\left( 1\right) ,
\]
\[
a_{7,v}=\frac{\gamma_{6}\mathrm{tr}\left( \mathbf{M}_{F}\mathbf{\odot M}
_{F}\mathbf{\odot M}_{F}\mathbf{\odot M}_{F}\right) }{v^{2}}=O\left(
\frac{1}{v}\right) ,
\]
\[
a_{8,v}=\frac{48\gamma_{1}^{2}\left[ \boldsymbol{\tau}_{T}^{\prime}\left(
\mathbf{I}_{T}\mathbf{\odot M}_{F}\right) \mathbf{M}_{F}\left(
\mathbf{I}_{T}\mathbf{\odot M}_{F}\right) \boldsymbol{\tau}_{T}\right]
}{v^{2}}=O\left( \frac{1}{\sqrt{v}}\right) ,
\]
\[
a_{9,v}=\frac{96\gamma_{1}^{2}\boldsymbol{\tau}_{T}^{\prime}\left(
\mathbf{M}_{F}\mathbf{\odot M}_{F}\right) \mathbf{M}_{F}\left(
\mathbf{I}_{T}\mathbf{\odot M}_{F}\right) \boldsymbol{\tau}_{T}}{v^{2}
}=O\left( \frac{1}{\sqrt{v}}\right) ,
\]
\[
a_{10,v}=\frac{96\gamma_{1}^{2}\mathrm{tr}\left[ \left( \mathbf{M}
_{F}\mathbf{\odot M}_{F}\right) \mathbf{M}_{F}\right] }{v^{2}}=O\left(
\frac{1}{v}\right) ,
\]
\[
a_{11,v}=\frac{3\gamma_{2}^{2}\left[ \mathrm{tr}\left( \mathbf{M}
_{F}\mathbf{\odot M}_{F}\right) \right] ^{2}}{v^{2}}=O\left( 1\right) ,
\]
\[
a_{12,v}=\frac{24\gamma_{2}^{2}\left[ \boldsymbol{\tau}_{T}^{\prime}\left(
\mathbf{I}_{T}\mathbf{\odot M}_{F}\right) \left( \mathbf{M}_{F}\mathbf{\odot
M}_{F}\right) \left( \mathbf{I}_{T}\mathbf{\odot M}_{F}\right)
\boldsymbol{\tau}_{T}\right] }{v^{2}}=O\left( \frac{1}{v}\right) ,
\]
\[
a_{13,v}=\frac{8\gamma_{2}^{2}\left[ \boldsymbol{\tau}_{T}^{\prime}\left(
\mathbf{M}_{F}\mathbf{\odot M}_{F}\mathbf{\odot M}_{F}\mathbf{\odot M}
_{F}\right) \boldsymbol{\tau}_{T}\right] }{v^{2}}=O\left( \frac{1}
{v}\right) ,
\]
\[
a_{14,v}=\frac{24\gamma_{1}\gamma_{3}\left[ \boldsymbol{\tau}_{T}^{\prime
}\left( \mathbf{I}_{T}\mathbf{\odot M}_{F}\right) \mathbf{M}_{F}\left(
\mathbf{I}_{T}\mathbf{\odot M}_{F}\mathbf{\odot M}_{F}\right)
\boldsymbol{\tau}_{T}\right] }{v^{2}}=O\left( \frac{1}{\sqrt{v}}\right) ,
\]
\[
a_{15,v}=\frac{32\gamma_{1}\gamma_{3}\left[ \boldsymbol{\tau}_{T}^{\prime
}\left( \mathbf{I}_{T}\mathbf{\odot M}_{F}\right) \left( \mathbf{M}
_{F}\mathbf{\odot M}_{F}\mathbf{\odot M}_{F}\right) \boldsymbol{\tau}
_{T}\right] }{v^{2}}=O\left( \frac{1}{v}\right) .
\]
Using these results in ((ref)) it now follows that $E\left( q_{v}
^{4}\right) =O\left( 1\right) $, as required.
lemmaSuppose the $T\times1$ vector $\boldsymbol{\varepsilon
}=\mathbf{(}\varepsilon_{1},\varepsilon_{2},...,\varepsilon_{T})^{\prime}$ is
$\boldsymbol{\varepsilon}\thicksim IID(\mathbf{0},\mathbf{I}_{T})$, $\sup
_{t}E\left( \left\vert \varepsilon_{t}\right\vert ^{8+s}\right) <C$ for some
small $s>0$, and $\mathbf{M}_{F}=\mathbf{I}_{T}-\mathbf{F(F}^{\prime
}\mathbf{F)}^{-1}\mathbf{F}^{\prime}$, where the $T\times m_{0}$ matrix
$\mathbf{F}$ is distributed independently of $\boldsymbol{\varepsilon}$.
Suppose there exists a finite integer $v_{0}$ such that for all $v>v_{0}$,
\begin{equation}
\frac{\boldsymbol{\varepsilon}^{\prime}\mathbf{M}_{F}\boldsymbol{\varepsilon}
}{v}>c>0.
\end{equation}
Then for $v>v_{0}$,
\begin{equation}
E\left[ \left( \frac{\boldsymbol{\varepsilon}^{\prime}\mathbf{M}
_{F}\boldsymbol{\varepsilon}}{v}\right) ^{1/2}\right] =1+O\left( \frac
{1}{v}\right) , E\left[ \left( \frac{\boldsymbol{\varepsilon
}^{\prime}\mathbf{M}_{F}\boldsymbol{\varepsilon}}{v}\right) ^{-s/2}\right]
=1+O\left( \frac{1}{v}\right) , for s=1,2,3,4,
\end{equation}
\begin{equation}
E\left[ \varepsilon_{t}^{2}\left( \frac{\boldsymbol{\varepsilon}^{\prime
}\mathbf{M}_{F}\boldsymbol{\varepsilon}}{v}\right) \right] =1+O\left(
\frac{1}{v}\right) ,
\end{equation}
\begin{equation}
E\left[ \varepsilon_{t}^{2}\left( \frac{\boldsymbol{\varepsilon}^{\prime
}\mathbf{M}_{F}\boldsymbol{\varepsilon}}{v}\right) ^{-1}\right] =1+O\left(
\frac{1}{v}\right) ,
\end{equation}
\begin{equation}
E\left[ \varepsilon_{t}^{2}\left( \frac{\boldsymbol{\varepsilon}^{\prime
}\mathbf{M}_{F}\boldsymbol{\varepsilon}}{v}\right) ^{-1/2}\right]
=1+O\left( \frac{1}{v}\right) ,
\end{equation}
\begin{equation}
E\left[ \varepsilon_{t}\varepsilon_{t^{\prime}}\left( \frac
{\boldsymbol{\varepsilon}^{\prime}\mathbf{M}_{F}\boldsymbol{\varepsilon}}
{v}\right) ^{-1/2}\right] =O\left( \frac{1}{v}\right) , for t\neq
t^{\prime},
\end{equation}
\begin{equation}
E\left[ \varepsilon_{t}\left( \frac{\boldsymbol{\varepsilon}^{\prime
}\mathbf{M}_{F}\boldsymbol{\varepsilon}}{v}\right) ^{-1/2}\right] =O\left(
\frac{1}{v}\right) .
\end{equation}
proofTo establish the results in ((ref)) we first note that
\begin{equation}
\frac{\boldsymbol{\varepsilon}^{\prime}\mathbf{M}_{F}\boldsymbol{\varepsilon}
}{v}=1+\frac{1}{\sqrt{v}}q_{v}
\end{equation}
where
\begin{equation}
q_{v}=\sqrt{v}\left( \frac{\boldsymbol{\varepsilon}^{\prime}\mathbf{M}
_{F}\boldsymbol{\varepsilon}}{v}-1\right) .
\end{equation}
Applying the Taylor Theorem to $\left( v^{-1}\boldsymbol{\varepsilon}
^{\prime}\mathbf{M}_{F}\boldsymbol{\varepsilon}\right) ^{1/2}$ we have
\begin{equation}
\left( \frac{\boldsymbol{\varepsilon}^{\prime}\mathbf{M}_{F}
\boldsymbol{\varepsilon}}{v}\right) ^{1/2}=1+\frac{1}{2}\frac{q_{v}}{\sqrt
{v}}-\frac{1}{8v}R_{v},
\end{equation}
where
\[
R_{v}=\left( 1+\frac{\bar{q}_{v}}{\sqrt{v}}\right) ^{^{-3/2}}q_{v}^{2},
\]
and $\bar{q}_{v}$ lies on the interval between $0$ and $q_{v}$. Since
$E\left( q_{v}\right) =0$ as shown by result ((ref)) of Lemma (ref),
taking expectations of both sides of ((ref)) yields \ \
\begin{equation}
E\left( \frac{\boldsymbol{\varepsilon}^{\prime}\mathbf{M}_{F}
\boldsymbol{\varepsilon}}{v}\right) ^{1/2}=1-\frac{1}{8v}E\left(
R_{v}\right) .
\end{equation}
It is, therefore, sufficient to show that $E\left( \left\vert R_{v}
\right\vert \right) <C$. By Cauchy-Schwarz inequality we have
\begin{equation}
E\left( \left\vert R_{v}\right\vert \right) \leq\left[ E\left( \left\vert
1+\frac{\bar{q}_{v}}{\sqrt{v}}\right\vert ^{-3}\right) \right] ^{1/2}\left[
E\left( q_{v}^{4}\right) \right] ^{1/2}.
\end{equation}
Consider $\left\vert 1+\frac{\bar{q}_{v}}{\sqrt{v}}\right\vert $\ and
distinguish the cases (a) $q_{v}\geq0$ or equivalently if $\frac
{\boldsymbol{\varepsilon}^{\prime}\mathbf{M}_{F}\boldsymbol{\varepsilon}}
{v}\geq1$ and (b) $q_{v}<0$ or equivalently if $\frac{\boldsymbol{\varepsilon
}^{\prime}\mathbf{M}_{F}\boldsymbol{\varepsilon}}{v}<1$. Under (a)$\ 0\leq$
$\bar{q}_{v}<q_{v}$, we have$\left\vert 1+\frac{\bar{q}_{v}}{\sqrt{v}
}\right\vert \geq1$. Under (b) $q_{v}<\bar{q}_{v}<0$, we have $\left\vert
1+\frac{\bar{q}_{v}}{\sqrt{v}}\right\vert >\left\vert 1+\frac{q_{v}}{\sqrt{v}
}\right\vert =\frac{\boldsymbol{\varepsilon}^{\prime}\mathbf{M}_{F}
\boldsymbol{\varepsilon}}{v}$, and under condition ((ref)) $\left\vert
1+\frac{\bar{q}_{v}}{\sqrt{v}}\right\vert >c>0$. Hence, irrespective of
whether $q_{v}\geq0$ or not,
\begin{equation}
\left\vert 1+\frac{\bar{q}_{v}}{\sqrt{v}}\right\vert >c>0,
\end{equation}
and we have $E\left( \left\vert 1+\frac{\bar{q}_{v}}{\sqrt{v}}\right\vert
^{-3}\right) <C$. Also it is established that $E\left( q_{v}^{4}\right)
=O(1)$ by result ((ref)) of Lemma (ref). Using these results in
((ref)) it follows that $E\left( \left\vert R_{v}\right\vert \right) <C$,
and given ((ref)) we can show$E\left[ \left( \boldsymbol{\varepsilon
}^{\prime}\mathbf{M}_{F}\boldsymbol{\varepsilon}/v\right) ^{1/2}\right]
=1+O\left( v^{-1}\right) . $ The other results in ((ref)) can also be
established similarly. Result ((ref)) follows (S.7) in Lemma 6 of Pesaran
and Yamagata (2024) by setting\ $\varepsilon_{t}^{2}
=\boldsymbol{\varepsilon}^{\prime}\mathbf{A}_{1}\boldsymbol{\varepsilon}$,
where $\mathbf{A}_{1}$ has only one none-zero element on its diagonal. Result
((ref)) can be established using a result due to Lieberman (1994) (see
Lemmas 5 and 21 in the online supplement of Pesaran and Yamagata (2024)). To
establish ((ref)), note that\ by applying the Taylor Theorem to $\left(
v^{-1}\boldsymbol{\varepsilon}^{\prime}\mathbf{M}_{F}\boldsymbol{\varepsilon
}\right) ^{-1/2}$,
\[
\varepsilon_{t}^{2}\left( \frac{\boldsymbol{\varepsilon}^{\prime}
\mathbf{M}_{F}\boldsymbol{\varepsilon}}{v}\right) ^{^{-1/2}}=\varepsilon
_{t}^{2}-\frac{1}{2}\frac{q_{v}\varepsilon_{t}^{2}}{\sqrt{v}}+\frac{3}
{8v}R_{e,v}
\]
where $R_{e,v}=\left( 1+\frac{\bar{q}_{v}}{\sqrt{v}}\right) ^{-5/2}q_{v}
^{2}\varepsilon_{t}^{2}, $ and taking expectations yields
\begin{equation}
E\left[ \varepsilon_{t}^{2}\left( \frac{\boldsymbol{\varepsilon}^{\prime
}\mathbf{M}_{F}\boldsymbol{\varepsilon}}{v}\right) ^{^{-1/2}}\right]
=E\left( \varepsilon_{t}^{2}\right) -\frac{1}{2}\frac{E\left(
q_{v}\varepsilon_{t}^{2}\right) }{\sqrt{v}}+\frac{3}{8v}E\left(
R_{e,v}\right) .
\end{equation}
$E(\varepsilon_{t}^{2})=1$, and using ((ref)) we have
\begin{equation}
E\left[ \varepsilon_{t}^{2}\left( \frac{q_{v}}{\sqrt{v}}\right) \right]
=E\left[ \varepsilon_{t}^{2}\left( \frac{\boldsymbol{\varepsilon}^{\prime
}\mathbf{M}_{F}\boldsymbol{\varepsilon}}{v}-1\right) \right] =O\left(
\frac{1}{v}\right) .
\end{equation}
By Cauchy-Schwarz inequality
\begin{align}
E\left( \left\vert R_{e,v}\right\vert \right) & \leq\left[ E\left[
\varepsilon_{t}^{4}\left( \left\vert 1+\frac{\bar{q}_{v}}{\sqrt{v}
}\right\vert ^{-5}\right) \right] \right] ^{1/2}\left[ E\left( q_{v}
^{4}\right) \right] ^{1/2}\nonumber\\
& \leq\left[ E\left( \left\vert 1+\frac{\bar{q}_{v}}{\sqrt{v}}\right\vert
^{-10}\right) \right] ^{1/4}\left[ E\left( \varepsilon_{t}^{8}\right)
\right] ^{1/4}\left[ E\left( q_{v}^{4}\right) \right] ^{1/2}.
\end{align}
Given ((ref)), it is easily seen that $E\left( \left\vert 1+\frac{\bar
{q}_{v}}{\sqrt{v}}\right\vert ^{-10}\right) <C$. Also $E\left(
\varepsilon_{t}^{8}\right) <C$ by assumption and $E\left( q_{v}^{4}\right)
<C$ by ((ref)). Hence, using these results in ((ref)) it follows
$E\left( \left\vert R_{e,v}\right\vert \right) <C$, which completes the
proof of ((ref)). To establish ((ref)), using ((ref)) note that
for $t\neq t^{\prime}$,
\[
E\left[ \varepsilon_{t}\varepsilon_{t^{\prime}}\left( \frac
{\boldsymbol{\varepsilon}^{\prime}\mathbf{M}_{F}\boldsymbol{\varepsilon}}
{v}\right) ^{-1/2}\right] =E\left( \varepsilon_{t}\varepsilon_{t^{\prime}
}\right) -\frac{1}{2}\frac{E\left( q_{v}\varepsilon_{t}\varepsilon
_{t^{\prime}}\right) }{\sqrt{v}}+\frac{3}{8v}E\left( R_{tt^{\prime}
,v}\right)
\]
where $R_{tt^{\prime},v}=\left( 1+\frac{\bar{q}_{v}}{\sqrt{v}}\right)
^{-5/2}q_{v}^{2}\varepsilon_{t}\varepsilon_{t^{\prime}}. $ Note that $E\left(
\varepsilon_{t}\varepsilon_{t^{\prime}}\right) =0$ for $t\neq t^{\prime}$ by
serial independence of $\varepsilon_{t}$. In addition, using definition of
$q_{v}$ in ((ref)) yields
\begin{align*}
\frac{E\left( q_{v}\varepsilon_{t}\varepsilon_{t^{\prime}}\right) }{\sqrt
{v}} & =E\left[ \left( \frac{\boldsymbol{\varepsilon}^{\prime}
\mathbf{M}_{F}\boldsymbol{\varepsilon}}{v}-1\right) \varepsilon
_{t}\varepsilon_{t^{\prime}}\right] =E\left[ \left( \frac
{\boldsymbol{\varepsilon}^{\prime}\mathbf{M}_{F}\boldsymbol{\varepsilon}}
{v}\right) \varepsilon_{t}\varepsilon_{t^{\prime}}\right] =\frac{1}
{2v}E\left[ \left( \boldsymbol{\varepsilon}^{\prime}\mathbf{A}
\boldsymbol{\varepsilon}\right) \left( \boldsymbol{\varepsilon}^{\prime
}\mathbf{B}\boldsymbol{\varepsilon}\right) \right] ,
\end{align*}
where $\mathbf{A}=\mathbf{M}_{F}$ and $\mathbf{B}=\mathbf{(}b_{tt^{\prime}})$
with $b_{tt^{\prime}}$ and $b_{t^{\prime}t}$ ($t\neq t^{\prime}$) being the
only non-zero elements. Now using (S.7) of Lemma 6 in Pesaran and Yamagata
(2024) it follows that $E\left( q_{v}\varepsilon_{t}\varepsilon_{t^{\prime}
}\right) =0$. Also by Cauchy-Schwarz inequality ,
\begin{align*}
E\left( \left\vert R_{tt^{\prime},v}\right\vert \right) & \leq\left[
E\left[ \varepsilon_{t}^{2}\varepsilon_{t^{\prime}}^{2}\left( \left\vert
1+\frac{\bar{q}_{v}}{\sqrt{v}}\right\vert ^{-5}\right) \right] \right]
^{1/2}\left[ E\left( q_{v}^{4}\right) \right] ^{1/2}\\
& \leq\left[ E\left( \left\vert 1+\frac{\bar{q}_{v}}{\sqrt{v}}\right\vert
^{-10}\right) \right] ^{1/4}\left[ E\left( \varepsilon_{t}^{4}\right)
E\left( \varepsilon_{t^{\prime}}^{4}\right) \right] ^{1/4}\left[ E\left(
q_{v}^{4}\right) \right] ^{1/2}<C,
\end{align*}
where the final inequality follows using the same line of argument used to
bound $E\left( \left\vert R_{e,v}\right\vert \right) $ in ((ref)).
Overall, result ((ref)) is established. Finally consider ((ref)) and
note that\
\[
E\left[ \varepsilon_{t}\left( \frac{\boldsymbol{\varepsilon}^{\prime
}\mathbf{M}_{F}\boldsymbol{\varepsilon}}{v}\right) ^{-1/2}\right] =E\left(
\varepsilon_{t}\right) -\frac{1}{2}\frac{E\left( q_{v}\varepsilon
_{t}\right) }{\sqrt{v}}+\frac{3}{8v}E\left( R_{t,v}\right) ,
\]
where $R_{t,v}=\left( 1+\frac{\bar{q}_{v}}{\sqrt{v}}\right) ^{-5/2}q_{v}
^{2}\varepsilon_{t}. $ We have $E\left( \varepsilon_{t}\right) =0$ and
\[
\frac{E\left( q_{v}\varepsilon_{t}\right) }{\sqrt{v}}=E\left(
\frac{\boldsymbol{\varepsilon}^{\prime}\mathbf{M}_{F}\boldsymbol{\varepsilon
}\varepsilon_{t}}{v}-\varepsilon_{t}\right) =E\left( \frac
{\boldsymbol{\varepsilon}^{\prime}\mathbf{M}_{F}\boldsymbol{\varepsilon
}\varepsilon_{t}}{v}\right) .
\]
Denote $\left\{ m_{jj^{\prime}}:j,j^{\prime}=1,2,\ldots,T\right\} $ as the
element of $\mathbf{M}_{F}$, such that based on part (c) of Assumption
(ref) $m_{jj^{\prime}}$ is independent from $\varepsilon_{t}$ and
$E\left( m_{jj^{\prime}}\right) =O\left( 1\right) $. Then it follows that
\begin{align*}
E\left( \frac{\boldsymbol{\varepsilon}^{\prime}\mathbf{M}_{F}
\boldsymbol{\varepsilon}\varepsilon_{t}}{v}\right) & =\frac{1}{v}\sum
_{j=1}^{v}\sum_{j^{\prime}=1}^{v}E\left( \varepsilon_{j}m_{jj^{\prime}
}\varepsilon_{j^{\prime}}\varepsilon_{t}\right) =\frac{1}{v}\sum_{j=1}
^{v}\sum_{j^{\prime}=1}^{v}E\left( \varepsilon_{j}\varepsilon_{j^{\prime}
}\varepsilon_{t}\right) E\left( m_{jj^{\prime}}\right) =\frac{1}{v}E\left(
\varepsilon_{t}^{3}\right) E\left( m_{tt}\right)
\end{align*}
which is $O\left( v^{-1}\right) $ where the second equation holds due to the
independence of $\varepsilon_{t}$ and $m_{jj^{\prime}},$ while the third
equation holds due to the serial independence of $\varepsilon_{t}$. Besides by
Cauchy-Schwarz inequality, \
\begin{align*}
E\left( \left\vert R_{t,v}\right\vert \right) & \leq\left[ E\left[
\varepsilon_{t}^{2}\left( \left\vert 1+\frac{\bar{q}_{v}}{\sqrt{v}
}\right\vert ^{-5}\right) \right] \right] ^{1/2}\left[ E\left( q_{v}
^{4}\right) \right] ^{1/2}\\
& \leq\left[ E\left( \left\vert 1+\frac{\bar{q}_{v}}{\sqrt{v}}\right\vert
^{-10}\right) \right] ^{1/4}\left[ E\left( \varepsilon_{t}^{4}\right)
\right] ^{1/4}\left[ E\left( q_{v}^{4}\right) \right] ^{1/2}<C,
\end{align*}
where the last inequality holds again using the same line of argument used to
bound $E\left( \left\vert R_{e,v}\right\vert \right) $ in ((ref)).
Overall, result ((ref)) is established.
lemmaConsider the latent factor model given by
((ref)) and ((ref)). $\widetilde{CD}$ and $CD$ statistics are
defined by ((ref)) and ((ref)). Suppose that Assumptions
(ref)-(ref) hold and $\left( n,T\right) \rightarrow
\infty$, such that $n/T\rightarrow\kappa\,,$ for $0<\kappa<\infty$. Then
\begin{equation}
CD=\widetilde{CD}+o_{p}\left( 1\right) .
\end{equation}
proofUsing ((ref)) and ((ref)) we first note that
\begin{equation}
\left( \sqrt{\frac{2\left( n-1\right) }{n}}\right) \left(
CD-\widetilde{CD}\right) =\frac{1}{\sqrt{T}}\sum_{t=1}^{T}\left[ \left(
\frac{1}{\sqrt{n}}\sum_{i=1}^{n}\frac{\hat{u}_{it}}{\hat{\sigma}_{i,T}
}\right) ^{2}-\left( \frac{1}{\sqrt{n}}\sum_{i=1}^{n}\frac{\hat{u}_{it}
}{\omega_{i,T}}\right) ^{2}\right] .
\end{equation}
Also note that
\begin{equation}
\frac{1}{\sqrt{n}}\sum_{i=1}^{n}\frac{\hat{u}_{it}}{\hat{\sigma}_{i,T}
}=h_{t,nT\ }+g_{t,nT}
\end{equation}
where (also see ((ref)))
\[
h_{t,nT\ }=\frac{1}{\sqrt{n}}\sum_{i=1}^{n}\frac{\hat{u}_{it}}{\omega_{i,T}
}=\frac{\mathbf{c}_{nT}^{\prime}\mathbf{\hat{u}}_{\circ t}}{\sqrt{n}}\text{,
and }g_{t,nT}=\frac{1}{\sqrt{n}}\sum_{i=1}^{n}\hat{u}_{it}\left( \frac
{1}{\hat{\sigma}_{i,T}}-\frac{1}{\omega_{i,T}}\right) =\frac{\mathbf{d}
_{nT}^{\prime}\mathbf{\hat{u}}_{\circ t}}{\sqrt{n}}
\]
$\mathbf{\hat{u}}_{\circ t}=(\hat{u}_{1t},\hat{u}_{2t},...,\hat{u}
_{nt})^{\prime}$, $\mathbf{c}_{nT}=(\omega_{1,T\ }^{-1},\omega_{2,T}
^{-1},...,\omega_{n,T}^{-1})^{\prime}$, $\mathbf{d}_{nT}=(d_{1T}
,d_{2T},...,d_{nT})^{\prime}\,$, and $d_{iT}=\hat{\sigma}_{i,T}^{-1}
-\omega_{i,T}^{-1}$. Then squaring both sides of ((ref)) and using the
result in ((ref)) we have
\begin{align}
\left( \sqrt{\frac{2\left( n-1\right) }{n}}\right) \left(
CD-\widetilde{CD}\right) & =\frac{1}{\sqrt{T}}\sum_{t=1}^{T}g_{t,nT}
^{2}+\frac{2}{\sqrt{T}}\sum_{t=1}^{T}h_{t,nT\ }g_{t,nT}\nonumber\\
& =\sqrt{\frac{T}{n}}\left( \frac{1}{\sqrt{n}}\mathbf{d}_{nT}^{\prime
}\mathbf{\hat{V}}_{T}\mathbf{d}_{nT}+\frac{2}{\sqrt{n}}\mathbf{c}_{nT}
^{\prime}\mathbf{\hat{V}}_{T}\mathbf{d}_{nT}\right) ,
\end{align}
where $\mathbf{\hat{V}}_{T}=T^{-1}\sum_{t=1}^{T}\mathbf{\hat{u}}_{\circ
t}\mathbf{\hat{u}}_{\circ t}^{\prime}$. Now using ((ref)), the error
vector $\mathbf{\hat{u}}_{\circ t}$ can be written as
\[
\mathbf{\hat{u}}_{\circ t}=\mathbf{u}_{\circ t}\left( \lambda_{T}\right)
-\mathbf{\Gamma}\left( \mathbf{\hat{f}}_{t}-\mathbf{f}_{t}\right) -\left(
\mathbf{\hat{\Gamma}}-\mathbf{\Gamma}\right) \mathbf{f}_{t}-\left(
\mathbf{\hat{\Gamma}}-\mathbf{\Gamma}\right) \left( \mathbf{\hat{f}}
_{t}-\mathbf{f}_{t}\right) .
\]
where $\mathbf{u}_{\circ t}\left( \lambda_{T}\right) =\left( \sigma
_{1}\varepsilon_{1t}\left( \lambda_{T}\right) ,\sigma_{2}\varepsilon
_{2t}\left( \lambda_{T}\right) ,\ldots,\sigma_{n}\varepsilon_{nt}\left(
\lambda_{T}\right) \right) ^{\prime}$. Using this expression we now have
\begin{align*}
\mathbf{\hat{V}}_{T} & =T^{-1}\sum_{t=1}^{T}\mathbf{u}_{\circ t}\left(
\lambda_{T}\right) \mathbf{u}_{\circ t}^{\prime}\left( \lambda_{T}\right)
+\mathbf{\Gamma}\left[ T^{-1}\sum_{t=1}^{T}\left( \mathbf{\hat{f}}
_{t}-\mathbf{f}_{t}\right) \left( \mathbf{\hat{f}}_{t}-\mathbf{f}
_{t}\right) ^{\prime}\right] \mathbf{\Gamma}^{\prime}\\
& +\left( \mathbf{\hat{\Gamma}-\Gamma}\right) \left( T^{-1}\sum_{t=1}
^{T}\mathbf{f}_{t}\mathbf{f}_{t}^{\prime}\right) \left( \mathbf{\hat{\Gamma
}-\Gamma}\right) ^{\prime}+\left( \mathbf{\hat{\Gamma}-\Gamma}\right)
\left[ T^{-1}\sum_{t=1}^{T}\left( \mathbf{\hat{f}}_{t}-\mathbf{f}
_{t}\right) \left( \mathbf{\hat{f}}_{t}-\mathbf{f}_{t}\right) ^{\prime
}\right] \left( \mathbf{\hat{\Gamma}-\Gamma}\right) ^{\prime}
\mathbf{\Gamma}^{\prime}\\
& -\left[ T^{-1}\sum_{t=1}^{T}\mathbf{u}_{\circ t}\left( \lambda
_{T}\right) \left( \mathbf{\hat{f}}_{t}-\mathbf{f}_{t}\right) ^{\prime
}\right] -\left[ T^{-1}\sum_{t=1}^{T}\mathbf{u}_{\circ t}\left( \lambda
_{T}\right) \mathbf{f}_{t}^{\prime}\right] \left( \mathbf{\hat{\Gamma
}-\Gamma}\right) ^{\prime}\\
& -\left[ T^{-1}\sum_{t=1}^{T}\mathbf{u}_{\circ t}\left( \lambda
_{T}\right) \left( \mathbf{\hat{f}}_{t}-\mathbf{f}_{t}\right) ^{\prime
}\right] \left( \mathbf{\hat{\Gamma}-\Gamma}\right) ^{\prime}
+\mathbf{\Gamma}\left[ T^{-1}\sum_{t=1}^{T}\left( \mathbf{\hat{f}}
_{t}-\mathbf{f}_{t}\right) \mathbf{f}_{t}^{\prime}\right] \left(
\mathbf{\hat{\Gamma}-\Gamma}\right) ^{\prime}\\
& +\mathbf{\Gamma}\left[ T^{-1}\sum_{t=1}^{T}\left( \mathbf{\hat{f}}
_{t}-\mathbf{f}_{t}\right) \left( \mathbf{\hat{f}}_{t}-\mathbf{f}
_{t}\right) ^{\prime}\right] \left( \mathbf{\hat{\Gamma}-\Gamma}\right)
^{\prime}+\left( \mathbf{\hat{\Gamma}-\Gamma}\right) \left[ T^{-1}
\sum_{t=1}^{T}\mathbf{f}_{t}\left( \mathbf{\hat{f}}_{t}-\mathbf{f}
_{t}\right) ^{\prime}\right] \left( \mathbf{\hat{\Gamma}-\Gamma}\right)
^{\prime},
\end{align*}
or in matrix forms
\begin{align*}
\mathbf{\hat{V}}_{T} & =\mathbf{V}_{T}\left( \lambda_{T}\right)
+\mathbf{\Gamma}\left[ T^{-1}\left( \mathbf{\hat{F}-F}\right) ^{\prime
}\left( \mathbf{\hat{F}-F}\right) \right] \mathbf{\Gamma}^{\prime}+\left(
\mathbf{\hat{\Gamma}-\Gamma}\right) \mathbf{\Sigma}_{T,ff}\left(
\mathbf{\hat{\Gamma}-\Gamma}\right) ^{\prime}\\
& +\left( \mathbf{\hat{\Gamma}-\Gamma}\right) \left[ T^{-1}\left(
\mathbf{\hat{F}-F}\right) ^{\prime}\left( \mathbf{\hat{F}-F}\right)
\right] \left( \mathbf{\hat{\Gamma}-\Gamma}\right) ^{\prime}-T^{-1}
\mathbf{U}^{\prime}\left( \lambda_{T}\right) \left( \mathbf{\hat{F}
-F}\right) \mathbf{\Gamma}^{\prime}\\
& -T^{-1}\mathbf{U}^{\prime}\left( \lambda_{T}\right) \mathbf{F}\left(
\mathbf{\hat{\Gamma}-\Gamma}\right) ^{\prime}-T^{-1}\mathbf{U}^{\prime
}\left( \lambda_{T}\right) \left( \mathbf{\hat{F}-F}\right) \left(
\mathbf{\hat{\Gamma}-\Gamma}\right) ^{\prime}+\mathbf{\Gamma}\left[
T^{-1}\mathbf{F}^{\prime}\left( \mathbf{\hat{F}-F}\right) \right] \left(
\mathbf{\hat{\Gamma}-\Gamma}\right) ^{\prime}\\
& +\mathbf{\Gamma}\left[ T^{-1}\left( \mathbf{\hat{F}-F}\right) ^{\prime
}\left( \mathbf{\hat{F}-F}\right) \right] \left( \mathbf{\hat{\Gamma
}-\Gamma}\right) ^{\prime}+\left( \mathbf{\hat{\Gamma}-\Gamma}\right)
\left[ T^{-1}\mathbf{F}^{\prime}\left( \mathbf{\hat{F}-F}\right) \right]
\left( \mathbf{\hat{\Gamma}-\Gamma}\right) ^{\prime},
\end{align*}
where $\mathbf{V}_{T}\left( \lambda_{T}\right) =T^{-1}\sum_{t=1}
^{T}\mathbf{u}_{\circ t}\left( \lambda_{T}\right) \mathbf{u}_{\circ
t}^{^{\prime}}\left( \lambda_{T}\right) $ and $\mathbf{U}\left( \lambda
_{T}\right) =\left( \mathbf{u}_{\circ1}\left( \lambda_{T}\right)
,\mathbf{u}_{\circ2}\left( \lambda_{T}\right) ,\ldots,\mathbf{u}_{\circ
T}\left( \lambda_{T}\right) \right) ^{\prime}$. By ((ref)) we have
$\left\Vert \mathbf{V}_{T}\left( \lambda_{T}\right) \right\Vert =\mu_{\max
}\left( \mathbf{V}_{T}\left( \lambda_{T}\right) \right) =O_{p}(\frac{n}
{T})$. By results in Lemma (ref) all other terms of the $\mathbf{\hat{V}
}_{T}$ are either $O_{p}(1)$ or of lower order, and we also have
$\mathbf{\hat{V}}_{T}=O_{p}\left( 1\right) $ since $n$ and $T$ are of the
same order. Consider now the terms in ((ref)) and note that
\[
\left( \sqrt{\frac{2\left( n-1\right) }{n}}\right) \left\vert
CD-\widetilde{CD}\right\vert <C\left\Vert \mathbf{\hat{V}}_{T}\right\Vert
\left[ \left( \frac{1}{\sqrt{n}}\left\Vert \mathbf{d}_{nT}\right\Vert
^{2}\right) +\left( \frac{2}{\sqrt{n}}\left\Vert \mathbf{c}_{nT}\right\Vert
\right) \left\Vert \mathbf{d}_{nT}\right\Vert \right] .
\]
where $\frac{1}{\sqrt{n}}\left\Vert \mathbf{c}_{nT}\right\Vert =\left(
n^{-1}\sum_{i=1}^{n}\omega_{i,T}^{-2}\right) ^{1/2}$ and $\left\Vert
\mathbf{d}_{nT}\right\Vert =\left( \sum_{i=1}^{n}\left( \hat{\sigma}
_{i,T}^{-1}-\omega_{i,T}^{-1}\right) ^{2}\right) ^{1/2}$. By part (c) of
Assumption (ref) $E\left( \omega_{i,T}^{-2}\right) <C<\infty$,
such that $\frac{1}{n}E\left\Vert \mathbf{c}_{nT}\right\Vert ^{2}=n^{-1}
\sum_{i=1}^{n}E\left( \omega_{i,T}^{-2}\right) <C$ and hence $n^{-1/2}
\left\Vert \mathbf{c}_{nT}\right\Vert =O_{p}\left( 1\right) $. Also by
((ref)) we have
\[
\sup_{i}\left\vert \hat{\sigma}_{i,T}^{-1}-\omega_{i,T}^{-1}\right\vert
^{2}=\left( \sup_{i}\left\vert \hat{\sigma}_{i,T}^{-1}-\omega_{i,T}
^{-1}\right\vert \right) ^{2}=O_{p}\left[ \left( \frac{\ln\left( n\right)
}{T}\right) ^{2}\right] ,
\]
therefore
\[
\left\Vert \mathbf{d}_{nT}\right\Vert ^{2}=\sum_{i=1}^{n}\left( \hat{\sigma
}_{i,T}^{-1}-\omega_{i,T}^{-1}\right) ^{2}\leq n\sup_{i}\left( \hat{\sigma
}_{i,T}^{-1}-\omega_{i,T}^{-1}\right) ^{2}=O_{p}\left[ n\left( \frac
{\ln\left( n\right) }{T}\right) ^{2}\right] =o_{p}\left( 1\right) ,
\]
recalling that $n$ and $T$ are of the same order. Hence, $\left\vert
CD-\widetilde{CD}\right\vert =o_{p}(1)$, as required.
lemmaSuppose the data are generated by the latent factor model given by
((ref)) and ((ref)). The latent factors, $\mathbf{f}_{t}$, and
their loadings, $\boldsymbol{\gamma}_{i}$, are estimated by principal
components, $\mathbf{\hat{f}}_{t}$ and $\boldsymbol{\hat{\gamma}}_{i}$, given
by ((ref)). Suppose that Assumptions (ref)-(ref) hold
and $\left( n,T\right) \rightarrow\infty$, such that $n/T\rightarrow
\kappa\,,$ for $0<\kappa<\infty$. Then
\begin{align}
\frac{1}{n}\sum_{i=1}^{n}\frac{\boldsymbol{\hat{\gamma}}_{i}
-\boldsymbol{\gamma}_{i}}{\sigma_{i}} & =O_{p}\left( \sqrt{\frac{\ln\left(
n\right) }{nT}}\right) ,\\
\frac{1}{n}\sum_{i=1}^{n}\left( \boldsymbol{\hat{\gamma}}_{i}
-\boldsymbol{\gamma}_{i}\right) \sigma_{i} & =O_{p}\left( \sqrt{\frac
{\ln\left( n\right) }{nT}}\right) ,\\
\frac{1}{n}\sum_{i=1}^{n}\sigma_{i}^{2}\left( \boldsymbol{\hat{\gamma}}
_{i}\boldsymbol{\hat{\gamma}}_{i}^{\prime}-\boldsymbol{\gamma}_{i}
\boldsymbol{\gamma}_{i}^{\prime}\right) & =O_{p}\left( \sqrt{\frac
{\ln\left( n\right) }{nT}}\right) ,\\
\frac{1}{n}\sum_{i=1}^{n}\hat{\sigma}_{i,T}\boldsymbol{\hat{\gamma}}_{i}
-\frac{1}{n}\sum_{i=1}^{n}\sigma_{i}\boldsymbol{\gamma}_{i} & =O_{p}\left(
\frac{\ln\left( n\right) }{T}\right) .
\end{align}
proofResults ((ref)) and ((ref)) follow directly from ((ref)) by
setting $b_{in}=\sigma_{i}^{-1}$ and $b_{in}=\sigma_{i}$, respectively. To
prove ((ref)), note that
\begin{align}
\frac{1}{n}\sum_{i=1}^{n}\sigma_{i}^{2}\left( \boldsymbol{\hat{\gamma}}
_{i}\boldsymbol{\hat{\gamma}}_{i}^{\prime}-\boldsymbol{\gamma}_{i}
\boldsymbol{\gamma}_{i}^{\prime}\right) & =\frac{1}{n}\sum_{i=1}^{n}
\sigma_{i}^{2}\left( \boldsymbol{\hat{\gamma}}_{i}-\boldsymbol{\gamma}
_{i}\right) \left( \boldsymbol{\hat{\gamma}}_{i}-\boldsymbol{\gamma}
_{i}\right) ^{\prime}\nonumber\\
& +\frac{1}{n}\sum_{i=1}^{n}\sigma_{i}^{2}\left( \boldsymbol{\hat{\gamma}
}_{i}-\boldsymbol{\gamma}_{i}\right) \boldsymbol{\gamma}_{i}^{\prime}
+\frac{1}{n}\sum_{i=1}^{n}\sigma_{i}^{2}\boldsymbol{\gamma}_{i}\left(
\boldsymbol{\hat{\gamma}}_{i}-\boldsymbol{\gamma}_{i}\right) ^{\prime}.
\end{align}
Since $\sigma_{i}^{2}$ is bounded, then
\begin{align*}
\left\Vert \frac{1}{n}\sum_{i=1}^{n}\sigma_{i}^{2}\left( \boldsymbol{\hat
{\gamma}}_{i}-\boldsymbol{\gamma}_{i}\right) \left( \boldsymbol{\hat{\gamma
}}_{i}-\boldsymbol{\gamma}_{i}\right) ^{\prime}\right\Vert & \leq\left(
\sup_{i}\sigma_{i}^{2}\right) \left\Vert \frac{1}{n}\sum_{i=1}^{n}\left(
\boldsymbol{\hat{\gamma}}_{i}-\boldsymbol{\gamma}_{i}\right) \left(
\boldsymbol{\hat{\gamma}}_{i}-\boldsymbol{\gamma}_{i}\right) ^{\prime
}\right\Vert \\
& \leq\left( \sup_{i}\sigma_{i}^{2}\right) \left( \frac{1}{n}\sum
_{i=1}^{n}\left\Vert \boldsymbol{\hat{\gamma}}_{i}-\boldsymbol{\gamma}
_{i}\right\Vert ^{2}\right) ,
\end{align*}
and using ((ref)) it follows that
\[
\frac{1}{n}\sum_{i=1}^{n}\sigma_{i}^{2}\left( \boldsymbol{\hat{\gamma}}
_{i}-\boldsymbol{\gamma}_{i}\right) \left( \boldsymbol{\hat{\gamma}}
_{i}-\boldsymbol{\gamma}_{i}\right) ^{\prime}=O_{p}\left( \frac{1}
{\delta_{nT}^{2}}\right) .
\]
Also, using ((ref)) and setting $b_{in}=\sigma_{i}$, we have
\[
\left\Vert \frac{1}{n}\sum_{i=1}^{n}\sigma_{i}^{2}\left( \boldsymbol{\hat
{\gamma}}_{i}-\boldsymbol{\gamma}_{i}\right) \boldsymbol{\gamma}_{i}^{\prime
}\right\Vert =\left\Vert \frac{1}{n}\sum_{i=1}^{n}\sigma_{i}^{2}
\boldsymbol{\gamma}_{i}\left( \boldsymbol{\hat{\gamma}}_{i}
-\boldsymbol{\gamma}_{i}\right) ^{\prime}\right\Vert =O_{p}\left(
\sqrt{\frac{\ln\left( n\right) }{nT}}\right) ,
\]
then ((ref)) is established based on ((ref)). Finally, consider
((ref)) and note that
\begin{align}
& \frac{1}{n}\sum_{i=1}^{n}\hat{\sigma}_{i,T}\boldsymbol{\hat{\gamma}}
_{i}-\frac{1}{n}\sum_{i=1}^{n}\sigma_{i}\boldsymbol{\gamma}_{i}\nonumber\\
& =\frac{1}{n}\sum_{i=1}^{n}\left[ \left( \hat{\sigma}_{i,T}-\omega
_{i,T}\right) +\omega_{i,T}\right] \left( \boldsymbol{\hat{\gamma}}
_{i}-\boldsymbol{\gamma}_{i}+\boldsymbol{\gamma}_{i}\right) -\frac{1}{n}
\sum_{i=1}^{n}\sigma_{i}\boldsymbol{\gamma}_{i}\nonumber\\
& =\frac{1}{n}\sum_{i=1}^{n}\boldsymbol{\gamma}_{i}\left( \omega
_{i,T}-\sigma_{i}\right) +\frac{1}{n}\sum_{i=1}^{n}\boldsymbol{\gamma}
_{i}\left( \hat{\sigma}_{i,T}-\omega_{i,T}\right) +\frac{1}{n}\sum_{i=1}
^{n}\sigma_{i}\left( \boldsymbol{\hat{\gamma}}_{i}-\boldsymbol{\gamma}
_{i}\right) \nonumber\\
& +\frac{1}{n}\sum_{i=1}^{n}\left( \omega_{i,T}-\sigma_{i}\right) \left(
\boldsymbol{\hat{\gamma}}_{i}-\boldsymbol{\gamma}_{i}\right) +\frac{1}{n}
\sum_{i=1}^{n}\left( \hat{\sigma}_{i,T}-\omega_{i,T}\right) \left(
\boldsymbol{\hat{\gamma}}_{i}-\boldsymbol{\gamma}_{i}\right) \nonumber\\
& =\mathbf{A}_{1,nT}+\mathbf{A}_{2,nT}+\mathbf{A}_{3,nT}+\mathbf{A}
_{4,nT}+\mathbf{A}_{5,nT}.
\end{align}
Recall also that under Assumptions (ref) and (ref)
$\sigma_{i}$ and $\boldsymbol{\gamma}_{i}$ are bounded and $\omega_{i,T}
^{2}=T^{-1}\sigma_{i}^{2}\boldsymbol{\varepsilon}_{i\circ}^{^{\prime}
}\mathbf{M}_{F}\boldsymbol{\varepsilon}_{i\circ}$, for $i=1,2,...,n$ are
distributed independently across $i,$ and from $\sigma_{i}$ and
$\boldsymbol{\gamma}_{i}$. Starting with $\mathbf{A}_{1,nT}$, by ((ref)) we
have
\[
E\left( \sqrt{nT}\mathbf{A}_{1,nT}\right) =\frac{\sqrt{nT}}{n}\sum_{i=1}
^{n}(\boldsymbol{\gamma}_{i}\sigma_{i})E\left[ \left( \frac
{\boldsymbol{\varepsilon}_{i\circ}^{^{\prime}}\mathbf{M}_{F}
\boldsymbol{\varepsilon}_{i\circ}}{T}\right) ^{1/2}-1\right] =O\left(
\frac{\sqrt{nT}}{T}\right) .
\]
Since $n$ and $T$ are assumed to be of the same order then $E\left( \sqrt
{nT}\mathbf{A}_{1,nT}\right) =O\left( 1\right) $. Also, using result
((ref))
\[
Var\left( \sqrt{nT}\mathbf{A}_{1,nT}\right) =\frac{nT}{n^{2}}\sum_{i=1}
^{n}\left( \sigma_{i}^{2}\boldsymbol{\gamma}_{i}\boldsymbol{\gamma}
_{i}^{\prime}\right) Var\left[ \left( \frac{\boldsymbol{\varepsilon
}_{i\circ}^{^{\prime}}\mathbf{M}_{F}\boldsymbol{\varepsilon}_{i\circ}}
{T}\right) ^{1/2}\right] =O\left( 1\right) .
\]
Therefore, $\sqrt{nT}\mathbf{A}_{1,nT}=O_{p}(1)$ and it follows that
$\mathbf{A}_{1,nT}=O_{p}\left[ \left( nT\right) ^{-1/2}\right] $. Further,
using ((ref)) and setting $b_{in}=\gamma_{ij}$, for $j=1,2,...,m_{0}$, it
follows that
\[
\mathbf{A}_{2,nT}=\frac{1}{n}\sum_{i=1}^{n}\boldsymbol{\gamma}_{i}\left(
\hat{\sigma}_{i,T}-\omega_{i,T}\right) =O_{p}\left( \frac{\ln\left(
n\right) }{T}\right) .
\]
Since $\mathbf{A}_{3,nT}$ is the same as the result in ((ref)), which is
already established, then $\mathbf{A}_{3,nT}=O_{p}\left( \sqrt{\ln\left(
n\right) /\left( nT\right) }\right) $. Using result ((ref)) it
follows that
\[
\mathbf{A}_{4,nT}=\frac{1}{n}\sum_{i=1}^{n}\left( \omega_{i,T}-\sigma
_{i}\right) \left( \boldsymbol{\hat{\gamma}}_{i}-\boldsymbol{\gamma}
_{i}\right) =O_{p}\left( \frac{1}{\delta_{nT}^{2}}\right) .
\]
Using result ((ref)) we have
\[
\mathbf{A}_{5,nT}=\frac{1}{n}\sum_{i=1}^{n}\left( \hat{\sigma}_{i,T}
-\omega_{i,T}\right) \left( \boldsymbol{\hat{\gamma}}_{i}-\boldsymbol{\gamma
}_{i}\right) =O_{p}\left[ \left( \frac{\ln\left( n\right) }{T}\right)
^{3/2}\right] .
\]
Result ((ref)) now follows straightforwardly based on ((ref)).
lemmaSuppose the data are generated by the latent factor
model given by ((ref)) and ((ref)). Denote $\boldsymbol{\varphi
}_{n}=n^{-1}\sum_{i=1}^{n}\boldsymbol{\gamma}_{i}/\sigma_{i}$,
$\boldsymbol{\varphi}_{nT}=n^{-1}\sum_{i=1}^{n}\boldsymbol{\gamma}_{i}
/\omega_{i,T}$ with $\omega_{i,T}=\left( T^{-1}\sigma_{i}^{2}
\boldsymbol{\varepsilon}_{i\circ}^{\prime}\mathbf{M}_{F}
\boldsymbol{\varepsilon}_{i\circ}\right) ^{1/2}$ where
$\boldsymbol{\varepsilon}_{i\circ}=\left( \varepsilon_{i1},\varepsilon
_{i2},\ldots,\varepsilon_{iT}\right) ^{\prime}$ and $\mathbf{M}
_{F}=\mathbf{I}_{T}-\mathbf{F}\left( \mathbf{F}^{\prime}\mathbf{F}\right)
^{-1}\mathbf{F}^{\prime}$. Suppose that Assumptions (ref)
-(ref) hold and $\left( n,T\right) \rightarrow\infty$, such that
$n/T\rightarrow\kappa\,,$ for $0<\kappa<\infty$. Then
\begin{align*}
g_{2,nT}\left( \lambda_{T}\right) & =\frac{1}{\sqrt{T}}\sum_{t=1}
^{T}\upsilon_{t,nT}^{2}\left( \lambda_{T}\right) =o_{p}\left( 1\right) ,\\
g_{3,nT}\left( \lambda_{T}\right) & =\sqrt{T}\left( \boldsymbol{\varphi
}_{nT}-\boldsymbol{\varphi}_{n}\right) ^{\prime}\left( \frac{\sum_{t=1}
^{T}\boldsymbol{\kappa}_{t,n}\left( \lambda_{T}\right) \boldsymbol{\kappa
}_{t,n}^{\prime}\left( \lambda_{T}\right) }{T}\right) \left(
\boldsymbol{\varphi}_{nT}-\boldsymbol{\varphi}_{n}\right) =o_{p}\left(
1\right) ,\\
g_{4,nT}\left( \lambda_{T}\right) & =\sqrt{T}\left( \boldsymbol{\varphi
}_{nT}-\boldsymbol{\varphi}_{n}\right) ^{\prime}\left( \frac{1}{T}\sum
_{t=1}^{T}\boldsymbol{\kappa}_{t,n}\left( \lambda_{T}\right) \upsilon
_{t,nT}\left( \lambda_{T}\right) \right) =o_{p}\left( 1\right) ,\\
g_{5,nT}\left( \lambda_{T}\right) & =\sqrt{T}\left( \boldsymbol{\varphi
}_{nT}-\boldsymbol{\varphi}_{n}\right) ^{\prime}\left( \frac{1}{T}\sum
_{t=1}^{T}\boldsymbol{\kappa}_{t,n}\left( \lambda_{T}\right) \xi
_{t,n}\left( \lambda_{T}\right) \right) =o_{p}\left( 1\right) ,
\end{align*}
where
\begin{align*}
\upsilon_{t,nT}\left( \lambda_{T}\right) & =\frac{1}{\sqrt{n}}\sum
_{i=1}^{n}\left( \frac{1}{\left( \boldsymbol{\varepsilon}_{i\circ}^{\prime
}\mathbf{M}_{F}\boldsymbol{\varepsilon}_{i\circ}/T\right) ^{1/2}}-1\right)
\varepsilon_{it}\left( \lambda_{T}\right) ,\\
\boldsymbol{\kappa}_{t,n}\left( \lambda_{T}\right) & =\frac{1}{\sqrt{n}
}\sum_{i=1}^{n}\boldsymbol{\gamma}_{i}\sigma_{i}\varepsilon_{it}\left(
\lambda_{T}\right) ,\\
\xi_{t,n}\left( \lambda_{T}\right) & =\frac{1}{\sqrt{n}}\sum_{i=1}
^{n}a_{i,n}\varepsilon_{it}\left( \lambda_{T}\right) ,
a_{i,n}=1-\sigma_{i}\boldsymbol{\varphi}_{n}^{\prime}\boldsymbol{\gamma}_{i}.
\end{align*}
proofStarting with the $g_{2,nT}\left( \lambda_{T}\right) $, note that
\begin{equation}
g_{2,nT}\left( \lambda_{T}\right) =\sqrt{T}\left( \frac{1}{T}\sum_{t=1}
^{T}\upsilon_{t,nT}^{2}\left( \lambda_{T}\right) \right) \leq\sqrt
{T}\left( \sup_{t}\upsilon_{t,nT}^{2}\left( \lambda_{T}\right) \right) .
\end{equation}
Meanwhile, we have
\[
\left\vert \upsilon_{t,nT}\left( \lambda_{T}\right) \right\vert \leq\frac
{1}{\sqrt{n}}\sum_{i=1}^{n}\left\vert \left( \frac{1}{\left(
\boldsymbol{\varepsilon}_{i\circ}^{\prime}\mathbf{M}_{F}
\boldsymbol{\varepsilon}_{i\circ}/T\right) ^{1/2}}-1\right) \right\vert
\left\vert \varepsilon_{it}\left( \lambda_{T}\right) \right\vert ,
\]
and
\begin{equation}
\sup_{t}\left\vert \upsilon_{t,nT}\left( \lambda_{T}\right) \right\vert
\leq\left[ \frac{1}{\sqrt{n}}\sum_{i=1}^{n}\left\vert \left( \frac
{1}{\left( \boldsymbol{\varepsilon}_{i\circ}^{\prime}\mathbf{M}
_{F}\boldsymbol{\varepsilon}_{i\circ}/T\right) ^{1/2}}-1\right) \right\vert
\right] \left( \sup_{i,t}\left\vert \varepsilon_{it}\left( \lambda
_{T}\right) \right\vert \right) .
\end{equation}
Consider the first term of the product in ((ref)) and note since
$\boldsymbol{\varepsilon}_{i\circ}$ is cross-sectionally independent
conditional then
\[
E\left[ \frac{1}{\sqrt{n}}\sum_{i=1}^{n}\left\vert \left( \frac{1}{\left(
\boldsymbol{\varepsilon}_{i\circ}^{\prime}\mathbf{M}_{F}
\boldsymbol{\varepsilon}_{i\circ}/T\right) ^{1/2}}-1\right) \right\vert
\right] ^{2}=\frac{1}{n}\sum_{i=1}^{n}E\left( \frac{1}{\left(
\boldsymbol{\varepsilon}_{i\circ}^{\prime}\mathbf{M}_{F}
\boldsymbol{\varepsilon}_{i\circ}/T\right) ^{1/2}}-1\right) ^{2}.
\]
Also, using ((ref)) we obtain
\begin{align*}
E\left( \frac{1}{\left( \boldsymbol{\varepsilon}_{i\circ}^{\prime}
\mathbf{M}_{F}\boldsymbol{\varepsilon}_{i\circ}/T\right) ^{1/2}}-1\right)
^{2} & =E\left( \frac{1}{\boldsymbol{\varepsilon}_{i\circ}^{\prime
}\mathbf{M}_{F}\boldsymbol{\varepsilon}_{i\circ}/T}\right) -2E\left(
\frac{1}{\left( \boldsymbol{\varepsilon}_{i\circ}^{\prime}\mathbf{M}
_{F}\boldsymbol{\varepsilon}_{i\circ}/T\right) ^{1/2}}\right) +1\\
& =\left[ E\left( \frac{1}{\boldsymbol{\varepsilon}_{i\circ}^{\prime
}\mathbf{M}_{F}\boldsymbol{\varepsilon}_{i\circ}/T}\right) -1\right]
-2\left[ E\left( \frac{1}{\left( \boldsymbol{\varepsilon}_{i\circ}^{\prime
}\mathbf{M}_{F}\boldsymbol{\varepsilon}_{i\circ}/T\right) ^{1/2}}\right)
-1\right] \\
& =O\left( \frac{1}{T}\right) .
\end{align*}
Therefore by Markov inequality, we have
\begin{equation}
\frac{1}{\sqrt{n}}\sum_{i=1}^{n}\left\vert \left( \frac{1}{\left(
\boldsymbol{\varepsilon}_{i\circ}^{\prime}\mathbf{M}_{F}
\boldsymbol{\varepsilon}_{i\circ}/T\right) ^{1/2}}-1\right) \right\vert
=O_{p}\left( \frac{1}{\sqrt{T}}\right) .
\end{equation}
Now consider the second term of of the product in ((ref)) and note
there exist $C_{\varepsilon,1}$, $C_{\varepsilon,2}$ and $r_{\varepsilon}>0$
such that
\[
\Pr\left( \sup_{i,t}\left\vert \varepsilon_{it}\left( \lambda_{T}\right)
\right\vert >a_{\varepsilon}\right) \leq nT\sup_{i,t}\Pr\left( \left\vert
\varepsilon_{it}\left( \lambda_{T}\right) \right\vert >a_{\varepsilon
}\right) \leq nTC_{\varepsilon,1}\exp\left( -C_{\varepsilon,2}\left(
a_{\varepsilon}\right) ^{r_{\varepsilon}}\right) .
\]
Hence, by letting $a_{\varepsilon}=\ominus\left( \ln\left( nT\right)
\right) $, we have
\[
\Pr\left( \sup_{i,t}\left\vert \varepsilon_{it}\left( \lambda_{T}\right)
\right\vert >a_{\varepsilon}\right) \leq C_{\varepsilon,1}\exp\left(
\ln\left( nT\right) -C_{\varepsilon,2}\left( a_{\varepsilon}\right)
^{r_{\varepsilon}}\right) =O\left( 1\right) ,
\]
which further implies $\sup_{i,t}\left\vert \varepsilon_{it}\left(
\lambda_{T}\right) \right\vert =O_{p}\left( \ln\left( nT\right) \right)
$. Invoking this result and ((ref)) in ((ref)) now yields
\begin{equation}
\sup_{t}\left\vert \upsilon_{t,nT}\left( \lambda_{T}\right) \right\vert
=O_{p}\left( \frac{\ln\left( nT\right) }{\sqrt{T}}\right) .
\end{equation}
Then consider ((ref)) and it follows
\begin{equation}
g_{2,nT}\left( \lambda_{T}\right) \leq\sqrt{T}\left( \sup_{t}
\upsilon_{t,nT}^{2}\left( \lambda_{T}\right) \right) \leq\sqrt{T}\left(
\sup_{t}\left\vert \upsilon_{t,nT}\left( \lambda_{T}\right) \right\vert
\right) ^{2}=O_{p}\left( \frac{\left[ \ln\left( nT\right) \right] ^{2}
}{\sqrt{T}}\right) =o_{p}\left( 1\right) ,
\end{equation}
as $n$ and $T$ are of the same order of magnitude. Consider $g_{3,nT}\left(
\lambda_{T}\right) $ and note that the result of Lemma (ref) holds such
that
\begin{equation}
\sqrt{T}\left( \boldsymbol{\varphi}_{n}-\boldsymbol{\varphi}_{nT}\right)
=O_{p}\left( n^{-1/2}\right) +O_{p}\left( T^{-1/2}\right) .
\end{equation}
Furthermore, we have
\begin{align*}
E\left\Vert \boldsymbol{\kappa}_{t,n}\left( \lambda_{T}\right) \right\Vert
^{2} & =\frac{1}{n}\sum_{i=1}^{n}\sum_{j=1}^{n}\boldsymbol{\gamma}
_{i}^{\prime}\boldsymbol{\gamma}_{j}\sigma_{i}\sigma_{j}E\left(
\varepsilon_{it}\left( \lambda_{T}\right) \varepsilon_{jt}\left(
\lambda_{T}\right) \right) \\
& \leq\frac{1}{n}\sum_{i=1}^{n}\sum_{j=1}^{n}\left\Vert \boldsymbol{\gamma
}_{i}\right\Vert \left\Vert \boldsymbol{\gamma}_{j}\right\Vert \left\vert
\sigma_{i}\sigma_{j}\right\vert \left\vert E\left( \varepsilon_{it}\left(
\lambda_{T}\right) \varepsilon_{jt}\left( \lambda_{T}\right) \right)
\right\vert \\
& \leq\left( \sup_{i}\left\Vert \boldsymbol{\gamma}_{i}\right\Vert
^{2}\right) \left( \sup_{i}\sigma_{i}^{2}\right) \left[ \frac{1}{n}
\sum_{i=1}^{n}\sum_{j=1}^{n}\left\vert E\left( \varepsilon_{it}\left(
\lambda_{T}\right) \varepsilon_{jt}\left( \lambda_{T}\right) \right)
\right\vert \right] .
\end{align*}
Given ((ref)) and boundedness of $\boldsymbol{\gamma}_{i}$ and
$\sigma_{i}$, it follows
\begin{equation}
E\left\Vert \boldsymbol{\kappa}_{t,n}\left( \lambda_{T}\right) \right\Vert
^{2}\leq\left( \sup_{i}\left\Vert \boldsymbol{\gamma}_{i}\right\Vert
^{2}\right) \left( \sup_{i}\sigma_{i}^{2}\right) \left[ \frac{1}{n}
\sum_{i=1}^{n}\sum_{j=1}^{n}\left\vert E\left( \varepsilon_{it}\left(
\lambda_{T}\right) \varepsilon_{jt}\left( \lambda_{T}\right) \right)
\right\vert \right] =O\left( 1\right)
\end{equation}
and $\boldsymbol{\kappa}_{t,n}\left( \lambda_{T}\right) $ is therefore
$O_{p}\left( 1\right) $. Also, under Assumption (ref)
$\boldsymbol{\kappa}_{t,n}\left( \lambda_{T}\right) $ is serially
independent and we have $T^{-1}\sum_{t=1}^{T}\boldsymbol{\kappa}_{t,n}\left(
\lambda_{T}\right) \boldsymbol{\kappa}_{t,n}^{\prime}\left( \lambda
_{T}\right) =O_{p}\left( 1\right) $. Using this result together with
((ref)) we then have
\begin{equation}
g_{3,nT}\left( \lambda_{T}\right) =o_{p}(1).
\end{equation}
Next consider $g_{4,nT}\left( \lambda_{T}\right) $ and note that by
Cauchy-Schwarz inequality
\[
\left\Vert \frac{1}{T}\sum_{t=1}^{T}\boldsymbol{\kappa}_{t,n}\left(
\lambda_{T}\right) \upsilon_{t,nT}\left( \lambda_{T}\right) \right\Vert
\leq\left( \frac{1}{T}\sum_{t=1}^{T}\left\Vert \boldsymbol{\kappa}
_{t,n}\left( \lambda_{T}\right) \right\Vert ^{2}\right) ^{1/2}\left(
\frac{1}{T}\sum_{t=1}^{T}\upsilon_{t,nT}^{2}\left( \lambda_{T}\right)
\right) ^{1/2}=O_{p}\left( \frac{\ln\left( nT\right) }{\sqrt{T}}\right)
\]
where the equation holds by ((ref))\ and ((ref)). Then using the
above results it also follows that
\begin{equation}
g_{4,nT}\left( \lambda_{T}\right) =\sqrt{T}\left( \boldsymbol{\varphi}
_{nT}-\boldsymbol{\varphi}_{n}\right) ^{\prime}\left( \frac{1}{T}\sum
_{t=1}^{T}\boldsymbol{\kappa}_{t,n}\left( \lambda_{T}\right) \upsilon
_{t,nT}\left( \lambda_{T}\right) \right) =o_{p}(1).
\end{equation}
Similarly, note $\xi_{t,n}\left( \lambda_{T}\right) =n^{-1/2}\sum_{i=1}
^{n}a_{i,n}\varepsilon_{it}\left( \lambda_{T}\right) $ and by Cauchy-Schwarz
inequality,
\[
\frac{1}{T}\sum_{t=1}^{T}\boldsymbol{\kappa}_{t,n}\left( \lambda_{T}\right)
\xi_{t,n}\left( \lambda_{T}\right) \leq\left( \frac{1}{T}\sum_{t=1}
^{T}\left\Vert \boldsymbol{\kappa}_{t,n}\left( \lambda_{T}\right)
\right\Vert ^{2}\right) ^{1/2}\left( \frac{1}{T}\sum_{t=1}^{T}\xi_{t,n}
^{2}\left( \lambda_{T}\right) \right) ^{1/2}
\]
where given ((ref)), $\sup_{i}a_{i,n}^{2}<C$ and $\sup_{i}\sigma
_{i}^{-2}<C$, it further follows that
\begin{align}
E\left( \xi_{t,n}^{2}\left( \lambda_{T}\right) \right) & =\frac{1}
{n}\sum_{i=1}^{n}\sum_{j=1}^{n}a_{i,n}a_{j,n}E\left( \varepsilon_{it}\left(
\lambda_{T}\right) \varepsilon_{jt}\left( \lambda_{T}\right) \right)
\nonumber\\
& <\left( \sup_{i}a_{i,n}^{2}\right) \frac{1}{n}\sum_{i=1}^{n}\sum
_{j=1}^{n}\left\vert E\left( \varepsilon_{it}\left( \lambda_{T}\right)
\varepsilon_{jt}\left( \lambda_{T}\right) \right) \right\vert <C.
\end{align}
Hence, $T^{-1}\sum_{t=1}^{T}\boldsymbol{\kappa}_{t,n}\left( \lambda
_{T}\right) \xi_{t,n}\left( \lambda_{T}\right) =O_{p}\left( 1\right) $,
and again using ((ref)) it follows that
\begin{equation}
g_{5,nT}\left( \lambda_{T}\right) =\sqrt{T}\left( \boldsymbol{\varphi}
_{nT}-\boldsymbol{\varphi}_{n}\right) ^{\prime}\left( \frac{1}{T}\sum
_{t=1}^{T}\boldsymbol{\kappa}_{t,n}\left( \lambda_{T}\right) \xi
_{t,n}\left( \lambda_{T}\right) \right) =o_{p}(1).
\end{equation}
lemmaSuppose the data are generated by the latent factor
model given by ((ref)) and ((ref)). Further denote
\[
w_{nT}\left( \lambda_{T}\right) =\frac{T^{-1/2}\sum_{t=1}^{T}\xi
_{t,n}\left( \lambda_{T}\right) \upsilon_{t,nT}\left( \lambda_{T}\right)
}{T^{-1}\sum_{t=1}^{T}\xi_{t,n}^{2}\left( \lambda_{T}\right) },
\]
where
\begin{align*}
\xi_{t,n}\left( \lambda_{T}\right) & =\frac{1}{\sqrt{n}}\sum_{i=1}
^{n}a_{i,n}\varepsilon_{it}\left( \lambda_{T}\right) ,a_{i,n}=1-\sigma
_{i}\boldsymbol{\varphi}_{n}^{\prime}\boldsymbol{\gamma}_{i},\\
\upsilon_{t,nT}\left( \lambda_{T}\right) & =\frac{1}{\sqrt{n}}\sum
_{i=1}^{n}\zeta_{it}\varepsilon_{it}\left( \lambda_{T}\right) ,\zeta
_{it}=\frac{1}{\left( T^{-1}\boldsymbol{\varepsilon}_{i\circ}^{\prime
}\mathbf{M}_{F}\boldsymbol{\varepsilon}_{i\circ}\right) ^{1/2}}-1,
\end{align*}
with $\boldsymbol{\varphi}_{n}=n^{-1}\sum_{i=1}^{n}\boldsymbol{\gamma}
_{i}/\sigma_{i}$, $\boldsymbol{\varepsilon}_{i\circ}=\left( \varepsilon
_{i1},\varepsilon_{i2},\ldots,\varepsilon_{iT}\right) ^{\prime}$,
$\mathbf{M}_{F}=\mathbf{I}_{T}-\mathbf{F}\left( \mathbf{F}^{\prime}
\mathbf{F}\right) ^{-1}\mathbf{F}^{\prime}$. Suppose that Assumptions
(ref)-(ref) hold and $(n,T)\rightarrow\infty$, such that
$n/T\rightarrow\kappa,$ and $0<\kappa<\infty$. Then $w_{nT}\left( \lambda
_{T}\right) =o_{p}\left( 1\right) $.
proofConsider first the denominator of $w_{nT}\left( \lambda_{T}\right) $ and
using ((ref)) note that
\[
E\left( \frac{1}{T}\sum_{t=1}^{T}\xi_{t,n}^{2}\left( \lambda_{T}\right)
\right) =\frac{1}{T}\sum_{t=1}^{T}E\left( \xi_{t,n}^{2}\left( \lambda
_{T}\right) \right) <C.
\]
Hence, by Markov inequality it is obvious that the denominator of
$w_{nT}\left( \lambda_{T}\right) $ is $O_{p}\left( 1\right) $.
Consider now the numerator of $w_{nT}\left( \lambda_{T}\right) $, which is
denoted as $r_{nT}\left( \lambda_{T}\right) $. For simplicity, let
\[
\xi_{t,n}=\frac{1}{\sqrt{n}}\sum_{i=1}^{n}a_{i,n}\varepsilon_{it}
,\upsilon_{t,nT}=\frac{1}{\sqrt{n}}\sum_{i=1}^{n}\zeta_{it}\varepsilon_{it},
\]
then
\begin{align*}
r_{nT}\left( \lambda_{T}\right) & =\frac{1}{\sqrt{T}}\sum_{t=1}^{T}
\xi_{t,n}\upsilon_{t,nT}+\frac{1}{\sqrt{T}}\sum_{t=1}^{T}\left[ \xi
_{t,n}\left( \lambda_{T}\right) -\xi_{t,n}\right] \upsilon_{t,nT}\left(
\lambda_{T}\right) \\
& +\frac{1}{\sqrt{T}}\sum_{t=1}^{T}\xi_{t,n}\left( \lambda_{T}\right)
\left[ \upsilon_{t,nT}\left( \lambda_{T}\right) -\upsilon_{t,nT}\right]
-\frac{1}{\sqrt{T}}\sum_{t=1}^{T}\left[ \xi_{t,n}\left( \lambda_{T}\right)
-\xi_{t,n}\right] \left[ \upsilon_{t,nT}\left( \lambda_{T}\right)
-\upsilon_{t,nT}\right] \\
& =r_{1,nT}+\sum_{j=2}^{4}r_{j,nT}\left( \lambda_{T}\right) .
\end{align*}
To bound $r_{1,nT}$, we firstly note $\xi_{t,n}\upsilon_{t,nT}=\frac{1}{n}
\sum_{i=1}^{n}\sum_{j=1}^{n}a_{in}\varepsilon_{it}\zeta_{jt},$ where for
$i\neq j$ $\varepsilon_{it}$ and $\zeta_{jt}$ are distributed independently by
parts (a) and (c) of Assumption (ref). Therefore,
\[
E\left( \xi_{t,n}\upsilon_{t,nT}\right) =\frac{1}{n}\sum_{i=1}^{n}
a_{in}E\left[ \varepsilon_{it}^{2}\left( \frac{1}{\left( T^{-1}
\boldsymbol{\varepsilon}_{i\circ}^{\prime}\mathbf{M}_{F}
\boldsymbol{\varepsilon}_{i\circ}\right) ^{1/2}}-1\right) \right] .
\]
Furthermore, since $a_{in}$ is bounded, and by result ((ref))
\[
E\left[ \varepsilon_{it}^{2}\left( \frac{1}{\left( T^{-1}
\boldsymbol{\varepsilon}_{i\circ}^{\prime}\mathbf{M}_{F}
\boldsymbol{\varepsilon}_{i\circ}\right) ^{1/2}}-1\right) \right] =O\left(
\frac{1}{T}\right) .
\]
Then $E\left( \xi_{t,n}\upsilon_{t,nT}\right) =O\left( T^{-1}\right) $,
and it follows that
\begin{align}
E\left( r_{1,nT}\right) & =\frac{1}{\sqrt{T}}\sum_{t=1}^{T}E\left(
\xi_{t,n}\upsilon_{t,nT}\right) =O\left( \frac{1}{\sqrt{T}}\right)
,\\
Var\left( r_{1,nT}\right) & =E\left( r_{1,nT}^{2}\right) -\left[
E\left( r_{1,nT}\right) \right] ^{2}=\frac{1}{T}\sum_{t=1}^{T}
\sum_{t^{\prime}=1}^{T}E\left( \xi_{t,n}\upsilon_{t,nT}\xi_{t^{\prime}
,n}v_{t^{\prime},nT}\right) +O\left( T^{-1}\right) .
\end{align}
Consider now the first term of $Var\left( r_{1,nT}\right) $, and using
((ref)) we have
\begin{align*}
& E\left( \xi_{t,n}\upsilon_{t,nT}\xi_{t^{\prime},n}v_{t^{\prime},nT}\right)
\\
& \equiv E\left[ \left( \frac{1}{n}\sum_{i=1}^{n}\sum_{j=1}^{n}a_{in}
a_{jn}\varepsilon_{it}\varepsilon_{jt^{\prime}}\right) \left( \frac{1}
{n}\sum_{r=1}^{n}\sum_{s=1}^{n}\zeta_{rt}\zeta_{st^{\prime}}\right) \right]
\\
& =\frac{1}{n^{2}}\sum_{i=1}^{n}\sum_{j=1}^{n}\sum_{r=1}^{n}\sum_{s=1}
^{n}a_{in}a_{jn}E\left\{ \left( \frac{1}{\left( T^{-1}
\boldsymbol{\varepsilon}_{r\circ}^{\prime}\mathbf{M}_{F}
\boldsymbol{\varepsilon}_{r\circ}\right) ^{1/2}}-1\right) \left( \frac
{1}{\left( T^{-1}\boldsymbol{\varepsilon}_{s\circ}^{\prime}\mathbf{M}
_{F}\boldsymbol{\varepsilon}_{s\circ}\right) ^{1/2}}-1\right) \varepsilon
_{it}\varepsilon_{jt^{\prime}}\varepsilon_{rt}\varepsilon_{st^{\prime}
}\right\} .
\end{align*}
Since by part (a) of Assumption (ref), $\varepsilon_{it}^{\prime}s$
are cross-sectionally independent and after some algebra we have
\begin{equation}
\frac{1}{T}\sum_{t=1}^{T}\sum_{t^{\prime}=1}^{T}E\left( \xi_{t,n}
\upsilon_{t,nT}\xi_{t^{\prime},n}v_{t^{\prime},nT}\right) =\sum_{l=1}
^{6}D_{l,nT},
\end{equation}
where
\[
D_{1,nT}=\frac{1}{Tn^{2}}\sum_{t=1}^{T}\sum_{t^{\prime}=1}^{T}\sum_{i=1}
^{n}a_{in}^{2}E\left[ \varepsilon_{it}^{2}\varepsilon_{it^{\prime}}
^{2}\left( \frac{1}{\left( T^{-1}\boldsymbol{\varepsilon}_{i\circ}^{\prime
}\mathbf{M}_{F}\boldsymbol{\varepsilon}_{i\circ}\right) ^{1/2}}-1\right)
^{2}\right] ,
\]
\begin{align*}
D_{2,nT} & =\frac{1}{Tn^{2}}\sum_{t=1}^{T}\sum_{t^{\prime}=1}^{T}\sum
_{i=1}^{n}\sum_{r\neq i}^{n}a_{in}^{2}E\left[ \varepsilon_{it}\varepsilon
_{it^{\prime}}^{2}\left( \frac{1}{\left( T^{-1}\boldsymbol{\varepsilon
}_{i\circ}^{\prime}\mathbf{M}_{F}\boldsymbol{\varepsilon}_{i\circ}\right)
^{1/2}}-1\right) \right] \times\\
& E\left[ \left( \frac{1}{\left( T^{-1}\boldsymbol{\varepsilon}_{r\circ
}^{\prime}\mathbf{M}_{F}\boldsymbol{\varepsilon}_{r\circ}\right) ^{1/2}
}-1\right) \varepsilon_{rt}\right] ,
\end{align*}
\begin{align*}
D_{3,nT} & =\frac{1}{Tn^{2}}\sum_{t=1}^{T}\sum_{t^{\prime}=1}^{T}\sum
_{i=1}^{n}\sum_{s\neq i}^{n}a_{in}^{2}E\left[ \varepsilon_{it}^{2}
\varepsilon_{it^{\prime}}\left( \frac{1}{\left( T^{-1}
\boldsymbol{\varepsilon}_{i\circ}^{\prime}\mathbf{M}_{F}
\boldsymbol{\varepsilon}_{i\circ}\right) ^{1/2}}-1\right) \right] \times\\
& E\left[ \left( \frac{1}{\left( T^{-1}\boldsymbol{\varepsilon}_{s\circ
}^{\prime}\mathbf{M}_{F}\boldsymbol{\varepsilon}_{s\circ}\right) ^{1/2}
}-1\right) \varepsilon_{st^{\prime}}\right] ,
\end{align*}
\[
D_{4,nT}=\frac{1}{Tn^{2}}\sum_{t=1}^{T}\sum_{t^{\prime}=1}^{T}\sum_{i=1}
^{n}\sum_{r\neq i}^{n}a_{in}^{2}E\left( \varepsilon_{it}\varepsilon
_{it^{\prime}}\right) E\left[ \left( \frac{1}{\left( T^{-1}
\boldsymbol{\varepsilon}_{r\circ}^{\prime}\mathbf{M}_{F}
\boldsymbol{\varepsilon}_{r\circ}\right) ^{1/2}}-1\right) ^{2}
\varepsilon_{rt}\varepsilon_{rt^{\prime}}\right] ,
\]
\begin{align*}
D_{5,nT} & =\frac{1}{Tn^{2}}\sum_{t=1}^{T}\sum_{t^{\prime}=1}^{T}\sum
_{i=1}^{n}\sum_{j\neq i}^{n}a_{in}a_{jn}E\left[ \varepsilon_{it}^{2}\left(
\frac{1}{\left( T^{-1}\boldsymbol{\varepsilon}_{i\circ}^{\prime}
\mathbf{M}_{F}\boldsymbol{\varepsilon}_{i\circ}\right) ^{1/2}}-1\right)
\right] \times\\
& E\left[ \varepsilon_{jt^{\prime}}^{2}\left( \frac{1}{\left(
T^{-1}\boldsymbol{\varepsilon}_{j\circ}^{\prime}\mathbf{M}_{F}
\boldsymbol{\varepsilon}_{j\circ}\right) ^{1/2}}-1\right) \right] ,
\end{align*}
and
\begin{align*}
D_{6,nT} & =\frac{1}{Tn^{2}}\sum_{t=1}^{T}\sum_{t^{\prime}=1}^{T}\sum
_{i=1}^{n}\sum_{j\neq i}^{n}a_{in}a_{jn}E\left[ \varepsilon_{it}
\varepsilon_{it^{\prime}}\left( \frac{1}{\left( T^{-1}
\boldsymbol{\varepsilon}_{i\circ}^{\prime}\mathbf{M}_{F}
\boldsymbol{\varepsilon}_{i\circ}\right) ^{1/2}}-1\right) \right] \times\\
& E\left[ \varepsilon_{jt}\varepsilon_{jt^{\prime}}\left( \frac{1}{\left(
T^{-1}\boldsymbol{\varepsilon}_{j\circ}^{\prime}\mathbf{M}_{F}
\boldsymbol{\varepsilon}_{j\circ}\right) ^{1/2}}-1\right) \right] .
\end{align*}
For $D_{1,nT}$ we have
\begin{align*}
D_{1,nT} & =\frac{1}{n^{2}}\sum_{i=1}^{n}a_{in}^{2}\frac{1}{T}\sum_{t=1}
^{T}E\left[ \varepsilon_{it}^{4}\left( \frac{1}{\left( T^{-1}
\boldsymbol{\varepsilon}_{i\circ}^{\prime}\mathbf{M}_{F}
\boldsymbol{\varepsilon}_{i\circ}\right) ^{1/2}}-1\right) ^{2}\right] \\
+ & \frac{1}{n^{2}}\sum_{i=1}^{n}a_{in}^{2}\frac{1}{T}\sum_{t^{\prime}\neq
t}^{T}E\left[ \varepsilon_{it}^{2}\varepsilon_{it^{\prime}}^{2}\left(
\frac{1}{\left( T^{-1}\boldsymbol{\varepsilon}_{i\circ}^{\prime}
\mathbf{M}_{F}\boldsymbol{\varepsilon}_{i\circ}\right) ^{1/2}}-1\right)
^{2}\right] =D_{1,1,nT}+D_{1,2,nT}.
\end{align*}
It is clear that $D_{1,1,nT}=O(n^{-1})$. Also by Cauchy-Schwarz inequality
\[
E\left[ \varepsilon_{it}^{2}\varepsilon_{it^{\prime}}^{2}\left( \frac
{1}{\left( T^{-1}\boldsymbol{\varepsilon}_{i\circ}^{\prime}\mathbf{M}
_{F}\boldsymbol{\varepsilon}_{i\circ}\right) ^{1/2}}-1\right) ^{2}\right]
\leq\left[ E\left( \varepsilon_{it}^{4}\varepsilon_{it^{\prime}}^{4}\right)
\right] ^{1/2}\left[ E\left( \frac{1}{\left( T^{-1}\boldsymbol{\varepsilon
}_{i\circ}^{\prime}\mathbf{M}_{F}\boldsymbol{\varepsilon}_{i\circ}\right)
^{1/2}}-1\right) ^{4}\right] ^{1/2},
\]
where by part (a) of Assumption (ref) we have $E\left(
\varepsilon_{it}^{4}\varepsilon_{it^{\prime}}^{4}\right) =E\left(
\varepsilon_{it}^{4}\right) E\left( \varepsilon_{it^{\prime}}^{4}\right)
=O\left( 1\right) $. Further, in view of ((ref))
\begin{align*}
E\left( \frac{1}{\left( T^{-1}\boldsymbol{\varepsilon}_{i\circ}^{\prime
}\mathbf{M}_{F}\boldsymbol{\varepsilon}_{i\circ}\right) ^{1/2}}-1\right)
^{4} & =E\left[ \frac{1}{\left( T^{-1}\boldsymbol{\varepsilon}_{i\circ
}^{\prime}\mathbf{M}_{F}\boldsymbol{\varepsilon}_{i\circ}\right) ^{2}
}\right] -4E\left[ \frac{1}{\left( T^{-1}\boldsymbol{\varepsilon}_{i\circ
}^{\prime}\mathbf{M}_{F}\boldsymbol{\varepsilon}_{i\circ}\right) ^{3/2}
}\right] +6E\left[ \frac{1}{T^{-1}\boldsymbol{\varepsilon}_{i\circ}^{\prime
}\mathbf{M}_{F}\boldsymbol{\varepsilon}_{i\circ}}\right] \\
& -4E\left[ \frac{1}{\left( T^{-1}\boldsymbol{\varepsilon}_{i\circ}
^{\prime}\mathbf{M}_{F}\boldsymbol{\varepsilon}_{i\circ}\right) ^{1/2}
}\right] +1=O\left( \frac{1}{T}\right) .
\end{align*}
Using this result it follows that $D_{1,2,nT}=$ $O\left( \sqrt{T}
n^{-1}\right) ,$ and overall we have
\begin{equation}
D_{1,nT}=O\left( \sqrt{T}n^{-1}\right) .
\end{equation}
Consider now $D_{2,nT}$ and note
\begin{align*}
\left\vert D_{2,nT}\right\vert & \leq\frac{1}{Tn^{2}}\sum_{t=1}^{T}
\sum_{t^{\prime}=1}^{T}\sum_{i=1}^{n}\sum_{r\neq i}^{n}a_{in}^{2}\left\vert
E\left[ \varepsilon_{it}\varepsilon_{it^{\prime}}^{2}\left( \frac{1}{\left(
T^{-1}\boldsymbol{\varepsilon}_{i\circ}^{\prime}\mathbf{M}_{F}
\boldsymbol{\varepsilon}_{i\circ}\right) ^{1/2}}-1\right) \right]
\right\vert \times\\
& \left\vert E\left[ \left( \frac{1}{\left( T^{-1}\boldsymbol{\varepsilon
}_{r\circ}^{\prime}\mathbf{M}_{F}\boldsymbol{\varepsilon}_{r\circ}\right)
^{1/2}}-1\right) \varepsilon_{rt}\right] \right\vert .
\end{align*}
By ((ref)) we have
\begin{equation}
E\left[ \left( \frac{1}{\left( T^{-1}\boldsymbol{\varepsilon}_{r\circ
}^{\prime}\mathbf{M}_{F}\boldsymbol{\varepsilon}_{r\circ}\right) ^{1/2}
}-1\right) \varepsilon_{rt}\right] =E\left[ \frac{\varepsilon_{rt}}{\left(
T^{-1}\boldsymbol{\varepsilon}_{r\circ}^{\prime}\mathbf{M}_{F}
\boldsymbol{\varepsilon}_{r\circ}\right) ^{1/2}}\right] =O\left( \frac
{1}{T}\right) .
\end{equation}
Under Assumption (ref) and given results ((ref)) and ((ref)),
it follows
\begin{align}
E\left[ \varepsilon_{it}^{2}\left( \frac{1}{\left( T^{-1}
\boldsymbol{\varepsilon}_{i\circ}^{\prime}\mathbf{M}_{F}
\boldsymbol{\varepsilon}_{i\circ}\right) ^{1/2}}-1\right) ^{2}\right] &
=E\left[ \frac{\varepsilon_{it}^{2}}{\left( T^{-1}\boldsymbol{\varepsilon
}_{i\circ}^{\prime}\mathbf{M}_{F}\boldsymbol{\varepsilon}_{i\circ}\right)
}\right] -2E\left[ \frac{\varepsilon_{it}^{2}}{\left( T^{-1}
\boldsymbol{\varepsilon}_{i\circ}^{\prime}\mathbf{M}_{F}
\boldsymbol{\varepsilon}_{i\circ}\right) ^{1/2}}\right] +1\nonumber\\
& =\left[ E\left[ \frac{\varepsilon_{it}^{2}}{\left( T^{-1}
\boldsymbol{\varepsilon}_{i\circ}^{\prime}\mathbf{M}_{F}
\boldsymbol{\varepsilon}_{i\circ}\right) }\right] -1\right] -2\left[
E\left[ \frac{\varepsilon_{it}^{2}}{\left( T^{-1}\boldsymbol{\varepsilon
}_{i\circ}^{\prime}\mathbf{M}_{F}\boldsymbol{\varepsilon}_{i\circ}\right)
^{1/2}}\right] -1\right] \nonumber\\
& =O\left( \frac{1}{T}\right) .
\end{align}
Further, by Cauchy-Schwarz inequality,
\[
\left\vert E\left[ \varepsilon_{it}\varepsilon_{it^{\prime}}^{2}\left(
\frac{1}{\left( T^{-1}\boldsymbol{\varepsilon}_{i\circ}^{\prime}
\mathbf{M}_{F}\boldsymbol{\varepsilon}_{i\circ}\right) ^{1/2}}-1\right)
\right] \right\vert \leq\left[ E\left( \varepsilon_{it^{\prime}}
^{4}\right) \right] ^{1/2}\left[ E\left[ \varepsilon_{it}^{2}\left(
\frac{1}{\left( T^{-1}\boldsymbol{\varepsilon}_{i\circ}^{\prime}
\mathbf{M}_{F}\boldsymbol{\varepsilon}_{i\circ}\right) ^{1/2}}-1\right)
^{2}\right] \right] ^{1/2},
\]
whi