Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.
Factor Network Autoregressions
abstractWe propose a factor network autoregressive (FNAR) model for time series with complex network structures. The coefficients of the model reflect many different types of connections between economic agents (“multilayer network"), which are summarized into a smaller number of network matrices (“network factors") through a novel tensor-based principal component approach.
We provide consistency and asymptotic normality results for the estimation of the factors, their loadings, and the coefficients of the FNAR, as the number of layers, nodes and time points diverges to infinity.
Our approach combines two different dimension-reduction techniques and can be applied to high-dimensional datasets.
Simulation results show the goodness of our estimators in finite samples.
In an empirical application, we use the FNAR to investigate the cross-country interdependence of GDP growth rates based on a variety of international trade and financial linkages. The model provides a rich characterization of macroeconomic network effects as well as good forecasts of GDP growth rates.
\noindentKeywords:\ Networks, factor models, principal components, VAR, tensor decomposition.
\def\spacingset#1
\spacingset{1}
\thispagestyle{empty}
\spacingset{1.7}
\refstepcounter{lettersection}
\oldsection{Introduction}
Network models are key tools for analyzing interconnected systems and are increasingly used in many areas of research. Typical applications involve social networks, where a high number of agents are connected along several dimensions (e.g., zhuetal17), or economic networks, which are usually applied to evaluate spillover effects between agents, geographical regions and economic sectors (e.g., acemogluetal15; dieboldyilmaz14; billioetal12).
From an empirical perspective, the analysis of networks calls for the development of novel statistical techniques that are able to account for complex interactions and handle large datasets, possibly recorded over time.
Several approaches have been proposed to integrate network structures into well-established time series models; the network autoregression (NAR) by zhuetal17 or the global vector autoregressive (GVAR) model by pesaranetal04 are two leading examples.
These models are generally designed to deal with one type of connection at a time between the agents or “nodes" of a network, e.g., import/export flows.
They characterize networks in the form of a single adjacency matrix; i.e., a matrix whose $ij$-th element represents a link between node $i$ and node $j$.
The assumption of a single adjacency matrix, however, appears in general restrictive. In fact, modern economies or social networks are complex systems, in which agents interact through many different channels. For instance, countries are simultaneously linked not only by international trade flows, but also by financial linkages, multinational firms' activities, migration flows, etc.
Accordingly, models of multilayer networks, i.e., networks where nodes are linked through multiple types of connections (kivelaet14), have important applications in many different areas, such as international economics and financial (systemic) risk assessment.
To state the problem we consider in this paper, let $y_t$ be an $N$-dimensional vector of stationary time series, e.g., GDP growth rates for different countries. To model the dynamics of $y_t$ when $N$ is large, we could consider a VAR model in which the dynamic relations between the $N$ elements of $y_t$ are mediated by $m$ time-varying networks, denoted as $W_{k,t}$, $k=1,\ldots,m$, e.g., measuring trade flows of different commodities or various financial linkages. All together these networks, which are $N\times N$ matrices, form the $m$ layers of a multilayer network. We could then consider a multilayer NAR:
align[align omitted — 176 chars of source]
where $b_k$, $k=1,\ldots, m$, $\varrho$, and $a$ are unknown scalars, and $\zeta_{t}:=(\zeta_{1t},\cdots, \zeta_{Nt})'$ is an $N$-dimensional vector of errors.
This model generalizes the NAR by zhuetal17 to the case of multilayer networks. However, in empirical applications, the number $m$ of layers can be very large (e.g., for trade data one may get to hundreds or even thousands of product categories); hence, least-squares estimation of (ref) can become quickly infeasible.
In this paper, we assume that the observed layers $W_{k,t}$, $k=1,\ldots,m$, of the multilayer network have a factor structure, where the common factors are unobserved networks, each represented by an $N\times N$ matrix, capturing the commonality between the layers. We call these latent factors network factors; once retrieved, we can use them in (ref) to model the dynamics of $y_t$ in place of all the $m$ layers, thus reducing the dimensionality of the problem. We call the resulting model a factor network autoregression (FNAR). To estimate the FNAR we first develop a novel, tensor-based, principal component analysis approach, which allows to estimate the network factors. We then use these network factors to estimate an $N$-dimensional FNAR.
Our model exploits two different types of dimension-reduction techniques.
First, by assuming, as in zhuetal17, that the network effects are homogeneous across nodes (or across a finite number of groups of nodes), the number of unknown parameters in our FNAR is at most linear in the number of networks considered, instead of being proportional to $N^2$ as is the case for ordinary vector autoregressions. Second, since the FNAR is based on the extraction of just few network factors common across the $m$ network layers, the number of parameters does not grow with the number of layers $m$, but with the number of network factors which we assume to be finite and independent of $N$ and $m$.
In terms of theory, we first provide sufficient conditions such that the FNAR admits a (weakly) stationary and causal solution. Second, we provide sufficient conditions such that our estimators
of the network factors and of the FNAR coefficients are consistent and asymptotically normal as the sample size $T$, the number of layers $m$, and the number of nodes $N$, diverge. Crucially, the estimator of the FNAR coefficients we propose is consistent even when the FNAR errors admit cross-sectional correlation as well as serial correlation, due to the presence of node specific factors. Finally, the number of network factors can be determined by an eigenvalue ratio approach, by adapting to the present context those recently proposed by hanetal20 and helitrapani2022.
In an empirical application, we use the FNAR model to study the dynamics of GDP growth rates in a multilayer network of $N=24$ countries with $m=25$ different layers, reflecting international trade flows for different categories of goods and services, and a variety of cross-border financial linkages. In a first step we retrieve 6 network factors common across the $m$ layers, which, once identified, have a clear economic meaning. These are then used to construct the regressors which are employed in a FNAR to study and predict GDP growth rates for all countries. A similar idea was explored by chen2022modeling where, however, the tensor data of trade flows is compressed also along the node (country) dimension, thus making the interpretation of the underlying factors less clear.
The paper is organized as follows. Section (ref) gives a summary of the related literature. Section (ref) illustrates the model. Section (ref) presents the estimation approach. Section (ref) introduces the assumptions and the asymptotic properties of the estimators.
Section (ref) evaluates the finite sample performance of our estimators by means of Monte Carlo simulations.
Section (ref) presents the empirical application and Section (ref) concludes. All proofs are in the online Supplementary Material, which also contains additional possible estimators, further simulation results and details on the empirical application.
Notation. The order of a tensor is the number of dimensions, also known as modes. The fibers of a given mode are defined by fixing every index but one. We deal only with order-3 tensors.
The mode-$q$ matricization, or unfolding, of a tensor $ \mathcal{T}$, denoted as $\text{mat}_{(q)} (\mathcal{T})$ or, equivalently, $ \mathcal{T}_{(q)} $, is a matrix having as columns its mode-$q$ fibers.
For a generic order-3 tensor $\mathcal{T}$ of dimensions $d_1 \times d_2 \times d_3$ with generic element $\mathcal T_{ijk}$, the mode-1 matricization is a $d_1 \times (d_2 d_3)$ matrix having as columns the $d_1$-dimensional vectors $\mathcal T_{\cdot jk}$, the mode-2 matricization is a $d_2 \times (d_1 d_3)$ matrix having as columns the $d_2$-dimensional vectors $\mathcal T_{i\cdot k}$, and the mode-3 matricization is a $d_3 \times (d_1 d_2)$ matrix having as columns the $d_3$-dimensional vectors $\mathcal T_{ij\cdot}$.
Finally,
we denote the mode-$q$ multiplication of tensor $\mathcal{T}$ by a matrix $X$ of size $p \times d_q$ as $\mathcal{T} \times_q X$, which is a tensor of size $p$ in the $q$-th dimension and the same size as $\mathcal{T}$ in the other dimensions.
\refstepcounter{lettersection}
\oldsection{Related literature}
Our modelling approach is especially related to the network autoregression (NAR) model by zhuetal17, the community network autoregression (CNAR) model by chenetal2020cnar and the group network autoregression (GNAR) model by zhu2022simultaneous. Differently from these papers, where just one network is considered, our main contribution is to integrate a large multilayer network into a VAR model by representing multilayer networks as a superposition of common network factors.
Our methodology for extracting the network factors is related to the recent statistical literature on factor analysis of tensor time series building on the tensor Tucker decomposition, which assumes the existence of low-rank tensor factors (chenyangzhang22; hanetal2020iter,hanetal20; helitrapani2022; zhangetal2022; chen2024rank). This approach differs from the canonical polyadic (CP) decomposition which always assumes the factors to be vectors chang2023modelling,hanetal2021.
Differently from all the above cited works, since our aim is to reduce the dimensionality of a multilayer network, our setting is a special case of the Tucker decomposition, where the factors are low-rank only along the layers' mode, while they retain all information along the nodes' two modes, i.e., they are $N\times N$ matrices which we can directly interpret as network factors.
Moreover, differently from some of the above cited works, we allow for serial dependence in the idiosyncratic tensor, reflecting the assumption, commonly made in the literature on factor modeling of economic data, that factors account for the cross-sectional (for us, cross-layer) variation of the data (see, e.g., bai2003inferential, in the vector case).
In Appendix F.5 we provide a simulation-based comparison between our estimates of the network factor loadings and those obtained by means of the approaches by chenyangzhang22 where no serial idiosyncratic correlation is allowed for.
Our paper is also related to a growing literature that uses tensor decomposition methods to estimate the coefficients of high-dimensional time series models.
First,
wangetal21 apply tensor decomposition to the order-3 tensor whose slices are the matrices of (unknown) VAR coefficients at different lags, which makes this approach an alternative to ours. However, unlike our approach, wangetal21 do not exploit data on observed networks to estimate the VAR.
We refer to Appendix G.2.4 for an empirical comparison between this approach and our FNAR.
Second, wangetal24 consider an autoregressive model for tensor-valued time series and use a Tucker decomposition to estimate the autoregressive coefficients, while chang2023modelling model matrix-valued time series using a tensor CP decomposition.
Both these last two approaches are more distant from ours, since we do not aim to estimate an AR model for a multilayer network itself, but we aim to use a multilayer network to estimate a vector autoregression in the same spirit as zhuetal17.
Finally, we also contribute to two further strands of literature. First, to the literature on factor and factor-augmented models (e.g., stockandwatson02, baiandng2006) by developing a new framework where factors are matrices rather than vectors, and enter a factor-augmented autoregression by multiplying (weighting) the lagged vector of endogenous variables, rather than being included directly as regressors. Second, to two important streams of the literature on network econometrics, that is: (i) works that investigate the properties of observed networks, such as production networks (see, e.g., acemogluetal12), trading networks in financial and interbank markets (denbeeetal21) and social networks (zhuetal17), and (ii) works concerned with the estimation of network links from the data (e.g., dieboldyilmaz14, billioetal12, barigozzibrownlees19). Our approach combines these two streams of research: on the one hand, we use data on a large number of observed economic networks; on the other hand, we estimate unobserved common network factors driving them.
\refstepcounter{lettersection}
\oldsection{Model}
We assume to observe $m$ networks, each represented by a matrix $W_{k,t}$, $k=1,\ldots, m$, of dimension $N\times N$. The networks have no self-loops, so the diagonal elements of the matrices are zero, and are weighted and {directed}, so in general the matrices have real entries and are not symmetric. In the present high-dimensional setting, it is convenient to assume that the weights are normalized in such a way that the elements in each row of $W_{k,t}$ sum to $N$. This can be done without loss of generality (see Section (ref) for specific comments on this aspect).
We then assume that each observed network can be written as a linear combination of $r$ network factors, $F_{k,t}$, $k=1,\ldots, r$, each of them being an $N\times N$ matrix and with $r\ll m$, plus a network $\mathcal E_{t\cdot\cdot k}$ idiosyncratic to the $k$-th layer, which is an $N\times N$ matrix. Specifically, we assume
equation[equation omitted — 159 chars of source]
where $u_{kh}$, $h=1,\ldots, r$, are the scalar factor loadings for network $k$.
If the factors are sufficiently “pervasive”, we can think of replacing the multilayer NAR in (ref) with a FNAR model, which is given by:
align[align omitted — 200 chars of source]
where $\beta_k$, $k=1,\ldots, r$, $\rho$, and $\alpha$ are unknown scalars, and $\nu_{t}:=(\nu_{1t},\cdots, \nu_{Nt})'$ is the $N$-dimensional vector of FNAR errors.
Notice that the network factors are rescaled by $N$ in agreement with our normalization assumption on the observed networks.
Furthermore, in order to allow also for factors common across the $N$ nodes, we assume each element of the FNAR errors $\nu_{it}$, $i=1,\ldots, N$, to have a factor structure:
align[align omitted — 153 chars of source]
where $G_{kt}$, $k=1,\ldots,q$, are node factors, $\lambda_{ik}$, $k=1,\ldots, q$, are the scalar factor loadings for node $i$, and $\epsilon_{it}$ is the node idiosyncratic component.
We call the term $\sum_{k=1}^{r} \beta_k N^{-1} F_{k,t-1}y_{t-1}$ the {\it network effect}, where the coefficients capture the strength of dynamic network effects between nodes, exerted through different types of relationships, which are summarized by means of few network factors. We call the term $\rho y_{t-1}$ the {\it momentum effect}, which captures the direct dynamic interaction of a node with itself. Last, we call the term $\alpha$ the {\it nodal effect}. This can be generalized to include a random effect by adding a term $Z_{t-1}\delta$, where $\delta$ is a $K$-dimensional parameter vector, common to all nodes, and $Z_{t}$ is a $N\times K$ matrix of node-specific exogenous variables. Because of the factor structure in the FNAR errors, the network nodes are correlated with each other not only through the network relationships, but also through the common factors which characterize the cross-sectional dependence at a global level.
By comparing the FNAR in (ref) with the multilayer NAR in (ref) and a standard VAR(1) for $y_t$ we see that, while the latter requires estimating $N^2+N$ parameters, the multilayer NAR requires estimating $m+2$ parameters, and the FNAR further reduces the number of parameters to $r+2$. Hence, the FNAR model provides two forms of dimension reduction, both along the layer direction, by means of the network factors in (ref) and along the node direction by means of the network mediated interactions in (ref).
\refstepcounter{lettersection}
\oldsection{Estimation}
Estimation of the FNAR model requires two distinct steps. The first one involves estimation of the $r$ network factors.
The second step involves fitting the FNAR equation (ref).
Furthermore, it is necessary to determine the number of network and node factors.
\noindentNetwork factors.
We start by discussing estimation of the network factors. To this aim, it is convenient to collect the $m$ weight matrices or layers of the multinetwork into a
weight tensor of order 3, which we denote as $\mathcal{W}_t$ and has size $N \times N \times m$. In our notation the $m$ network matrices are the frontal slices of the tensor.
The factor network structure in (ref) is then rewritten as:
equation[equation omitted — 116 chars of source]
where $\mathcal{F}_t $ is a $N \times N \times r$ tensor containing, as frontal slices, the $r$ network factors $F_{k,t}$, $k=1,\ldots, r$, each of dimensions $N\times N$, which are common across all layers, $U$ is a $m \times r$ matrix of factor loadings, with entries $u_{kh}$, $k=1,\ldots, m$, $h=1,\ldots, r$, determining how each layer of the original network loads on the network factors, and $\mathcal{E}_t$ is an $N \times N \times m$ tensor containing, as frontal slices, the {idiosyncratic networks}, $\mathcal E_{t\cdot\cdot k}$, $k=1,\ldots, m$. The elements of $\mathcal{E}_t$ are allowed to be
(i) weakly correlated in all three modes, and (ii) autocorrelated over time (see Section (ref) for the specific assumptions).
Finally, from (ref) the mode-3 matricization of $\mathcal W_t$, denoted as $\mathcal{W}_{(3)t}$, is a $m\times N^2$ matrix such that:
equation[equation omitted — 122 chars of source]
where $\mathcal{F}_{(3)t}$ is $r\times N^2$ and $\mathcal{E}_{(3)t}$ is $m\times N^2$. This expression resembles a conventional factor model, with the major difference that each of the $r$ factors is no longer a scalar but a vector of size $N^2$, containing up to $N(N-1)$ non-zero elements (as there are no self-interactions in the network).
To estimate $\mathcal{F}_t$, or, equivalently, $\mathcal{F}_{(3)t}$, first, we compute the sample $m\times m$ outer product of $\mathcal{W}_{(3)t}$:
align[align omitted — 131 chars of source]
Second, letting $\widehat{V}^{\mathcal{W}} $ be the $m \times r$ matrix whose $j$-th column is the normalized eigenvector corresponding to the $j$-th largest eigenvalue, $ \widehat{\mu}_j^{\mathcal{W}}$, of $ \widehat{\Gamma}^{\mathcal{W}}$, and $\widehat{M}^{\mathcal{W}}$ be the $r \times r$ diagonal matrix with $ \widehat{\mu}_j^{\mathcal{W}}$ as its $j$-th diagonal entry,
we estimate the $m\times r$ loadings matrix $U$ as:
equation[equation omitted — 126 chars of source]
Third, we estimate the mode-3 matricization of the network factors $\mathcal{F}_{(3)t}$ as the principal components (PCs) of $\mathcal{W}_{(3)t}$, i.e., by linear projection of $\mathcal W_{(3)t}$ onto $\widehat U$:
$
\widehat{\mathcal{F}}_{(3)t} := (\widehat{U}'\widehat{U})^{-1}\widehat{U}'\mathcal{W}_{(3)t}= N ( \widehat{M}^{\mathcal{W}} )^{-1/2}\widehat{V}^{\mathcal{W}'} \mathcal W_{(3)t},
$
which is an $r\times N^2$ matrix; once folded into an order-3 tensor, it gives the estimated $N\times N\times r$ tensor of network factors:
equation[equation omitted — 243 chars of source]
having as layers $\widehat F_{k,t}:=\text{mat}_{(1)}( \widehat{\mathcal{F}}_{\cdot ,\cdot ,k, t} )$, $k=1,\ldots, r$, which are the $N\times N$ matrices of estimated network factors.
Like in ordinary principal component analysis (PCA), the key intuition is that by exploiting the cross-sectional variation we can estimate the space spanned by the factors. In this case, the relevant cross-sectional dimension is the dimension $m$ of the layers in a network; i.e., the third dimension of the weight tensor $\mathcal{W}_{t}$. The rescaling by $N$ in estimating the loadings reflects the fact that no dimension reduction is applied to the first and second modes of $\mathcal W_t$.
Notice that the estimated loadings and factor tensor are such that they satisfy the identifying conditions:
$m^{-1}\widehat{U}'\widehat U=m^{-1}N^{-2} \widehat{M}^{\mathcal{W}}$,
which is a diagonal matrix by construction, and
equation[equation omitted — 293 chars of source]
so that the $r$ rows of $\widehat{\mathcal{F}}_{(3)t}$ are the classical normalized PCs of the $m$ rows of ${\mathcal{W}}_{(3)t}$.
There are two main differences between our approach and the one proposed by chenyangzhang22 for estimating common tensor factors from tensor times series admitting a Tucker decomposition. First, we estimate factors using only contemporaneous sample second moments, thus allowing the idiosyncratic tensor to be also autocorrelated.
The second difference is that we extract factors along a single dimension of the tensor, namely the dimension of the network layers.
This implies that the extracted factors still have a network interpretation; i.e., they are common network factors. Indeed, by construction, the estimated factor matrices $\widehat F_{k,t} $, $k=1,\dots, r$, have zeros along the main diagonal, as the observed network matrices, and have a scale fixed by means of (ref).
To preserve the network properties of the factors, we do not standardize nor demean the elements in $\mathcal{W}_{t} $ along the time dimension. Indeed, our variables are all expressed in the same unit of measurement and standardization would eliminate the fundamental interpretation of $\mathcal{W}_{t}$ as a network, as centering would affect the zero diagonal entries. The same strategy is adopted by chenyangzhang22.
Our estimators in (ref) and (ref) are defined consistently with the assumption that the observed networks have rows summing to $N$. This implies that the estimated network factors in (ref) have variance growing with $N$, as shown in (ref), and they must be rescaled before being used in the FNAR defined in (ref) to ensure that the scale of the estimated network coefficients $\beta_k$ does not depend on $N$.
Clearly, we could equivalently work with row-normalized observed networks and then no rescaling by $N$ would be needed anywhere, although this would imply that, as the number of nodes $N$ grows, the entries of $\mathcal W_t$ would have to become smaller and smaller.
\noindentFNAR coefficients.
To describe the estimation of the FNAR, it is convenient to introduce some further notation.
Hereafter, with reference to (ref),
let $y := (y_1,\cdots ,y_T)' = (y_1,\cdots ,y_N)$ be the $T\times N$ matrix of observed data (with $y_t$, $t=1,\ldots, T$, being $N$-dimensional, and $y_i$, $i=1,\ldots, N$, being $T$-dimensional), and define also
the $r$-dimensional vector $\beta:=(\beta_1,\cdots,\beta_r)'$ and the $(r+2)$-dimensional vector $\theta := (\beta', \rho, \alpha )'$ of the FNAR coefficients.
Moreover, with reference to the node factor structure in (ref), we let
$\epsilon:=(\epsilon_1,\cdots ,\epsilon_T)'=(\epsilon_1,\cdots ,\epsilon_N)$ be the $T\times N$ matrix of node idiosyncratic components (with $\epsilon_t$, $t=1,\ldots, T$, being $N$-dimensional, and $\epsilon_i$, $i=1,\ldots, N$, being $T$-dimensional), and we let $G_t:=(G_{1t},\cdots, G_{qt})'$ and $\Lambda_i=(\lambda_{i1},\cdots,\lambda_{iq})'$, so that
$G:=(G_1,\cdots ,G_T)'$ is the $T\times q$ matrix of node factors, and $\Lambda:=(\Lambda_1,\cdots ,\Lambda_N)'$ is the $N\times q$ matrix of loadings.
Finally, define the order-3 tensor $\mathcal X$ of dimensions $T\times N\times (r+2)$ having in each of the first $r$ layers one of the $r$ matrices
$
\mathsf F_k:=\left(N^{-1} F_{k,0} y_{0},\cdots, N^{-1} F_{k,T-1} y_{T-1}\right)^\prime
$, $k=1,\ldots, r$, each of size $T\times N$,
in layer $(r+1)$ the $T\times N$ matrix $y_{(-1)}:=(y_0,\cdots,y_{T-1})^\prime$, and in layer $(r+2)$ a $T\times N$ matrix of ones, denoted as $1_{T\times N}$. It follows that $\mathcal X_{(1)}:=\text{mat}_{(1)}(\mathcal X)$ is a $T\times N(r+2)$ matrix having as $t$-th row $\text{vec}(X_t)^\prime$ and $\mathcal X_{(2)}:=\text{mat}_{(2)}(\mathcal X)$ is a $N\times T(r+2)$ matrix having as $i$-th row $\text{vec}(X_i)^\prime$, where we define
$X_t := (\mathsf F_{1,t-1,\cdot}',\cdots, \mathsf F_{r,t-1,\cdot}', y_{t-1}, \iota_N)$
and $X_i:=(\mathsf F_{1,\cdot,i},\cdots, \mathsf F_{r,\cdot,i}, y_{i}, \iota_T)$, with $\iota_N$ and $\iota_T$ being $N$- and $T$-dimensional vectors of ones, respectively, with $\mathsf F_{k,t-1,\cdot}$ and $\mathsf F_{k,\cdot,i}$ being the $t$th row and $i$th column of $\mathsf F_{k}$, respectively.
According to the above notation, the FNAR in (ref), jointly with the node factor structure in its errors given in (ref), can be equivalently rewritten as:
equation[equation omitted — 187 chars of source]
which, by stacking all $T$ or $N$ equations, can also be written as:
align[align omitted — 232 chars of source]
Since both the network factors (contained in $\mathcal X$) and the node factors are unobserved, estimation of (ref) is infeasible.
To make estimation feasible, we start by substituting $\mathcal X$ with $\widehat{\mathcal X}$, which, in turn, is obtained by replacing the network factors $F_{k,t}$ with their estimates $\widehat F_{k,t}$, $k=1,\ldots, r$, given by (ref).
An infeasible OLS estimator of $\theta$ would then be obtained by applying the Frisch-Waugh theorem to partial out the effect of either $G$ or $\Lambda$:
align[align omitted — 415 chars of source]
where $M_G:= I_T-G(G'G)^{-1}G'$ and $M_\Lambda:=I_N-\Lambda(\Lambda'\Lambda)^{-1}\Lambda'$ are the $T\times T$ and $N\times N$ linear projectors onto the spaces orthogonal to the factors and loadings spaces, respectively.
Likewise, given $\theta$, the estimators of $\Lambda$ and $G$ would be the usual PC estimators applied to the FNAR residuals $\widehat{\nu} := y- \text{mat}_{(1)}(\widehat{\mathcal X}\times_3\theta')$.
Following this reasoning, two equivalent estimators of $\theta$ are given by:
equation[equation omitted — 527 chars of source]
where, letting $\widehat \nu^\dag:=y-\text{mat}_{(1)}(\widehat{\mathcal X}\times_3\widehat{\theta}^{\dag'})$ and $\widehat \nu^*:=y-\text{mat}_{(1)}(\widehat{\mathcal X}\times_3\widehat{\theta}^{*'})$, we defined (see also Appendix A)
align[align omitted — 186 chars of source]
with $\widehat V^{\widehat\nu^\dag}$ being the $T\times q$ matrix of normalized eigenvectors of $N^{-1}\widehat \nu^{\dag}\widehat \nu^{\dag'}$, and
$\widehat M^{\widehat\nu^*}$ being the $q\times q$ diagonal matrix of eigenvalues of $T^{-1}\widehat \nu^{*'}\widehat \nu^{*}$ with corresponding normalized eigenvectors given by the columns of the $N\times q$ matrix $\widehat V^{\widehat\nu^*}$.
By iterating between (ref) and (ref), we solve such minimization and compute the two final estimators of $\theta$. As explained at the end of Section (ref), the two are asymptotically equivalent and they might differ numerically just because of the iterative approaches used to compute them.
Finally, notice that, obviously, once we have computed $\widehat{\theta}^\dag$ we can also estimate the loadings $\Lambda$ by linear projection of $\widehat{G}^\dag$ onto $\widehat \nu^{\dag}$, and once we have computed $\widehat{\theta}^*$ we can also estimate the node factors $G$ by linear projection of $\widehat{\Lambda}^*$ onto $\widehat \nu^{*}$. These estimates, however, are not needed for estimating $\theta$.
The outlined estimators are robust to the presence of autocorrelated node factors $G_t$, and, as a consequence, the FNAR errors $\nu_t$ are allowed to be both cross-sectionally and serially correlated. The adopted algorithm is similar to the one proposed by chenetal2020cnar, who in turn adapted the approach by bai2009panel to the NAR setting. Under the assumption of no autocorrelation in the node factors, the OLS and GLS estimators are also valid estimators of $\theta$ (see Appendices B and C, respectively).
\noindentNumber of factors.
Letting $\widehat{\mu}_j^{\mathcal W}$, $j=1,\ldots, m$, be the $j$-th largest eigenvalue of $\widehat{\Gamma}^{\mathcal W}$ defined in (ref), we estimate the number of network factors, $r$,
by means of the eigenvalue ratio criterion:
equation[equation omitted — 155 chars of source]
where $r_{\max}$ is a predefined maximum number of network factors such that $r_{\max}<\min\{m,T,N^2\}$. This is the criterion proposed by helitrapani2022, which generalizes the approach proposed by hanetal20 to the tensor factor model with autocorrelated idiosyncratic components.
Likewise, letting $\widehat{\mu}_j^{\widehat \nu}$, $j=1,\ldots, N$, be the $j$-th largest eigenvalue of the sample covariance matrix of the FNAR residuals,
the number of node factors, $q$, is determined by means of the criterion:
equation[equation omitted — 159 chars of source]
where $q_{\max}$ is a predefined maximum number of node factors such that $q_{\max}<\min\{T,N\}$; see ahn2013eigenvalue.
In practice, since the FNAR residuals $\widehat{\nu}_t$ depend on the chosen value of $q$, we can adopt an iterative procedure, which starts by over-estimating $q$ in the first stage, identical to the one described by bai2009panel.
Alternative approaches to estimating $r$ and $q$ are possible; see, e.g., the information criterion by baing02, or the randomized test by trapani2018randomized which is based on the divergence rates of the eigenvalues. Importantly, the latter could also be used to test for no factors.
\refstepcounter{lettersection}
\oldsection{Theory}
Assumptions
In the following, we allow for the number of layers $m$ to grow to infinity. Therefore, all our assumptions are stated for the infinite sequence of $N\times N$ networks
$W_{i,t}:=\text{mat}_{(1)}( \mathcal{W}_{\cdot ,\cdot ,i, t} )$ with $i\in\mathbb N$. Equivalently, we could state the assumptions for $i=1,\ldots, m$ with $m\in\mathbb N$. Moreover, all assumptions are stated contemplating the possibility that also the number of nodes $N$ and the sample size $T$ grow to infinity.
assumption[Common component of multilayer network]\
\begin{enumerate}[label=(\roman*)]
• $\lim_{m \rightarrow \infty} \norm{m^{-1} U'U - \Sigma_U} = 0$
where $\Sigma_U$ is $r \times r$ finite and positive definite, and, for all $k \in \mathbb{N}$, $\norm{U_{k\cdot}} \leq M_U$
for some finite $M_U$
independent of $k$.
• For all $t \in \mathbb{Z}$ and all $N \in \mathbb{N}$,
$ \Gamma^{\mathcal F}:=\mathbb{E} [ \mathcal{F}_{(3)t} \mathcal{F}_{(3) t}' ] $ is $r \times r$ positive definite, and such that
$\norm{N^{-2} \Gamma^{\mathcal F}}\le M_{\mathcal F}$ for some finite $M_{\mathcal F}$
independent of $N$.
• For all $N\in\mathbb N$ and all $t \in \mathbb{Z}$,
$\mathbb E\left[\norm{N^{-1}\mathcal{F}_{(3)t}}^4\right]\le K_{\mathcal F}$ for some finite $K_{\mathcal{F}}$ independent of $t$ and $N$.
• For all $i,j=1,\dots,r$ and all $T, N\in\mathbb N$,
$
\mathbb{E} \left[
\abs{ \frac{1}{\sqrt{T} N } \sum_{t=1}^{T}
\left\{
\mathcal{F}_{(3)t i \cdot} \mathcal{F}_{(3)t j \cdot}^\prime
-
\mathbb{E} \left[ \mathcal{F}_{(3)t i \cdot} \mathcal{F}_{(3)t j \cdot}^\prime \right]
\right\}
}^2
\right]
\leq C_{\mathcal{F}}
$
for some finite $C_{\mathcal{F}}$ independent of $i,j,T$, and $N$.
• There exists an integer $\overline{M}$ such that for all $m>\overline{M}$, $r$ is a finite positive integer, independent of $m$.
• For all $h\in\mathbb N$, all $s\in\mathbb Z$, and all $l=1,\ldots,r$,
$\mathbb E\left[\left\vert
\mathcal{F}_{(3)slh}\right\vert\right] \le C_{\mathcal F}^\prime$, for some finite $C_{\mathcal{F}}^\prime$ independent of $h,s,$ and $l$.
\end{enumerate}
Assumptions (ref)(ref) and (ref)(ref) imply that we consider only pervasive, or strong, factors. In other words, the network factors are loaded by most or all the network layers. Notice that the factor tensor has dimension $N\times N\times r$, hence the rescaling introduced in Assumption (ref)(ref), which accounts for its first two modes having diverging dimensions.
Furthermore, under Assumption (ref)(ref), the 2nd order moments of the process $\{N^{-1}\text{vec}(\mathcal F_{(3)t}), t\in\mathbb Z\}$ are finite and independent of time, for all $N\in\mathbb N$.
Assumptions (ref)(ref) and (ref)(ref) imply that, given the factors, we can consistently estimate $\Gamma^{\mathcal F} := \mathbb{E} [ \mathcal{F}_{(3)t} \mathcal{F}_{(3)t}' ]$, as proved in Lemma E.5(i).
Therefore, given the loadings, we can also consistently estimate $\Gamma^\chi:= U\mathbb E [ \mathcal{F}_{(3)t} \mathcal{F}_{(3)t}' ] U' $.
Assumption (ref)(ref) simply states that the number of factors is finite for all $m,N\in\mathbb N$ and that in order to find such factors we need $m$ to be large enough. Finally, Assumption (ref)(ref) is very mild as it simply requires each element of the network factors to have finite first moment.
assumption[Idiosyncratic component of multilayer network]\
\begin{enumerate}[label=(\roman*)]
• For all $m, N \in \mathbb{N}$ and all $ t \in \mathbb{Z}$, $\mathbb{E}\left[ \mathcal{E}_{(3)t} \right]= 0_{m \times N^2}$ and $\Gamma^{\mathcal{E}} := \mathbb{E} [\mathcal{E}_{(3)t} \mathcal{E}_{(3)t}']$ is $m \times m$ positive definite.
• For all $N \in \mathbb{N}$, all $t,s \in \mathbb{Z}$, and all $i,j=1,\ldots N^2$,
$
{N^{-2}}
\sum_{h=1}^{N^2} \sum_{k=1}^{N^2}
\abs{
\mathbb{E} \left[ \mathcal{E}_{(3) t i h} \mathcal{E}_{(3) s j k} \right]
}
\leq \rho_{\mathcal{E}}^{\abs{t-s}} M_{ij}
$
and, for all $N \in \mathbb{N}$, all $t,s \in \mathbb{Z}$, and all $i,j,k=1, \dots, N^2$,
$
{N^{-2}}
\sum_{h=1}^{N^2}
\abs{
\mathbb{E} \left[ \mathcal{E}_{(3) t i h} \mathcal{E}_{(3) s j k} \right]
}
\leq \rho_{\mathcal{E}}^{\abs{t-s}} M_{ij}
$
for
some finite $\rho_{\mathcal{E}}$ and $M_{ij}$ independent of $t,s,k$ and $N$ such that
$0 \leq \rho_{\mathcal{E}} <1 $, $\sum_{i=1, i\ne j}^m M_{ij} \leq M_{\mathcal{E}}$ and $\sum_{j=1,j\ne i}^m M_{ij} \leq M_{\mathcal{E}}$, for some finite $ M_{\mathcal{E}}$ independent of $i,j$ and $m$.
• For all $i,j\in\mathbb N$ and all $t \in \mathbb{Z}$,
$\mathbb E[\vert \mathcal{E}_{(3)t ij}\vert^4]\le K_{\mathcal E}$ for some finite $K_{\mathcal{E}}$ independent of $i,j,t$.
• For all $m,T, N \in \mathbb{N}$ and all $j=1,\ldots, N^2$ and all $s=1,\ldots, T$,
\begin{equation*}
\mathbb{E} \left[ \abs{ \frac{1}{\sqrt{mT} N^{2} } \sum_{i=1}^{m} \sum_{t=1}^{T} \sum_{h=1}^{N^2} \sum_{k=1}^{N^2}
\left\{
\mathcal{E}_{(3)t i h} \mathcal{E}_{(3)t j k}
-
\mathbb{E} \left[ \mathcal{E}_{(3)t i h} \mathcal{E}_{(3)t j k} \right]
\right\} }^2 \right]
\leq C_{\mathcal{E}}
\end{equation*}
and
\begin{equation*}
\mathbb{E} \left[ \abs{ \frac{1}{\sqrt{mT} N^{2} } \sum_{i=1}^{m} \sum_{t=1}^{T} \sum_{h=1}^{N^2} \sum_{k=1}^{N^2}
\left\{
\mathcal{E}_{(3)t i h} \mathcal{E}_{(3)s i k}
-
\mathbb{E} \left[ \mathcal{E}_{(3)t i h} \mathcal{E}_{(3)s i k} \right]
\right\} }^2 \right]
\leq C_{\mathcal{E}}
\end{equation*}
for some finite $C_{\mathcal{E}}$ independent of $j, s, m, T, N$.
• For all $m,T, N \in \mathbb{N}$ and all $j=1,\ldots, N^2$ and all $s=1,\ldots, T$,
\begin{equation*}
\mathbb{E} \left[ \abs{ \frac{1}{\sqrt{mT} N} \sum_{i=1}^{m} \sum_{t=1}^{T} \sum_{h=1}^{N^2}
\left\{
\mathcal{E}_{(3)t i h} \mathcal{E}_{(3)t j h}
-
\mathbb{E} \left[ \mathcal{E}_{(3)t i h} \mathcal{E}_{(3)t j h} \right]
\right\} }^2 \right]
\leq C_{\mathcal{E}}
\end{equation*}
and
\begin{equation*}
\mathbb{E} \left[ \abs{ \frac{1}{\sqrt{mT} N} \sum_{i=1}^{m} \sum_{t=1}^{T} \sum_{h=1}^{N^2}
\left\{
\mathcal{E}_{(3)t i h} \mathcal{E}_{(3)s i h}
-
\mathbb{E} \left[ \mathcal{E}_{(3)t i h } \mathcal{E}_{(3)s i h} \right]
\right\} }^2 \right]
\leq C_{\mathcal{E}}
\end{equation*}
for some finite $C_{\mathcal{E}}$ independent of $j, s, m, T, N$.
\end{enumerate}
This assumption controls the serial and cross-sectional dependence of the entries of the idiosyncratic tensor. In particular, according to Assumption (ref)(ref), the covariance across time and layers is controlled in a standard way; hence, the classical summability conditions hold (see Lemma E.1 and bai2003inferential). The covariances across the elements of the $N^2$-dimensional vector $\mathcal E_{(3)t j\cdot}$ are of order $N^4$ for any given $t=1,\ldots, T$ and $j=1,\ldots, m$, and we require to scale their sum by $N^2$, thus assuming a standard summability condition.
Assumptions (ref)(ref), (ref)(ref), and (ref)(ref) imply that,
given the idiosyncratic tensor, we can consistently estimate $m^{-1}N^{-2}\Gamma^{\mathcal E}$, for any $m,N\in\mathbb N$, as proved in Lemma E.5(ii).
assumption[Moment conditions - part 1]
For all $i,j\in\mathbb N$, all $k=1,\ldots,r$, and all $t\in\mathbb Z$, $\mathbb E[\mathcal{F}_{(3)tkj}\mathcal{E}_{(3)tij}]=0$,
and, for all $m,N,T\in\mathbb N$ and all $t=1,\ldots, T$,
\begin{align}
&\mathbb{E} \left[\frac 1{m N^{2}} \sum_{i=1}^{m} \norm{ \frac 1{\sqrt T} \sum_{t=1}^{T} \mathcal{F}_{(3)t} \mathcal{E}_{(3)t i\cdot}^\prime }_F^2 \right] \leq C_{\mathcal{F}\mathcal{E}},\nonumber\\
&\mathbb{E}\left[\frac 1{ N^{4} }\left\Vert\frac 1{\sqrt{mT}}\sum_{i=1}^{m} \sum_{s=1}^{T}\mathcal{F}_{(3)s}\left\{\mathcal{E}_{(3)s i \cdot}^\prime \mathcal{E}_{(3)t i \cdot}-\mathbb E[\mathcal{E}_{(3)s i \cdot}^\prime \mathcal{E}_{(3)t i \cdot}]\right\} \right\Vert_F^2\right]\le C_{\mathcal{F}\mathcal{E}}^\prime,\nonumber
\end{align}
for some finite $C_{\mathcal{F}\mathcal{E}}$ and $C_{\mathcal{F}\mathcal{E}}^\prime$ independent of $t$, $m, N$, and $T$.
Uncorrelatedness of the processes $\{\mathcal F_{(3)t}\}$ and $\{\mathcal E_{(3)t}\}$ is a natural assumption, while the moment conditions controlling higher-order dependence are standard and are direct extensions of what typically assumed in the vector case bai2003inferential.
Define the $m\times m$ matrix $ \Gamma^{\mathcal W}:=\mathbb E[\mathcal W_{(3)t}\mathcal W_{(3)t}']$ and let $\Gamma^\chi$ be as previously defined. Denote the eigenvalues of $\Gamma^\chi$ as $\mu_j^\chi$, $j=1,\ldots, r$, in decreasing order. Then, as established by Lemmas E.1(i) and E.1(iv), we have that $\mu_j^\chi\asymp mN^2$ and $\norm{N^{-2}\Gamma^{\mathcal{E}}}$ is finite for all $N\in\mathbb N$.
As a consequence, by Assumption (ref) and Weyl's inequality, the matrix $\Gamma^{\mathcal W}=\Gamma^\chi+\Gamma^{\mathcal E}$ is characterized by an eigengap between the $r$-th and the $(r+1)$-th largest eigenvalues which widens as $m\to\infty$.
This property allows us to identify the number of network factors $r$ and it is the rationale for the eigenvalue ratio criterion for estimating $r$ defined in (ref). This also means that the network factor model in (ref) is always identified as long as $m\to\infty$.
In general, the network factors and their loadings are not separately identified unless we impose further restrictions. To this end we make the following assumption.
assumption[Identification of network factors and loadings]\
\begin{enumerate}[label=(\roman*)]
• For all $m \in \mathbb{N}$, $m^{-1}U'U$ is diagonal with distinct entries.
• For all $N, T \in \mathbb{N}$, ${N^{-2}T^{-1}}\sum_{t=1}^{T} \mathcal{F}_{(3)t} \mathcal{F}_{(3)t}' = I_r$.
\end{enumerate}
Under Assumption (ref) the columns of the loadings matrix $U$ and the layers of the tensor factor $\mathcal F_t$ are identified up to a sign multiplication.
This identification scheme is a classical one adopted for example by bai2009panel in the vector factor model case.
assumption[CLTs]\
\begin{enumerate}[label=(\roman*)]
• For any given $i=1,\ldots, m$,
as $N,T\to\infty$,
$
\frac{1}{N\sqrt {T}} \sum_{t=1}^{T} \mathcal{F}_{(3)t} \mathcal{E}_{(3)ti\cdot}' \overset{d}{\to}\mathcal N(0_r,\Phi_i),
$
where
$
\Phi_i:=\lim_{N,T\to\infty} \mathbb E\left[\left (\frac 1{N\sqrt {T }}\sum_{t=1}^{T} \mathcal{F}_{(3)t} \mathcal{E}_{(3)t i \cdot}' \right)
\left(\frac 1{N\sqrt { T }}\sum_{t=1}^{T} \mathcal{F}_{(3)t} \mathcal{E}_{(3)t i \cdot}' \right)'
\right].
$
• For any given $t=1,\ldots, T$ and $j=1,\ldots, N^2$, as $m\to\infty$,
$
\frac{ 1}{\sqrt{m}}\sum_{i=1}^m u_i \mathcal E_{(3)tij}\overset{d}{\to}\mathcal N(0_{r},\Pi_{tj}),
$
where
$\Pi_{tj}:=\lim_{m\to\infty}\mathbb E\left[\left(
\frac{1}{\sqrt{m}}\sum_{i=1}^m u_i \mathcal {E}_{(3)t i j}
\right)
\left(
\frac{1}{\sqrt{m}}\sum_{i=1}^m u_i \mathcal {E}_{(3)t i j}
\right)'
\right]$
and $u_i'$ is the $i$-th row of $U$.
\end{enumerate}
Assumption (ref)(ref) is standard in the vector factor model case; i.e., when $N=1$, where it is satisfied for example by strong-mixing processes bai2003inferential.
In Assumption
(ref)(ref)
we directly assume a cross-sectional CLT, which is a standard approach in the vector factor model case bai2003inferential.
The FNAR errors $\nu_t$ follow a factor model given in (ref), characterized by the following assumption.
assumption[FNAR errors]\
\begin{enumerate}[label=(\roman*)]
• $
\lim_{N \to \infty}
\norm{ N^{-1} \Lambda' \Lambda - \Sigma_{\Lambda} } = 0
$
where $\Sigma_{\Lambda}$ is $q \times q$ finite and positive definite, and, for all $i \in \mathbb{N}$,
$\norm{\Lambda_{i\cdot}} \leq M_{\Lambda}$
for some finite $M_{\Lambda}$ independent of $i$.
• For all $t\in\mathbb Z$,
$\mathbb{E}[G_t] = 0_q $,
$\Gamma^{G} := \mathbb{E}\left[G_t G_t' \right]$ is $q \times q$ positive definite,
and such that
$\norm{\Gamma^{G}} \leq M_G$
for some
finite $ M_G$ independent of $t$.
• For all $t\in\mathbb Z$, $\mathbb E[\norm{G_t}^4]\le K_G$ for some finite $K_G$ independent of $t$.
• For all $i,j = 1, \dots, q$ and all $T \in \mathbb{N}$,
$
\mathbb{E} \left[ \abs{ \frac{1}{\sqrt{T}} \sum_{t=1}^{T}
\left\{
G_{it} G_{jt}
-
\mathbb{E} \left[ G_{it} G_{jt} \right]
\right\} }^2 \right]
\leq C_{G}
$
for some finite $ C_{G}$ independent of $i,j$, and $T$.
• There exists an integer $\underline{N}$ such that for all $N > \underline{N}$, $q$ is a finite positive integer, independent of $N$.
• For all $i \in \mathbb{N}$ and all $t\in\mathbb Z$,
$\mathbb{E}[\epsilon_{it}] = 0$,
$\mathbb E[\epsilon_{it}^2]= \sigma_i^2$ such that $\sigma_i^2 \ge \underline M_{\epsilon}$ and $\sigma_i^2\le \overline M_{\epsilon}$ for some finite $\underline M_{\epsilon}$ and
$\overline M_{\epsilon}$ independent of $i$ and $t$.
• For all $i,j \in \mathbb{N}$ and all $t,s \in \mathbb{Z}$, $\mathbb{E} \left[ \epsilon_{ it} \epsilon_{js} \right]=0$ if $i\ne j$ and
$\mathbb{E} \left[ \epsilon_{ it} \epsilon_{is} \right]=0$ if $t\ne s$.
• For all $i \in \mathbb N$ and all $t \in \mathbb{Z}$,
$\mathbb{E}[\epsilon_{it}^4] \leq K_{\epsilon}$
for some finite $K_{\epsilon}$ independent of $i$ and $t$.
• For all $N, T \in \mathbb{N}$,
$
\mathbb{E} \left[ \abs{ \frac{1}{\sqrt{NT}} \sum_{i=1}^{N} \sum_{t=1}^{T}
\left\{
\epsilon_{it}^2
-
\mathbb{E} \left[ \epsilon_{it}^2 \right]
\right\} }^2 \right]
\leq C_{\epsilon}
$
for some finite $C_{\epsilon}$ independent of $N$ and $T$.
\end{enumerate}
This assumption is similar to the usual set of assumptions for the vector factor model bai2003inferential. The comments to these assumptions are analogous to those previously made for the network factors and are therefore omitted.
Concerning the idiosyncratic components, we follow chenetal2020cnar and assume zero correlations both in time and across units. Given that we are considering a factor structure for the FNAR errors, this assumption is not very restrictive, as most of the correlations are likely to be already captured by the lagged terms in the FNAR and by the common factors $G_t$. Nevertheless, it is possible to develop the following asymptotic theory by allowing for the usual kind of weak cross- and autocorrelations between the components of $\epsilon_t$.
assumption[Independence of network and node factors]
For all $N\in\mathbb N$, the processes $\{\mathcal F_t,\,t\in\mathbb Z\}$, $\{\epsilon_{t},\, t \in \mathbb{Z}\}$,
and $\{G_{t},\, t \in \mathbb{Z}\}$ are
mutually independent.
Assumption (ref) is taken from
bai2009panel
and is made just to simplify the proof.
In principle it could be relaxed to allow for weak dependence between $\{G_t\}$ and $\{\epsilon_{t}\}$
by means of a condition similar to those required in Assumption (ref), which, in turn, derives from bai2003inferential.
Because of Assumptions (ref)(ref), (ref)(ref),
(ref)(ref),
(ref)(ref),
and (ref), the FNAR errors have covariance matrix $V= \Lambda\Gamma^G\Lambda' + S$, with $S=\mathbb{E} \left[ \epsilon_{t} \epsilon_{t}' \right]$, which is positive definite for all $N\in\mathbb N$. This also implies that $V^{-1}$ is finite for all $N\in\mathbb N$. Moreover, $V$ has the usual eigengap property; i.e., its largest $q$ eigenvalues diverge at rate $N$, while the remaining $N-q$ stay bounded for all $N\in\mathbb N$. This implies that the factor model in (ref), and therefore the number of factors $q$, is always identified as $N\to\infty$. This is the rationale for the eigenvalue ratio criterion for estimating $q$ defined in (ref).
assumption[Identification of node factors and loadings]\
\begin{enumerate}[label=(\roman*)]
• For all $N \in \mathbb{N}$, ${N}^{-1}\Lambda'\Lambda$ is diagonal with distinct entries.
• For all $T \in \mathbb{N}$, ${T}^{-1}\sum_{t=1}^{T} G_tG_t' = I_q$.
\end{enumerate}
Under Assumption (ref) the columns of the loadings matrix $\Lambda$ and the factors $G_t$ are identified up to a sign multiplication.
Turning to the FNAR defined in (ref), we make the following assumption.
assumption[Stability of FNAR]\
\begin{enumerate}[label=(\roman*)]
• For all $t\in\mathbb Z$ and all $N\in\mathbb N$,
$\det(I_N-\rho I_N- N^{-1}\sum_{j=1}^r\beta_j\mathbb E[F_{j,t}])\ne 0$.
• For all $t\in\mathbb Z$ and all $N\in\mathbb N$,
$\det ({\rho^2 I_{N^2}+ {N^{-2}}\sum_{j=1}^r \beta_j^2 \mathbb E [ F_{j,t}\otimes F_{j,t}]}- zI_{N^2}) = 0$ has roots $z_j^*\in\mathbb C$, $j=1,\ldots,N^2$, such that $|z_j^*|\le C_S$ for some finite $C_S<1$ independent of $j,t$, and $N$.
\end{enumerate}
This assumption is a generalization to the case of random multivariate AR models of the usual stability conditions for a VAR.
As shown below it implies, together with Assumptions (ref)(ref) and (ref), that the FNAR has a stationary solution for all $N\in\mathbb N$. Notice that Assumption (ref)(ref) is stated for the general case in which $\mathbb E[F_{j,t}]\ne 0$, otherwise the condition needed to ensure the existence of the mean is simply $|\rho|<1$.
assumption[Moment conditions - part 2]
For all $m,N,T\in\mathbb N$,
\begin{align}
&\mathbb E\left[\norm{\frac 1{\sqrt {mT}N^{2}}\sum_{t=1}^T\sum_{i=1}^m u_i \mathcal E_{(3)ti\cdot} (y_{t-1}\otimes X_t)}^2 \right]\le\mathfrak K_1,\nonumber\\
&\mathbb E\left[\norm{\frac 1{\sqrt {mT}N^{2}}\sum_{t=1}^T\sum_{i=1}^m u_i \mathcal E_{(3)ti\cdot} (y_{t-1}\otimes \nu_t)}^2 \right]\le\mathfrak K_2,\nonumber
\end{align}
for some finite $\mathfrak K_1$ and $\mathfrak K_2$ independent of $m,N$, and $T$.
To get an intuition of this assumption, consider the $m\times (r+2)$ matrix process $\{\mathcal E_{(3)t} (y_{t-1}\otimes X_t)\}$. We are saying that this process is weakly correlated along the time dimension, which is a standard requirement, but it is also weakly correlated across its $m$ rows. The latter requirement is fulfilled by the idiosyncratic terms $\mathcal E_{(3)t}$ via Assumption (ref)(ref), and here is extended to the case in which $\mathcal E_{(3)t}$ is multiplied by $y_{t-1}\otimes X_t$ which is weakly dependent of $\mathcal E_{(3)t}$ because of Assumption (ref).
assumption[CLT for FNAR]
Let $Z_i:={
M_{ G}X_i-\frac 1N\sum_{k=1}^N(\Lambda_i'\left(\frac{\Lambda'\Lambda} N)^{-1}\Lambda_k \right)M_{ G}X_k
}$ such that $Z_i :=(Z_{i1} \cdots Z_{iT})'$ is $T\times (r+2)$
and $W_t:={
M_{ \Lambda}X_t-\frac 1T\sum_{s=1}^T(G_t'G_s )M_{ \Lambda}X_s
}$ such that $W_t :=(W_{1t} \cdots W_{Nt})'$ is $N\times (r+2)$.
Then, as $N,T\to\infty$,
\begin{enumerate}[label=(\roman*)]
• $
\frac{1}{\sqrt{NT}} \sum_{i=1}^{N} Z_i' \epsilon_{i} \overset{d}{\to} N(0,D_1)
$,
where $D_1 := \lim_{N \to \infty} \frac{1}{N} \sum_{i=1}^N \mathbb{E} \left[ Z_{it}Z_{it}'\right] \sigma_i^2$ is an $(r+2)\times (r+2)$ positive definite matrix.
• $\frac{1}{NT} \sum_{i=1}^{N} Z_i' Z_{i} \overset{p}{\to}\Sigma_{ZZ}$,
where
$\Sigma_{ZZ}:=\lim_{N\to\infty}\frac 1N \sum_{i=1}^{N} \mathbb E[Z_{it} Z_{it}']$ is an $(r+2)\times (r+2)$ positive definite matrix.
• $
\frac{1}{\sqrt{NT}} \sum_{t=1}^{T} W_t' \epsilon_{t} \overset{d}{\to} N(0,D_2)
$,
where $D_2 := \lim_{N \to \infty} \frac{1}{N} \sum_{i=1}^N \mathbb{E} \left[ W_{it}W_{it}'\right] \sigma_i^2$ is an $(r+2)\times (r+2)$ positive definite matrix.
• $\frac{1}{NT} \sum_{t=1}^{T} W_t' W_{t} \overset{p}{\to}\Sigma_{WW}$,
where
$\Sigma_{WW}:=\lim_{N\to\infty}\frac 1N \sum_{i=1}^{N} \mathbb E[W_{it} W_{it}']$ is an $(r+2)\times (r+2)$ positive definite matrix.
\end{enumerate}
Notice that we do not rule out autocorrelation in the node factors, as only the node specific idiosyncratic components are required to have no correlation for the above CLTs to hold, and this is ensured by Assumption (ref)(ref). These assumptions are similar to the conditions in bai2009panel.
Stationarity
In order to develop the theory for the FNAR, we first discuss under which conditions equation (ref) admits a stationary causal solution. Given the difficulty of the problem we limit ourselves to consider weakly stationary solutions. This poses two issues. First, the FNAR is defined for an $N$-dimensional vector $y_t$ where we allow $N\to\infty$. Second, the FNAR is an autoregressive model with stochastic time-varying coefficients. Regarding the former issue, we adopt the definition proposed by zhuetal17.
definitionLet $\{y_t\}$ be an $N$-dimensional stochastic process with $N\in\mathbb N$. Let $W:=\{\omega:=(\omega_1\cdots\omega_N)' \in\mathbb R^N\,:\, \sum_{i=1}^N \abs{\omega_i}<\infty , N\in\mathbb N\}$. Then, $\{y_t\}$ is weakly stationary if for all $N\in\mathbb N$ and any given $\omega\in W$, $y_t^\omega:=\lim_{N\to\infty} \omega'y_t$ exists almost surely and $\{y_t^w\}$ is weakly stationary and causal.
Turning to the second problem, we have the following result.
proposition[Stationarity of FNAR]
Under Assumptions (ref)(ref), (ref)(ref), (ref), and (ref), for all $N\in\mathbb N$, the FNAR has a unique weakly stationary and causal solution.
In general, one might object that when $N\to\infty$ a meaningful concept of stationarity cannot be stated, as no causal solution to the FNAR can exist since Assumption (ref) will break down, see, e.g., the remark by zhou2020network, in a similar context. What we mean by Definition (ref) and Proposition (ref) is that we are implicitly assuming that there exists a space where the causal solution is well-defined even when $N \to\infty$. This is the same approach adopted by zhuetal17. An interesting implication of this definition is that, under our assumptions, we can ensure that any finite linear combination of the elements of $\{y_t\}$ satisfies a finite dimensional FNAR with a causal solution.
Asymptotic properties of network factors and FNAR coefficients
Consistency and asymptotic normality of the estimated network factor loadings are given next.
theorem[Consistency and asymptotic normality of loadings]\
\begin{enumerate}[label=(\roman*)]
• Under Assumptions (ref)-(ref),
as $m,N,T \to \infty$,
\[
\norm{ \frac{\widehat{U} - U J}{\sqrt{m}} }
= O_p \left(
\max \left( \frac{1}{N \sqrt{T}}, \frac{1}{ m } \right) \right),
\]
where $J$ is a $r \times r$ diagonal matrix whose diagonal entries are equal to $\pm 1$.
• Under Assumptions (ref)-(ref), for any given $i=1,\ldots, m$,
as $m,N,T\to\infty$, if $N\sqrt T/m \to 0$,
\[
N\sqrt T \left(\widehat u_i'-u_i' J \right) \overset{d}{\to}\mathcal N\left(0_r, \Phi_i\right),
\]
where $\widehat u_i'$ and $u_i'$ are the $i$-th rows of $\widehat U$ and $U$, respectively, $\Phi_i$ is defined in Assumption (ref)(ref), and $J$ is defined in part (ref).
\end{enumerate}
Theorem (ref) shows that, when applying PCA to a given mode of the tensor $\mathcal W_t$, the dimensions of all other modes contribute to a faster convergence rate, hence allowing for more degrees of freedom. This is an advantage with respect to the vector case, since even for moderately small values of $T$ we can still have good estimates of the loadings matrix and therefore of the network factors. In particular,
we see that the estimated loadings vector $\widehat u_i$ has a consistency rate $\min(m,N\sqrt T)$ and is asymptotically normal if $N\sqrt T/m\to 0$. This is the generalization to the multilayer network case (i.e., to order-3 tensors) of the usual vector case, which corresponds to setting $N=1$ (see bai2003inferential).
Next we prove consistency and asymptotic normality of the estimated network factors.
theorem[Consistency and asymptotic normality of network factors]\
\begin{enumerate}[label=(\roman*)]
• Under Assumptions (ref)-(ref),
for any given $t=1,\ldots, T$, as $m,N,T \to \infty$,
\[
\norm{ \frac{\widehat{\mathcal{F}}_{(3)t} - J \mathcal{F}_{(3)t}}{N} } =
O_p \left(\max
\left(
\frac 1{N^{2} T},
\frac{1}{\sqrt{m} }
\right) \right),
\]
where
$J$ is defined in Theorem (ref)(ref).
• Under Assumptions (ref)-(ref), for any given $t=1,\ldots, T$ and $j=1,\ldots,N^2$, as $m,N,T\to\infty$,
if $\sqrt m/(N^2T)\to 0$,
\[
\sqrt{m}\left(\widehat {\mathcal F}_{(3)t\cdot j}-J {\mathcal F}_{(3)t\cdot j}
\right)\overset{d}{\to}\mathcal N\left(0_{r}, \Sigma_U^{-1}\Pi_{tj}\Sigma_U^{-1} \right),
\]
where $\Pi_{tj}$ is defined in Assumption (ref)(ref) and $J$ is defined in Theorem (ref)(ref).
\end{enumerate}
Theorem (ref)(ref) proves consistency of the whole network factor tensor.
Theorem (ref)(ref) proves asymptotic normality of any given column of $\widehat {\mathcal F}_{(3)t}$, which is equivalent to asymptotic normality of any of the $N^2$ entries of each of the $r$ layers of the multilayer network factor $\widehat {\mathcal F}_t$.
This is the natural generalization to the multinetwork case of the usual vector case, i.e., when $N=1$ (see bai2003inferential).
We then turn to the asymptotic properties of the estimated FNAR coefficients.
Hereafter, let
equation[equation omitted — 158 chars of source]
with $J$ as in Theorem (ref)(ref).
Then, we analyze the properties of the estimators $\widehat{\theta}^\dag$ and $\widehat{\theta}^*$ by noticing that
equation[equation omitted — 821 chars of source]
where $u_i:= (X_i\bar J-\widehat X_i)\bar J\theta$ and $u_t:= (X_t\bar J-\widehat X_t)\bar J\theta$, and recall that $\theta := (\beta', \rho, \alpha )'$.
The following theorem holds.
theorem[CLT for FNAR coefficients estimated by iterative OLS]
Under Assumptions (ref)-(ref) and (ref), and if ${ \sqrt {NT}}/ m \to 0$ and ${ {N}}/ {m} \to 0$, as $m,N,T\to\infty$, then:
\begin{enumerate}[label=(\roman*)]
• if $T/N\to 0$ and $\sqrt N/T\to 0$, we have
\begin{equation}
\sqrt {NT}(\widehat{\theta}^{\dag}-\bar J\theta)
\overset{d}{\to}
\mathcal{N} ( 0_{r+2},
\Sigma_{ZZ}^{-1} D_1\Sigma_{ZZ}^{-1}
),\nonumber
\end{equation}
where $D_1$ and $\Sigma_{ZZ}$ are defined in Assumptions (ref)(ref) and (ref)(ref), and $\bar J$ is defined in (ref);
• if $\sqrt T/N\to 0$ and $\sqrt N/T\to 0$ and $\sigma_i^2=\sigma^2$ for all $i=1,\ldots, N$, we have
\begin{equation}
\sqrt {NT}(\widehat{\theta}^{\dag}-\bar J\theta)
\overset{d}{\to}
\mathcal{N} ( 0_{r+2},\sigma^2
\Sigma_{ZZ}^{-1}
),\nonumber
\end{equation}
where $\Sigma_{ZZ}$ is defined in Assumption (ref)(ref), and $\bar J$ is defined in (ref);
• if $T/N\to 0$ and $\sqrt N/T\to 0$, we have
\begin{equation}
\sqrt {NT}(\widehat{\theta}^{*}-\bar J\theta)
\overset{d}{\to}
\mathcal{N} ( 0_{r+2},
\Sigma_{WW}^{-1} D_2\Sigma_{WW}^{-1}
),\nonumber
\end{equation}
where $D_2$ and $\Sigma_{WW}$ are defined in Assumptions (ref)(ref) and (ref)(ref), and $\bar J$ is defined in (ref);
• if $\sqrt T/N\to 0$ and $\sqrt N/T\to 0$ and $\sigma_i^2=\sigma^2$ for all $i=1,\ldots, N$, we have
\begin{equation}
\sqrt {NT}(\widehat{\theta}^{*}-\bar J\theta)
\overset{d}{\to}
\mathcal{N} ( 0_{r+2},\sigma^2
\Sigma_{WW}^{-1}
),\nonumber
\end{equation}
where $\Sigma_{WW}$ is defined in Assumption (ref)(ref), and $\bar J$ is defined in (ref);
\end{enumerate}
Parts (ref) and (ref) extend Theorem 2 in bai2009panel to the FNAR case. The interesting cases are parts (ref) and (ref), where we do not impose homoskedastic idiosyncratic components in the FNAR errors. Notice that the network coefficients, $\beta_j$, $j=1,\ldots, r$, which are the first $r$ elements of $\theta$, are consistently estimated only up to a sign, due to the indeterminacy in the identification of the network factors.
Estimators of the asymptotic variance-covariance matrix of $\widehat \theta^\dag$ and $\widehat\theta^*$ under the assumptions in parts (ref) and (ref) are (see also bai2009panel):
align[align omitted — 825 chars of source]
where
$
\widehat Z_{i} :=
M_{\widehat G^\dag}\widehat{X}_i-\frac 1N\sum_{k=1}^N \left(\widehat\Lambda_i^{\dag'}\left(\frac{\widehat{\Lambda}^{\dag'}\widehat\Lambda^\dag} N\right)^{-1}\widehat\Lambda_k^\dag\right)M_{\widehat G^\dag}\widehat X_k
,
$
and
$
\widehat W_t:=
M_{\widehat \Lambda^*}\widehat X_t-\frac 1T\sum_{s=1}^T (\widehat G_t^{*'}\widehat G_s^* )M_{\widehat \Lambda^*}\widehat X_s
.
$
Four important comments about this result follow.
First, the proof is based on showing that if ${ \sqrt {NT}}/ m \to 0$ and $N/m\to 0$, as $m,N,T\to\infty$, then the network factors can be treated as observed; i.e., the generated regressors bias is asymptotically negligible (see Proposition C.1).
Under these conditions, and if also $T/N\to 0$ and $\sqrt N/T\to 0$, the iterative estimators are $\sqrt{NT}$-consistent.
Second, if the node factors are not autocorrelated we can also compare the iterative estimators with the OLS and the GLS estimators studied in Appendix B and C, respectively. In this case, the GLS is also $\sqrt{NT}$-consistent since, similarly to the iterative estimator, it rescales $X_t$ by the FNAR error covariance matrix $V$, which is $O(N)$ by Assumption (ref). For the same reason the OLS estimator is just $\sqrt T$-consistent since it does not control for the FNAR error covariance, and it would be $\sqrt{NT}$-consistent only if $V$ were a diagonal matrix; i.e., when no node factor is present, as assumed by zhuetal17. The same results on OLS and GLS are obtained by chenetal2020cnar for the case of observed networks.
Third, the two estimators are asymptotically equivalent and thus also equally efficient. To see this notice that $\widehat{\theta}^\dag$ and $\widehat{\theta}^*$ are such that they solve
align[align omitted — 476 chars of source]
so the two losses are identical and must have the same minimum. Once we fix the identification constraints as in Assumption (ref), the only difference is then about the implementation of the minimizations. For $\widehat{\theta}^{\dag}$, by replacing $\Lambda_i= (G^\prime G)^{-1} G^\prime(y_i-X_i\theta)=T^{-1}G^\prime(y_i-X_i\theta)$ in (ref), we can solve for $G$ and $\theta$ only. For $\widehat{\theta}^{*}$, by replacing $G_t=(\Lambda^\prime\Lambda)^{-1}\Lambda'(y_t-X_t\theta)$ in (ref), we can solve for $\Lambda$ and $\theta$ only. These two procedures lead to the two solutions given in (ref). However, since in practice the solutions are obtained by iteration, the two estimates might not coincide exactly, although they will be very similar. And, as expected, the estimated standard errors will coincide (see the results in Section (ref) and Appendix F.4).
Fourth, and last, we should view Theorem (ref) as giving the asymptotic distribution of the theoretical estimator minimizing (ref) or (ref). This is the same point of view adopted by bai2009panel. In practice, it might be important to investigate how the initialization of the algorithm affects such convergence. In the simpler case of a panel regression having errors with a factor structure, jiang2021recursive show that any initial estimator could still lead to a consistent iterated estimator, depending on the structure of the regressors which can be quite general. However, the estimator computed in practice might have a slower convergence rate if the initial estimator is not consistent. We do not explore this aspect further here, but we limit ourselves to notice that in our numerical exercises of Sections (ref) and (ref), convergence is always achieved in few steps and the iterated estimator works well even in presence of weak serial correlation of the node factors.
\refstepcounter{lettersection}
\oldsection{Monte Carlo simulations}
To evaluate the finite sample performance of the proposed estimators, we generate artificial time series of $y_t$ and $\mathcal{W}_t$, for $t=1, \dots, T$, according to the model equations (ref), (ref), and (ref).
We fix $r=1$ and $q=1$; i.e., one network factor and one node-specific factor. We also fix
the values of FNAR parameters $\beta=0.5$, $\rho=0.3$, and $\alpha=0.2$. We consider $N \in \{10,20,50,100,200\}$ nodes, $m \in \{20,50,100\}$ layers, and $T \in \{50,100\}$ time periods.
Also, for each value of the pair ($N ,m$), we randomly generate the entries of the loading vectors $U$ and $\Lambda$ once (and independently)
from $\mathcal{N}(1,1)$, and then we keep them fixed across Monte Carlo (MC) iterations (see chenetal2020cnar). All other quantities are generated at each MC iteration. All results are based on 500 iterations. Full details on the data generating process are in Appendix F.1.
table[table omitted — 2,362 chars of source]
table[table omitted — 2,360 chars of source]
Tables (ref)-(ref) report the {RMSE} and Relative RMSE ({ReRMSE}) of the estimates.
Here we report results under case II, which corresponds to idiosyncratic terms $\mathcal E_t$ having serial and cross-layer correlation, and using the iterative estimator $\widehat{\theta}^\dag$ in (ref). Additional results for case I of uncorrelated idiosyncratic terms are in Appendix F.2.
As predicted by the theory, the accuracy of estimates for $\beta$, $\rho$ and $\alpha$ improves with both $N$ and $T$, and the RMSE of network factors and loadings decreases when the number of layers $m$ increases. Furthermore, as shown in Appendix F.3, the MC distributions of the estimated network coefficient are all strongly centered around the true value $\beta=0.5$ and become narrower as $T$ and $N$ increase.
Last, in Appendix F.5 we compare our estimates of the loadings $U$ with those obtained using the TOPUP and TIPUP estimation methods proposed by chenyangzhang22. As expected our approach improves over those estimators in presence of serial idiosyncratic correlation.
\refstepcounter{lettersection}
\oldsection{Empirical application}
In this section, we present an application of the FNAR for studying cross-country macroeconomic interdependence determined by global trade flows and cross-border financial relationships.
\noindentData.
For a sample of $N=24$ countries, we use $m=25$ networks constructed using bilateral import/export flows for different good (layers 1--9) and services (10--19) categories, bilateral financial positions for different types of financial claims (20--23) and cross-border mergers and acquisitions classified by sector of economic activity (24--25). The list of countries and network layers, including details on how the networks are built, are given in Appendix G.1.
\noindentNetwork factors. Due to data limitations in the time series of financial positions, we collect data for the networks at the annual frequency from 2001 to 2019, so the factor analysis is conducted on a sample of length $T_1=19$.\footnote{We have few missing values over this period. In these cases, we use the previous year's value or the closest available year's value.} Although this is a short time span, we recall that in tensor PCA the effective sample size when estimating the loadings space is $N^{2}T_1$
(see Theorem (ref)).
We then extract the common network factors from the 25 layers of the network, and we set $\widehat r=6$ network factors, as in chenyangzhang22.
From Figures G.9-G.12 in Appendix G.2 we can interpret the six network factors as follows. The first network factor conveys approximately the average country weights across all layers of the network. In particular, the factor values are very close to the average weights (scaled by a constant), and the loading coefficients are almost the same for all layers. The countries with the largest factor weights for the US are its major economic partners: Canada, UK, Mexico, China, Japan and Germany.
The second factor captures a difference between financial relationships and trade in goods. The factor loadings for financial layers have opposite sign (positive) compared to the loadings for trade-in-goods layers (negative). Recall that factors are identified up to a sign.
The largest positive weights are assigned to economies having relatively large financial sectors with global reach: UK, US, and Hong Kong. In the case of the US connections, a large positive weight is assigned to the UK, whose tight economic links with the US are mostly concentrated in the financial sector, and large negative weights are assigned to Canada, Mexico, and China, i.e., the US biggest trade partners.
The third factor distinguishes between equity and debt relationships, being the only factor where equity, on the one hand, and debt, on the other hand, show loadings with opposite signs. The fourth factor is strongly associated with M&A relationships.
The fifth factor is mainly driven by agricultural/extractive goods (positive weights, especially for vegetable fuels, oils, fats, and waxes). It also loads on trade in manufacturing goods (negative weights). Positive weights are assigned to countries with strong trade links with the US in non-manufacturing sectors, such as Canada, Saudi Arabia, and Italy, while negative weights are associated with large manufacturing partners, like China. This factor also distinguishes between stocks of portfolio holdings and flows associated with M&A deals and banking. Finally, the sixth factor captures a distinction between goods-sector M&A integration and services-sector integration.
Next, in line with conventional PCA, we evaluate the fraction of variance in network layers explained by each factor, denoted as $v^{(k)}$, $k=1,\ldots, 6$ (computed as in Appendix G.2). We have $v^{(1)}=0.68$, $v^{(2)}=0.07$, $v^{(3)}=0.03$, $v^{(4)}=0.03$, $v^{(5)}=0.02$, and $v^{(6)}=0.02$. Thus, overall the 6 factors explain about 85% of the total variance of $\mathcal W$. However, the importance of different factors varies greatly across countries; see Table G.11 in Appendix G.2.
\noindentFNAR coefficients.
The endogenous vector $y_t$, $t=1,\ldots, T$, collects (quarterly) real GDP growth rates for all $N$ considered countries and for the sample 2001Q1-2019Q4; i.e., $T_2=76$. To address heterogeneity of nodal and momentum effects, we split the countries into two groups: (1) advanced economies ($N_1=15$), and (2) emerging economies ($N_2=9$) and the vector $y_t$ is partitioned accordingly as $y_t=(y_t^{(1)'}; y_t^{(2)'})'$.
We consider the following FNAR, for $t=1,\ldots, T$,
align[align omitted — 471 chars of source]
where $\widetilde F_{j,t} = \widehat F_{j,\tau}$ for $4(\tau-1)+1\le t \le 4\tau$, $\tau=1,\ldots, T_1$. In other words, the network factors $F_{k,t}$, $t=1,\ldots, T_1$, which are computed on a yearly basis, are treated as constant throughout all quarters of a given year.
Hereafter, we let $\theta:=(\beta', \rho^{(1)},\rho^{(2)},\alpha^{(1)}, \alpha^{(2)})'$.
By means of the criterion defined in (ref), we find evidence of one common node factor, i.e., $\widehat q=1$. We then estimate the model by GLS as described in Appendix C. Last, we consider the iterative estimators $\widehat\theta^\dag$ or $\widehat\theta^*$ defined in (ref) and we initialize the algorithm by using the GLS estimator and the estimated node loadings, $\widehat \Lambda$, and factor, $\widehat G_t$, computed by PCA on the GLS residuals as described in Appendix A.
Since these residuals do not display significant autocorrelation, we are confident that the GLS estimator is $\sqrt{NT}$-consistent and, based on the results of jiang2021recursive we conjecture that Theorem (ref) holds for our iterated estimators. Convergence is reached in 8 or 4 iterations for $\widehat\theta^\dag$ or $\widehat\theta^*$, respectively.
Table (ref) reports the estimated coefficients and their standard errors with significance reported according to the usual $Z$-test. The coefficients on $N^{-1}\widehat F_{1,t-1}y_{t-1}$ and $N^{-1}\widehat F_{5,t-1}y_{t-1}$ are always strongly significant, while there is mixed evidence regarding the coefficients on $N^{-1}\widehat F_{2,t-1}y_{t-1}$, $N^{-1}\widehat F_{4,t-1}y_{t-1}$, and $N^{-1}\widehat F_{6,t-1}y_{t-1}$ which are mildly significant and not for all estimates.
table[table omitted — 1,901 chars of source]
Based on the interpretation of the first network factor, the coefficient on $N^{-1}\widehat F_{1,t-1}y_{t-1}$ captures a general network effect operating through aggregate economic linkages.
Given the loadings of factor 5 in Figure G.10, the coefficient on $N^{-1}\widehat F_{5,t-1}y_{t-1}$ indicates that trade in mineral fuels and in animal and vegetable oils (major inputs of chemical industry) has the main impact on GDP growth.
Apart from this, trade linkages tend to generate larger spillovers in manufacturing sectors (layers 6-9) than in non-manufacturing sectors (layers 1-3), and financial linkages tend to generate larger spillovers when they take the form of M&A or flows of banking assets (rather than portfolio holdings).
Finally, based on these estimates, we can approximate the network effects associated with the original layers of the network, by appropriately rescaling the estimated network effect coefficients $\beta$. Specifically, given the definition of estimated loadings in (ref) and the properties of tensor multiplication, and letting $\widehat{\mathcal{W}}_{t} := \widehat{\mathcal{F}}_{t} \times_3 \widehat{U}$, for a given estimate $\widehat{\beta}$ we have that:
align[align omitted — 422 chars of source]
Thus, $N^2 \widehat{U} (\widehat{M}^{\mathcal{W}})^{-1}\widehat{\beta}$ is the vector of network effects in terms of the row-normalized tensor $\widehat{\mathcal{W}}_{t-1} / N$ and its entries are shown in Figure (ref), when computed using the iterated estimator $\widehat{\theta}^*$.
The figure confirms a substantial heterogeneity of effects across layers, reflecting their different loadings on the network factors.
\noindentForecasting. We conclude by studying the performance of our FNAR when producing 1-quarter-ahead forecasts of GDP growth rates based on a recursive window exercise from 2002Q1 to 2019Q4. We consider the following competitors: the tensor-based estimators MLR and SHORR by wangetal21;
a NAR estimated either using the $m$ layers of the common component (denoted as TUCKER COMMON and further regularized via LASSO or Ridge due to collinearity of the regressors) or just $r$ network factors (denoted as TUCKER FACTORS) both obtained from a full Tucker factor decomposition, estimated via TOPUP as in chenyangzhang22 (see Appendix H for more details);
a multilayer NAR estimated via LASSO or Ridge; and an ordinary VAR. Details on the adopted forecasting scheme and on the implementation of the alternative estimation methods are in Appendix G.2.4.
Table (ref) reports the root mean squared forecast errors (RMSFE) in terms of percentage points of GDP growth, for each country considered. In the last two rows of Table (ref), we report the average (across all countries) RMFSE and relative RMSFE (ReRMSFE) with respect to the FNAR, for all forecasting methods (values larger than one indicate a better performance of the FNAR). For most countries, and on average, our approach delivers the forecasts with smallest RMSFEs. In particular, we outperform the approaches based on a full Tucker decomposition which are our most natural competitors. Indeed, contrary to the case of factors extracted by means of a full Tucker decomposition, our network factors still contain terms which are idiosyncratic to the network nodes, i.e., countries, which are potentially relevant for predicting country specific GDP growth rates.
figure[figure omitted — 665 chars of source]
table[table omitted — 3,690 chars of source]
\refstepcounter{lettersection}
\oldsection{Conclusions}
In this paper, we have introduced a factor network autoregression (FNAR) for time series characterized by multiple network effects. Estimation is based on two steps. First, we extract few network factors common across the layers of the underlying multilayer network. Second, we estimate a factor-augmented NAR or FNAR where the network effects are determined by the latent network factors. The FNAR errors are allowed to have an underlying factor structure capturing common correlations across the network nodes. We prove consistency and asymptotic normality of the proposed estimators as the number of layers, nodes and time observations diverges to infinity.
The results of an empirical application show that, by accounting for cross-country economic and financial linkages, the model provides a rich description of the dynamics of GDP growth rates and produces accurate forecasts.
We outline three possible extensions of this work, which we leave for further research. First, by adapting the works by chenetal2020cnar and zhu2022simultaneous to the FNAR framework, we could consider a FNAR with momentum and nodal coefficients which are group specific, where the group structure is unknown and the number of groups $K$ can grow with the number of nodes $N$.
Second, by generalizing to the tensor setting the three-pass regression filter by kelly2015three,
we could improve the performance of our estimator by accounting also for the information contained in the vector of dependent variables $y_t$ when extracting the network factors.
Third, by extending to the tensor case an approach similar to the one proposed by wu2020adaptive, we could allow for time-varying network factor loadings under the standard assumption of local stationarity.
center[center omitted — 216 chars of source]
\newcounter{lettersection}
\setcounter{lettersection}{0}
\let\oldsection\section
\setcounter{figure}{0}
\setcounter{table}{0}
This supplemental material contains several appendices to our paper.
In Appendix (ref) we describe how to estimate the node factors by Principal Components Analysis (PCA) on the FNAR residuals.
In Appendices (ref) and (ref), we introduce the OLS estimator and a GLS-type estimator of the FNAR coefficients. In Appendix (ref), we provide the proofs for Propositions and Theorems. In Appendix (ref), we provide and prove all auxiliary results.
In Appendix (ref) we provide additional simulation results.
In Appendix (ref) we provide additional information on the data used in our empirical application, as well as additional empirical results.
Finally, in Appendix (ref) we consider alternative approaches for the estimation of a multilayer NAR, based on full Tucker decompositions.
\refstepcounter{lettersection}
\oldsection{Estimation of node factors}
Once we compute the OLS estimator, let $\widehat\nu:=(\widehat\nu_1\cdots \widehat\nu_T)'$ be the $T\times N$ matrix of residuals of the FNAR such that
$\widehat \nu_t:=y_t -\widehat X_t \widehat{\theta}^{\text{\upshape\tiny OLS}}$, $t=1,\ldots, T$. Then, the node factors $G_t$ and their loadings $\Lambda$ can be estimated by PCA in two equivalent ways.
Specifically, consider either the $N\times N$ or $T\times T$ sample covariance matrices
equation[equation omitted — 191 chars of source]
Then, letting $\widehat G:=(\widehat G_1,\cdots, \widehat G_T)'$ and $\widetilde G:=(\widetilde G_1,\cdots, \widetilde G_T)'$ be the estimated $T\times q$ matrices of factors, the PC estimators are given by:
align[align omitted — 471 chars of source]
where $\widehat M^{\widehat\nu}$ is the $q\times q$ diagonal matrix of eigenvalues of $\widehat {\Gamma}^{\widehat \nu}$ with corresponding normalized eigenvectors being the columns of the $N\times q$ matrix $\widehat V^{\widehat\nu}$, and $\widetilde V^{\widehat\nu}$ is the $T\times q$ matrix of normalized eigenvectors of $\widetilde {\Gamma}^{\widehat \nu}$. It is easy to verify that, regardless of the choice made in (ref), $\widehat G\widehat \Lambda'=\widetilde G\widetilde \Lambda'$.
Notice that the estimated loadings and factors are such that $\widehat \Lambda'\widehat \Lambda /N$ and $\widetilde \Lambda'\widetilde \Lambda /N$ are diagonal and
\[
\frac 1T\sum_{t=1}^T \widehat G_t\widehat G_t'= I_q,\quad \frac 1T\sum_{t=1}^T \widetilde G_t\widetilde G_t'= I_q
\]
so that the estimated factors are orthonormal.
\refstepcounter{lettersection}
\oldsection{OLS estimator and its asymptotic properties}
The OLS estimator of $\theta$ is given by:
equation[equation omitted — 229 chars of source]
This is the estimator proposed by zhuetal17 and chenetal2020cnar for the case in which the network is observed and, respectively, when the FNAR errors are either uncorrelated or have a factor structure with serially uncorrelated node factors.
We start by making the following standard assumption.
assumption[CLTs for FNAR - OLS]
\
\begin{enumerate}[label=(\roman*)]
• For all $t,s\in\mathbb Z$ with $t\ne s$, $\mathbb E[G_tG_s']=0$.
• As $N,T\to\infty$,
$
\frac{1}{N\sqrt{T}} \sum_{t=1}^{T} X_t' \nu_{t} \overset{d}{\to} N(0_{r+2},\Omega_0)
$,
where
$\Omega_0 := \lim_{N \to \infty} \frac{1}{N^2} \mathbb{E} \left[ X_t' V X_t \right]$ is an $(r+2)\times (r+2)$ positive definite matrix.
• As $N,T\to\infty$,
$
\frac{1}{NT} \sum_{t=1}^{T} X_t' X_{t} \overset{p}{\to}\Sigma_{XX}
$,
where
$\Sigma_{XX}:=\lim_{N\to\infty}\frac 1N \mathbb E[X_t' X_{t}]$ is an $(r+2)\times (r+2)$ positive definite matrix.
\end{enumerate}
Because of Assumptions (ref)(ref), (ref)(ref), and (ref), we have that the FNAR errors $\{\nu_t\}$ are not autocorrelated. This is necessary for the CLT in the next part of this assumption to hold. Assumption (ref)(ref) is also found in baiandng2006. In fact, Assumptions (ref)(ref) and (ref)(ref) are made for simplicity and could be proved in a similar way as in chenetal2020cnar.
The OLS estimator in (ref) satisfies:
align[align omitted — 356 chars of source]
The asymptotic properties of the terms of (ref) are given in the following Proposition.
propositionUnder Assumptions (ref)-(ref) and (ref), as $m,N,T\to\infty$,
\begin{enumerate}[label=(\roman*)]
• $\frac1{NT}\sum_{t=1}^T\widehat X_{t}' u_t = O_p \left( \max\left(\frac 1{N^{2}T},\frac 1{m},\frac 1{\sqrt {mT}}\right)\right) $.
• $\frac1{NT}\sum_{t=1}^T\widehat X_{t}' \nu_t = O_p\left(\max\left(\frac 1{\sqrt {T}},\frac 1{N^{2}T},\frac 1{m},\frac 1{\sqrt {mT}}\right)\right)$.
• $ \norm {\frac 1{NT}\sum_{t=1}^T \widehat{X}_{t}'\widehat X_{t}-\frac 1{NT}\sum_{t=1}^T\bar J {X}_{t}' X_{t}\bar J} = O_p \left( \max\left(\frac 1{N^{2}T},\frac 1{m},\frac 1{\sqrt {mT}}\right)\right)$.
\end{enumerate}
The next theorem follows.
theorem[CLT for FNAR coefficients estimated by OLS]
Under Assumptions (ref)-(ref) and (ref),
if ${ \sqrt T}/ m \to 0$, as $m,N,T\to\infty$,
\begin{equation}
\sqrt T(\widehat{\theta}^{\tiny OLS}-\bar J\theta)
\overset{d}{\to}
\mathcal{N} ( 0_{r+2},
\Sigma_{XX}^{-1}
\Omega_0
\Sigma_{XX}^{-1}
),\nonumber
\end{equation}
where $\Omega_0$ and $\Sigma_{XX}$ are defined in Assumptions (ref)(ref) and (ref)(ref), respectively, and
$\bar J$ is defined in (ref).
From Proposition (ref), we see that if $\sqrt{ T}/ m \to 0$, as $m,N,T\to\infty$, then the network factors can be treated as observed and Theorem (ref) follows. In particular, by virtue of Assumptions (ref)(ref) and (ref)(ref), the OLS estimator is $\sqrt T$-consistent and asymptotically normal. Notice that
the requirement $\sqrt T/m\to 0$ for Theorem (ref) to hold is analogous to the one assumed in the vector case by baiandng2006.
An estimator of the asymptotic variance-covariance matrix of the OLS estimator is then given by (see baiandng2006):
\[
\widehat{\operatorname*{Avar}}\left[ \sqrt{T} (\widehat{\theta}^{\text{\upshape\tiny OLS}} -\bar J\theta) \right]
=
\left(\frac{1}{NT} \sum_{t=1}^{T} \widehat{X}_{t}^' \widehat{X}_t \right)^{-1}
\left(\frac{1}{N^2T} \sum_{t=1}^{T} \widehat{X}_t^'
\widehat \nu_t \widehat \nu_t'
\widehat{X}_t \right)
\left( \frac{1}{NT} \sum_{t=1}^{T} \widehat{X}_{t}^' \widehat{X}_t \right)^{-1}.
\]
\refstepcounter{lettersection}
\oldsection{GLS estimator and its asymptotic properties}
A GLS extension of the OLS estimator of the FNAR considered in Appendix (ref) can be computed by means of the following procedure, initially proposed by chenetal2020cnar for the special case where the network is observed.
Once we estimate the node factors and their loadings as described in Appendix (ref), let $\widehat{\epsilon}:=\widehat \nu-\widehat G\widehat \Lambda' = \widehat \nu-\widetilde G\widetilde \Lambda'$ and $\widehat S$ be the diagonal matrix with entries the diagonal entries of $T^{-1}\widehat{\epsilon}'\widehat{\epsilon}$. We can estimate the covariance matrix $V$ of $\nu_t$ as $\widehat V :=\widehat \Lambda\widehat \Lambda'+\widehat S$ and, by applying the Sherman-Morrison-Woodbury formula, its inverse as:
equation[equation omitted — 200 chars of source]
where $\widehat S^{-1}$ is a diagonal matrix and hence easy to compute.
The GLS estimator of $\theta$ is then given by:
equation[equation omitted — 238 chars of source]
To study the asymptotic properties of the GLS estimator, we make two more assumptions. First, we extend Assumptions (ref)(ref) and (ref)(ref) to the following.
assumption[CLTs for FNAR - GLS]
\
\begin{enumerate}[label=(\roman*)]
• For all $t,s\in\mathbb Z$ with $t\ne s$, $\mathbb E[G_tG_s']=0$.
• As $N,T\to\infty$,
$
\frac{1}{\sqrt{NT}} \sum_{t=1}^{T} X_t' V^{-1}\nu_{t} \overset{d}{\to} N(0_{r+2},\Omega_1)
$, where
$\Omega_1 := \lim_{N \to \infty} \frac{1}{N} \mathbb{E} \left[ X_t' V^{-1} X_t \right]$ is an $(r+2)\times (r+2)$ positive definite matrix.
• As $N,T\to\infty$,
$
\frac{1}{NT} \sum_{t=1}^{T} X_t' V^{-1} X_{t}\overset{p}{\to}\Omega_1
$,
where
$\Omega_1$ is defined in part (ref).
\end{enumerate}
Because of Assumptions (ref)(ref), (ref)(ref), and (ref), we have that the FNAR errors $\{\nu_t\}$ are not autocorrelated. This is necessary for the CLT in the next part of this assumption to hold.
Assumptions (ref)(ref) and (ref)(ref) follow directly from (ref)(ref) and (ref)(ref) since we know that $V^{-1}$ is finite for all $N\in\mathbb N$. Notice, however, the different role of $N$ in the definition of $\Omega_0$ and $\Omega_1$, indeed, as $N\to\infty$ we have $X_t'VX_t=O_p(N^2)$, but $X_t' V^{-1} X_t=O_p(N)$, since $X_t=O_p(\sqrt N)$, $V=O(N)$ and $V^{-1}=O(1)$.
Second, to study the properties of the GLS estimator (ref) we need to prove consistency of the estimated inverse of the FNAR errors covariance $\widehat V^{-1}$ defined in (ref). This is not an easy task, for at least three reasons: first, the FNAR errors are estimated and not observed; second, the matrix $V$ is $N\times N$ so it is a high-dimensional one; third, to study $\widehat V^{-1}$, we need uniform consistency over all $N^2$ entries of the estimated covariance $\widehat V$. These difficulties are reduced if we assume that all considered random variables are sub-Gaussian, which is a classical assumption in high-dimensional statistics vershynin2018high.
assumption[Sub-Gaussianity]\
\begin{enumerate}[label=(\roman*)]
• For all $i\in\mathbb N$, all $j=1,\ldots, r+2$, and all $t\in\mathbb Z$, $\text{\upshape P}(\abs{X_{ijt}-\mathbb E[X_{ijt}]}>s)\le 2\exp(-s^2/c_1^2)$
for some finite $c_1$ independent of $i,j$, and $t$.
• For all $i,j,k\in\mathbb N$ and all $t\in\mathbb Z$, $\text{\upshape P}(\abs{\mathcal E_{tijk}}>s)\le 2\exp(-s^2/c_2^2)$ for some finite $c_2$ independent of $i,j,k$, and $t$.
• For all $i\in\mathbb N$ and all $t\in\mathbb Z$, $\text{\upshape P}(\abs{\epsilon_{it}}>s)\le 2\exp(-s^2/c_3^2)$ for some finite $c_3$ independent of $i$ and $t$.
\end{enumerate}
This approach is similar to the one adopted in a vector factor model context by fan2013large. Instead, chenetal2020cnar assume a set of moment conditions on the regressors matrix $X_t$, which amount to bounding up the 8th order cross-cumulants (in addition, they also make use of Hanson-Wright concentration inequalities which are based
on the assumption of sub-Gaussianity).
The GLS estimator in (ref) satisfies:
align[align omitted — 360 chars of source]
The asymptotic properties of the terms of (ref) are then given in the following Proposition.
propositionUnder Assumptions (ref)-(ref), (ref), and (ref), if $\sqrt T/m\to 0$, as $m,N,T\to\infty$,
\begin{enumerate}[label=(\roman*)]
• $\frac1{NT}\sum_{t=1}^T\widehat X_{t}'\widehat V^{-1} u_t = O_p \left( \max\left(\frac 1{N^{2}T},\frac 1{m},\frac 1{\sqrt {mT}}\right)\right)$.
• $\frac1{NT}\sum_{t=1}^T\widehat X_{t}' \widehat V^{-1}\nu_t = O_p \left( \max\left(\frac 1{\sqrt{NT}},\frac 1{N^{2}T},\frac 1{m},\frac 1{\sqrt {mT}}\right)\right)$.
• $ \norm {\frac 1{NT}\sum_{t=1}^T \widehat{X}_{t}'\widehat V^{-1}\widehat X_{t}-\frac 1{NT}\sum_{t=1}^T\bar J {X}_{t}' V^{-1} X_{t}\bar J} = O_p\left(\max\left(\frac 1{\sqrt N},\sqrt{\frac{\log N}{T}},\frac 1{m},\frac 1{\sqrt {mT}}\right)\right)$.
\end{enumerate}
The next theorem follows.
theorem[CLT for FNAR coefficients estimated by GLS]
Under Assumptions (ref)-(ref), (ref), and (ref), if ${ \sqrt {NT}}/ m \to 0$ and $N/m \to 0$, as $m,N,T\to\infty$,
\begin{equation}
\sqrt {NT}(\widehat{\theta}^{\tiny GLS}-\bar J\theta)
\overset{d}{\to}
\mathcal{N} ( 0_{r+2},
\Omega_1^{-1}
),\nonumber
\end{equation}
where $\Omega_1$ is defined in Assumptions (ref)(ref) and
$\bar J$ is defined in (ref).
For observed network factors, the GLS estimator has a faster rate of convergence than the OLS estimator and it is more efficient; see also chenetal2020cnar. The different rates depend on the different scaling needed for Assumptions (ref)(ref) and (ref)(ref) to hold. Indeed, on the one hand $\mathbb{E} \left[ X_t' V X_t \right]=O(N^2)$, while, on the other hand $\mathbb{E} \left[ X_t' V^{-1} X_t \right]=O(N)$. This is because, by Assumption (ref), $\norm{V}=O(N)$ but $\norm{V^{-1}}=O(1)$.
Now, from Proposition (ref), we see that the conditions ${ \sqrt {NT}}/ m \to 0$ and $N/m\to 0$ allow us to treat the factors as observed. These conditions are stronger than in the OLS case.
An estimator of the asymptotic variance-covariance matrix of the GLS estimator is then given by:
\[
\widehat{\operatorname*{Avar}}\left[ \sqrt{NT} (\widehat{\theta}^{\text{\upshape\tiny GLS}} -\bar J\theta) \right]
=
\left(\frac{1}{NT} \sum_{t=1}^{T} \widehat{X}_{t}^' \widehat{V}^{-1} \widehat{X}_t \right)^{-1},
\]
where $ \widehat{V}^{-1}$ is defined in (ref). Alternatively, to address possible residual cross-correlation of the node idiosyncratic components, we could use:
\[
\widehat{\operatorname*{Avar}}\left[ \sqrt{NT} (\widehat{\theta}^{\text{\upshape\tiny GLS}} -\bar J\theta) \right]
=
\left(\frac{1}{NT} \sum_{t=1}^{T} \widehat{X}_{t}^' \widehat{V}^{-1} \widehat{X}_t \right)^{-1}
\left(\frac{1}{NT} \sum_{t=1}^{T} \widehat{X}_{t}^' \widehat{V}^{-1} \widehat\nu_t\widehat\nu_t' \widehat{V}^{-1} \widehat{X}_t \right)
\left(\frac{1}{NT} \sum_{t=1}^{T} \widehat{X}_{t}^' \widehat{V}^{-1} \widehat{X}_t \right)^{-1}.
\]
\refstepcounter{lettersection}
\oldsection{Proofs of the main results}
General statements of Assumptions (ref),
(ref),
(ref),
and (ref)
All the following results in Appendices (ref) and (ref) are proved under a more general version of the assumptions in the main text. Namely, we replace Assumptions
(ref)(ref),
(ref)(ref),
(ref)(ref),
(ref),
(ref)(ref),
and (ref) with:
Assumption 2 (General statement).
{\it
enumerate• \vskip -.3cm
For all $N \in \mathbb{N}$, all $t,s \in \mathbb{Z}$, and all $i,j=1,\ldots N^2$,
$
{N^{-\gamma}}
\sum_{h=1}^{N^2} \sum_{k=1}^{N^2}
\abs{
\mathbb{E} \left[ \mathcal{E}_{(3) t i h} \mathcal{E}_{(3) s j k} \right]
}
\leq \rho_{\mathcal{E}}^{\abs{t-s}} M_{ij}
$
and, for all $N \in \mathbb{N}$, all $t,s \in \mathbb{Z}$, and all $i,j,k=1, \dots, N^2$,
$
{N^{-\gamma}}
\sum_{h=1}^{N^2}
\abs{
\mathbb{E} \left[ \mathcal{E}_{(3) t i h} \mathcal{E}_{(3) s j k} \right]
}
\leq \rho_{\mathcal{E}}^{\abs{t-s}} M_{ij}
$
for some $\gamma \in [0,2]$ and some finite $\rho_{\mathcal{E}}$ and $M_{ij}$ independent of $t,s,k$ and $N$ such that
$0 \leq \rho_{\mathcal{E}} <1 $, $\sum_{i=1, i\ne j}^m M_{ij} \leq M_{\mathcal{E}}$ and $\sum_{j=1,j\ne i}^m M_{ij} \leq M_{\mathcal{E}}$, for some finite $ M_{\mathcal{E}}$ independent of $i,j$ and $m$.
•
For all $m,T, N \in \mathbb{N}$ and all $j=1,\ldots, N^2$ and all $s=1,\ldots, T$,
\begin{equation*}
\mathbb{E} \left[ \abs{ \frac{1}{\sqrt{mT} N^{\gamma} } \sum_{i=1}^{m} \sum_{t=1}^{T} \sum_{h=1}^{N^2} \sum_{k=1}^{N^2}
\left\{
\mathcal{E}_{(3)t i h} \mathcal{E}_{(3)t j k}
-
\mathbb{E} \left[ \mathcal{E}_{(3)t i h} \mathcal{E}_{(3)t j k} \right]
\right\} }^2 \right]
\leq C_{\mathcal{E}}
\end{equation*}
and
\begin{equation*}
\mathbb{E} \left[ \abs{ \frac{1}{\sqrt{mT} N^{\gamma} } \sum_{i=1}^{m} \sum_{t=1}^{T} \sum_{h=1}^{N^2} \sum_{k=1}^{N^2}
\left\{
\mathcal{E}_{(3)t i h} \mathcal{E}_{(3)s i k}
-
\mathbb{E} \left[ \mathcal{E}_{(3)t i h} \mathcal{E}_{(3)s i k} \right]
\right\} }^2 \right]
\leq C_{\mathcal{E}}
\end{equation*}
for some finite $C_{\mathcal{E}}$ independent of $j, s, m, T, N$ and some $\gamma \in [0,2]$.
•
For all $m,T, N \in \mathbb{N}$ and all $j=1,\ldots, N^2$ and all $s=1,\ldots, T$,
\begin{equation*}
\mathbb{E} \left[ \abs{ \frac{1}{\sqrt{mT} N^{\gamma/2} } \sum_{i=1}^{m} \sum_{t=1}^{T} \sum_{h=1}^{N^2}
\left\{
\mathcal{E}_{(3)t i h} \mathcal{E}_{(3)t j h}
-
\mathbb{E} \left[ \mathcal{E}_{(3)t i h} \mathcal{E}_{(3)t j h} \right]
\right\} }^2 \right]
\leq C_{\mathcal{E}}
\end{equation*}
and
\begin{equation*}
\mathbb{E} \left[ \abs{ \frac{1}{\sqrt{mT} N^{\gamma/2} } \sum_{i=1}^{m} \sum_{t=1}^{T} \sum_{h=1}^{N^2}
\left\{
\mathcal{E}_{(3)t i h} \mathcal{E}_{(3)s i h}
-
\mathbb{E} \left[ \mathcal{E}_{(3)t i h } \mathcal{E}_{(3)s i h} \right]
\right\} }^2 \right]
\leq C_{\mathcal{E}}
\end{equation*}
for some finite $C_{\mathcal{E}}$ independent of $j, s, m, T, N$ and some $\gamma \in [0,2]$.
}
Assumption 3 (General statement).
{\it
For all $i,j\in\mathbb N$, all $k=1,\ldots,r$, and all $t\in\mathbb Z$, $\mathbb E[\mathcal{F}_{(3)tkj}\mathcal{E}_{(3)tij}]=0$,
and, for all $m,N,T\in\mathbb N$ and all $t=1,\ldots, T$,
align[align omitted — 563 chars of source]
for some finite $C_{\mathcal{F}\mathcal{E}}$ and $C_{\mathcal{F}\mathcal{E}}^\prime$ independent of $t$, $m, N$, and $T$ and some $\gamma \in [0,2]$.}
Assumption 5 (General statement).
{\it
enumerate• \vskip -.3cm For any given $i=1,\ldots, m$,
as $N,T\to\infty$,
$
\frac{1}{\sqrt {N^\gamma T}} \sum_{t=1}^{T} \mathcal{F}_{(3)t} \mathcal{E}_{(3)ti\cdot}' \overset{d}{\to}\mathcal N(0_r,\Phi_i),
$
where
$
\Phi_i:=\lim_{N,T\to\infty} \mathbb E\left[\left (\frac 1{\sqrt {N^{\gamma} T }}\sum_{t=1}^{T} \mathcal{F}_{(3)t} \mathcal{E}_{(3)t i \cdot}' \right)
\left(\frac 1{\sqrt {N^{\gamma} T }}\sum_{t=1}^{T} \mathcal{F}_{(3)t} \mathcal{E}_{(3)t i \cdot}' \right)'
\right],
$
for some $\gamma \in [0,2]$.
}
Assumption 10 (General statement).
{\it For all $m,N,T\in\mathbb N$,
align[align omitted — 359 chars of source]
for some finite $\mathfrak K_1$ and $\mathfrak K_2$ independent of $m,N$, and $T$ and some $\gamma\in[0,2]$.}
These assumptions all depend on a generic $\gamma\in[0,2]$, thus are more general than those stated in the main text, which correspond to the case $\gamma=2$. All Theorems and Propositions in the main text and in Appendices (ref) and (ref) are stated in the case $\gamma=2$, which is the least favorable one, meaning it is the case giving the slowest possible convergence rates.
Proof of Proposition (ref)
proofFirst, let $A_{t}:= N^{-1}\sum_{j=1}^r\beta_j F_{j,t-1}+\rho$ and
notice that (ref) can be rewritten as
\begin{equation}
y_t = \alpha + A_{t} y_{t-1}+\nu_t.
\end{equation}
Thus, for any given $N\in\mathbb N$,
letting $y_{-\infty}=0$, if there exists a causal solution it is given by:
\begin{align}
y_t
&=\left\{\prod_{k=0}^\ell(\alpha+A_{t-k})\right\}y_{t-\ell-1}+\sum_{j=1}^\ell\left\{\prod_{k=0}^{j-1}(\alpha+A_{t-k}) \right\}\nu_{t-j}+\nu_t
\nonumber\\
&=\sum_{j=1}^\infty\left\{\prod_{k=0}^{j-1}(\alpha+A_{t-k}) \right\}\nu_{t-j}+\nu_t.
\end{align}
To this end first notice that since by Assumption (ref), $\{\mathcal F_t\}$ is independent of $\{\nu_t\}$, for any given $N\in\mathbb N$,
\[
\mathbb E[y_t]= \alpha+\rho \mathbb E[ y_{t-1}]+\frac 1N \sum_{j=1}^r\beta_j\mathbb E[F_{j,t-1}] \mathbb E[ y_{t-1}]
\]
hence, a stationary solution must have mean
\[
\mathbb E[y_t]=\left(I_N-\rho I_N- \frac 1N\sum_{j=1}^r\beta_j\mathbb E[F_{j,t}]\right)^{-1}\alpha,
\]
which is finite and independent of $t$ because of Assumptions (ref)(ref) and (ref)(ref). Clearly if $\alpha=0$ then $ \mathbb E[y_t]=0$ and vice versa.
Let then $\alpha=0$ for simplicity and define $\Sigma_{t,s} := \mathbb E[y_ty_s']$ and recall that $V:=\mathbb E[\nu_t\nu_t']$, then
\begin{align}
vec(\Sigma_{t,t})&= \rho^2 vec(\Sigma_{t-1,t-1}) +\frac 1{N^{2}}\sum_{j=1}^r\beta_j^2 \mathbb E [ F_{j,t-1}\otimes F_{j,t-1}] vec(\Sigma_{t-1,t-1})+ vec(V)\nonumber\\
&= \left\{\rho^2 I_{N^2}+ \frac1 {N^2}\sum_{j=1}^r \beta_j^2 \mathbb E [ F_{j,t-1}\otimes F_{j,t-1}]\right\} vec(\Sigma_{t-1,t-1})+ vec(V)\nonumber\\
& = \sum_{k=0}^\ell \left\{\rho^2 I_{N^2}+ \frac1 {N^2}\sum_{j=1}^r\beta_j^2 \mathbb E [ F_{j,t-1}\otimes F_{j,t-1}]\right\}^k \text{vec}(V)\nonumber\\
&+ \left\{\rho^2 I_{N^2}+ \frac 1{N^2}\sum_{j=1}^r\beta_j^2 \mathbb E [ F_{j,t-1}\otimes F_{j,t-1}]\right\}^{\ell+1}\text{vec}(\Sigma_{t-\ell-1,t-\ell-1}).
\end{align}
Notice also that $V$ is positive definite, indeed its smallest eigenvalue is such that, by Weyl's inequality,
\[
\mu_N(V) \ge \mu_N(\Lambda\Lambda')+\mu_N(S) = \min_{i=1,\ldots, N}\mathbb E[\epsilon_{it}^2]\ge \underline M_{\epsilon},
\]
because of Assumption (ref)(ref) and where $\mu_N(\Lambda\Lambda')$, $\mu_N(V)$, and $\mu_N(S)$ are the smallest eigenvalues of $\Lambda\Lambda'$, $V$, and $S$, respectively, and $\mu_N(\Lambda\Lambda')=0$.
Hence, letting $\ell\to\infty$ we see that $\text{vec}(\Sigma_{t,t})$ is finite and independent of $t$, because of Assumptions (ref)(ref) and (ref)(ref), and since $V$ is positive definite, see also Theorem 2.1 in nicholls1981multiple.
To show that the above arguments imply that also $y_t^\omega=\lim_{N\to\infty} \omega'y_t$ exists almost surely and is stationary and causal it is enough to follow the same steps as in Theorem 2 by zhuetal17.
Proof of Theorem (ref)
proofThe proof of part (ref) follows directly from Lemmas (ref)(ref) and (ref).
For part (ref), from (ref) in the proof of Lemma (ref)(ref) and Lemma (ref)(ref), if $N^{\gamma/2}\sqrt T/m \to 0$,
as $m,N,T\to\infty$, we have
\begin{align}
N^{2-\gamma/2}\sqrt{T}\left( \widehat{u}_i' - u_i' J \right) &=
\frac{N^{2-\gamma/2}\sqrt{T}}{N^2 T} \sum_{t=1}^{T} \mathcal{E}_{(3)ti\cdot} \mathcal{F}_{(3)t}' \left(\frac 1m \sum_{j=1}^{m} u_j u_j'
\right)
\left( \frac{\widehat{M}^{\mathcal{W}}}{m N^2} \right)^{-1}J +o_p(1)\nonumber\\
&=\frac 1{\sqrt{N^{\gamma}T}}\sum_{t=1}^{T} \mathcal{E}_{(3)ti\cdot} \mathcal{F}_{(3)t}' \Sigma_U
\left( \frac{{M}^{\chi}}{m N^2} \right)^{-1}J +o_p(1)\nonumber\\
&=\frac 1{\sqrt{N^{\gamma}T}}\sum_{t=1}^{T} \mathcal{E}_{(3)ti\cdot} \mathcal{F}_{(3)t}' J +o_p(1)\nonumber\\
&\overset{d}{\to}\mathcal N(0_r, J_0\Phi_i J_0),\nonumber
\end{align}
because of Assumption (ref)(ref) and Lemma (ref)(ref), Assumption (ref)(ref), and Slutsky's Theorem, and where $J_0=\operatorname*{plim}_{m,N,T\to\infty} J$. Notice that $J\Phi_i J=\Phi_i$. This proves part (ref).
The final statement follows by setting $\gamma=2$.
Proof of Theorem (ref)
proofFor part (ref), consider the estimated factors $\widehat{\mathcal{F}}_{(3)}$. We have that:
\begin{align}
\widehat{\mathcal{F}}_{t} &= \mathcal{W}_{t} \times_3 \left( \widehat{U}'\widehat{U} \right)^{-1} \widehat{U}' \nonumber\\
&= \left( \mathcal{F}_{t} \times_3 U + \mathcal{E}_t \right) \times_3 \left( \widehat{U}'\widehat{U} \right)^{-1} \widehat{U}' \nonumber \\
&= \mathcal{F}_{t} \times_3 \left( \widehat{U}'\widehat{U} \right)^{-1} \widehat{U}' \left( U - \widehat{U}J + \widehat{U} J \right) \nonumber
+ \mathcal{E}_t \times_3 \left( \widehat{U}'\widehat{U} \right)^{-1} \left( \widehat{U} - \widehat{U}J + \widehat{U} J \right)',\nonumber
\end{align}
and
\begin{align}
\widehat{\mathcal{F}}_{(3)t} &=
\left( \widehat{U}'\widehat{U} \right)^{-1}
\left[
\widehat{U}' \left(U - \widehat{U} J + \widehat{U} J \right) \mathcal{F}_{(3)t}
\right]
+ \left(\widehat{U} - UJ +UJ \right)' \mathcal{E}_{(3)t}\nonumber \\
&= \left( \frac{\widehat{M}^{\mathcal{W}}}{m N^2} \right)^{-1}
\left[
\frac{ \widehat{U}' \left(U - \widehat{U} J \right) }{m} \mathcal{F}_{(3)t}
+ \frac{ \widehat{U}'\widehat{U}}{m} J \mathcal{F}_{(3)t}
+ \frac{ \left(\widehat{U} - U J \right)' }{m} \mathcal{E}_{(3)t}
+ \frac{ J U' \mathcal{E}_{(3)t}}{m}
\right].\nonumber
\end{align}
Thus, because of Lemma (ref)(ref),
\begin{align}
\frac{\widehat{\mathcal{F}}_{(3)t} - J \mathcal{F}_{(3)t}}{N}
&= \left( \frac{\widehat{M}^{\mathcal{W}}}{m N^2} \right)^{-1}
\left[
\frac{ \widehat{U}' \left(U - \widehat{U} J \right) }{mN} \mathcal{F}_{(3)t}
+ \frac{ \left(\widehat{U} - U J \right)' }{mN} \mathcal{E}_{(3)t}
+ \frac{ J U' \mathcal{E}_{(3)t}}{mN}
\right] \nonumber \\
&= O_p(1) \left[ A + B + C \right].
\end{align}
For term (A), given (ref) in the proof of Lemma (ref) when $\widehat{H}=J$ and Lemma (ref)(ref)
\begin{equation}
\norm{\frac{ \widehat{U}' \left(U - \widehat{U} J \right) }{mN} \mathcal{F}_{(3)t}}
\leq
\norm{ \frac{ \widehat{U}' \left(U - \widehat{U} J \right) }{m} }\
\norm{\frac{\mathcal{F}_{(3)t}}{N}}
= O_p \left(\frac{1}{\xi}\right),
\end{equation}
where $\xi =
\min \left(
\sqrt{mT} N^{2-\gamma/2},
N^{3-\gamma/2} T,
m \sqrt{T} N^{2-\gamma},
m N^{2-\gamma}
\right)
$.
For term (B), from (ref) in the proof of Lemma (ref)(ref) with $\widehat{H}=J$, we have
\begin{align}
\frac{ \left(\widehat{U} - U J \right)' \mathcal{E}_{(3)t} }{mN}
=& \left(\frac{\widehat{M}^{\mathcal{W}}}{mN^2}\right)^{-1} J
\frac{U'\mathcal{E}_{(3)} \mathcal{F}_{(3)}' U' \mathcal{E}_{(3)t} }{m^2N^3T}
+
\left(\frac{\widehat{M}^{\mathcal{W}}}{mN^2}\right)^{-1} J
\frac{U'U \mathcal{F}_{(3)} \mathcal{E}_{(3)}' \mathcal{E}_{(3)t} }{m^2N^3T}\nonumber\\
&+ \left(\frac{\widehat{M}^{\mathcal{W}}}{mN^2}\right)^{-1} J
\frac{U' \mathcal{E}_{(3)} \mathcal{E}_{(3)}' \mathcal{E}_{(3)t} }{m^2N^3T}\nonumber\\
&+
\left(\frac{\widehat{M}^{\mathcal{W}}}{mN^2}\right)^{-1} (\widehat{U} - UJ)'
\frac{ \mathcal{E}_{(3)} \mathcal{F}_{(3)}' U' \mathcal{E}_{(3)t} }{m^2N^3T}
+
\left(\frac{\widehat{M}^{\mathcal{W}}}{mN^2}\right)^{-1} (\widehat{U} - UJ)'
\frac{ U \mathcal{F}_{(3)} \mathcal{E}_{(3)}' \mathcal{E}_{(3)t} }{m^2N^3T}\nonumber\\
&+
\left(\frac{\widehat{M}^{\mathcal{W}}}{mN^2}\right)^{-1} (\widehat{U} - UJ)'
\frac{ \mathcal{E}_{(3)} \mathcal{E}_{(3)}' \mathcal{E}_{(3)t} }{m^2N^3T}\nonumber\\
=& III_a + III_b + III_c + III_d + III_e + III_f.\nonumber
\end{align}
Then, because of Lemma (ref)(ref), (ref)(ref), and (ref)(ref), and using (ref) and $\norm{J}=O(1)$,
\begin{align}
\norm{III_a}
&\leq
\norm{\left( \frac{\widehat{M}^{\mathcal{W}}}{mN^2} \right)^{-1}} \
\norm{J} \
\norm{ \frac{U' \mathcal{E}_{(3)} \mathcal{F}_{(3)}'}{m N^2 T} } \
\norm{\frac{U}{\sqrt{m}}} \
\norm{\frac{\mathcal{E}_{(3)t}}{\sqrt{m}N}}\nonumber \\
&= O_p(1) O(1) O_p\left(\frac{1}{\sqrt{mT}N^{2-\gamma/2}}\right)
O_p(1) O\left(\frac{1}{N^{1-\gamma/2}}\right)\nonumber
\\
&= O \left( \frac{1}{\sqrt{mT}N^{3-\gamma}} \right).\nonumber
\end{align}
Turning to $III_b$ we have
\begin{equation}
\norm{\frac{\mathcal{F}_{(3)}\mathcal{E}_{(3)}' \mathcal{E}_{(3)t} }{m N^3 T}}\le \norm{\frac{\mathcal{F}_{(3)}\left\{\mathcal{E}_{(3)}' \mathcal{E}_{(3)t}-\mathbb {E}\left[\mathcal{E}_{(3)}' \mathcal{E}_{(3)t}\right]\right\}}{mN^3 T}}+\norm{\frac{\mathcal{F}_{(3)}\mathbb{E}\left[\mathcal{E}_{(3)}' \mathcal{E}_{(3)t}\right] }{m N^3 T}}.
\end{equation}
By Assumption (ref), we have
\begin{align}
\mathbb{E}&\left[
\norm{\frac{\mathcal{F}_{(3)}\left\{\mathcal{E}_{(3)}' \mathcal{E}_{(3)t}-\mathbb {E}\left[\mathcal{E}_{(3)}' \mathcal{E}_{(3)t}\right]\right\}}{mN^3 T}}^2
\right]\le \mathbb{E}\left[
\norm{\frac{\mathcal{F}_{(3)}\left\{\mathcal{E}_{(3)}' \mathcal{E}_{(3)t}-\mathbb {E}\left[\mathcal{E}_{(3)}' \mathcal{E}_{(3)t}\right]\right\}}{mN^3 T}}^2_F\right]\nonumber\\
&\le \frac{1}{m^2 N^6 T^2}\mathbb E\left[\left\Vert\sum_{i=1}^{m} \sum_{s=1}^{T}\mathcal{F}_{(3)s}\left\{\mathcal{E}_{(3)s i \cdot}^\prime \mathcal{E}_{(3)t i \cdot}-\mathbb E[\mathcal{E}_{(3)s i \cdot}^\prime \mathcal{E}_{(3)t i \cdot}]\right\} \right\Vert^2\right]\le \frac{C_{\mathcal{F}\mathcal{E}}^\prime }{m N^{6-2\gamma}T}.
\end{align}
Moreover, by Assumption (ref)(ref) and Lemma (ref)(ref),
\begin{align}
\mathbb E&\left[\norm{\frac{\mathcal{F}_{(3)}\mathbb{E}\left[\mathcal{E}_{(3)}' \mathcal{E}_{(3)t}\right] }{m N^3 T}}\right]\le \mathbb E\left[\norm{\frac{\mathcal{F}_{(3)}\mathbb{E}\left[\mathcal{E}_{(3)}' \mathcal{E}_{(3)t}\right] }{m N^3 T}}_F\right]\nonumber\\
&=\frac 1{mN^3 T} \mathbb E\left[\sqrt{ \sum_{l=1}^r \sum_{k=1}^{N^2} \left(\sum_{s=1}^T\sum_{h=1}^{N^2}\sum_{i=1}^m
\mathcal{F}_{(3)slh} \mathbb E\left[\mathcal{E}_{(3)sih}\mathcal{E}_{(3)tik}\right]
\right)^2 }\right]\nonumber\\
&\le\frac 1{mN^3 T} \mathbb E\left[ \sum_{l=1}^r \sum_{k=1}^{N^2} \left\vert \sum_{s=1}^T\sum_{h=1}^{N^2}\sum_{i=1}^m
\mathcal{F}_{(3)slh} \mathbb E\left[\mathcal{E}_{(3)sih}\mathcal{E}_{(3)tik}\right]
\right\vert \right]\nonumber\\
&\le\frac 1{mN^3 T} \sum_{l=1}^r \sum_{k=1}^{N^2} \sum_{s=1}^T\sum_{h=1}^{N^2}\sum_{i=1}^m \mathbb E\left[\left\vert
\mathcal{F}_{(3)slh} \mathbb E\left[\mathcal{E}_{(3)sih}\mathcal{E}_{(3)tik}\right]
\right\vert \right]\nonumber\\
&\le\frac r{mN^3 T} \max_{l=1,\ldots,r} \max_{s=1,\ldots, T}\max_{h=1,\ldots, N^2} \mathbb E\left[\left\vert
\mathcal{F}_{(3)slh}\right\vert\right] \sum_{k=1}^{N^2} \sum_{s=1}^T\sum_{h=1}^{N^2}\sum_{i=1}^m
\left\vert\mathbb E\left[\mathcal{E}_{(3)sih}\mathcal{E}_{(3)tik}\right]
\right\vert \nonumber\\
&\le \frac {r C_{\mathcal F}^\prime N^\gamma m }{mN^3 T}.
\end{align}
From (ref), (ref), and (ref),
\begin{align}
\norm{III_b}
&\leq
\norm{\left(\frac{\widehat{M}^{\mathcal{W}}}{mN^2}\right)^{-1}} \
\norm{J} \
\norm{\frac{U'U }{m}} \
\norm{\frac{\mathcal{F}_{(3)}\mathcal{E}_{(3)}' \mathcal{E}_{(3)t} }{m N^3 T}}
= O_p\left( \frac{1}{ \sqrt{mT} N^{3-\gamma} }\right)+O_p\left( \frac{1}{ T N^{3-\gamma} }\right).\nonumber
\end{align}
Because of Lemma (ref)(ref), (ref)(ref), and Lemma (ref)(ref), (ref)(iv), and since $\norm{J}=O(1)$,
\begin{align}
\norm{III_c}
&\leq
\norm{\left(\frac{\widehat{M}^{\mathcal{W}}}{mN^2}\right)^{-1}} \
\norm{J} \
\norm{\frac{U \mathcal{E}_{(3)} \mathcal{E}_{(3)}' }{m^{3/2} N^2T}}
\norm{\frac{\mathcal{E}_{(3)t} }{\sqrt{m} N}}
= O_p \left( \max \left( \frac{1}{\sqrt{mT}N^{3-\gamma}}, \frac{1}{mN^{3(1-\gamma/2)}} \right) \right).\nonumber
\end{align}
Next, notice that, by Assumption (ref)(ref)
and Lemma (ref)(ref)
\begin{align}
\mathbb{E} \left[
\norm{\frac{U' \mathcal{E}_{(3)t}}{mN}}^2
\right]
&\leq
\mathbb{E} \left[
\norm{\frac{U' \mathcal{E}_{(3)t}}{mN}}^2_F
\right]
=
\frac{1}{m^2 N^2} \sum_{k=1}^{r}
\mathbb{E} \left[
\norm{u_k' \mathcal{E}_{(3)t}}^2
\right] \nonumber\\
&= \frac{1}{m^2 N^2} \sum_{k=1}^{r}
\mathbb{E} \left[
\left(\sum_{i=1}^{m} \sum_{h=1}^{N^2} U_{ik} \mathcal{E}_{(3)t i h}
\right)^2
\right]\nonumber \\
&\leq
\frac{r}{m^2 N^2}
\max_{k=1, \dots, r}
\sum_{i=1}^{m} \sum_{j=1}^{m}
\sum_{h=1}^{N^2} \sum_{l=1}^{N^2}
\abs{ U_{ik} } \abs{ U_{jk} }
\abs{\mathbb{E}\left[ \mathcal{E}_{(3)t i h} \mathcal{E}_{(3)t j l} \right]}\nonumber \\
&\leq
\frac{r M_U^2}{m^2 N^2}
\sum_{i=1}^{m} \sum_{j=1}^{m}
\sum_{h=1}^{N^2} \sum_{l=1}^{N^2}
\abs{\mathbb{E}\left[ \mathcal{E}_{(3)t i h}
\mathcal{E}_{(3)t j l} \right]}
\leq
\frac{r M_U^2 M_{\mathcal{E}}}{m N^{2-\gamma}}
\end{align}
since $M_{\mathcal{E}}$ does not depend on $t$.
Therefore, by Theorem (ref)(ref), Lemma (ref)(ref), (ref)(ref), and
using (ref)
and $\norm{J}=O(1)$
\begin{align}
\norm{III_d}
&\leq
\norm{\left(\frac{\widehat{M}^{\mathcal{W}}}{mN^2}\right)^{-1}} \
\norm{\frac{\widehat{U} - UJ}{\sqrt{m}}} \
\norm{\frac{ \mathcal{E}_{(3)} \mathcal{F}_{(3)}'}{\sqrt{m}N^2T}}
\norm{\frac{ U' \mathcal{E}_{(3)t} }{m N }}\nonumber\\
&= O_p(1)
O_p \left(\max \left( \frac{1}{N^{2-\gamma/2} \sqrt{T}}, \frac{1}{ N^{2-\gamma} m} \right) \right)
O_p \left(\frac{1}{\sqrt{T}N^{2-\gamma/2}}\right)
O_p \left(\frac{1}{\sqrt{m}N^{1-\gamma/2}}\right) \nonumber\\
&=
O_p \left(\max \left( \frac{1}{N^{5-3\gamma/2} T \sqrt{m}}, \frac{1}{ N^{5-2\gamma} m^{3/2} \sqrt{T}} \right) \right) .\nonumber
\end{align}
Note that the term $III_d$ is clearly dominated by $III_a$. Analogously, $III_e$ and $III_f$ are dominated by $III_b$ and $III_c$, respectively. Thus, (B) is $O_p(1/(\sqrt{mT}N^{3-\gamma}, mN^{3-3\gamma/2}, TN^{3-\gamma}))$, hence it is dominated by term (A).
For term (C), using (ref)
\begin{equation}
\norm{\frac{ J U' \mathcal{E}_{(3)t} }{mN}}
\leq
\norm{ J} \norm{\frac{U' \mathcal{E}_{(3)t}}{mN}}
=O_p \left(\frac{1}{\sqrt{m}N^{1-\gamma/2}}\right).
\end{equation}
By noticing that
\[
\max\left(\frac 1\xi, \frac{1}{\sqrt{m}N^{1-\gamma/2}}\right)=\max\left(\frac 1{N^{3-\gamma/2} T},\frac 1{\sqrt{m}N^{1-\gamma/2}}\right),
\]
we prove part (ref).
For part (ref), for any given $j=1,\ldots, N^2$, following the same reasoning as in part (i), we have
\[
{\widehat{\mathcal{F}}_{(3)t\cdot j} - J \mathcal{F}_{(3)t\cdot j}} =N \left(\frac{\widehat{M}^{\mathcal{W}}}{m N^2} \right)^{-1}J \left[\frac{U' \mathcal{E}_{(3)t\cdot j}}{mN}+ O_p\left(\frac 1{N^3T}\right)\right].
\]
Moreover,
\[
\norm{\frac{U' \mathcal{E}_{(3)t\cdot j}}{m}}=O_p\left(\frac 1{\sqrt m}\right).
\]
Therefore, if $\sqrt m/(N^2T)\to 0$ as $m,N,T\to\infty$,
\begin{align}
\sqrt{m} \left({\widehat{\mathcal{F}}_{(3)t\cdot j} - J \mathcal{F}_{(3)t\cdot j}} \right)
&=\left(\frac{\widehat{M}^{\mathcal{W}}}{m N^2} \right)^{-1} J
\left[\frac{U' \mathcal{E}_{(3)t\cdot j}}{\sqrt m}\right] +o_p(1)\nonumber\\
&=
\Sigma_U^{-1} J
\left[ \frac{ 1}{\sqrt{m}}\sum_{i=1}^m u_i \mathcal E_{(3)tij} \right] + o_p(1)\nonumber\\
&\overset{d}{\to}\mathcal N\left (0_{r}, \Sigma_U^{-1}J_0\Pi_{tj} J_0\Sigma_U^{-1}\right),\nonumber
\end{align}
because of Assumption (ref)(ref), Lemma (ref)(ref),
Assumption (ref)(ref), and Slutsky's theorem, and $J_0=\operatorname*{plim}_{m,N,T\to\infty} J$. Notice that $ \Sigma_U^{-1}J_0\Pi_{tj} J_0\Sigma_U^{-1}= \Sigma_U^{-1}\Pi_{tj}\Sigma_U^{-1}$. This proves part (ref).
Proof of Theorem (ref)
proofFirst, notice that
\begin{equation}
\norm{M_{\widehat G^\dag}} = \norm{I_T-\widehat G^\dag\widehat G^{\dag'}/T}= O(1),\qquad \norm{M_{\widehat \Lambda^*}} = \norm{I_N-\widehat V^{\widehat\nu^*}\widehat V^{\widehat{\nu}^{*'}} }= O(1)
\end{equation}
since $\norm{\widehat G^\dag}=O_p(\sqrt T)$ and eigenvectors are normalized.
Then, because of (ref) and by the same arguments used in the proof of Proposition (ref)(ref), we have
\begin{align}
\norm{\frac 1{NT}\sum_{i=1}^N\widehat{X}_i'M_{\widehat G^\dag} u_i }=&\, O_p \left( \max\left(\frac 1{N^{3-\gamma/2}T},\frac 1{mN^{2-\gamma}},\frac 1{\sqrt {mT}N^{1-\gamma/2}}\right)\right),\\
\norm{\frac 1{NT}\sum_{t=1}^T\widehat{X}_t'M_{\widehat \Lambda^*} u_t }=&\, O_p \left( \max\left(\frac 1{N^{3-\gamma/2}T},\frac 1{mN^{2-\gamma}},\frac 1{\sqrt {mT}N^{1-\gamma/2}}\right)\right).
\end{align}
Moreover,
\begin{align}
&\norm{\frac 1{NT}\sum_{i=1}^N \widehat{X}_{i}' M_{\widehat G^\dag} \widehat X_{i}-\frac 1{NT} \sum_{i=1}^N \bar J \widehat{X}_{i}' M_{\widehat G^\dag} \widehat X_{i}\bar J}= O_p \left( \max\left(\frac 1{N^{3-\gamma/2}T},\frac 1{mN^{2-\gamma}},\frac 1{\sqrt {mT}N^{1-\gamma/2}}\right)\right),\\
&\norm{\frac 1{NT}\sum_{t=1}^T \widehat{X}_{t}' M_{\widehat \Lambda^*} \widehat X_{t}-\frac 1{NT} \sum_{t=1}^T \bar J \widehat{X}_{t}' M_{\widehat \Lambda^*} \widehat X_{t}\bar J}= O_p \left( \max\left(\frac 1{N^{3-\gamma/2}T},\frac 1{mN^{2-\gamma}},\frac 1{\sqrt {mT}N^{1-\gamma/2}}\right)\right),
\end{align}
because of (ref) and following Proposition (ref)(ref).
And also,
\begin{align}
{\frac 1{NT}\sum_{i=1}^N\widehat{X}_{i}'M_{\widehat G^\dag}\, \left[G\Lambda'+\epsilon\right]} =&\, {\frac 1{NT}\sum_{i=1}^N\bar J X_i' M_{\widehat G^\dag} \left\{G\Lambda_i+\epsilon_i\right\}}\nonumber\\
&+{\frac 1{NT}\sum_{i=1}^N\left(\widehat X_i'-\bar J X_i'\right) M_{\widehat G^\dag} \left\{G\Lambda_i+\epsilon_i\right\}} =: A +B,\\
{\frac 1{NT}\sum_{t=1}^T\widehat{X}_{t}'M_{\widehat \Lambda^*}\, \left[\Lambda G_t+\epsilon_t\right]} =&\, {\frac 1{NT}\sum_{t=1}^T\bar J X_t' M_{\widehat \Lambda^*} \left\{\Lambda G_t+\epsilon_t\right\}}\nonumber\\
&+{\frac 1{NT}\sum_{t=1}^T\left(\widehat X_t'-\bar J X_t'\right) M_{\widehat \Lambda^*} \left\{\Lambda G_t+\epsilon_t\right\}} =: C +D,.
\end{align}
Then, because of (ref), (ref), (ref), and (ref)
\begin{align}
\norm {B} &=O_p \left( \max\left(\frac 1{N^{3-\gamma/2}T},\frac 1{mN^{2-\gamma}},\frac 1{\sqrt {mT}N^{1-\gamma/2}}\right)\right),\\
\norm {D} &=O_p \left( \max\left(\frac 1{N^{3-\gamma/2}T},\frac 1{mN^{2-\gamma}},\frac 1{\sqrt {mT}N^{1-\gamma/2}}\right)\right).
\end{align}
Consider first part (ref). If ${ \sqrt {NT}}/ (mN^{2-\gamma}) \to 0$ and ${ \sqrt {N}}/ (\sqrt{m}N^{1-\gamma/2}) \to 0$, as $m,N,T\to\infty$,
by substituting (ref) into (ref), from (ref) we get:
\begin{align}
\sqrt{NT}\left(\widehat\theta^\dag-\bar J\theta\right) =&\, \left(\frac 1{NT} \sum_{i=1}^N \bar J{X}_{i}' M_{\widehat G^\dag} X_{i}\bar J\right)^{-1} \left(\sqrt{NT}
\, A\right) + o_p(1) =: I+o_p(1).
\end{align}
By similar arguments to those used in (ref),
\begin{align}
&\norm{\frac 1{NT} \sum_{i=1}^N \left(y_i-\widehat X_i\widehat\theta^{\dag}\right) \left(y_i-\widehat X_i\widehat\theta^{\dag}\right)'-
\frac 1{NT} \sum_{i=1}^N \left(y_i- X_i\bar J\widehat\theta^{\dag}\right) \left(y_i- X_i\bar J\widehat\theta^{\dag}\right)'}\nonumber\\
&\phantom{cippirimerlo}= O_p \left( \max\left(\frac 1{N^{3-\gamma/2}T},\frac 1{mN^{2-\gamma}},\frac 1{\sqrt {mT}N^{1-\gamma/2}}\right)\right).
\end{align}
Thus, from (ref) and the definition of $\widehat{G}^\dag$ in (ref) it is clear that, if ${ \sqrt {NT}}/ (mN^{2-\gamma}) \to 0$ and ${ \sqrt {N}}/ (\sqrt{m}N^{1-\gamma/2}) \to 0$, as $m,N,T\to\infty$, it solves
\begin{align}
&\left\{\frac 1{NT} \sum_{i=1}^N \left(y_i- X_i\bar J\widehat\theta^{\dag}\right) \left(y_i- X_i\bar J\widehat\theta^{\dag}\right)' + o_p\left (\frac 1{\sqrt {NT}}\right) \right\}\widehat G^\dag = \widehat G^\dag\frac{\widehat M^{\widehat\nu^\dag}}{T},
\end{align}
where $\widehat M^{\widehat\nu^\dag}$ is the $q\times q$ diagonal matrix of eigenvalues of $N^{-1}\widehat \nu^{\dag}\widehat \nu^{\dag'}$.
Now, define the $T\times (r+2)$ matrices:
\begin{align}
\widehat Z_i^\dag &:= \left\{
M_{\widehat G^\dag}X_i-\frac 1N\sum_{k=1}^N\left(\Lambda_i'\left(\frac{\Lambda'\Lambda} N\right)^{-1}\Lambda_k \right)M_{\widehat G^\dag}X_k
\right\},\nonumber\\
Z_i &:= \left\{
M_{ G}X_i-\frac 1N\sum_{k=1}^N\left(\Lambda_i'\left(\frac{\Lambda'\Lambda} N\right)^{-1}\Lambda_k \right)M_{ G}X_k
\right\}.\nonumber
\end{align}
Then, following the same steps as in bai2009panel, if ${ \sqrt {NT}}/ (mN^{2-\gamma}) \to 0$ and ${ \sqrt {N}}/ (\sqrt{m}N^{1-\gamma/2}) \to 0$, as $m,N,T\to\infty$, we have that $I$ in (ref) is such that
\begin{align}
I =&\,\left(\frac 1{NT}\sum_{i=1}^N\bar J\widehat Z_i^{\dag'}\widehat Z_i^\dag\bar J\right)^{-1}
\left\{ \frac 1{\sqrt {NT}}\sum_{i=1}^N
\bar J \widehat Z_i^\dag \epsilon_i\right.\nonumber\\
&\left.-\sqrt{\frac NT}\left[\frac 1{NT}\sum_{i=1}^N \bar J X_i'M_{\widehat G^\dag}\left(\frac 1N \sum_{k=1}^N \mathbb E[\epsilon_k\epsilon_k']\right)\widehat G^\dag \left(\frac{G'\widehat G^\dag}{T}\right)^{-1} \left(\frac{\Lambda'\Lambda}{N}\right)^{-1} \Lambda_i\right]\right\}+ o_p(1)\nonumber\\
=&\,\left(\frac 1{NT}\sum_{i=1}^N\bar J \widehat{Z}_i^{\dag'} \widehat{Z}_i^{\dag}\bar J\right)^{-1}
\left\{ \frac 1{\sqrt {NT}}\sum_{i=1}^N
\bar J Z_i \epsilon_i\right.\nonumber\\
&\left.-\sqrt{\frac NT}\left[\frac 1{NT}\sum_{i=1}^N \bar J X_i'M_{\widehat G^\dag}\left(\frac 1N \sum_{k=1}^N \mathbb E[\epsilon_k\epsilon_k']\right)\widehat G^\dag \left(\frac{G'\widehat G^\dag}{T}\right)^{-1} \left(\frac{\Lambda'\Lambda}{N}\right)^{-1}\Lambda_i\right]\right.\nonumber\\
&\left.- \sqrt {\frac TN} \left[\frac 1{NT}\sum_{i=1}^N\bar J\left(X_i-\frac 1N\sum_{k=1}^N\left(\Lambda_i'\left(\frac{\Lambda'\Lambda}{N}\right)^{-1}\Lambda_k \right) X_k\right)' G \left(\frac{\Lambda'\Lambda}{N}\right)^{-1}\sum_{j=1}^N\Lambda_j \left(\frac 1T\sum_{t=1}^T\epsilon_{jt}\epsilon_{it}\right)\right]\right\}+o_p(1)\nonumber\\
=&\,\left(\frac 1{NT}\sum_{i=1}^N\bar J Z_i' Z_i\bar J\right)^{-1}
\left\{ \frac 1{\sqrt {NT}}\sum_{i=1}^N
\bar J Z_i \epsilon_i\right.\nonumber\\
&\left.-\sqrt{\frac NT}\left[\frac 1{NT}\sum_{i=1}^N \bar J X_i'M_{G}\left(\frac 1N \sum_{k=1}^N \mathbb E[\epsilon_k\epsilon_k']\right)G \left(\frac{\Lambda'\Lambda}{N}\right)^{-1}\Lambda_i\right]\right.\nonumber\\
&\left.- \sqrt {\frac TN} \left[\frac 1{NT}\sum_{i=1}^N\bar J\left(X_i-\frac 1N\sum_{k=1}^N\left(\Lambda_i'\left(\frac{\Lambda'\Lambda}{N}\right)^{-1}\Lambda_k \right) X_k\right)' G\left(\frac{\Lambda'\Lambda}{N}\right)^{-1}\sum_{j=1}^N\Lambda_j \left(\frac 1T\sum_{t=1}^T\mathbb E[\epsilon_{jt}\epsilon_{it}]\right)\right]\right\}\nonumber\\
&+ O_p\left(\frac {\sqrt T}{N}\right)+O_p\left(\frac {\sqrt N}{T}\right)+o_p(1)\nonumber\\
=:&\,\left(\frac 1{NT}\sum_{i=1}^N\bar J Z_i' Z_i\bar J\right)^{-1}
\left\{ \frac 1{\sqrt {NT}}\sum_{i=1}^N
\bar J Z_i \epsilon_i +\sqrt{\frac NT} I_a+\sqrt{\frac TN}I_b \right\}\nonumber\\
&+ O_p\left(\frac {\sqrt T}{N}\right)+O_p\left(\frac {\sqrt N}{T}\right)+o_p(1).
\end{align}
Since by Assumptions (ref)(ref) and (ref)(ref), $\mathbb E[\epsilon_k\epsilon_k'] = \mathbb E[\epsilon_{kt}^2] I_T=\sigma^2_k I_T$ and since $M_{G} G=0_{T\times q}$, we have $I_a=0_{r+2}$. Moreover, since we assumed $T/N\to 0$ and $\sqrt N/T\to 0$, as $N,T\to\infty$, by using (ref) into (ref)
we have
\begin{align}
\sqrt{NT}\left(\widehat\theta^\dag-\bar J\theta\right) &= \left(\frac 1{NT}\sum_{i=1}^N\bar J Z_i' Z_i\bar J\right)^{-1}\left(\frac 1{\sqrt {NT}}\sum_{i=1}^N
\bar JZ_i \epsilon_i\right) + o_p(1)\nonumber\\
&\overset{p}{\to}\mathcal N\left(0_{r+2}, \bar J_0 \Sigma_{ZZ}^{-1} \bar J_0 \bar J_0D_1 \bar J_0\bar J_0\Sigma_{ZZ}^{-1}\bar J_0\right),
\end{align}
because of Assumptions (ref)(ref) and (ref)(ref), and Slutsky's theorem and where $\bar J_0:=\operatorname*{plim}_{m,N,T\to\infty} \bar J$. We complete the proof by noticing that $\bar J_0 \Sigma_{ZZ}^{-1} \bar J_0 \bar J_0D_1 \bar J_0\bar J_0\Sigma_{ZZ}^{-1}\bar J_0:= \Sigma_{ZZ}^{-1}D_1 \Sigma_{ZZ}^{-1}$.
For part (ref), notice that if $\sigma_i^2=\sigma^2$ for all $i=1,\ldots, N$, then in (ref) we have
\begin{align}
I_b=&\, \frac 1{NT}\sum_{i=1}^N\bar JX_i' G\left(\frac{\Lambda'\Lambda}{N}\right)^{-1}\Lambda_i \sigma^2
-\frac 1{NT}\sum_{i=1}^N\bar J\frac 1N\sum_{k=1}^N\left(\Lambda_i'\left(\frac{\Lambda'\Lambda}{N}\right)^{-1}\Lambda_k \right) X_k' G\left(\frac{\Lambda'\Lambda}{N}\right)^{-1}\Lambda_i\sigma^2\nonumber\\
=&\,\frac 1{NT}\sum_{i=1}^N\bar JX_i' G\left(\frac{\Lambda'\Lambda}{N}\right)^{-1}\Lambda_i \sigma^2
-\frac 1{NT}\sum_{k=1}^N\bar J X_k'G\left(\frac{\Lambda'\Lambda}{N}\right)^{-1}\frac 1N\sum_{i=1}^N \Lambda_i\Lambda_i'\left(\frac{\Lambda'\Lambda}{N}\right)^{-1}\Lambda_k \sigma^2\nonumber\\
=&\,\frac 1{NT}\sum_{i=1}^N\bar JX_i' G\left(\frac{\Lambda'\Lambda}{N}\right)^{-1}\Lambda_i \sigma^2
-\frac 1{NT}\sum_{k=1}^N\bar J X_k'G\left(\frac{\Lambda'\Lambda}{N}\right)^{-1}\left(\frac{\Lambda'\Lambda}{N}\right)\left(\frac{\Lambda'\Lambda}{N}\right)^{-1}\Lambda_k \sigma^2\nonumber\\
=&\, 0_{r+2}.\nonumber
\end{align}
The proof then follows as in part (ref) and by noticing that in this case $ \bar J_0 \Sigma_{ZZ}^{-1} \bar J_0 \bar J_0D_1 \bar J_0\bar J_0\Sigma_{ZZ}^{-1}\bar J_0= \sigma^2\Sigma_{ZZ}^{-1}$.
For part (ref). By similar arguments to those used in (ref),
\begin{align}
&\norm{\frac 1{NT} \sum_{t=1}^T \left(y_t-\widehat X_t\widehat\theta^{*}\right) \left(y_t-\widehat X_t\widehat\theta^{*}\right)'-
\frac 1{NT} \sum_{t=1}^T \left(y_t- X_t\bar J\widehat\theta^{*}\right) \left(y_t- X_t\bar J\widehat\theta^{*}\right)'}\nonumber\\
&\phantom{cippirimerlo}= O_p \left( \max\left(\frac 1{N^{3-\gamma/2}T},\frac 1{mN^{2-\gamma}},\frac 1{\sqrt {mT}N^{1-\gamma/2}}\right)\right).
\end{align}
Thus, from (ref) and the definition of $\widehat{\Lambda}^*$ in (ref) it is clear that, if ${ \sqrt {NT}}/ (mN^{2-\gamma}) \to 0$ and ${ \sqrt {N}}/ (\sqrt{m}N^{1-\gamma/2}) \to 0$, as $m,N,T\to\infty$, it solves
\begin{align}
&\left\{\frac 1{NT} \sum_{t=1}^T \left(y_t- X_t\bar J\widehat\theta^{*}\right) \left(y_t- X_t\bar J\widehat\theta^{*}\right)' + o_p\left (\frac 1{\sqrt {NT}}\right) \right\}\widehat \Lambda^* = \widehat \Lambda^*\frac{\widehat M^{\widehat\nu^*}}{N},
\end{align}
where $\widehat M^{\widehat\nu^*}$ is the $q\times q$ diagonal matrix of eigenvalues of $T^{-1}\widehat \nu^{*'}\widehat \nu^{*}$.
Moreover, since $y_t- X_t\bar J\widehat\theta^{*}=X_t\bar J(\bar J\theta-\widehat{\theta}^*)+\Lambda G_t+\epsilon_t$, from (ref) we get
\begin{align}
\widehat \Lambda^*\frac{\widehat M^{\widehat\nu^*}}{N}=&\,\frac 1{NT}\sum_{t=1}^TX_t\bar J(\bar J\theta-\widehat{\theta}^*)(\bar J\theta-\widehat{\theta}^*)'\bar JX_t'\widehat{\Lambda}^*+\frac 1{NT}\sum_{t=1}^TX_t\bar J(\bar J\theta-\widehat{\theta}^*)G_t'\Lambda'\widehat{\Lambda}^*\nonumber\\
&+\frac 1{NT}\sum_{t=1}^TX_t\bar J(\bar J\theta-\widehat{\theta}^*)\epsilon_t'\widehat{\Lambda}^*
+\frac 1{NT}\sum_{t=1}^T\Lambda G_t(\bar J\theta-\widehat{\theta}^*)' \bar JX_t'\widehat{\Lambda}^*
\nonumber\\
&+\frac 1{NT}\sum_{t=1}^T\epsilon_t(\bar J\theta-\widehat{\theta}^*)' \bar JX_t'\widehat{\Lambda}^*+\frac 1{NT}\sum_{t=1}^T\Lambda G_t\epsilon_t'\widehat{\Lambda}^*+\frac 1{NT}\sum_{t=1}^T\epsilon_tG_t'\Lambda'\widehat{\Lambda}^*\nonumber\\
&+\frac 1{NT}\sum_{t=1}^T\epsilon_t\epsilon_t'\widehat{\Lambda}^*+\frac 1{NT}\sum_{t=1}^T \Lambda G_tG_t'\Lambda'\widehat{\Lambda}^*\nonumber\\
=:&\, I_1+I_2+I_3+I_4+I_5+I_6+I_7+I_8+ \Lambda\frac{\Lambda'\widehat{\Lambda}^*}{N} + o_p\left (\frac 1{\sqrt {NT}}\right).\nonumber
\end{align}
since $T^{-1}G'G=I_q$ by Assumption (ref)(ref). So,
\begin{align}
\widehat \Lambda^*\frac{\widehat M^{\widehat\nu^*}}{N}\left(\frac{\Lambda'\widehat{\Lambda}^*}{N}\right)^{-1}-\Lambda=
\sum_{j=1}^8 I_j
\left(\frac{\Lambda'\widehat{\Lambda}^*}{N}\right)^{-1}+o_p\left (\frac 1{\sqrt {NT}}\right),
\end{align}
and notice that
\begin{align}
\left\Vert\left (\frac{\Lambda'\widehat{\Lambda}^*}N\right)^{-1}\right\Vert=O_p(1),
\end{align}
because of (ref) in Lemma (ref) (where $\widehat{\Lambda}^*$ is simply denoted as $\widehat{\Lambda}$) and Assumption (ref)(ref).
It follows that,
\begin{align}
\frac 1{\sqrt N}
\left\Vert
\widehat \Lambda^*\frac{\widehat M^{\widehat\nu^*}}{N}\left(\frac{\Lambda'\widehat{\Lambda}^*}{N}\right)^{-1}-\Lambda
\right\Vert \le \frac 1{\sqrt N}\sum_{j=1}^8\Vert I_j\Vert\,\left\Vert\left(\frac{\Lambda'\widehat{\Lambda}^*}{N}\right)^{-1}\right\Vert +
o_p\left (\frac 1{\sqrt {NT}}\right).
\end{align}
Now,
\begin{align}
\frac 1{\sqrt N}
\left\Vert I_1\right\Vert \le\left\Vert \frac {\widehat \Lambda^*}{\sqrt N}\right\Vert \frac 1T\sum_{t=1}^T\frac{\Vert X_t\Vert ^2}{N}\left\Vert\widehat{\theta}^*-\bar J\theta\right\Vert^2= o_p\left(\Vert\widehat{\theta}^*-\bar J\theta\Vert\right),
\end{align}
because of (ref) in Lemma (ref) and Assumption (ref)(ref), and because
\begin{align}
\mathbb E\left[\left( \frac 1T\sum_{t=1}^T\frac{\Vert X_t\Vert ^2}{N}\right)^2\right]&=\frac 1{T^2N^2}\sum_{t=1}^T\sum_{s=1}^T\mathbb E\left[\Vert X_t\Vert ^2\Vert X_s\Vert ^2\right] \le \frac 1{N^2} \mathbb E\left[\Vert X_t\Vert ^4\right] = \frac 1{N^2}\mathbb E\left[\left( \sum_{i=1}^N X_{it}^2\right)^2\right]\nonumber\\
& = \frac 1{N^2}\sum_{i=1}^N\sum_{j=1}^N \mathbb E[X_{it}^2X_{jt}^2] \le \mathcal K_X,
\end{align}
for some finite $\mathcal K_X$ independent of $N$, $i$, and $t$, due to Assumptions (ref)(ref), (ref)(ref), and (ref)(ref).
Similarly to (ref) we have
\begin{align}
\frac 1{\sqrt N}\left\Vert I_2\right\Vert
= \frac 1{\sqrt N}\left\Vert I_3\right\Vert
= \frac 1{\sqrt N}\left\Vert I_4\right\Vert
= \frac 1{\sqrt N}\left\Vert I_5\right\Vert
=
O_p\left(\Vert\widehat{\theta}^*-\bar J\theta\Vert\right),
\end{align}
Furthermore, by barigozzi2022
\begin{align}
\frac 1{\sqrt N}\left\Vert I_6\right\Vert
= \frac 1{\sqrt N}\left\Vert I_7\right\Vert
= \frac 1{\sqrt N}\left\Vert I_8\right\Vert
=
O_p\left(\max\left(\frac 1{\sqrt N},\frac 1{\sqrt T}\right)\right),
\end{align}
and, by barigozzi2022
\begin{align}
\left\Vert\frac{\widehat M^{\widehat\nu^*}}{N}\left(\frac{\Lambda'\widehat{\Lambda}^*}{N}\right)^{-1}-J\right\Vert = O_p\left(\max\left(\frac 1{\sqrt {NT}},\frac 1{ T}\right)\right).
\end{align}
By substituting (ref), (ref), (ref), (ref), and (ref) into (ref) and since we assumed $\sqrt N/T\to 0$ as $N,T\to\infty$,
\begin{align}
\frac 1{\sqrt N}
\left\Vert
\widehat \Lambda^*-\Lambda J
\right\Vert = O_p\left(\Vert\widehat{\theta}^*-\bar J\theta\Vert\right)+ O_p\left(\max\left(\frac 1{\sqrt N},\frac 1{\sqrt T}\right)\right).
\end{align}
From (ref) and (ref), it follows also that
\begin{align}
\frac 1{NT}\sum_{t=1}^T\bar JX_t' M_{\widehat{\Lambda}^*} (\Lambda-\widehat \Lambda^* J) = O_p\left(\widehat{\theta}^*-\bar J\theta\right)+ O_p\left(\max\left(\frac 1{\sqrt N},\frac 1{\sqrt T}\right)\right).
\end{align}
Now, from (ref) and (ref), and noticing that $M_{\widehat{\Lambda}^*}\widehat{\Lambda}^*=0_{r\times r}$, it follows that:
\begin{align}
\frac 1{NT}\sum_{t=1}^T \bar JX_t'M_{\widehat{\Lambda}^*}\Lambda G_t =&\,\frac 1{NT}\sum_{t=1}^T \bar JX_t'M_{\widehat{\Lambda}^*}(\Lambda-\widehat{\Lambda}^*J) G_t +O_p\left(\max\left(\frac 1{\sqrt {NT}},\frac 1{ T}\right)\right) \nonumber\\
=&\,- \frac 1{NT}\sum_{t=1}^T \bar JX_t'M_{\widehat{\Lambda}^*}
\left\{I_1+I_2+I_3+I_4+I_5+I_6+I_7+I_8\right\}\left(\frac{\Lambda'\widehat{\Lambda}^*}{N}\right)^{-1}G_t\nonumber\\
&+o_p\left (\frac 1{\sqrt {NT}}\right)
+o_p\left(\max\left(\frac 1{\sqrt {NT}},\frac 1{ T}\right)\right)\nonumber\\
=:&\, J_1+J_2+J_3+J_4+J_5+J_6+J_7+J_8+O_p\left(\max\left(\frac 1{\sqrt {NT}},\frac 1{T}\right)\right).
\end{align}
Then, for $J_1$ by (ref) we have
\begin{align}
\Vert J_1\Vert &=\left\Vert- \frac 1{NT}\sum_{t=1}^T \bar JX_t'M_{\widehat{\Lambda}^*}I_1 \left(\frac{\Lambda'\widehat{\Lambda}^*}{N}\right)^{-1}G_t\right\Vert\le \frac 1{T}\sum_{t=1}^T \frac{\Vert \bar JX_t'M_{\widehat{\Lambda}^*}\Vert}{\sqrt N} \,\frac{\Vert I_1\Vert}{\sqrt N}\left\Vert\left(\frac{\Lambda'\widehat{\Lambda}^*}{N}\right)^{-1}\right\Vert\, \Vert G_t\Vert\nonumber\\
&= O_p(1) \left\Vert\widehat{\theta}^*-\bar J\theta\right\Vert^2= o_p(1) \left\Vert\widehat{\theta}^*-\bar J\theta\right\Vert,
\end{align}
since $\Vert J\Vert=1$, $\Vert M_{\widehat{\Lambda}^*}\Vert=1$ (it is a projector), $\Vert G_t\Vert=O_p(1)$ due to Assumption (ref)(ref), and because of (ref) and (ref).
For $J_2$ we have
\begin{align}
J_2&=-\frac 1{NT}\sum_{t=1}^T \bar JX_t'M_{\widehat{\Lambda}^*}I_2 \left(\frac{\Lambda'\widehat{\Lambda}^*}{N}\right)^{-1}G_t\nonumber\\
&= -\frac 1{NT}\sum_{t=1}^T \bar JX_t'M_{\widehat{\Lambda}^*}\left[
\frac 1{T}\sum_{s=1}^TX_s\bar J(\bar J\theta-\widehat{\theta}^*)G_s'\frac{\Lambda'\widehat{\Lambda}^*}N
\right] \left(\frac{\Lambda'\widehat{\Lambda}^*}{N}\right)^{-1}G_t\nonumber\\
&= \frac 1{NT^2}\sum_{t=1}^T\sum_{s=1}^T\bar JX_t'M_{\widehat{\Lambda}^*}X_s\bar J G_s'G_t(\widehat{\theta}^*-\bar J\theta).
\end{align}
For $J_3$ we have
\begin{align}
J_3&=-\frac 1{NT}\sum_{t=1}^T \bar JX_t'M_{\widehat{\Lambda}^*}I_3 \left(\frac{\Lambda'\widehat{\Lambda}^*}{N}\right)^{-1}G_t\nonumber\\
&= -\frac 1{NT}\sum_{t=1}^T \bar JX_t'M_{\widehat{\Lambda}^*}\left[
\frac 1{NT}\sum_{s=1}^TX_s\bar J(\bar J\theta-\widehat{\theta}^*)\epsilon_s'\widehat{\Lambda}^*
\right] \left(\frac{\Lambda'\widehat{\Lambda}^*}{N}\right)^{-1}G_t\nonumber\\
&=\frac 1{NT^2}\sum_{t=1}^T\sum_{s=1}^T\bar JX_t'M_{\widehat{\Lambda}^*}X_s\bar J\left[\frac{\epsilon_s'\widehat{\Lambda}^*}{N}\left(\frac{\Lambda'\widehat{\Lambda}^*}{N}\right)^{-1}\right](\widehat{\theta}^*-\bar J\theta)\nonumber\\
&= o_p(1)(\widehat{\theta}^*-\bar J\theta),
\end{align}
since, by Assumption (ref)(ref) and (ref),
\begin{align}
\frac{\epsilon_s'\widehat{\Lambda}^*}{N}=\frac{\epsilon_s'{\Lambda}J}{N}+\frac{\epsilon_s'(\widehat{\Lambda}^*-\Lambda J)}{N} = O_p\left(\frac 1{\sqrt N}\right)+
O_p\left(\Vert\widehat{\theta}^*-\bar J\theta\Vert\right)+ O_p\left(\max\left(\frac 1{\sqrt N},\frac 1{\sqrt T}\right)\right).\nonumber
\end{align}
For $J_4$, because of (ref), we have
\begin{align}
J_4&=-\frac 1{NT}\sum_{t=1}^T \bar JX_t'M_{\widehat{\Lambda}^*}I_4 \left(\frac{\Lambda'\widehat{\Lambda}^*}{N}\right)^{-1}G_t\nonumber\\
&= - \frac 1{NT}\sum_{t=1}^T \bar JX_t'M_{\widehat{\Lambda}^*}\left[
\frac 1{NT}\sum_{s=1}^T\Lambda G_s(\bar J\theta-\widehat{\theta}^*)'\bar JX_s'\widehat{\Lambda}^*
\right] \left(\frac{\Lambda'\widehat{\Lambda}^*}{N}\right)^{-1}G_t\nonumber\\
&=-\frac 1{\sqrt NT^2}\sum_{t=1}^T\sum_{s=1}^T\bar JX_t'M_{\widehat{\Lambda}^*}\frac{(\Lambda-\widehat{\Lambda}^*J)}{\sqrt N}G_s(\bar J\theta-\widehat{\theta}^*)'\bar J\frac{X_s'\widehat{\Lambda}^*}{N}
\left(\frac{\Lambda'\widehat{\Lambda}^*}{N}\right)^{-1}G_t\nonumber\\
&= o_p(1)(\widehat{\theta}^*-\bar J\theta).
\end{align}
Likewise, because of (ref), for $J_5$ we have
\begin{align}
J_5&=-\frac 1{NT}\sum_{t=1}^T \bar JX_t'M_{\widehat{\Lambda}^*}I_5 \left(\frac{\Lambda'\widehat{\Lambda}^*}{N}\right)^{-1}G_t
= o_p(1)(\widehat{\theta}^*-\bar J\theta).
\end{align}
Then, for $J_6$ we have
\begin{align}
J_6&= -\frac 1{NT}\sum_{t=1}^T \bar JX_t'M_{\widehat{\Lambda}^*}I_6 \left(\frac{\Lambda'\widehat{\Lambda}^*}{N}\right)^{-1}G_t\nonumber\\
&= -\frac 1{NT}\sum_{t=1}^T \bar JX_t'M_{\widehat{\Lambda}^*}\left[\frac 1{NT}\sum_{s=1}^T\Lambda G_s\epsilon_s'\widehat{\Lambda}^*\right] \left(\frac{\Lambda'\widehat{\Lambda}^*}{N}\right)^{-1}G_t\nonumber\\
&= -\frac 1{NT}\sum_{t=1}^T\bar JX_t'M_{\widehat{\Lambda}^*}(\Lambda-\widehat{\Lambda}^*J)\frac 1T \sum_{s=1}^T G_s\frac{\epsilon_s'\widehat{\Lambda}^*}{N}
\left(\frac{\Lambda'\widehat{\Lambda}^*}{N}\right)^{-1}G_t\nonumber\\
&= o_p(\widehat{\theta}^*-\bar J\theta)+o_p\left(\frac 1{\sqrt{NT}}\right),
\end{align}
because of (ref) and since from barigozzi2022 and (ref) we have
\begin{align}
\frac 1T \sum_{s=1}^T G_s\frac{\epsilon_s'\widehat{\Lambda}^*}{N}&=\frac 1T \sum_{s=1}^T G_s\frac{\epsilon_s'{\Lambda}J}{N}+
\frac 1T \sum_{s=1}^T G_s\frac{\epsilon_s'(\widehat{\Lambda}^*-\Lambda J)}{N}\nonumber\\
&= O_p\left(\frac 1{\sqrt{NT}}\right)+O_p\left(\frac 1{\sqrt{NT}}\right)\left\{O_p(\widehat{\theta}^*-\bar J\theta)+O_p\left(\max\left(\frac 1{\sqrt N},\frac 1{\sqrt T}\right)\right)\right\}.\nonumber
\end{align}
For $J_7$ we have
\begin{align}
J_7&= -\frac 1{NT}\sum_{t=1}^T \bar JX_t'M_{\widehat{\Lambda}^*}I_7 \left(\frac{\Lambda'\widehat{\Lambda}^*}{N}\right)^{-1}G_t\nonumber\\
&= -\frac 1{NT}\sum_{t=1}^T \bar JX_t'M_{\widehat{\Lambda}^*}\left[\frac 1T\sum_{s=1}^T\epsilon_s G_s'\frac{\Lambda'\widehat{\Lambda}^*}{N}\right] \left(\frac{\Lambda'\widehat{\Lambda}^*}{N}\right)^{-1}G_t\nonumber\\
&= -\frac 1{NT^2}\sum_{t=1}^T\sum_{s=1}^T G_s'G_t\bar JX_t'M_{\widehat{\Lambda}^*}\epsilon_s.
\end{align}
Finally, for $J_8$ we have
\begin{align}
J_8=&\, -\frac 1{NT}\sum_{t=1}^T \bar JX_t'M_{\widehat{\Lambda}^*}I_8 \left(\frac{\Lambda'\widehat{\Lambda}^*}{N}\right)^{-1}G_t\nonumber\\
=&\, -\frac 1{NT}\sum_{t=1}^T \bar JX_t'M_{\widehat{\Lambda}^*}\left[
\frac 1{NT}\sum_{s=1}^T\left\{\epsilon_s\epsilon_s'-\mathbb E[\epsilon_s\epsilon_s']+\mathbb E[\epsilon_s\epsilon_s']\right\}\widehat{\Lambda}^*
\right]\left(\frac{\Lambda'\widehat{\Lambda}^*}{N}\right)^{-1}G_t\nonumber\\
=&\, -\frac 1{N^2T^2}\sum_{t=1}^T\sum_{s=1}^T \bar JX_t'M_{\widehat{\Lambda}^*}
\mathbb E[\epsilon_s\epsilon_s']\widehat{\Lambda}^*\left(\frac{\Lambda'\widehat{\Lambda}^*}{N}\right)^{-1}G_t\nonumber\\
& -\frac 1{NT}\sum_{t=1}^T \bar JX_t'M_{\widehat{\Lambda}^*}\left[
\frac 1{NT}\sum_{s=1}^T\left\{\epsilon_s\epsilon_s'-\mathbb E[\epsilon_s\epsilon_s']\right\}\widehat{\Lambda}^*
\right]\left(\frac{\Lambda'\widehat{\Lambda}^*}{N}\right)^{-1}G_t\nonumber\\
=&\, A_{NT}-\frac 1{NT}\sum_{t=1}^T \bar JX_t'M_{\widehat{\Lambda}^*}\left[
\frac 1{NT}\sum_{s=1}^T\left\{\epsilon_s\epsilon_s'-\mathbb E[\epsilon_s\epsilon_s']\right\}\widehat{\Lambda}^*
\right]\left(\frac{\Lambda'\widehat{\Lambda}^*}{N}\right)^{-1}G_t\nonumber\\
=&\, A_{NT}+o_p(\widehat{\theta}^*-\bar J\theta)+o_p\left(\frac 1{\sqrt{NT}}\right),
\end{align}
since, by barigozzi2022 and (ref),
\begin{align}
\frac 1{NT}\sum_{s=1}^T\left\{\epsilon_s\epsilon_s'-\mathbb E[\epsilon_s\epsilon_s']\right\}\widehat{\Lambda}^*&=
\frac 1{NT}\sum_{s=1}^T\left\{\epsilon_s\epsilon_s'-\mathbb E[\epsilon_s\epsilon_s']\right\}{\Lambda}J +
\frac 1{NT}\sum_{s=1}^T\left\{\epsilon_s\epsilon_s'-\mathbb E[\epsilon_s\epsilon_s']\right\}(\widehat{\Lambda}^*-\Lambda J)\nonumber\\
&=O_p\left( \frac 1{\sqrt T}\frac 1{\sqrt{NT}}\right) + O_p\left(\frac 1{\sqrt{NT}}\right)O_p(\widehat{\theta}^*-\bar J\theta).\nonumber
\end{align}
Therefore, by substituting (ref), (ref), (ref), (ref), (ref), (ref), (ref), and (ref)
into (ref) we get
\begin{align}
\frac 1{NT}\sum_{t=1}^T \bar JX_t'M_{\widehat{\Lambda}^*}\Lambda G_t = J_2+J_7+A_{NT}+o_p(\widehat{\theta}^*-\bar J\theta)+
O_p\left(\max\left(\frac 1{\sqrt {NT}},\frac 1{ T}\right)\right).
\end{align}
Now, if ${ \sqrt {NT}}/ (mN^{2-\gamma}) \to 0$ and ${ \sqrt {N}}/ (\sqrt{m}N^{1-\gamma/2}) \to 0$, as $m,N,T\to\infty$,
by substituting (ref) into (ref), from (ref) we get:
\begin{align}
\left(\widehat\theta^*-\bar J\theta\right) =&\, \left(\frac 1{NT} \sum_{t=1}^T \bar J{X}_{t}' M_{\widehat \Lambda^*} X_{t}\bar J\right)^{-1}
\, C + o_p\left(\frac 1{\sqrt{NT}}\right).
\end{align}
Thus, from (ref) and (ref) and given the definition of $C$ in (ref),
\begin{align}
\left(\frac 1{NT} \sum_{t=1}^T \bar J{X}_{t}' M_{\widehat \Lambda^*} X_{t}\bar J+o_p(1)\right)&\left(\widehat\theta^*-\bar J\theta\right) -J_2\nonumber\\
&=
\frac 1{NT}\sum_{t=1}^T \bar JX_t'M_{\widehat{\Lambda}^*}\epsilon_t +J_7 +A_{NT}+O_p\left(\max\left(\frac 1{\sqrt {NT}},\frac 1{ T}\right)\right).
\end{align}
Then, define the $T\times (r+2)$ matrices:
\begin{align}
\widehat W_t^* &:= \left\{
M_{\widehat \Lambda^*}X_t-\frac 1T\sum_{s=1}^T \left(G_t'G_s \right)M_{\widehat \Lambda^*}X_s
\right\},\nonumber\\
W_t &:= \left\{
M_{ \Lambda}X_t-\frac 1T\sum_{s=1}^T\left(G_t'G_s \right)M_{ \Lambda}X_s.
\right\}.\nonumber
\end{align}
Using these definitions, by multiplying (ref) by $\sqrt{NT}$, from (ref), (ref) we have
\begin{align}
\left(\frac 1{NT} \sum_{t=1}^T \bar J\widehat{W}_{t}^{*'} \widehat{W}_{t}^{*}\bar J+o_p(1)\right)\sqrt{NT}\left(\widehat\theta^*-\bar J\theta\right)=
\frac 1{\sqrt{NT}}\sum_{t=1}^T \bar J\widehat{W}_t^{*'}\epsilon_t+
\sqrt{NT} A_{NT}+ o_p(1),
\nonumber
\end{align}
which implies
\begin{align}
\sqrt{NT}\left(\widehat\theta^*-\bar J\theta\right)&=: II+o_P(1).
\end{align}
Now, by the definition of $A_{NT}$ in (ref) term $II$ is such that:
\begin{align}
II=&\, \left(\frac 1{NT} \sum_{t=1}^T \bar J\widehat{W}_{t}^{*'} \widehat{W}_{t}^{*}\bar J\right)^{-1} \left\{\frac 1{\sqrt{NT}}\sum_{t=1}^T \bar J\widehat{W}_t^{*'}\epsilon_t\right.\nonumber\\
&\left.-\sqrt{\frac TN}\left[\frac {1}{NT}\sum_{t=1}^T \bar JX_t'M_{\widehat{\Lambda}^*}\left(\frac 1T\sum_{s=1}^T
\mathbb E[\epsilon_s\epsilon_s']\right)\widehat{\Lambda}^*\left(\frac{\Lambda'\widehat{\Lambda}^*}{N}\right)^{-1}G_t\right]
\right\}+o_p(1)\nonumber\\
=&\, \left(\frac 1{NT} \sum_{t=1}^T \bar J\widehat{W}_{t}^{*'} \widehat{W}_{t}^{*}\bar J\right)^{-1} \left\{\frac 1{\sqrt{NT}}\sum_{t=1}^T \bar J{W}_t^{'}\epsilon_t\right.\nonumber\\
&-\sqrt{\frac TN}\left[\frac {1}{NT}\sum_{t=1}^T \bar JX_t'M_{\widehat{\Lambda}^*}\left(\frac 1T\sum_{s=1}^T
\mathbb E[\epsilon_s\epsilon_s']\right)\widehat{\Lambda}^*\left(\frac{\Lambda'\widehat{\Lambda}^*}{N}\right)^{-1}G_t\right]\nonumber\\
&\left.-\sqrt{\frac NT} \left[\frac 1{NT}\sum_{t=1}^T\bar J\left(X_t-\frac 1T\sum_{s=1}^T (G_t'G_s)X_s\right)'\Lambda\left(\frac{\Lambda'\Lambda}{N}\right)^{-1}\sum_{u=1}^T G_u\left(\frac 1N\sum_{i=1}^N\epsilon_{iu}\epsilon_{it} \right)
\right]
\right\}+o_p(1)\nonumber\\
=&\, \left(\frac 1{NT} \sum_{t=1}^T \bar J{W}_{t}^{'} {W}_{t}\bar J\right)^{-1} \left\{\frac 1{\sqrt{NT}}\sum_{t=1}^T \bar J{W}_t^{'}\epsilon_t\right.\nonumber\\
&-\sqrt{\frac TN}\left[\frac {1}{NT}\sum_{t=1}^T \bar JX_t'M_{{\Lambda}}\left(\frac 1T\sum_{s=1}^T
\mathbb E[\epsilon_s\epsilon_s']\right){\Lambda}\left(\frac{\Lambda'{\Lambda}}{N}\right)^{-1}G_t\right]\nonumber\\
&\left.-\sqrt{\frac NT} \left[\frac 1{NT}\sum_{t=1}^T\bar J\left(X_t-\frac 1T\sum_{s=1}^T (G_t'G_s)X_s\right)'\Lambda\left(\frac{\Lambda'\Lambda}{N}\right)^{-1}\sum_{u=1}^T G_u\left(\frac 1N\sum_{i=1}^N\mathbb E[\epsilon_{iu}\epsilon_{it}] \right)
\right]
\right\}\nonumber\\
&+O_p\left(\frac{\sqrt N}{T}\right)+O_p\left(\frac{\sqrt T}{N}\right)+o_p(1)\nonumber\\
=:&\, \left(\frac 1{NT} \sum_{t=1}^T \bar J{W}_{t}^{'} {W}_{t}\bar J\right)^{-1} \left\{\frac 1{\sqrt{NT}}\sum_{t=1}^T \bar J{W}_t^{'}\epsilon_t
+\sqrt{\frac TN} II_a+\sqrt{\frac NT} II_b\right\}
\nonumber\\
&+O_p\left(\frac{\sqrt N}{T}\right)+O_p\left(\frac{\sqrt T}{N}\right)+o_p(1).
\end{align}
To derive (ref) in the first step we used Lemma (ref) and the fact that $\sqrt NO_p(\Vert\widehat{\theta}^*-\bar J\theta \Vert^2)$ and $O_p(\Vert\widehat{\theta}^*-\bar J\theta \Vert)$ are dominated by $\sqrt {NT} O_p(\Vert\widehat{\theta}^*-\bar J\theta \Vert)$, and $\sqrt NO_p\left(\max\left(\frac 1N, \frac 1T\right)\right)=o_p(1)$ since $\sqrt N/T\to 0$, as $N,T\to\infty$, by assumption. Furthermore, in the second step we used Lemma (ref)(ref) for the denominator, Lemma (ref)(ref) for the second term at the numerator and Lemma (ref)(ref) for the third term at the numerator.
Since by Assumptions (ref)(ref) and (ref)(ref), $\mathbb E[\epsilon_{iu}\epsilon_{it}] = \sigma_i^2 \mathbb I(t=u)$ letting $\bar {\sigma}^2=N^{-1}\sum_{i=1}^N \sigma_i^2$, we get
\begin{align}
II_b& = - \frac 1{NT}
\sum_{t=1}^T
\bar J\left(X_t-\frac 1T\sum_{s=1}^T (G_t'G_s)X_s\right)'
\Lambda\left(\frac{\Lambda'\Lambda}{N}\right)^{-1}G_t\bar{\sigma}^2\nonumber\\
&=- \frac 1{NT}
\sum_{t=1}^T
\bar JX_t'
\Lambda\left(\frac{\Lambda'\Lambda}{N}\right)^{-1}G_t\bar{\sigma}^2+
- \frac 1{NT}
\sum_{t=1}^T
\bar J\left(\frac 1T\sum_{s=1}^T X_s\right)'
\Lambda\left(\frac{\Lambda'\Lambda}{N}\right)^{-1}G_t'G_sG_t\bar{\sigma}^2\nonumber\\
&=- \frac 1{NT}
\sum_{t=1}^T
\bar JX_t'
\Lambda\left(\frac{\Lambda'\Lambda}{N}\right)^{-1}G_t\bar{\sigma}^2+
- \frac 1{NT}
\sum_{s=1}^T
\bar J X_s' \Lambda\left(\frac{\Lambda'\Lambda}{N}\right)^{-1}
\frac 1T\sum_{t=1}^T
G_tG_t'G_s\bar{\sigma}^2\nonumber\\
&=- \frac 1{NT}
\sum_{t=1}^T
\bar JX_t'
\Lambda\left(\frac{\Lambda'\Lambda}{N}\right)^{-1}G_t\bar{\sigma}^2+
- \frac 1{NT}
\sum_{s=1}^T
\bar J X_s' \Lambda\left(\frac{\Lambda'\Lambda}{N}\right)^{-1}
G_s\bar{\sigma}^2 = 0_{r+2}.\nonumber
\end{align}
Moreover, since we assumed $T/N\to 0$ and $\sqrt N/T\to 0$, as $N,T\to\infty$, by using (ref) into (ref) we have
\begin{align}
\sqrt{NT}\left(\widehat\theta^*-\bar J\theta\right) &= \left(\frac 1{NT}\sum_{t=1}^T\bar J W_t' W_t\bar J\right)^{-1}\left(\frac 1{\sqrt {NT}}\sum_{t=1}^T
\bar JW_t \epsilon_t\right) + o_p(1)\nonumber\\
&\overset{p}{\to}\mathcal N\left(0_{r+2}, \bar J_0 \Sigma_{WW}^{-1} \bar J_0 \bar J_0D_2 \bar J_0\bar J_0\Sigma_{WW}^{-1}\bar J_0\right),
\end{align}
because of Assumptions (ref)(ref) and (ref)(ref), and Slutsky's theorem and where $\bar J_0:=\operatorname*{plim}_{m,N,T\to\infty} \bar J$. We complete the proof by noticing that $\bar J_0 \Sigma_{WW}^{-1} \bar J_0 \bar J_0D_2 \bar J_0\bar J_0\Sigma_{WW}^{-1}\bar J_0= \Sigma_{WW}^{-1}D_2 \Sigma_{WW}^{-1}$.
For part (ref), if $\sigma_i^2=\sigma^2$ for all $i=1,\ldots, N$, then in (ref) we have $\mathbb E[\epsilon_s\epsilon_s'] = \sigma^2 I_N$ and since $M_{\Lambda} \Lambda=0_{N\times q}$, we have $II_a=0_{r+2}$. The proof then follows as in part (ref) and by noticing that in this case $ \bar J_0 \Sigma_{WW}^{-1} \bar J_0 \bar J_0D_2 \bar J_0\bar J_0\Sigma_{WW}^{-1}\bar J_0= \sigma^2\Sigma_{WW}^{-1}$. This completes the proof.
Proof of Proposition (ref)
proofFirst, consider part (ref).
We have
\begin{align}
\frac1{NT}\sum_{t=1}^T\widehat X_t' u_t
=&
\frac1{NT}\sum_{t=1}^T \bar J X_t' u_t
+\frac1{NT}\sum_{t=1}^T \left(\widehat X_t' - \bar J X_t' \right) u_t = a + b.
\end{align}
Then, notice that we can write
\[
u_t =(X_t\bar J-\widehat X_t)\bar J\theta= \text{mat}_1 \left( N^{-1}(\mathcal F_{t-1}\times_3 J -\widehat {\mathcal F}_{t-1})\times_2 y'_{t-1}\times_3 \beta' J \right).
\]
For term $a$, we have that
\begin{align}
\norm{a}
&=
\norm{ \bar J X_t' mat_1 \left( \frac{1}{N^2T} \sum_{t=1}^T \left(\widehat{\mathcal{F}}_{t-1} - \mathcal{F}_{t-1} \times_3 J \right) \times_2 y_{t-1}' \times_3 \beta' J \right)}\nonumber \\
&=
\norm{ mat_1 \left( \frac{1}{N^2T} \sum_{t=1}^T \left(\widehat{\mathcal{F}}_{t-1} - \mathcal{F}_{t-1} \times_3 J \right) \times_1 \bar J X_t' \times_2 y_{t-1}' \times_3 \beta' J \right)}\nonumber \\
&=
\norm{ mat_3 \left( \frac{1}{N^2T} \sum_{t=1}^T \left(\widehat{\mathcal{F}}_{t-1} - \mathcal{F}_{t-1} \times_3 J \right) \times_1 \bar J X_t' \times_2 y_{t-1}' \times_3 \beta' J \right)^\prime}\nonumber \\
&=
\norm{ \beta' J mat_3 \left( \frac{1}{N^2T} \sum_{t=1}^T \left(\widehat{\mathcal{F}}_{t-1} - \mathcal{F}_{t-1} \times_3 J \right) \times_1 \bar J X_t' \times_2 y_{t-1}' \right)}\nonumber \\
&\leq \norm{\beta J} \cdot \norm{mat_3 \left( \frac{1}{N^2T} \sum_{t=1}^T \left(\widehat{\mathcal{F}}_{t-1} - \mathcal{F}_{t-1} \times_3 J \right) \times_1 \bar J X_t' \times_2 y_{t-1}' \right)}\nonumber \\
&= \norm{a_1} \cdot \norm{a_2}.
\end{align}
Clearly, $\norm{a_1} = O(1)$.
As for $\norm{a_2}$, let us first define $z_t := y_{t-1} \otimes X_t \bar J$.
Then, recall that for a generic tensor $\mathcal{Z}$ and matrices $A,B$, and $C$ such that $\mathcal{Z} = \mathcal{X} \times_1 A \times_2 B \times_3 C$, we have
$\text{mat}_3(\mathcal{Z}) = C \text{mat}_3(\mathcal{X}) \left( B \otimes A \right)'$. Therefore, by using also (ref), we have that:
\begin{align}
a_2 = & mat_3 \left( \frac{1}{N^2T} \sum_{t=1}^T \left(\widehat{\mathcal{F}}_{t-1} - \mathcal{F}_{t-1} \times_3 J \right) \times_1 \bar J X_t' \times_2 y_{t-1}' \right)\nonumber\\
&= \frac{1}{N^2T} \sum_{t=1}^T \text{mat}_3 \left(\widehat{\mathcal{F}}_{t-1} - \mathcal{F}_{t-1} \times_3 J \right) \left( y_{t-1}' \otimes \bar J X_t' \right)'
\nonumber\\
&=
\frac{1}{N^2T} \sum_{t=1}^T \left( \widehat{\mathcal{F}}_{(3)t-1} - J \mathcal{F}_{(3)t-1} \right) \left( y_{t-1} \otimes X_t\bar J \right)\nonumber \\
&=
\frac{1}{N^2T} \sum_{t=1}^T \left( \widehat{\mathcal{F}}_{(3)t-1} - J \mathcal{F}_{(3)t-1} \right) z_t\nonumber\nonumber \\
&= \left( \frac{\widehat{M}^{\mathcal{W}}}{mN^2} \right)^{-1}
\Bigg\{
\frac{1}{N^2T} \sum_{t=1}^T \frac{ \widehat{U}' \left(U - \widehat{U}J \right) }{m} \mathcal{F}_{(3)t} z_t \nonumber\\
&+ \frac{1}{N^2T} \sum_{t=1}^T \frac{ \left(\widehat{U} - U J\right)' }{m} \mathcal{E}_{(3)t} z_t \nonumber\\
&+ \frac{1}{N^2T} \sum_{t=1}^T \frac{ J' U'}{m} \mathcal{E}_{(3)t} z_t \Bigg\} \nonumber \\
&= \left( \frac{\widehat{M}^{\mathcal{W}}}{mN^2} \right)^{-1} \left\{ a_I + a_{II} + a_{III} \right\}.
\end{align}
Now, because of (ref) in the proof of Lemma (ref) when $\widehat{H}=J$, using Cauchy-Schwarz inequality
\begin{align}
\norm{a_I} &= \norm{\frac{1}{N^2T} \sum_{t=1}^T \frac{ \widehat{U}' \left(U - \widehat{U}J \right) }{m} \mathcal{F}_{(3)t} z_t}_F\nonumber\\
&\le \left(\frac{1}{N^2T} \sum_{t=1}^T \norm{\frac{ \widehat{U}' \left(U - \widehat{U}J \right) }{m} \mathcal{F}_{(3)t}}_F^2 \right)^{1/2}
\left(\frac{1}{N^2T} \sum_{t=1}^T \norm{z_t}_F^2\right)^{1/2}\nonumber\\
&\le \sqrt r \norm{\frac{ \widehat{U}' \left(U - \widehat{U}J \right) }{m} }\left(\frac{1}{N^2T} \sum_{t=1}^T \norm{\mathcal{F}_{(3)t}}_F^2 \right)^{1/2}
\left(\frac{1}{N^2T} \sum_{t=1}^T \norm{z_t}_F^2\right)^{1/2}
&= O_p\left(\frac 1{\xi}\right),
\end{align}
where $\xi =
\min \left(
\sqrt{mT} N^{2-\gamma/2},
N^{3-\gamma/2} T,
m \sqrt{T} N^{2-\gamma},
m N^{2-\gamma}
\right)
$,
and since
\begin{align}
\frac{1}{N^2T} \sum_{t=1}^T
\mathbb{E} \left[ \norm{\mathcal{F}_{(3)t}}_F^2 \right]
&=
\frac{1}{N^2T} \sum_{t=1}^T \sum_{i=1}^{N^2} \sum_{j=1}^{r}
\mathbb{E} \left[ \mathcal{F}_{(3)tji}^2 \right] \leq
r
\max_{t=1,\dots,T} \max_{i=1, \dots, N^2}
\max_{j=1, \dots, r}
\mathbb{E} \left[ \mathcal{F}_{(3)tji}^2 \right]=r O(1),\nonumber
\end{align}
because of Assumption (ref)(ref), and
\begin{align}
\frac{1}{N^2T} \sum_{t=1}^T
\mathbb{E} \left[ \norm{z_t }_F^2 \right]
&=
\frac{1}{N^2T} \sum_{t=1}^T \sum_{i=1}^{N^2} \sum_{j=1}^{r+2}
\mathbb{E} \left[ z_{ijt}^2 \right] \leq
\frac{1}{N^2T} \sum_{t=1}^T
\sum_{i,j=1}^{N}
\sum_{k=1}^{r+2}
\sqrt{\mathbb{E} \left[ y_{it}^4 \right]} \sqrt{\mathbb{E} \left[ X_{jkt}^4 \right]}\nonumber \\
&\leq
(r+2)
\max_{t=1,\dots,T} \max_{i,j=1, \dots, N}
\max_{k=1, \dots, r+2}
\sqrt{\mathbb{E} \left[ y_{it}^4 \right]} \sqrt{\mathbb{E} [ X_{jkt}^4 ]} \nonumber\\
&=
(r+2) O(1)O(1),\nonumber
\end{align}
following from Assumptions (ref)(ref), the MA representation (ref) of the FNAR, and Assumptions (ref)(ref) and (ref)(ref).
Moreover,
\begin{align}
\norm {a_{III}}&\le \norm{J} \norm{\frac 1{mN^2T}\sum_{t=1}^T\sum_{i=1}^m u_i \mathcal E_{(3)ti\cdot} z_t} = O_p\left(\frac 1{\sqrt {mT}N^{1-\gamma/2}}\right),
\end{align}
since
\[
\mathbb E\left[\norm{\frac 1{mN^2T}\sum_{t=1}^T\sum_{i=1}^m u_i \mathcal E_{(3)ti\cdot} z_t}^2\right]\le \frac{\mathfrak K_1N^{2+\gamma}}{mTN^4} = \frac{\mathfrak K_1}{mTN^{2-\gamma}},
\]
because of Assumption (ref). Finally, term $a_{II}$ is dominated by term $a_{III}$ because of Theorem (ref)(ref).
Therefore, by using (ref) and (ref) into (ref), and since by Lemma (ref)(ref), $\norm { \left( \frac{\widehat{M}^{\mathcal{W}}}{mN^2} \right)^{-1}}=O_p(1)$, we have
\[
\norm{a}=O_p \left( \max\left(\frac 1{\xi},\frac 1{\sqrt {mT}N^{1-\gamma/2}}\right)\right) = O_p \left( \max\left(\frac 1{N^{3-\gamma/2}T},\frac 1{mN^{2-\gamma}},\frac 1{\sqrt {mT}N^{1-\gamma/2}}\right)\right) .
\]
As for term $b$, by analogy with term $a$ above, we have
\begin{align}
\norm{b} \leq& \norm{\beta J} \cdot \norm{\text{mat}_3 \left( \frac{1}{N^2T} \sum_{t=1}^T \left(\widehat{\mathcal{F}}_{t-1} - \mathcal{F}_{t-1} \times_3 J \right) \times_1 \left(\widehat X_{t}' - \bar J X'_t \right) \times_2 y_{t-1}' \right)}\nonumber \\
=&
\norm{\beta J} \cdot \norm{\frac{1}{N^2T} \sum_{t=1}^T \left( \widehat{\mathcal{F}}_{(3)t-1} - J \mathcal{F}_{(3)t-1} \right) \left( y_{t-1} \otimes \left(\widehat X_{t}' - \bar J X'_t \right) \right)}\nonumber \\
=&
\norm{\beta J} \cdot \norm{\frac{1}{N^2T} \sum_{t=1}^T \left( \widehat{\mathcal{F}}_{(3)t-1} - J \mathcal{F}_{(3)t-1} \right) \left(\widehat z_{t} - z_t \right)} \nonumber\\
=& \norm{b_1} \cdot \norm{b_2}.
\end{align}
Now, $b_1$ is the same as $a_1$ defined before, while $b_2$ is like $a_2$ but with
$\widehat z_{t} - z_t = y_{t-1} \otimes (\widehat X_{t}- X_t\bar J )$ replacing $z_t$.
Also,
\begin{align}
\widehat X_{t} - X_t \bar J
=&
\left( \text{mat}_1 \left( \frac{1}{N} (\widehat{\mathcal{F}}_{t-1} -\mathcal{F}_{t-1} \times_3 J ) \times_2 y_{t-1}' \right), 0_{N}, 0_{N} \right) \nonumber \\
=&
\left( \text{mat}_3 \left(\frac{1}{N} (\widehat{\mathcal{F}}_{t-1} -\mathcal{F}_{t-1} \times_3 J) \times_2 y_{t-1}' \right)', 0_{N \times 2} \right)\nonumber\\
=&
\left( \frac{1}{N} \left( y_{t-1} \otimes I_N \right)' \left( \widehat{\mathcal{F}}_{(3)t-1} - J \mathcal{F}_{(3)t-1} \right)', 0_{N \times 2} \right).
\end{align}
Thus, from term $a_{III}$ in (ref) we see that the leading term in $b_2$ is given by
\begin{equation}
\frac 1{mN^2T}\sum_{t=1}^T\sum_{i=1}^m u_i \mathcal E_{(3)ti\cdot} \left(y_{t-1}\otimes\left\{y_{t-1}\otimes I_N\right\}' \frac{\mathcal E_{(3)t-1}'U}{mN}\right),
\end{equation}
which is dominated by term $a_{III}$ because of (ref) in the proof of Theorem (ref)(ref). Therefore, $b_2$ is dominated by $a_2$.
By combining (ref), (ref), and (ref), we complete the proof of part (ref).
Next, consider part (ref).
We have
\begin{align}
\frac1{TN}\sum_{t=1}^T\widehat X_{t}' \nu_t
=&
\frac1{TN}\sum_{t=1}^T \bar JX_t' \nu_t
+\frac1{TN}\sum_{t=1}^T \left(\widehat X_{t}' - \bar J X'_t \right) \nu_t = A + B.
\end{align}
By Assumption (ref)(ref),
\[
\norm{A}=O_p\left(\frac 1{\sqrt T}\right).
\]
As for term $B$, we have
\begin{align}
\norm{B}=\norm{\frac{1}{NT} \sum_{t=1}^{T} \left(\widehat X_{t}' - \bar J X'_t \right) \nu_t }
&=\norm{ \frac{1}{N^2T} \sum_{t=1}^{T} \left( \text{mat}_1 ( (\widehat{\mathcal{F}}_t -\mathcal{F}_t \times_3 J ) \times_2 y_{t-1}' ), 0_{N \times N}, 0_{N \times N} \right)' \nu_t}\nonumber \\
&=\norm{ \frac{1}{N^2T} \sum_{t=1}^{T} \left( \text{mat}_1 \left( (\widehat{\mathcal{F}}_t -\mathcal{F}_t \times_3 J) \times_1 \nu'_t \times_2 y_{t-1}' \right), 0_{N \times 2N} \right)' }
\nonumber\\
&= \norm{\frac{1}{N^2T} \sum_{t=1}^{T} \left( \text{mat}_3 \left( (\widehat{\mathcal{F}}_t -\mathcal{F}_t \times_3 J ) \times_1 \nu'_t \times_2 y_{t-1}' \right)' , 0_{N \times 2N} \right)' }
\nonumber\\
& = O_p \left( \max\left(\frac 1{N^{3-\gamma/2}T},\frac{1}{mN^{2-\gamma}},\frac{1}{\sqrt{mT} N^{1-\gamma/2}} \right) \right).\nonumber
\end{align}
The proof is similar to that of term $a_2$ defined in part (ref), with $\nu_t$ replacing $X_t$.
Finally, consider part (ref).
We have that
\begin{align}
\norm{\frac{1}{TN} \sum_{t=1}^T \widehat{X}_{t}'\widehat X_{t} - \frac{1}{TN} \sum_{t=1}^T \bar J X_{t}'X_{t}\bar J } \nonumber
&\leq
2\norm{ \frac{1}{TN} \sum_{t=1}^T (\widehat{X}_{t}' - \bar J X_t') X_{t}\bar J }
+
\norm{ \frac{1}{TN} \sum_{t=1}^T (\widehat{X}_{t}' - \bar J X_t') (\widehat{X}_{t}' - \bar J X_{t}) }
\nonumber\\
&= I + II.
\end{align}
Now,
\begin{align}
\norm{I} & =2 \norm{
\frac{1}{N^2T} \sum_{t=1}^{T} \left( \text{mat}_3 \left( (\widehat{\mathcal{F}}_t -\mathcal{F}_t \times_3 J ) \times_1 \bar J X'_t \times_2 y_{t-1}' \right)', 0_{N \times 2N} \right)'
}\nonumber \\
& = O_p \left( \max\left(\frac 1{N^{3-\gamma/2}T},\frac{1}{mN^{2-\gamma}},\frac{1}{\sqrt{mT} N^{1-\gamma/2}} \right) \right),\nonumber
\end{align}
following the same steps as in the proof for term $a_2$ in part (ref).
Moreover, $II$ is clearly dominated by $I$. This completes the proof.
Proof of Theorem (ref)
proofFrom (ref), Proposition (ref), Assumptions (ref)(ref) and (ref)(ref), and Slutsky's theorem:
\begin{align}
\sqrt T \left(\widehat{\theta}^{\tiny OLS}-\bar J\theta\right) =&\, \left(\frac 1{NT} \sum_{t=1}^T \widehat{X}_{t}'\widehat X_{t}\right)^{-1} \left\{\left(\frac1{N\sqrt T}\sum_{t=1}^T\widehat X_{t}' \nu_t\right)+ \left(\frac1{N\sqrt T}\sum_{t=1}^T\widehat X_{t}' u_t\right)\right\}\nonumber\\
=&\, \left(\frac 1{NT} \sum_{t=1}^T \widehat{X}_{t}'\widehat X_{t}\right)^{-1} \left\{\left(\frac1{N\sqrt T}\sum_{t=1}^T \bar J X_{t}' \nu_t\right)+\left(\frac1{N\sqrt T}\sum_{t=1}^T(\widehat X_{t}-\bar J X_t)' \nu_t\right)\right.\nonumber\\
&\left.+ \left(\frac1{N\sqrt T}\sum_{t=1}^T\widehat X_{t}' u_t\right)\right\}\nonumber\\
=&\, \left(\frac 1{NT} \sum_{t=1}^T\bar J {X}_{t}' X_{t}\bar J \right)^{-1} \left(\frac1{N\sqrt T}\sum_{t=1}^T \bar J X_{t}' \nu_t\right) + O_p\left(\frac {\sqrt T}{ mN^{2-\gamma}}\right)+o_p(1)\nonumber\\
&\overset{d}{\to} \mathcal N\left(0_{r+2},\bar J_0 \Sigma_{XX}^{-1}\bar J_0\bar J_0\Omega_0\bar J_0\bar J_0\Sigma_{XX}^{-1}\bar J_0\right),\nonumber
\end{align}
since we assumed $\sqrt { T}/ (mN^{2-\gamma}) \to 0$ and where and $\bar J_0:=\operatorname*{plim}_{m,N,T\to\infty} \bar J$. Notice, finally that
$\bar J_0 \Sigma_{XX}^{-1}\bar J_0\bar J_0\Omega_0\bar J_0\bar J_0\Sigma_{XX}^{-1}\bar J_0= \Sigma_{XX}^{-1}\Omega_0\Sigma_{XX}^{-1}$. This completes the proof.
Proof of Proposition (ref)
proofFirst of all, notice that from Assumption (ref)(ref) and
by Weyl's inequality,
\begin{align}
\norm{V^{-1}} = \frac 1{\mu_N(V)} \le \frac 1{\mu_N(\Lambda\Lambda')+\mu_N(S)}=\frac 1{ \min_{i=1,\ldots, N}\mathbb E[\epsilon_{it}^2]}\le \frac 1{\underline M_{\epsilon}},
\end{align}
where $\mu_N(\Lambda\Lambda')=0$, $\mu_N(V)$, and $\mu_N(S)$ are the smallest eigenvalues of $\Lambda\Lambda'$, $V$, and $S$, respectively, and $\mu_N(\Lambda\Lambda')=0$.
Consider part (ref). We have
\begin{align}
\frac1{NT}\sum_{t=1}^T\widehat X_t'\widehat V^{-1} u_t
=&\,
\frac1{NT}\sum_{t=1}^T \bar J X_t' V^{-1} u_t
+\frac1{NT}\sum_{t=1}^T \left(\widehat X_t' - \bar J X_t' \right) V^{-1} u_t \nonumber\\
&+\frac1{NT}\sum_{t=1}^T \bar J X_t'(\widehat V^{-1}-V^{-1}) u_t
+\frac1{NT}\sum_{t=1}^T \left(\widehat X_t' - \bar J X_t' \right)(\widehat V^{-1}-V^{-1}) u_t\nonumber\\
=&\, a + b +c+d.\nonumber
\end{align}
Because of (ref), term $a$ behaves like term $a$ in (ref) in the proof of Proposition (ref)(ref), thus:
\[
\norm{a}= O_p \left( \max\left(\frac 1{N^{3-\gamma/2}T},\frac 1{mN^{2-\gamma}},\frac 1{\sqrt {mT}N^{1-\gamma/2}}\right)\right) .
\]
Likewise, term $b$ behaves like term $b$ in the proof of Proposition (ref)(ref), thus term $b$ is dominated by term $a$. Moreover, because of Lemma (ref), terms $c$ and $d$ are dominated by terms $a$ and $b$, respectively. This proves part (ref).
For part (ref), we have:
\begin{align}
\frac1{NT}\sum_{t=1}^T\widehat X_{t}'\widehat V^{-1} \nu_t =&\,
\frac1{NT}\sum_{t=1}^T\widehat X_{t}' V^{-1} \nu_t
+\frac1{NT}\sum_{t=1}^T\widehat X_{t}'(\widehat V^{-1}-V^{-1}) \nu_t\nonumber\\
=&\, \frac1{NT}\sum_{t=1}^T \bar J X_{t}' V^{-1} \nu_t+\frac1{NT}\sum_{t=1}^T(\widehat X_{t}-X_t\bar J)' V^{-1} \nu_t\nonumber\\
&+\frac1{NT}\sum_{t=1}^T \bar J X_{t}'(\widehat V^{-1}-V^{-1}) \nu_t
+\frac1{NT}\sum_{t=1}^T(\widehat X_{t}-X_t\bar J)'(\widehat V^{-1}-V^{-1}) \nu_t.\nonumber\\
=&\, A+B+C+D.\nonumber
\end{align}
By Assumption (ref)(ref):
\[
\norm{A}=O_p\left(\frac1{\sqrt{NT}}\right).
\]
Term $B$, because of (ref), behaves like term $B$ in (ref) in the proof of Proposition (ref)(ref), i.e., it behaves like term $A$ in part (ref):
\[
\norm{B}= O_p \left( \max\left(\frac 1{N^{3-\gamma/2}T},\frac 1{mN^{2-\gamma}},\frac 1{\sqrt {mT}N^{1-\gamma/2}}\right)\right) .
\]
Terms $C$ and $D$ are dominated by terms $A$ and $B$, respectively, because of Lemma (ref). This proves part (ref).
For part (ref), we have
\begin{align}
\norm{\frac 1{NT}\sum_{t=1}^T \widehat X_t'\widehat V^{-1}\widehat X_t- \frac 1{NT}\sum_{t=1}^T \bar J X_t' V^{-1} X_t\bar J}\le&\,
2\norm{ \frac{1}{TN} \sum_{t=1}^T (\widehat{X}_{t}' - \bar J X_t') V^{-1}X_{t}\bar J } \nonumber\\
&+\norm{ \frac{1}{TN} \sum_{t=1}^T (\widehat{X}_{t}' - \bar J X_t')V^{-1} (\widehat{X}_{t}' - \bar J X_{t}) } \nonumber\\
&+\norm{\frac{1}{TN} \sum_{t=1}^T \bar J X_t' (\widehat V^{-1}-V^{-1})X_t\bar J } \nonumber\\
&+2\norm{\frac{1}{TN} \sum_{t=1}^T (\widehat{X}_{t}' - \bar J X_t')(\widehat V^{-1}-V^{-1})X_t\bar J} \nonumber\\
&+\norm{\frac{1}{TN} \sum_{t=1}^T (\widehat{X}_{t}' - \bar J X_t')(\widehat V^{-1}-V^{-1})(\widehat{X}_{t}' - \bar J X_t')} \nonumber\\
=&\, I+II+III+IV+V.\nonumber
\end{align}
Because of (ref), terms $I$ and $II$ behave like terms $I$ and $II$ in (ref) in the proof of Proposition (ref)(ref), thus
\[
\norm{I}=O_p \left( \max\left(\frac 1{N^{3-\gamma/2}T},\frac{1}{mN^{2-\gamma}},\frac{1}{\sqrt{mT} N^{1-\gamma/2}} \right) \right),
\]
while $II$ is dominated by $I$. Then, by Lemma (ref), we have
\[
\norm{III}=O_p\left(\max\left(\frac 1{\sqrt N},\sqrt{\frac{\log N}{T}}\right)\right).
\]
Finally, terms $IV$ and $V$ are dominated by terms $I$ and $II$, respectively, by Lemma (ref). This completes the proof.
Proof of Theorem (ref)
proofFrom (ref), Proposition (ref), Assumptions (ref)(ref) and (ref)(ref), and Slutsky's theorem:
\begin{align}
\sqrt {NT} \left(\widehat{\theta}^{\tiny GLS}-\bar J\theta\right) =&\, \left(\frac 1{NT} \sum_{t=1}^T \widehat{X}_{t}'\widehat{V}^{-1}\widehat X_{t}\right)^{-1} \left\{\left(\frac1{\sqrt{N T}}\sum_{t=1}^T\widehat X_{t}' \widehat{V}^{-1}\nu_t\right)+ \left(\frac1{\sqrt{N T}}\sum_{t=1}^T\widehat X_{t}'\widehat{V}^{-1} u_t\right)\right\}\nonumber\\
=&\, \left(\frac 1{NT} \sum_{t=1}^T\bar J {X}_{t}' {V}^{-1} X_{t}\bar J \right)^{-1} \left(\frac1{\sqrt{N T}}\sum_{t=1}^T \bar J X_{t}'{V}^{-1} \nu_t\right)\nonumber\\
& + O_p\left(\frac {\sqrt{N T}}{ mN^{2-\gamma}}\right)+O_p\left(\frac {\sqrt{N }}{ \sqrt mN^{1-\gamma/2}}\right)+o_p(1)\nonumber\\
&\overset{d}{\to} \mathcal N\left(0_{r+2},\bar J_0 \Omega_1^{-1}\bar J_0\bar J_0\Omega_1\bar J_0\bar J_0\Omega_1^{-1}\bar J_0\right),\nonumber
\end{align}
since we assumed $\sqrt { {NT}}/ (mN^{2-\gamma}) \to 0$ and $ {\sqrt{N }}/{ \sqrt mN^{1-\gamma/2}}\to 0$ and where and $\bar J_0:=\operatorname*{plim}_{m,N,T\to\infty} \bar J$. Notice that $\bar J_0 \Omega_1^{-1}\bar J_0\bar J_0\Omega_1\bar J_0\bar J_0\Omega_1^{-1}\bar J_0=\Omega_1^{-1}$. This completes the proof.
\refstepcounter{lettersection}
\oldsection{Auxiliary lemmata}
We start with some notation. Let $\chi_t:=U\mathcal{F}_{(3)t}$ and recall the definition of the $m\times m$ matrices:
align[align omitted — 435 chars of source]
with $j$ largest eigenvalues $\widehat{\mu}_j^{\mathcal{W}}$, ${\mu}_j^{\mathcal{W}}$, ${\mu}_j^{\chi}$, and ${\mu}_j^{\mathcal{E}}$, respectively.
lemmaUnder Assumptions (ref) and (ref):
\begin{enumerate}[label=(\roman*)]
• for all $j=1, \dots, r$,
$\underline{C}_j < \underset{m, N \to \infty }{\text{lim inf }} \frac{\mu_j^{\chi}}{m N^2} \leq
\underset{m, N \to \infty }{\text{lim sup }} \frac{\mu_j^{\chi}}{m N^2} < \overline{C}_{j}$
for some finite $\underline{C}_{j}$ and $\overline{C}_{j}$ independent of $m$ and $N$.
• For all $m \in \mathbb{N}$, $N \in \mathbb{N}$ and
$T \in \mathbb{N}$,
\[ \frac{1}{m N^\gamma T} \sum_{i,j=1}^{m} \sum_{h,k=1}^{N^2} \sum_{t,s=1}^{T}
\abs{
\mathbb{E} \left[ \mathcal{E}_{(3) t i h} \mathcal{E}_{(3) s j k} \right]
} \leq M_{2 \mathcal{E}}
\]
for some finite $M_{2 \mathcal{E}}$ independent of $m, N, T$ and some $\gamma \in [0,2]$.
• For all $m \in \mathbb{N}$ and $N \in \mathbb{N}$,
\[ \frac{1}{m N^\gamma} \sum_{i,j=1}^{m} \sum_{h,k=1}^{N^2}
\abs{
\mathbb{E} \left[ \mathcal{E}_{(3) t i h} \mathcal{E}_{(3) s j k} \right]
} \leq M_{3 \mathcal{E}}
\]
for some finite $M_{3 \mathcal{E}}$ independent of $m, N, T$ and some $\gamma \in [0,2]$.
• Let $\mu_1^{\mathcal{E}} $ denote the largest eigenvalue of $\Gamma^{\mathcal{E}}$.
Then, $\frac{\mu_1^{\mathcal{E}} }{N^{\gamma}} \leq M_{\mathcal{E}} $ for some $\gamma \in [0,2]$ and some finite $ M_{\mathcal{E}}$.
\end{enumerate}
proofFor part (ref), by merikoskikumar2004 (Theorem 7), for all $j=1, \dots, r$, we have
\begin{equation}
\frac{\mu_r (U'U)}{m} \frac{\mu_j(\Gamma^{\mathcal{F}})}{N^2}
\leq \frac{\mu_j^{\chi}}{m N^2}
\leq \frac{\mu_j (U'U)}{m} \frac{\mu_1(\Gamma^{\mathcal{F}})}{N^2}, \nonumber
\end{equation}
where $\Gamma^{\mathcal{F}} = \mathbb{E} [\mathcal{F}_{(3)t}' \mathcal{F}_{(3)t}]$ and $\mu_j (\cdot)$ denotes the $j$-th eigenvalues of the matrix in parenthesis.
The proof follows from Assumption (ref)(ref) which, by continuity of the eigenvalues, implies that, for any $j=1, \dots, r$, as $m \to \infty$
\begin{equation}
\lim_{m \to \infty} \frac{\mu_j (U'U)}{m} = \mu_j (\Sigma_U),\nonumber
\end{equation}
with
\begin{equation}
0 < m_U^2 \leq \mu_r(\Sigma_U) \leq \mu_1 (\Sigma_U) \leq M^2_U < \infty, \nonumber
\end{equation}
and by Assumption (ref)(ref), which implies that
$\frac{\mu_r(\Gamma^{\mathcal{F}})}{N^2}$ and $\frac{\mu_1(\Gamma^{\mathcal{F}})}{N^2}$ are both finite and bounded away from zero.
For part (ref), by Assumption (ref)(ref) we have
\begin{align}
\frac{1}{m N^\gamma T} \sum_{i,j=1}^{m} \sum_{h,k=1}^{N^2} \sum_{t,s=1}^{T}
\abs{
\mathbb{E} \left[ \mathcal{E}_{(3) t i h} \mathcal{E}_{(3) s j k} \right]
}
=&
\frac{1}{m} \sum_{i,j=1}^{m} \sum_{h,k=1}^{N^2}
\sum_{l=-(T-1)}^{T-1} \left(1 - \frac{\abs{l}}{T}\right) \frac{\abs{
\mathbb{E} \left[ \mathcal{E}_{(3) t i h} \mathcal{E}_{(3) s j k} \right] }}{N^\gamma}\nonumber \\
\leq&
\max_{i=1, \dots, m} \sum_{j=1}^{m} \sum_{l=-\infty}^{\infty}
\rho_{\mathcal{E}}^{\abs{l}} M_{ij}
\leq \frac{M_{\mathcal{E}} (1+\rho_{\mathcal{E}})}{1-\rho_{\mathcal{E}}}.\nonumber
\end{align}
Similarly, for part (ref),
\begin{align}
\frac{1}{m N^\gamma } \sum_{i,j=1}^{m} \sum_{h,k=1}^{N^2}
\abs{
\mathbb{E} \left[ \mathcal{E}_{(3) t i h} \mathcal{E}_{(3) t j k} \right]
}
\leq&
\max_{i=1, \dots, m} \sum_{j=1}^{m}
M_{ij}
\leq M_{\mathcal{E}}.\nonumber
\end{align}
For part (ref),
\begin{align}
\frac{1}{N^{\gamma}} \mu_1^{\mathcal{E}} =& \frac{1}{N^\gamma} \norm{\mathbb{E}\left[\mathcal{E}_{(3) t} \mathcal{E}_{(3) t}' \right]} \\
\leq&
\frac{1}{N^\gamma}
\max_{i=1, \dots, m} \sum_{j=1}^{m}
\abs{\mathbb{E}\left[\mathcal{E}_{(3) t i \cdot}' \mathcal{E}_{(3) t \cdot j}' \right] }\nonumber \\
\leq&
\frac{1}{N^\gamma}
\max_{i=1, \dots, m} \sum_{j=1}^{m}
\abs{ \sum_{h=1}^{N^2} \mathbb{E}\left[\mathcal{E}_{(3) t i h} \mathcal{E}_{(3) t h j}' \right] }\nonumber \\
\leq&
\frac{1}{N^\gamma}
\max_{i=1, \dots, m} \sum_{j=1}^{m} \sum_{h=1}^{N^2}
\abs{ \mathbb{E}\left[\mathcal{E}_{(3) t i h} \mathcal{E}_{(3) t h j}' \right] }\nonumber\\
\leq& M_{\mathcal{E}},\nonumber
\end{align}
following from Assumption (ref)(ref)
lemmaUnder Assumption (ref),
\begin{equation}
\frac{1}{\sqrt{mT}N^{\gamma/2}}
\norm{ \sum_{t=1}^{T} \mathcal{F}_{(3)t} \mathcal{E}_{(3)t}' }
=
O_p (1).\nonumber
\end{equation}
proofFrom Assumption (ref),
\begin{align}
\mathbb{E} \left[
\frac{1}{mTN^{\gamma}}
\norm{ \sum_{t=1}^{T} \mathcal{F}_{(3)t} \mathcal{E}_{(3)t }' }^2
\right]
&\leq
\mathbb{E} \left[
\frac{1}{mTN^{\gamma}}
\norm{ \sum_{t=1}^{T} \mathcal{F}_{(3)t} \mathcal{E}_{(3)t }' }^2_F
\right]\nonumber \\
&=
\mathbb{E} \left[\frac 1{m N^{\gamma} T} \sum_{i=1}^{m} \norm{ \sum_{t=1}^{T} \mathcal{F}_{(3)t} \mathcal{E}_{(3)t i\cdot}^\prime }^2 \right]
\leq
C_{\mathcal{FE}}.\nonumber
\end{align}
lemmaUnder Assumptions (ref) and (ref),
\begin{enumerate}[label=(\roman*)]
• $U = V^{\chi} (M^{\chi})^{1/2} N^{-1}$.
• $\mathcal{F}_{(3)t} = N (M^{\chi})^{-\frac{1}{2}} V^{\chi}{}' \chi_t$.
\end{enumerate}
where $ V^{\chi}$ is the $m \times r$ matrix whose $j$-th column is the normalized eigenvector corresponding to the $j$-th largest eigenvalue $ \mu_j^{\chi}$ of the matrix $\Gamma^{\chi}=E(\chi_t \chi_t')$,
and $M^{\chi}$ is the $r \times r$ diagonal matrix with $ \mu_j^{\chi}$ as its entry $(j,j)$.
proofFor part (ref),
Assumption (ref)(ref) implies $\frac{\Gamma^{\chi}}{N^2} = U U' $. Therefore, since the non-zero eigenvalues of $\frac{\Gamma^{\chi}}{m N^2}$ are the same as the $r$ eigenvalues of $\frac{U'U}{m}$, which is diagonal by Assumption (ref)(ref).
Then, we must have, for all $m \in \mathbb{N}$,
\begin{equation}
\frac{U'U}{m} = \frac{M^{\chi}}{mN^2}.\nonumber
\end{equation}
Since $\Gamma^{\chi} = V^{\chi} M^{\chi} V^{\chi}{}'$, it must be that
\begin{equation}
U K = \frac{V^{\chi} (M^{\chi})^{1/2}}{N},\nonumber
\end{equation}
for some $r \times r$ invertible $K$. By multiplying on the right both sides by their transposed:
\begin{equation}
K' U' U K = \frac{M^{\chi}}{N^2},\nonumber
\end{equation}
since eigenvectors are normalized. Thus, we must have $K = I_r$.
This proves part (ref).
For part (ref), since $\chi = U \mathcal{F}_{(3)t}$, then by linear projection of $U$ onto $\chi_t$, and using part (i), for $t=1, \dots, T$,
\begin{equation}
\mathcal{F}_{(3)t} = \left( U' U \right)^{-1} U' \chi_t =
N (M^{\chi})^{-\frac{1}{2}} V^{\chi}' \chi_t.\nonumber
\end{equation}
This proves part (ref).
lemmaUnder Assumptions (ref) and (ref), for all $t=1, \dots, T$ and all $m,N,T \in \mathbb{N}$,
\begin{enumerate}[label=(\roman*)]
• $\norm{ \frac{U}{\sqrt{m}} } = O(1)$.
• $\norm{\frac{\mathcal{F}_{(3)t}}{N}} = O_p(1)$ and
$\norm{\frac{\mathcal{F}_{(3)}}{\sqrt{T}N}} = O_p(1)$.
• $\norm{\frac{\mathcal{E}_{(3)t}}{\sqrt{m}N}} = O_p\left(\frac{1}{N^{1-\gamma/2}}\right)$ and
$\norm{\frac{\mathcal{E}_{(3)}}{\sqrt{mT}N}} = O_p\left(\frac{1}{N^{1-\gamma/2}}\right)$ .
\end{enumerate}
proofBy Assumption (ref)(ref), which holds for all $m \in \mathbb{N}$,
\begin{equation}
\sup_{m \in \mathbb{N}} \norm{\frac{U}{\sqrt{m}}}^2
\leq
\sup_{m \in \mathbb{N}} \norm{\frac{U}{\sqrt{m}}}^2_F
=
\sup_{m \in \mathbb{N}} \frac{1}{m} \sum_{j=1}^{r} \sum_{i=1}^{m}
u_{ij}^2
\leq \sup_{m \in \mathbb{N}} \max_{i=1, \dots, m} \norm{u_i}^2
\leq M_U^2,\nonumber
\end{equation}
since $M_U$ is independent of $i$. This proves part (ref).
For part (ref), by Assumption (ref)(ref)
\begin{align}
\sup_{N,T \in \mathbb{N}} \max_{t=1, \dots, T}
\mathbb{E} \left[ \norm{\frac{\mathcal{F}_{(3)t}}{N}}^2 \right]
&\leq \sup_{N,T \in \mathbb{N}} \max_{t=1, \dots, T} \mathbb{E} \left[ \norm{\frac{\mathcal{F}_{(3)t}}{N}}^2_F \right]\nonumber \\
&=
\sup_{N,T \in \mathbb{N}} \max_{t=1, \dots, T}
\mathbb{E} \left[\tr\left( \frac{\mathcal{F}_{(3)t}\mathcal{F}_{(3)t}'}{N^2}\right) \right]\nonumber\\
&= \sup_{N,T \in \mathbb{N}} \max_{t=1, \dots, T}
\tr( \frac{1}{N^2} \mathbb{E} \left[ \mathcal{F}_{(3)t}\mathcal{F}_{(3)t}' \right])
= r.\nonumber
\end{align}
Therefore,
\begin{align}
\sup_{N,T \in \mathbb{N}}
\mathbb{E} \left[ \norm{\frac{\mathcal{F}_{(3)}}{\sqrt{T}N}}^2 \right]
&\leq \sup_{N,T \in \mathbb{N}} \mathbb{E} \left[ \norm{\frac{\mathcal{F}_{(3)}}{\sqrt{T}N}}^2_F \right] \nonumber \\
&= \sup_{N,T \in \mathbb{N}}
\frac{1}{T} \sum_{t=1}^{T} \mathbb{E} \left[ \norm{\frac{\mathcal{F}_{(3)t}}{N}}^2_F \right] \nonumber \\
&=\sup_{N,T \in \mathbb{N}} \max_{t=1, \dots, T}
\mathbb{E} \left[ \norm{\frac{\mathcal{F}_{(3)t}}{N}}^2_F \right]
= r.\nonumber
\end{align}
For part (ref), by Lemma (ref)(ref)
\begin{align}
\sup_{N,T \in \mathbb{N}} \max_{t=1, \dots, T}
\mathbb{E} \left[ \norm{\frac{\mathcal{E}_{(3)t}}{\sqrt{m}N}}^2 \right]
&\leq \sup_{N,T \in \mathbb{N}} \max_{t=1, \dots, T} \mathbb{E} \left[ \norm{\frac{\mathcal{E}_{(3)t}}{\sqrt{m}N}}^2_F \right] \nonumber \\
&=
\sup_{N,T \in \mathbb{N}} \max_{t=1, \dots, T}
\mathbb{E} \left[\tr\left( \frac{\mathcal{E}_{(3)t}\mathcal{E}_{(3)t}'}{mN^2}\right) \right]\nonumber\\
&= \sup_{N,T \in \mathbb{N}} \max_{t=1, \dots, T}
\frac{1}{mN^2} \tr( \Gamma^{\mathcal{E}})\nonumber\\
&\leq \frac{ M_{\mathcal{E}} N^{\gamma}}{N^{2-\gamma}} .\nonumber
\end{align}
Thus,
\begin{align}
\sup_{N,T \in \mathbb{N}}
\mathbb{E} \left[ \norm{\frac{\mathcal{E}_{(3)}}{\sqrt{mT}N}}^2 \right]
&\leq \sup_{N,T \in \mathbb{N}} \mathbb{E} \left[ \norm{\frac{\mathcal{E}_{(3)}}{\sqrt{mT}N}}^2_F \right]\nonumber \\
&=
\sup_{N,T \in \mathbb{N}}
\frac{1}{T} \sum_{t=1}^T \mathbb{E} \left[ \norm{\frac{\mathcal{E}_{(3)t}}{\sqrt{m}N}}^2_F \right]\nonumber \\
&= \sup_{N,T \in \mathbb{N}} \max_{t=1, \dots, T}
\mathbb{E} \left[ \norm{\frac{\mathcal{E}_{(3)t}}{\sqrt{m}N}}^2_F \right]
\leq \frac{ M_{\mathcal{E}} N^{\gamma}}{N^{2-\gamma}} ,\nonumber
\end{align}
which completes the proof.
lemmaUnder Assumptions (ref)-(ref), for all $m,N,T \in \mathbb{N}$,
\begin{enumerate}[label=(\roman*)]
• $
\norm{ \frac{1}{N^2 T} \sum_{t=1}^{T} \mathcal{F}_{(3)t} \mathcal{F}_{(3)t}' - \frac{\Gamma^{\mathcal{F}}}{N^2}}
= O_p \left(\frac{1}{N \sqrt{T}}\right).
$
• $
\norm{ \frac{1}{ m N^2 T} \sum_{t=1}^{T} \mathcal{E}_{(3)t}\mathcal{E}_{(3)t}' - \frac{\Gamma^{\mathcal{E}}}{m N^2} } = O_p \left(\frac{1}{ N^{2-\gamma/2} \sqrt{T}}\right).
$
\end{enumerate}
proofFor part (ref), because of Assumption (ref)(ref),
\begin{align}
\mathbb{E} \left[ \norm{ \frac{1}{N^2 T} \sum_{t=1}^{T} \mathcal{F}_{(3)t} \mathcal{F}_{(3)t}' - \frac{\Gamma^{\mathcal{F}}}{N^2}}^2 \right]
&\leq
\mathbb{E} \left[ \norm{ \frac{1}{N^2 T} \sum_{t=1}^{T} \mathcal{F}_{(3)t} \mathcal{F}_{(3)t}' - \frac{\Gamma^{\mathcal{F}}}{N^2}}^2_F \right]\nonumber\\
&= \frac{1}{N^2 T} \sum_{j=1}^{r} \sum_{i=1}^{r}
\mathbb{E} \left[ \left(
\frac{1}{N \sqrt{T}}
\sum_{t=1}^{T}
\{
\mathcal{F}_{(3)t j \cdot} \mathcal{F}_{(3)t i \cdot}'
-
\mathbb{E} [ \mathcal{F}_{(3)t j \cdot} \mathcal{F}_{(3)t i \cdot}' ]
\}
\right)^2 \right]\nonumber \\
&\leq \frac{r^2 C_{\mathcal{F}}}{N^2 T} .\nonumber
\end{align}
This proves part (ref).
For part (ref),
because of Assumption (ref)(ref),
\begin{align}
&\mathbb{E} \left[ \norm{ \frac{1}{m N^2 T} \sum_{t=1}^{T} \mathcal{E}_{(3)t} \mathcal{E}_{(3)t}' - \frac{\Gamma^{\mathcal{E}}}{m N^2}}^2 \right]\nonumber
\\
&\leq \frac{1}{m N^{4- \gamma} T}
\sum_{j=1}^{m}
\mathbb{E} \left[ \left(
\frac{1}{\sqrt{mT} N^{\gamma/2} }
\sum_{i=1}^{m} \sum_{t=1}^{T}
\left\{
\mathcal{E}_{(3)t i \cdot} \mathcal{E}_{(3)t j \cdot}'
-
\mathbb{E} \left[ \mathcal{E}_{(3)t i \cdot} \mathcal{E}_{(3)t j \cdot}' \right]
\right\}
\right)^2 \right] \nonumber\\
&\leq \frac{C_{\mathcal{E}}}{N^{4 - \gamma} T}, \nonumber
\end{align}
which completes the proof.
lemmaUnder Assumptions (ref)-(ref), for all $m,T, N \in \mathbb{N}$
\begin{enumerate}[label=(\roman*)]
• $\frac{1}{m N^2} \norm{ \widehat{\Gamma}^{\mathcal{W}} - \Gamma^{\mathcal{W}} } = O_p \left(\frac{1}{N \sqrt{T}}\right)$.
• $\frac{1}{m N^2} \norm{ \widehat{\Gamma}^{\mathcal{W}} - \Gamma^{\chi} } = O_p \left( \max \left( \frac{1}{N \sqrt{T}}, \frac{1}{N^{2-\gamma} m}\right) \right)$.
• $\frac{1}{m N^2} \left| \widehat{\mu}_j^{\mathcal{W}} - \mu_j^{\chi} \right| = O_p \left(\max \left(\frac{1}{N \sqrt{T}},\frac{1}{N^{2-\gamma} m} \right) \right)$.
for all $j$
• $\norm{\widehat{V}^{\mathcal{W}} - V^{\chi} J} = O_p \left(\max \left(\frac{1}{N \sqrt{T}},\frac{1}{N^{2-\gamma} m} \right) \right)$,
where $\widehat{V}^{\mathcal{W}} $ is the $m \times r$ matrix whose $j$-th column is the normalized (unit-modulus) eigenvector corresponding to the $j$-th largest eigenvalue of $ \widehat{\Gamma}^{\mathcal{W}}$, $V^{\chi}$ is the $m \times r$ matrix whose $j$-th column is the normalized eigenvector corresponding to the $j$-th largest eigenvalue of $ \Gamma^{\chi} $, and $J$ is a $r \times r$ diagonal matrix whose diagonal entries are $\pm 1$.
\end{enumerate}
proofFor part (ref), we have that
\begin{align}
\norm{\frac{1}{m N^2} \left( \widehat{\Gamma}^{\mathcal{W}} - \Gamma^{\mathcal{W}} \right) }
&\leq
\norm{
\frac{1}{mN^2}
\left\{
U \left( \frac{1}{T} \sum_{t=1}^{T} \mathcal{F}_{(3)t} \mathcal{F}_{(3)t}' - \Gamma^{\mathcal{F}} \right) U'
+ \frac{1}{T} \sum_{t=1}^{T} \mathcal{E}_{(3)t} \mathcal{E}_{(3)t}' - \Gamma^{\mathcal{E}}
\right\} }\nonumber \\
&+ \norm{\frac{2}{mN^2 T} \sum_{t=1}^{T} \mathcal{F}_{(3)t} \mathcal{E}_{(3)t}^\prime}\nonumber\\
&\leq
\norm{ \frac{U}{\sqrt{m}} }^2 \norm{ \frac{1}{T N^2} \sum_{t=1}^{T} \mathcal{F}_{(3)t} \mathcal{F}_{(3)t}' - \frac{\Gamma^{\mathcal{F}}}{N^2}}
+ \norm{ \left( \frac{1}{m T N^2} \sum_{t=1}^{T} \mathcal{E}_{(3)t} \mathcal{E}_{(3)t}' - \frac{\Gamma^{\mathcal{E}} }{m N^2}\right) } \nonumber\\
&+ \norm{\frac{2}{mN^2 T} \sum_{t=1}^{T} \mathcal{F}_{(3)t} \mathcal{E}_{(3)t}^\prime}.\nonumber
\end{align}
By Lemma (ref)(ref), (ref)(ref) and (ref)(ref), the first two terms are
$O_p \left( \frac{1}{N \sqrt{T}} \right)$.
By Lemma (ref),
$$
\mathbb{E} \left[\frac 1{m^2 N^4 T^2} \norm{ \sum_{t=1}^T \mathcal{F}_{(3)t} \mathcal{E}_{(3)t}^\prime }^2 \right] =
O \left( \frac{1}{ m T N^{4-\gamma}} \right),
$$
so
$$
\norm{ \frac 1{m N^2 T} \sum_{t=1}^T \mathcal{F}_{(3)t} \mathcal{E}_{(3)t}^\prime } =
O_p\left( \frac{1}{\sqrt
{mT} N^{2-\gamma/2}} \right).
$$
This proves part (ref).
For part (ref), under Assumption (ref), we have that
\begin{equation}
\widehat{\Gamma}^{\mathcal{W}} - \Gamma^{\chi} = \widehat{\Gamma}^{\mathcal{W}} - \Gamma^{\mathcal{W}} + \Gamma^{\mathcal{E}} .\nonumber
\end{equation}
Hence,
\begin{align}
\frac{1}{m N^2} \norm{ \widehat{\Gamma}^{\mathcal{W}} - \Gamma^{\chi} } &= \frac{1}{m N^2} \norm{ \widehat{\Gamma}^{\mathcal{W}} - \Gamma^{\mathcal{W}} } + \frac{1}{m N^2} \norm{ \Gamma^{\mathcal{E}} }\nonumber \\
&= O_p \left(\frac{1}{N \sqrt{T}}\right) + O_p\left(\frac{1}{m N^{2-\gamma}}\right) ,\nonumber
\end{align}
following from part (i) and Lemma (ref)(ref), since $\norm{\Gamma^{\mathcal{E}}}$ is bounded by $\mu_1^{\mathcal{E}}$.
For part (ref), given Weyl's inequality,
\begin{equation}
\left| \widehat{\mu}_j^{\mathcal{W}} - \mu_j^{\chi} \right|
\leq \mu_1 \left( \widehat{\Gamma}^{\mathcal{W}} - \Gamma^{\chi} \right)
=
\norm{ \widehat{\Gamma}^{\mathcal{W}} - \Gamma^{\chi} },\nonumber
\end{equation}
so, given part (ref),
\begin{equation}
\frac{1}{mN^2} \left| \widehat{\mu}_j^{\mathcal{W}} - \mu_j^{\chi} \right| = O_p \left(\max \left(\frac{1}{N \sqrt{T}},\frac{1}{m N^{2-\gamma}} \right) \right).\nonumber
\end{equation}
Last, for part (ref), given part (ref), Lemma (ref)(ref), and Theorem 2 in yu2015useful, which is a special case of Davis-Kahn Theorem
\begin{equation*}
\norm{\widehat{V}^{\mathcal{W}} - V^{\chi} J}
\leq
\frac{ \norm{ \widehat{\Gamma}^{\mathcal{W}} - \Gamma^{\chi} } }{\mu^{\chi}_r}
=
\frac{N^2 m O_p\left( \frac{1}{N\sqrt{T}}, \frac{1}{N^{2-\gamma}m} \right) }{O(N^2 m)} = O_p\left( \frac{1}{N\sqrt{T}}, \frac{1}{N^{2-\gamma}m} \right),\nonumber
\end{equation*}
and this completes the proof.
lemmaUnder Assumptions (ref)-(ref), for all $m,T, N \in \mathbb{N}$
\begin{enumerate}[label=(\roman*)]
• $\norm{ \frac{M^{\chi}}{m N^2}} = O(1)$.
• $\norm{ \left(\frac{M^{\chi}}{m N^2} \right)^{-1}} = O(1)$.
• $\norm{ \frac{\widehat{M}^{\mathcal{W}}}{m N^2}} = O_p(1)$.
• $\norm{ \left( \frac{\widehat{M}^{\mathcal{W}}}{m N^2} \right)^{-1}} = O_p(1)$.
\end{enumerate}
proofParts (ref) and (ref) follow directly from Lemma (ref), indeed,
\begin{equation}
\norm{\frac{M^{\chi}}{m N^2}} = \frac{\mu_1^{\chi}}{m N^2} \leq \overline{C}_1,\nonumber
\end{equation}
and
\begin{equation}
\norm{ \left( \frac{M^{\chi}}{m N^2} \right)} = \frac{mN^2}{\mu_r^{\chi}} \leq C_r.\nonumber
\end{equation}
Both statements hold for all $m \in \mathbb{N}$ since the eigenvalues are an increasing sequence in $m$ and $N$.
For part (ref), because of Lemma (ref)(ref),
\begin{equation}
\norm{ \frac{\widehat{M}^{\mathcal{W}}}{m N^2}}
\leq
\norm{ \frac{M^{\chi}}{m N^2} }
+
\norm{ \frac{\widehat{M}^{\mathcal{W}}}{m N^2} - \frac{M^{\chi}}{m N^2} }
\leq
\overline{C}_1 + O_p\left( \max \left( \frac{1}{N \sqrt{T}}, \frac{1}{N^{2-\gamma} m } \right) \right).\nonumber
\end{equation}
For part (ref), just notice that, because of Lemma (ref)(ref) and part (ref), then $\frac{\widehat{M}^{\mathcal{W}}}{m N^2} $ is positive definite with probability tending to one as $m, N, T \to \infty$. This completes the proof.
lemmaUnder Assumptions (ref)-(ref), for any given $i=1,\ldots ,m$ and
$m,T \to \infty$
\begin{enumerate}[label=(\roman*)]
• $\frac{1}{m N^2} \norm{\varepsilon_i' \left(\widehat{\Gamma}^{\mathcal{W}} - \Gamma^{\chi}\right) } = O_p \left( \max \left( \frac{1}{N \sqrt{T}}, \frac{1}{\sqrt{N^{2-\gamma} m}}\right) \right)$ .
• $\norm{\sqrt {m} v_i^\chi}=O_p(1)$.
• $\norm{\sqrt {m}\, \widehat v_i^{\mathcal W'}-\sqrt{m} v_i^{\chi'}J }=O_p\left( \frac{1}{N\sqrt{T}}, \frac{1}{N^{2-\gamma}m} \right)$.
\end{enumerate}
proofPart (ref) follows directly from Lemma (ref)(ref).
For part (ref) notice that for all $i=1,\ldots, m$, since $N^{-2}\Gamma^{\mathcal F}=I_r$ because of Assumption (ref)(ref), we must have
\begin{equation}
\frac{Var(\chi_{it})}{N^2}= u_i'u_i \le M_U^2,
\end{equation}
which is finite for all $i$ and $t$. Then, by Lemma (ref)(ref)
\begin{align}
\lim\inf_{m\to\infty}\max_{i=1,\ldots, m} Var(\chi_{it})&= \lim\inf_{m\to\infty}\max_{i=1,\ldots, m}\sum_{j=1}^r\mu_j^\chi \abs{V_{ij}^\chi}^2
\ge \lim\inf_{m\to\infty}\max_{i=1,\ldots, m}\mu_r^\chi \sum_{j=1}^r\abs{V_{ij}^\chi}^2\nonumber\\
&\ge \underline C_r mN^2\max_{i=1,\ldots, m}\norm {v_i^\chi}^2,\nonumber
\end{align}
and by (ref) we must have
\[
\underline C_r mN^2\max_{i=1,\ldots, m}\norm {v_i^\chi}^2\le M_U^2 N^2,
\]
which implies
\[
m\max_{i=1,\ldots, m}\norm {v_i^\chi}^2\le \frac{M_U^2}{\underline C_r }.
\]
This proves part (ref).
For part (ref) we follow the same approach as in the proof of Lemma (ref)(ref) but using part (ref).
lemmaUnder Assumptions (ref)-(ref), for $m,T \to \infty$
\begin{enumerate}[label=(\roman*)]
• $\frac{1}{\sqrt{m}}\norm{ \widehat{U} - U J } = O_p \left(\max \left(\frac{1}{N \sqrt{T}},\frac{1}{N^{2-\gamma} m} \right) \right)$.
• for any given $i=1,\ldots ,m$, $\norm{ \widehat{u_i}' - u_i' J } = O_p \left(\max \left(\frac{1}{N \sqrt{T}},\frac{1}{\sqrt{N^{2-\gamma} m}} \right) \right)$.
\end{enumerate}
proofFor part (ref), first notice that $\text{rk} \left(\frac{U}{\sqrt{m}} \right) = r$ for all $m$, since $\text{rk} \left( \frac{\Gamma^{\mathcal{F}}}{N^2} \right) = r$ by Assumption (ref)(ref) and $\text{rk} \left(\frac{\Gamma^{\chi}}{m} \right) = r$ by Lemma (ref)(ref).
Indeed,
$\text{rk} \left( \frac{\Gamma^{\mathcal{F}}}{m N^2} \right) \leq
\min
\left(
\text{rk} \left( \frac{\Gamma^{\mathcal{F}}}{N^2} \right),
\text{rk} \left(\frac{U}{\sqrt{m}} \right)
\right)$.
This holds for all $m>\overline{m}$ and since eigenvalues are an increasing sequence in $m$.
Therefore, $\left( \frac{U'U}{m}\right)^{-1}$ is well defined for all $m$ and $U$ admits a left inverse.
Now, because of Lemma (ref)(ref), (ref)(ref), (ref)(i), using (ref)
\begin{align}
\norm{ \frac{\widehat{U} - U J}{\sqrt{m}} }
&=
\norm{\widehat{V}^{\mathcal{W}} \left( \frac{\widehat{M}^{\mathcal{W}}}{m N^2} \right)^{1/2}
-
V^{\chi} \left( \frac{M^{\chi}}{m N^2} \right)^{1/2} J } \nonumber\\
&\leq \norm{\widehat{V}^{\mathcal{W}} - V^{\chi} J} \norm{\frac{M^{\chi}}{m N^2}}
+ \norm{ \frac{1}{\sqrt{m} N} \left\{ \left( \widehat{M}^{\mathcal{W}} \right)^{1/2} - \left( M^{\chi} \right)^{1/2} \right\} } \norm{V^{\chi}} \nonumber \\
&+ \norm{\widehat{V}^{\mathcal{W}} - V^{\chi} J} \norm{ \frac{1}{\sqrt{m} N} \left\{ \left( \widehat{M}^{\mathcal{W}} \right)^{1/2} - \left( M^{\chi} \right)^{1/2} \right\} }\nonumber \\
& = O_p \left(\max \left( \frac{1}{N \sqrt{T}}, \frac{1}{m N^{2-\gamma}} \right) \right).\nonumber
\end{align}
For part (ref), because of Lemmas
(ref)(ref),
(ref)(ref),
(ref)(ref), and
(ref)(ref)
\begin{align}
\norm{\widehat{u}_i'-u_i'J} &=
\norm{\sqrt {m}\, \widehat v_i^{\mathcal W'}\left(\frac{\widehat M^{\mathcal W}}{mN^2}\right)^{1/2}-\sqrt{m} v_i^{\chi'}J
\left(\frac{M^\chi}{mN^2}\right)^{1/2}}\nonumber\\
&\le \norm{\sqrt {m}\, \widehat v_i^{\mathcal W'}-\sqrt{m} v_i^{\chi'}J }\,\norm{\frac{\widehat M^{\mathcal W}}{mN^2}}
+\norm{\frac 1{\sqrt{m}} \left\{\left(\widehat M^{\mathcal W}\right)^{1/2}-\left( M^{\chi}\right)^{1/2}\right\} }\,\norm{\sqrt {m} v_i^\chi}\nonumber\\
&+\norm{\sqrt {m}\, \widehat v_i^{\mathcal W'}-\sqrt{m} v_i^{\chi'}J }\,\norm{\frac 1{\sqrt{mN^2}} \left\{\left(\widehat M^{\mathcal W}\right)^{1/2}-\left( M^{\chi}\right)^{1/2}\right\} }\nonumber\\
&=O_p \left(\max \left( \frac{1}{N \sqrt{T}}, \frac{1}{\sqrt {m N^{2-\gamma}}} \right) \right).\nonumber
\end{align}
lemmaUnder Assumptions (ref)-(ref), for all $m, N, T \in \mathbb{N}$
\begin{enumerate}[label=(\roman*)]
• $
\norm{ \frac{ \mathcal{F}_{(3)} \mathcal{E}_{(3)}' }{\sqrt{m} N^2T} }
= O_p\left( \frac{1}{ \sqrt{T} N^{2-\gamma/2}} \right) $.
• $
\norm{ \frac{ \mathcal{E}_{(3)} \mathcal{E}_{(3)}' }{mN^2T} }
= O_p \left( \max \left( \frac{1}{N^{2-\gamma/2}\sqrt{T}}, \frac{1}{mN^{2-\gamma}} \right) \right)$
and
$\norm{\frac{U' \mathcal{E}_{(3)} \mathcal{E}_{(3)}' }{m^{3/2} N^2T}} =
O_p \left( \max \left( \frac{1}{\sqrt{mT}N^{2-\gamma/2}}, \frac{1}{mN^{2-\gamma}} \right) \right)$.
• $ \norm{\frac{ U'U }{m}} = O\left( 1 \right) $.
\end{enumerate}
proofPart (ref) follows directly from Assumption (ref).
For part (ref), by Lemma (ref)(ref) and Lemma (ref)(ref)
\begin{align}
\norm{ \frac{ \mathcal{E}_{(3)} \mathcal{E}_{(3)}' }{mN^2T} }
\leq& \norm{ \frac{ \mathcal{E}_{(3)} \mathcal{E}_{(3)}' }{mN^2T}
- \frac{\Gamma^{\mathcal{E}}}{mN^2} }
+ \norm{ \frac{\Gamma^{\mathcal{E}}}{mN^2} }\nonumber \\
&= O_p \left( \frac{1}{N^{2-\gamma/2}\sqrt{T}}\right) + O\left(\frac{1}{mN^{2-\gamma}} \right).\nonumber
\end{align}
Similarly, by Lemma (ref)(ref) and Lemma (ref)(ref)
\begin{align}
\norm{\frac{U' \mathcal{E}_{(3)} \mathcal{E}_{(3)}' }{m^{3/2} N^2T}}
&\leq
\norm{\frac{U' \mathcal{E}_{(3)} \mathcal{E}_{(3)}' }{m^{3/2} N^2T} - \frac{U \Gamma^{\mathcal{E}}}{m^{3/2}N^2}}
+
\norm{ \frac{U' }{\sqrt{m}}}
\norm{ \frac{\Gamma^{\mathcal{E}}}{m N^2}} \nonumber\\
&=\norm{\frac{U \mathcal{E}_{(3)} \mathcal{E}_{(3)}' }{m^{3/2} N^2T} - \frac{U \Gamma^{\mathcal{E}}}{m^{3/2}N^2}}
+
O\left(\frac{1}{mN^{2-\gamma}} \right).\nonumber
\end{align}
Then, because of Assumption (ref)(ref),
\begin{align}
\mathbb{E} \left[\norm{\frac{U' \mathcal{E}_{(3)} \mathcal{E}_{(3)}' }{m^{3/2} N^2T} - \frac{U \Gamma^{\mathcal{E}}}{m^{3/2}N^2}}^2 \right]
\leq&
\mathbb{E} \left[\norm{\frac{U' \mathcal{E}_{(3)} \mathcal{E}_{(3)}' }{m^{3/2} N^2T} - \frac{U \Gamma^{\mathcal{E}}}{m^{3/2}N^2}}^2_F \right]\nonumber \\
=&
\sum_{k=1}^{r} \sum_{j=1}^{m}
\mathbb{E} \left[
\abs{
\frac{1}{m^{3/2} N^2T}
\sum_{i=1}^{m}
\sum_{t=1}^{T}
U_{ik} \mathcal{E}_{(3)t i \cdot} \mathcal{E}_{(3)t j \cdot}'
- U_{ik} \mathbb{E} \left[ \mathcal{E}_{(3)t i \cdot} \mathcal{E}_{(3)t j \cdot}' \right]
}^2
\right] \nonumber\\
\leq&
\frac{r M_U^2 m}{m^2N^{4-\gamma} T}
\max_{j=1, \dots, m}
\mathbb{E} \left[
\abs{
\frac{1}{\sqrt{mT} N^{\gamma/2}}
\sum_{i=1}^{m}
\sum_{t=1}^{T}
\mathcal{E}_{(3)t i \cdot} \mathcal{E}_{(3)t j \cdot}'
- \mathbb{E} \left[ \mathcal{E}_{(3)t i \cdot} \mathcal{E}_{(3)t j \cdot}' \right]
}^2
\right]\nonumber\\
\leq&
\frac{r M_U^2 C_{\mathcal{E}}}{mN^{4-\gamma} T}.\nonumber
\end{align}
This proves part (ref).
Part ((ref) follows from Lemma (ref)(ref), since
\begin{equation}
\norm{\frac{ U'U }{m}} \leq \norm{\frac{ U}{\sqrt{m}}}^2 \leq M_U^2 .\nonumber
\end{equation}
lemma[Consistency of loadings]
Under Assumptions (ref)-(ref)
\begin{enumerate}[label=(\roman*)]
• $
\norm{ \widehat{u}_i' - u_i' \widehat{H} }
= O_p \left(
\max \left( \frac{1}{N^{2-\gamma/2} \sqrt{T}}, \frac{1}{m N^{2-\gamma}} \right) \right),
$
where $\widehat H = \frac{\mathcal{F}_{(3)} \mathcal{F}_{(3)}'}{N^2 T} \frac{U'\widehat{U}}{m} \left( \frac{\widehat{M}^{\mathcal{W}}}{m N^2} \right)^{-1}$.
• $
\norm{ \frac{\widehat{U} - U\widehat{H}}{\sqrt{m}} }
= O_p \left(
\max \left( \frac{1}{N^{2-\gamma/2} \sqrt{T}}, \frac{1}{m N^{2-\gamma}} \right) \right).
$
\end{enumerate}
proofSubstituting
\begin{equation}
\mathcal{W}_{(3)}\mathcal{W}_{(3)}' =
U \mathcal{F}_{(3)} \mathcal{F}_{(3)}' U' + U\mathcal{F}_{(3)} \mathcal{E}_{(3)}' +
\mathcal{E}_{(3)} \mathcal{F}_{(3)}'U' +\mathcal{E}_{(3)} \mathcal{E}_{(3)}'\nonumber
\end{equation}
into
\begin{equation}
\frac{\mathcal{W}_{(3)}\mathcal{W}_{(3)}'}{m N^2 T} \widehat{U} = \widehat{U} \frac{\widehat{M}^{\mathcal{W}}}{m N^2},\nonumber
\end{equation}
where $\widehat{M}^{\mathcal{W}}$ are the eigenvalues of $\mathcal{W}_{(3)}\mathcal{W}_{(3)}'/T$, we get:
\begin{equation}
\frac{U \mathcal{F}_{(3)} \mathcal{F}_{(3)}' U'\widehat{U}}{m N^2 T}
+ \frac{U\mathcal{F}_{(3)} \mathcal{E}_{(3)}'\widehat{U}}{m N^2 T}
+
\frac{\mathcal{E}_{(3)} \mathcal{F}_{(3)}'U'\widehat{U}}{m N^2 T} +\frac{\mathcal{E}_{(3)} \mathcal{E}_{(3)}'\widehat{U}}{m N^2 T} = \widehat{U} \frac{\widehat{M}^{\mathcal{W}}}{m N^2}. \nonumber
\end{equation}
Define
\begin{equation}
\widehat{H} := \frac{\mathcal{F}_{(3)} \mathcal{F}_{(3)}'}{N^2 T} \frac{U'\widehat{U}}{m} \left( \frac{\widehat{M}^{\mathcal{W}}}{m N^2} \right)^{-1}.
\end{equation}
Then,
\begin{align}
\widehat{U} - U \widehat{H} =&
\left(
\frac{U \mathcal{F}_{(3)} \mathcal{E}_{(3)}' \widehat{U}}{m N^2 T} +
\frac{\mathcal{E}_{(3)} \mathcal{F}_{(3)}' U'\widehat{U}}{m N^2 T} +
\frac{\mathcal{E}_{(3)} \mathcal{E}_{(3)}' \widehat{U}}{m N^2 T}
\right)
\left( \frac{\widehat{M}^{\mathcal{W}}}{m N^2} \right)^{-1}\nonumber\\
=& \left(
\frac{U \mathcal{F}_{(3)} \mathcal{E}_{(3)}' U}{m N^2 T} +
\frac{\mathcal{E}_{(3)} \mathcal{F}_{(3)}' U'U}{m N^2 T} +
\frac{\mathcal{E}_{(3)} \mathcal{E}_{(3)}' U}{m N^2 T}
\right)
J
\left( \frac{\widehat{M}^{\mathcal{W}}}{m N^2} \right)^{-1}\nonumber \\
& + \left(
\frac{U \mathcal{F}_{(3)} \mathcal{E}_{(3)}' }{m N^2 T} +
\frac{\mathcal{E}_{(3)} \mathcal{F}_{(3)}' U'}{m N^2 T} +
\frac{\mathcal{E}_{(3)} \mathcal{E}_{(3)}' }{m N^2 T}
\right)
\left( \widehat{U} - U J \right)
\left( \frac{\widehat{M}^{\mathcal{W}}}{m N^2} \right)^{-1}
\end{align}
Taking the $i$-th row of (ref), we have that:
\begin{align}
&\left( \widehat{u}_i' - u_i' \widehat{H} \right) \nonumber
\\
&=
\left(
\frac{1}{m N^2 T} u_i' \sum_{t=1}^{T} \sum_{j=1}^{m} \mathcal{F}_{(3)t} \mathcal{E}_{(3)tj\cdot}' u_j'
+ \frac{1}{m N^2 T} \sum_{t=1}^{T} \mathcal{E}_{(3)ti\cdot} \mathcal{F}_{(3)t}' \sum_{j=1}^{m} u_j u_j'
+ \frac{1}{m N^2 T} \sum_{t=1}^{T} \sum_{j=1}^{m} \mathcal{E}_{(3)ti\cdot} \mathcal{E}_{(3)tj\cdot}' u_j'
\right) \nonumber\\
&\times
J
\left( \frac{\widehat{M}^{\mathcal{W}}}{m N^2} \right)^{-1}\nonumber \\
&+ \Bigg(
\frac{1}{m N^2 T} u_i' \sum_{t=1}^{T} \sum_{j=1}^{m} \mathcal{F}_{(3)t} \mathcal{E}_{(3)tj\cdot}' (\widehat{u}_j' - u_j' J)
+ \frac{1}{m N^2 T} \sum_{t=1}^{T} \mathcal{E}_{(3)ti\cdot} \mathcal{F}_{(3)t}' \sum_{j=1}^{m} u_j (\widehat{u}_j' - u_j' J)\nonumber\\
&+ \frac{1}{m N^2 T} \sum_{t=1}^{T} \sum_{j=1}^{m} \mathcal{E}_{(3)ti\cdot} \mathcal{E}_{(3)tj\cdot}' (\widehat{u}_j' - u_j' J)
\Bigg)
\left( \frac{\widehat{M}^{\mathcal{W}}}{m N^2} \right)^{-1}\nonumber\\
&= (a+b+c) J
\left( \frac{\widehat{M}^{\mathcal{W}}}{m N^2} \right)^{-1}
+ (d+e+f)
\left( \frac{\widehat{M}^{\mathcal{W}}}{m N^2} \right)^{-1}.
\end{align}
First consider part (ref).
For term $a$, by Assumption (ref)(ref),
\begin{equation}
\norm{
\frac{1}{m N^2 T} u_i' \sum_{t=1}^{T} \sum_{j=1}^{m} \mathcal{F}_{(3)t} \mathcal{E}_{(3)tj\cdot}' u_j' }
\leq
\norm{u_i}
\norm{
\frac{1}{m N^2 T} \sum_{t=1}^{T} \sum_{j=1}^{m} \mathcal{F}_{(3)t} \mathcal{E}_{(3)tj\cdot}' u_j' }
\leq
M_U
\norm{
\frac{1}{m N^2 T} \sum_{t=1}^{T} \sum_{j=1}^{m} \mathcal{F}_{(3)t} \mathcal{E}_{(3)tj\cdot}' u_j' },\nonumber
\end{equation}
for any $i=1, \dots, m$.
Then, by Assumption (ref)(ref) and Lemma (ref)
\begin{align}
\mathbb{E}
\left[
\norm{
\frac{1}{m N^2 T} \sum_{t=1}^{T} \sum_{j=1}^{m} \mathcal{F}_{(3)t} \mathcal{E}_{(3)tj\cdot}' u_j' }^2
\right]
&\leq
\mathbb{E}
\left[
\norm{
\frac{1}{m N^2 T} \sum_{t=1}^{T} \sum_{j=1}^{m} \mathcal{F}_{(3)t} \mathcal{E}_{(3)tj\cdot}' u_j' }^2_F
\right]\nonumber\\
&=
\frac{1}{m^2 N^4 T^2}
\sum_{h=1}^{r}
\mathbb{E}
\left[
\norm{
\sum_{t=1}^{T} \sum_{j=1}^{m} \mathcal{F}_{(3)t} \mathcal{E}_{(3)tj\cdot}' [U_{jh}] }^2_F
\right]\nonumber\\
&=
\max_{h=1, \dots, r}
\frac{r}{m^2 N^4 T^2}
\mathbb{E}
\left[
\norm{
\sum_{t=1}^{T} \sum_{j=1}^{m} \mathcal{F}_{(3)t} \mathcal{E}_{(3)tj\cdot}' [U_{jh}] }^2_F
\right]\nonumber\\
&=
\frac{r M_U^2}{m^2 N^4 T^2}
\mathbb{E}
\left[
\norm{
\sum_{t=1}^{T} \sum_{j=1}^{m} \mathcal{F}_{(3)t} \mathcal{E}_{(3)tj\cdot}' }^2_F
\right]\nonumber\\
&=
\frac{r M_U^2}{m N^{4-\gamma} T}
\mathbb{E} \left[\frac 1{m N^{\gamma} T} \sum_{i=1}^{m} \norm{ \sum_{t=1}^{T} \mathcal{F}_{(3)t} \mathcal{E}_{(3)t i\cdot}^\prime }^2 \right]\nonumber\\
&\leq
\frac{r M_U^2 C_{\mathcal{FE}}}{m N^{4-\gamma} T}
=
O_p \left(\frac{1}{mTN^{4-\gamma}}\right).
\end{align}
Therefore, term $a$ is $O_p (\frac{1}{\sqrt{mT}N^{2-\gamma/2}})$.
For term $b$, for any $i=1, \dots, m$, because of Lemma (ref)(ref)
\begin{align}
\norm{\frac{1}{m N^2 T} \sum_{t=1}^{T} \mathcal{E}_{(3)ti\cdot} \mathcal{F}_{(3)t}' \sum_{j=1}^{m} u_j u_j'} &=
\norm{ \frac{1}{m N^2 T} \sum_{t=1}^{T} \mathcal{E}_{(3)ti\cdot} \mathcal{F}_{(3)t}' (U' U)}\nonumber\\
&\leq \norm{ \frac{1}{N^2 T} \sum_{t=1}^{T} \mathcal{E}_{(3)ti\cdot} \mathcal{F}_{(3)t}'}
\norm{ \frac{U}{\sqrt{m}}}^2\nonumber\\
&\leq \norm{ \frac{1}{N^2 T} \sum_{t=1}^{T} \mathcal{E}_{(3)ti\cdot} \mathcal{F}_{(3)t}'} M_U^2.
\end{align}
Then, by Assumption (ref) with $m=1$,
\begin{align}
\norm{ \frac{1}{N^2 T} \sum_{t=1}^{T} \mathcal{E}_{(3)ti\cdot} \mathcal{F}_{(3)t}'}
&= \frac{\sqrt{m}}{N^{2-\gamma/2} \sqrt{T}}
\norm{ \frac{1}{ \sqrt{mT} N^{\gamma/2}} \sum_{t=1}^{T} \mathcal{E}_{(3)ti\cdot} \mathcal{F}_{(3)t}'}\nonumber \\
&= \frac{\sqrt{m}}{N^{2-\gamma/2} \sqrt{T}} O_p \left( \frac{1}{\sqrt{m}}\right).
\end{align}
By substituting (ref) into (ref), we prove that term $b$ is $O_p (\frac{1}{\sqrt{T} N^{2-\gamma/2}})$
For term $c$, for any $i=1, \dots, m$, because of Assumption (ref)(ref),
\begin{align}
\norm{\frac{1}{m N^2 T} \sum_{t=1}^{T} \sum_{j=1}^{m} \mathcal{E}_{(3)ti\cdot} \mathcal{E}_{(3)tj\cdot}' u_j'}
&=
\left\{ \sum_{t=1}^{T}\sum_{j=1}^{m} \left( \frac{1}{m N^2 T}
\sum_{t=1}^{T} \sum_{j=1}^{m} \mathcal{E}_{(3)ti\cdot} \mathcal{E}_{(3)tj\cdot}' U_{jk}
\right)^2 \right\}^{1/2}\nonumber\\
&\leq \sqrt{r} M_U \abs{ \frac{1}{m N^2 T} \sum_{t=1}^{T} \sum_{j=1}^{m} \mathcal{E}_{(3)ti\cdot} \mathcal{E}_{(3)tj\cdot}' }\nonumber \\
&\leq
\sqrt{r} M_U
\abs{\frac{1}{m N^2 T} \sum_{t=1}^{T} \sum_{j=1}^{m} \mathbb{E} \left[ \mathcal{E}_{(3)ti\cdot} \mathcal{E}_{(3)tj\cdot}' \right] }\nonumber \\
&+
\sqrt{r} M_U
\abs{ \frac{1}{m N^2 T} \sum_{t=1}^{T} \sum_{j=1}^{m}
\left\{ \mathcal{E}_{(3)ti\cdot} \mathcal{E}_{(3)tj\cdot}' - \mathbb{E} \left[ \mathcal{E}_{(3)ti\cdot} \mathcal{E}_{(3)tj\cdot}' \right]
\right\} }.\nonumber
\end{align}
By Assumption (ref)(ref),
\begin{align}
\abs{\frac{1}{m N^2 T} \sum_{t=1}^{T} \sum_{j=1}^{m} \mathbb{E} \left[ \mathcal{E}_{(3)ti\cdot} \mathcal{E}_{(3)tj\cdot}' \right] }
&\leq
\frac{1}{m N^2 T} \sum_{t=1}^{T} \sum_{j=1}^{m} \abs{\mathbb{E} \left[ \mathcal{E}_{(3)ti\cdot} \mathcal{E}_{(3)tj\cdot}' \right] }
\nonumber\\
&\leq
\frac{1}{m N^2 T} \sum_{t=1}^{T} \sum_{j=1}^{m} M_{ij} N^{\gamma}
\leq \frac{M_{\mathcal{E}}}{mN^{2-\gamma}}.\nonumber
\end{align}
Moreover, by Assumption (ref)(ref)
\begin{equation}
\mathbb{E} \left[
\abs{\frac{1}{m N^2 T} \sum_{t=1}^{T} \sum_{j=1}^{m}
\left\{ \mathcal{E}_{(3)ti\cdot} \mathcal{E}_{(3)tj\cdot}' - \mathbb{E} \left[ \mathcal{E}_{(3)ti\cdot} \mathcal{E}_{(3)tj\cdot}' \right]
\right\}}^2
\right]
\leq
\frac{C_{\mathcal{E}}}{mN^{4-\gamma}T}.\nonumber
\end{equation}
Hence, term $c$ is $O_p \left( \frac{1}{mN^{2-\gamma}} , \frac{1}{\sqrt{mT} N^{2-\gamma/2}} \right)$.
Given Lemma (ref)(ref), the terms $d$, $e$, and $f$ are dominated, since they are similar to $a$, $b$, and $c$, but with $(\widehat{u}'_j-u_j'J)$ replacing $u_j'$. This completes the proof of part (ref).
Second, consider part (ref).
From (ref)
\begin{align}
\norm{\frac{ \widehat{U} - U \widehat{H} }{\sqrt{m}}}
\leq&
\left(
\norm{\frac{\mathcal{F}_{(3)} \mathcal{E}_{(3)}' }{\sqrt{m}N^2 T}}
\norm{\frac{U}{\sqrt{m}}}^2
+
\norm{\frac{\mathcal{E}_{(3)} \mathcal{F}_{(3)}' }{\sqrt{m}N^2 T}}
\norm{\frac{U'U}{m}}
+
\norm{\frac{\mathcal{E}_{(3)} \mathcal{E}_{(3)}' }{mN^2 T}}
\norm{\frac{U}{\sqrt{m}}}
\right)
\norm{J} \norm{\left(\frac{\widehat{M}^{\mathcal{W}}}{mN^2} \right)^{-1}}\nonumber\\
&+
\left(
\norm{\frac{\mathcal{F}_{(3)} \mathcal{E}_{(3)}' }{\sqrt{m}N^2 T}}
\norm{\frac{U}{\sqrt{m}}}
+
\norm{\frac{\mathcal{E}_{(3)} \mathcal{F}_{(3)}' }{\sqrt{m}N^2 T}}
\norm{\frac{U}{\sqrt{m}}}
+
\norm{\frac{\mathcal{E}_{(3)} \mathcal{E}_{(3)}' }{mN^2 T}}
\right)
\norm{\frac{ \widehat{U} - U J }{\sqrt{m}}} \norm{\left(\frac{\widehat{M}^{\mathcal{W}}}{mN^2} \right)^{-1}}\nonumber\\
=&
O_p \left(\max \left(
\frac{1}{\sqrt{T}N^{2-\gamma/2}},
\frac{1}{mN^{2-\gamma}}
\right) \right),\nonumber
\end{align}
following from Lemma (ref)(ref), (ref), (ref)(ref), (ref)(ref), (ref)(ref), (ref)(ref) and $\norm{J} = O(1)$.
This completes the proof.
lemmaUnder Assumptions (ref)-(ref), as $m, N, T \to \infty$, $ \norm{\widehat{H}} = O(1)$.
proofBy Lemmas (ref), (ref) and (ref)
\begin{align}
\norm{\widehat{H}}
&\leq
\norm{\frac{\mathcal{F}_{(3)} \mathcal{F}_{(3)}'}{N^2 T}} \norm{\frac{U'\widehat{U}}{m}}
\norm{\left( \frac{\widehat{M}^{\mathcal{W}}}{m N^2} \right)^{-1} }
\leq
\norm{\frac{\mathcal{F}_{(3)}}{N \sqrt{T}}}^2 \norm{\frac{U}{\sqrt{m}}}
\norm{\frac{\widehat U}{\sqrt{m}}}
\norm{\left( \frac{\widehat{M}^{\mathcal{W}}}{m N^2} \right)^{-1} }\nonumber \\
&\leq
\norm{\frac{\mathcal{F}_{(3)}}{N \sqrt{T}}}^2 \norm{\frac{U}{\sqrt{m}}}
\norm{J}
\norm{\left( \frac{\widehat{M}^{\mathcal{W}}}{m N^2} \right)^{-1} }
+
\norm{\frac{\mathcal{F}_{(3)}}{N \sqrt{T}}}^2 \norm{\frac{U}{\sqrt{m}}}
\norm{\frac{ \widehat{U} - U \widehat{H} }{\sqrt{m}}}
\norm{\left( \frac{\widehat{M}^{\mathcal{W}}}{m N^2} \right)^{-1} }\nonumber \\
&= O_p(1) + O_p \left(\max \left(
\frac{1}{\sqrt{T}N^{2-\gamma/2}},
\frac{1}{mN^{2-\gamma}}
\right) \right).\nonumber
\end{align}
lemmaUnder Assumptions (ref)-(ref), as $m,N,T \to \infty$
\begin{enumerate}[label=(\roman*)]
• $ \norm{\widehat{H} -J} = O_p\left(\frac{1}{\xi}\right) = o_p \left( \max \left( \frac{1}{\sqrt{T}N^{2-\gamma/2}}, \frac{1}{\sqrt mN^{2-\gamma/2}} \right)\right)$,
• $\norm{\widehat{H}^{-1} -J} = O_p\left(\frac{1}{\xi}\right) = o_p \left( \max \left( \frac{1}{\sqrt{T}N^{2-\gamma/2}}, \frac{1}{\sqrt mN^{1-\gamma/2}} \right)\right)$,
\end{enumerate}
where $\xi =
\min \left(
\sqrt{mT} N^{2-\gamma/2},
N^{3-\gamma/2} T,
m \sqrt{T} N^{2-\gamma},
m N^{2-\gamma}
\right)
$.
proofFor part (ref), using (ref) we have
\begin{align}
\frac{\widehat{U}' U \widehat{H}}{m}
&=
\frac{\widehat{U}' \widehat{U}}{m}
+
\frac{\widehat{U}' (U \widehat{H}- \widehat{U} ) }{m}\nonumber\\
&= \frac{\widehat{M}^{\mathcal{W}}}{mN^2}
+
\frac{\widehat{U}' (U \widehat{H}- \widehat{U} ) }{m} .\nonumber
\end{align}
Now,
\begin{equation}
\frac{\widehat{U}' (U \widehat{H}- \widehat{U} ) }{m}
=
\frac{ (\widehat{U} - U \widehat{H} )' U \widehat{H} }{m} + \frac{(\widehat{U} - U \widehat{H} )' (\widehat{U} - U \widehat{H} )}{m}.\nonumber
\end{equation}
Then, \begin{align}
\norm{ \frac{\widehat{U}' (U \widehat{H}- \widehat{U} ) }{m} }
&\leq \norm{ \frac{ (\widehat{U} - U \widehat{H} )' U \widehat{H} }{m}}
+ \norm{ \frac{(\widehat{U} - U \widehat{H} )' (\widehat{U} - U \widehat{H} )}{m} }= I + II.\nonumber
\end{align}
First, consider $I$.
From (ref)
\begin{align}
I =& \frac{ (\widehat{U} - U \widehat{H} )' U \widehat{H} }{m}\nonumber\\
=& \left( \frac{\widehat{M}^{\mathcal{W}}}{mN^2} \right)^{-1} J \frac{U'U}{m} \frac{\mathcal{F}_{(3)} \mathcal{E}_{(3)}' U \widehat{H}}{m N^2 T}
+ \left( \frac{\widehat{M}^{\mathcal{W}}}{mN^2} \right)^{-1} J \frac{U' \mathcal{E}_{(3)} \mathcal{F}_{(3)}' }{m N^2 T}\frac{U'U \widehat{H}}{m}
+ \left( \frac{\widehat{M}^{\mathcal{W}}}{mN^2} \right)^{-1} J \frac{U' \mathcal{E}_{(3)} \mathcal{E}_{(3)}' U \widehat{H}}{m^2 N^2 T} \nonumber \\
&+ \left( \frac{\widehat{M}^{\mathcal{W}}}{mN^2} \right)^{-1} \frac{(\widehat{U} - UJ)'}{\sqrt{m}} \frac{ \mathcal{E}_{(3)} \mathcal{F}_{(3)}' U'}{mN^2T} \frac{U \widehat{H}}{\sqrt{m}}
+ \left( \frac{\widehat{M}^{\mathcal{W}}}{mN^2} \right)^{-1} \frac{(\widehat{U} - UJ)'}{\sqrt{m}} \frac{U \mathcal{F}_{(3)} \mathcal{E}_{(3)}'}{mN^2T} \frac{U \widehat{H}}{\sqrt{m}}\nonumber \\
&+ \left( \frac{\widehat{M}^{\mathcal{W}}}{mN^2} \right)^{-1} \frac{(\widehat{U} - UJ)'}{\sqrt{m}} \frac{\mathcal{E}_{(3)} \mathcal{E}_{(3)}'}{mN^2T} \frac{U \widehat{H}}{\sqrt{m}}\nonumber \\
=& I_a + I_b + I_c + I_d + I_e + I_f.\nonumber
\end{align}
Then, given (ref),
\begin{equation}
\mathbb{E} \left[ \norm{ \frac{\mathcal{F}_{(3)} \mathcal{E}_{(3)}' U}{m N^2 T}}^2 \right]
= O \left( \frac{1}{mTN^{4-\gamma}} \right).
\end{equation}
Therefore, by Assumption (ref)(i), Lemma (ref)(iv), Lemma (ref) and using $\norm{J} = O(1)$ and (ref),
we get
\begin{gather}
\norm{I_a} \leq
\norm{\left( \frac{\widehat{M}^{\mathcal{W}}}{mN^2} \right)^{-1}} \
\norm{J} \ \norm{\frac{U'U}{m}} \
\norm{ \frac{\mathcal{F}_{(3)} \mathcal{E}_{(3)}'U }{m N^2 T}} \
\norm{\widehat{H}}
= O_p \left( \frac{1}{\sqrt{mT} N^{2-\gamma/2}} \right) \nonumber \\
\norm{I_b} \leq
\norm{\left( \frac{\widehat{M}^{\mathcal{W}}}{mN^2} \right)^{-1}} \
\norm{J} \
\norm{ \frac{U' \mathcal{E}_{(3)} \mathcal{F}_{(3)}'}{m N^2 T}} \
\norm{\frac{U'U}{m}} \
\norm{\widehat{H}}
= O_p \left( \frac{1}{\sqrt{mT} N^{2-\gamma/2}} \right) .\nonumber
\end{gather}
Moreover, because of Lemma (ref)(ref), (ref)(ref),
Lemma (ref)(ref),
Lemma (ref) and $\norm{J} = O(1)$,
\begin{equation}
\norm{I_c}
\leq
\norm{\left( \frac{\widehat{M}^{\mathcal{W}}}{mN^2} \right)^{-1}} \
\norm{J} \
\norm{\frac{U}{\sqrt{m}}} \
\norm{ \frac{U'\mathcal{E}_{(3)} \mathcal{E}_{(3)}'}{m^{3/2} N^2 T} } \
\norm{\widehat{H}}
= O_p \left( \max \left( \frac{1}{\sqrt{mT}N^{2-\gamma/2}}, \frac{1}{mN^{2-\gamma}} \right) \right).\nonumber
\end{equation}
Similarly, because of Lemma (ref), Lemma (ref)(ref), (ref)(ref) and (ref), Lemma (ref)(ref) and (ref),
\begin{align}
\norm{I_d} \leq&
\norm{\left( \frac{\widehat{M}^{\mathcal{W}}}{mN^2} \right)^{-1}} \
\norm{\frac{\widehat{U} - UJ}{\sqrt{m}}} \
\norm{\frac{ \mathcal{E}_{(3)} \mathcal{F}_{(3)}' }{\sqrt{m} N^2T}} \
\norm{\frac{U'U}{m}} \
\norm{\widehat{H}} \nonumber\\
&= O_p(1) \ O_p \left(\max \left( \frac{1}{N \sqrt{T}}, \frac{1}{m N^{2-\gamma}} \right) \right) \
O_p\left( \frac{1}{ \sqrt{T} N^{2-\gamma/2}} \right) O_p(1) O_p(1)\nonumber\\
&= O_p \left(\max \left(
\frac{1}{N^{3-\gamma/2} T},
\frac{1}{m \sqrt{T} N^{2-\gamma}}
\right) \right),\nonumber\\
\norm{I_e} \leq&
\norm{\left( \frac{\widehat{M}^{\mathcal{W}}}{mN^2} \right)^{-1}} \
\norm{\frac{\widehat{U} - UJ}{\sqrt{m}}} \
\norm{\frac{ \mathcal{E}_{(3)} \mathcal{F}_{(3)}' }{\sqrt{m} N^2T}} \
\norm{\frac{U}{\sqrt{m}}}^2 \
\norm{\widehat{H}}\nonumber \\
&= O_p \left(\max \left(
\frac{1}{N^{3-\gamma/2} T},
\frac{1}{m \sqrt{T} N^{2-\gamma}}
\right) \right),\nonumber
\end{align}
and, because of Lemma (ref), Lemma (ref)(ref), (ref)(ref), Lemma (ref)(ref) and (ref)
\begin{align}
\norm{I_f} \leq&
\norm{\left( \frac{\widehat{M}^{\mathcal{W}}}{mN^2} \right)^{-1}} \
\norm{\frac{\widehat{U} - UJ}{\sqrt{m}}} \
\norm{\frac{ \mathcal{E}_{(3)} \mathcal{E}_{(3)}' }{m N^2T}} \
\norm{\frac{U}{\sqrt{m}}} \
\norm{\widehat{H}}\nonumber \\
&= O_p(1) \ O_p \left(\max \left( \frac{1}{N \sqrt{T}}, \frac{1}{m N^{2-\gamma}} \right) \right) \
O_p\left( \frac{1}{ \sqrt{mT} N^{2-\gamma/2}} \right) O_p(1) O_p(1)\nonumber\\
&= O_p \left(\max \left(
\frac{1}{\sqrt{m} N^{3-\gamma/2} T},
\frac{1}{m^{3/2} \sqrt{T} N^{2-\gamma}}
\right) \right).\nonumber
\end{align}
Therefore,
\begin{equation}
\norm{I} = O_p \left(
\max \left(
\frac{1}{\sqrt{mT} N^{2-\gamma/2}},
\frac{1}{N^{3-\gamma/2} T},
\frac{1}{m \sqrt{T} N^{2-\gamma}}
\right)
\right).
\end{equation}
Second, consider $II$. From Lemma (ref)(ref)
\begin{equation}
\norm{II} \leq \frac{1}{m} \norm{\widehat{U}- U \widehat{H}}^2
= O_p \left(
\max \left(
\frac{1}{N^{4-\gamma} T},
\frac{1}{m^2 N^{4-2\gamma}},
\right)
\right).\nonumber
\end{equation}
Therefore,
\begin{equation}
\norm{ \frac{ \widehat{U}' \left(U - \widehat{U} \widehat{H} \right) }{m} }
\leq \norm{I} + \norm{II}
= O_p \left( \frac{1}{\xi}\right).
\end{equation}
Recall that $\xi =
\min \left(
\sqrt{mT} N^{2-\gamma/2},
N^{3-\gamma/2} T,
m \sqrt{T} N^{2-\gamma},
m N^{2-\gamma}
\right)
$.
Because of (ref), using (ref),
\begin{align}
\frac{\widehat{U}' U \widehat{H}}{m}
&= \frac{\widehat{M}^{\mathcal{W}}}{mN^2}
+ O_p \left(\frac{1 }{\xi}\right).\nonumber
\end{align}
Or, equivalently,
\begin{align}
\left( \frac{\widehat{M}^{\mathcal{W}}}{mN^2} \right)^{-1}
\frac{\widehat{U}' {U}}{m}
=
\widehat{H}^{-1}
+ O_p \left(\frac{1 }{\xi}\right).\nonumber
\end{align}
Using this in (ref) and using Assumption (ref) we have
\begin{equation}
\widehat{H}'
=
\left( \frac{\widehat{M}^{\mathcal{W}}}{mN^2} \right)^{-1}
\frac{\widehat{U}' {U}}{m}
=
\widehat{H}^{-1}
+ O_p \left(\frac{1 }{\xi}\right).\nonumber
\end{equation}
Therefore, as $m, T, N \to \infty$, $\widehat{H}$ is an $r \times r$ orthogonal matrix, thus it has eigenvalues $\pm 1$.
Moreover, because of (ref)
\begin{align}
\frac{\widehat{U}' U }{m}
=
\frac{(\widehat{U}' - U \widehat{H} + U \widehat{H} )' U }{m}
&= \frac{\widehat{H}' \widehat{U}' U }{m} + O_p \left(\frac{1 }{\xi}\right).
\end{align}
And from (ref) it follows that
\begin{align}
\left( \frac{\widehat{M}^{\mathcal{W}}}{mN^2} \right)
\widehat{H}'
=
\widehat{H}'
\frac{U' U}{m}
+ O_p \left(\frac{1 }{\xi}\right).\nonumber
\end{align}
So, because of (eq. above), as $m, N, T \to \infty$, the columns of $\widehat{H}$ are the eigenvectors of $\frac{U' U}{m}$ with eigenvalues $\frac{\widehat{M}^{\mathcal{W}}}{mN^2} $.
The eigenvectors are normalized since $\widehat{H}$ is orthogonal, as $m, N, T \to \infty$.
Moreover, under Assumption (ref) (ref), $\frac{U' U}{m}$ is diagonal so, as $m, N, T \to \infty$, $\widehat{H}$ must be diagonal with eigenvalues $\pm 1$.
This proves part (i).
Part (ii) follows from the fact that $\widehat{H}$ is orthogonal, as $m, N, T\to \infty$. This completes the proof.
lemmaUnder Assumptions (ref)-(ref), (ref), and (ref), if $\sqrt{ T}/ (mN^{2-\gamma}) \to 0$, as $m,N,T\to\infty$
$$
\norm{\widehat V^{-1}-V^{-1}}= O_p\left(\max\left(\frac 1{\sqrt N},\sqrt{\frac{\log N}{T}}\right)\right).
$$
proofFirst, we have
\begin{align}
\widehat \nu_t := y_t-\widehat X_t\widehat{\theta}^{\tiny OLS} &= y_t -\left(\widehat X_t-X_t\bar J+X_t\bar J \right)\left(\widehat{\theta}^{\tiny OLS}-\bar J\theta+\bar J\theta \right)\nonumber\\
&= y_t - X_t\bar J\bar J\theta - X_t\bar J\left(\widehat{\theta}^{\tiny OLS}-\bar J\theta\right)-\left(\widehat X_t-X_t\bar J\right)\bar J\theta -\left(\widehat X_t-X_t\bar J\right)\left(\widehat{\theta}^{\tiny OLS}-\bar J\theta\right)\nonumber\\
&= \nu_t - X_t\bar J\left(\widehat{\theta}^{\tiny OLS}-\bar J\theta\right)-\left(\widehat X_t-X_t\bar J\right)\bar J\theta -\left(\widehat X_t-X_t\bar J\right)\left(\widehat{\theta}^{\tiny OLS}-\bar J\theta\right)\nonumber\\
& = \nu_t +\delta_t+\eta_t+\zeta_t.
\end{align}
Moreover,
\begin{align}
\norm{\frac{\widehat X_t-X_t\bar J}{\sqrt N}}&=\frac 1{\sqrt N}\norm{\left( \frac{1}{N} \left( y_{t-1} \otimes I_N \right)' \left( \widehat{\mathcal{F}}_{(3)t-1} - J \mathcal{F}_{(3)t-1} \right)', 0_{N \times 2N} \right)'}\nonumber\\
& \le \norm{\frac{y_{t-1}}{\sqrt N}} \norm{\frac{ \widehat{\mathcal{F}}_{(3)t-1} - J \mathcal{F}_{(3)t-1}}{N}} = O_p\left(\max\left(\frac 1{N^{3-\gamma/2}T},\frac 1{\sqrt m N^{1-\gamma/2}}\right)\right),
\end{align}
because of Theorem (ref)(ref) and since
\[
\mathbb E\left[\norm{\frac{y_{t-1}}{\sqrt N}}^2\right] =\frac 1N\sum_{i=1}^N \mathbb E[y_{it}^2] = O(1),
\]
because of stationarity (see Proposition (ref)). Likewise $\norm{X_t}= O_p(\sqrt N)$. Obviously $\norm{\theta}=O(1)$ and $\norm {\bar J}=1$. Thus, because of Theorem (ref) and (ref):
\begin{equation}
\norm{\frac {\delta_t}{\sqrt N}} = O_p\left(\frac 1{\sqrt T}\right), \quad \norm{\frac {\eta_t}{\sqrt N}} = O_p\left(\max\left(\frac 1{N^{3-\gamma/2}T},\frac 1{\sqrt m N^{1-\gamma/2}}\right)\right).
\end{equation}
From (ref) and (ref), it follows that,
\[
\norm{\frac{\widehat \nu_t-\nu_t}{\sqrt N}} = O_p\left(\max\left(\frac 1{\sqrt T},\frac 1{N^{3-\gamma/2}T},\frac 1{\sqrt m N^{1-\gamma/2}}\right)\right).
\]
Now, from (ref) and (ref), we also have that
\begin{equation}
\widehat \nu_t = \Lambda G_t +\epsilon_t + \delta_t+\eta_t+\zeta_t.
\end{equation}
Recall, the PC estimators $\widehat \Lambda$ and $\widehat G_t$ or $\widetilde \Lambda$ and $\widetilde G_t$ defined in barigozzi2022 or bai2003inferential, depending on the choice of the sample covariance matrix (see (ref)) and for simplicity of notation write $\widehat \Lambda$ and $\widehat G_t$ to indicate both sets of estimators. Then, since
\[
\norm{\eta_t} = o_p(\norm{\delta_t})\;\text{ and }\; \norm{\zeta_t} = o_p(\norm{\delta_t}),
\]
because $\sqrt{ T}/ (mN^{2-\gamma}) \to 0$, from barigozzi2022 or bai2020simpler, we have
\begin{align}
\norm{\frac{\widehat\Lambda-\Lambda \mathfrak J}{\sqrt N}} = O_p\left(\max\left(\frac1 {N},\frac 1{\sqrt T}\right)\right),
\end{align}
and for any given $i=1,\ldots, N$ (see barigozzi2022 or bai2003inferential)
\begin{align}
\norm{\widehat\Lambda_{i\cdot}-\Lambda_{i\cdot} \mathfrak J} = O_p\left(\max\left(\frac1 {N},\frac 1{\sqrt T}\right)\right),
\end{align}
where $\mathfrak J$ is a $q\times q$ diagonal matrix with diagonal entries $\pm 1$.
Moreover, from barigozzi2022 or bai2003inferential, for any given $t=1,\ldots, T$, we have
\begin{align}
\norm{\widehat G_{t}- \mathfrak JG_{t}} = O_p\left(\max\left(\frac1 {\sqrt N},\frac 1{\sqrt T}\right)\right).
\end{align}
Notice the term $1/\sqrt T$, coming from $\delta_t$, which is slower than the usual $1/T$ due to estimation of $\nu_t$. By steps analogous to baiandng2006, from (ref) and (ref), we get
\begin{align}
\frac 1T\sum_{t=1}^T \norm{\widehat G_{t}- \mathfrak JG_{t}}^2 = O_p\left(\max\left(\frac1 { N},\frac 1{ T}\right)\right).
\end{align}
And, from (ref) and chenetal2020cnar,
\begin{align}
\max_{i=1,\ldots, N}\norm{\widehat\Lambda_{i\cdot}-\Lambda_{i\cdot} \mathfrak J} = O_p\left(\max\left(\frac1 {\sqrt N},\sqrt{\frac {\log N}{T}}\right)\right),
\end{align}
where the $\sqrt{\log N}$ terms is due to Assumption (ref) of sub-Gaussian tails and it appears when taking the max over $i=1,\ldots, N$ in (ref) and applying the Bonferroni inequality.
Let, $\widehat{\epsilon}_{it}= \widehat{\nu}_{it}- \widehat\Lambda_{i\cdot}\widehat G_t$ and recall that $\epsilon_{it}=\nu_{it}-\Lambda_{i\cdot}G_t$. Hence, from (ref)
\begin{equation}
\abs{\widehat{\epsilon}_{it}-\epsilon_{it}}= \abs{\Lambda_{i\cdot}G_t-\widehat\Lambda_{i\cdot}\widehat G_t+\delta_{it}+\eta_{it}+\zeta_{it}}.
\end{equation}
From (ref) it follows that (see also chenetal2020cnar):
\begin{align}
\max_{i=1,\ldots, N}\frac 1T\sum_{t=1}^T \abs{\widehat{\epsilon}_{it}-\epsilon_{it}}^2\le&\, \max_{i=1,\ldots, N}\frac 2T\sum_{t=1}^T\abs{\Lambda_{i\cdot}G_t-\widehat\Lambda_{i\cdot}\widehat G_t}^2\nonumber\\
&+\max_{i=1,\ldots, N}\frac 2T\sum_{t=1}^T\abs{\delta_{it}}^2+\max_{i=1,\ldots, N}\frac 2T\sum_{t=1}^T\abs{\eta_{it}}^2+\max_{i=1,\ldots, N}\frac 2T\sum_{t=1}^T\abs{\zeta_{it}}^2\nonumber\\
&\le 8 \max_{i=1,\ldots, N}\norm{\Lambda_{i\cdot}\mathfrak J} \frac 1T\sum_{t=1}^T \norm{\widehat G_{t}- \mathfrak JG_{t}}^2+ 8 \max_{i=1,\ldots, N}\norm{\widehat\Lambda_{i\cdot}-\Lambda_{i\cdot} \mathfrak J} \frac 1T\sum_{t=1}^T \norm{\mathfrak J G_t}^2\nonumber\\
&+\max_{i=1,\ldots, N}\frac 2T\sum_{t=1}^T\abs{\delta_{it}}^2+\max_{i=1,\ldots, N}\frac 2T\sum_{t=1}^T\abs{\eta_{it}}^2+\max_{i=1,\ldots, N}\frac 2T\sum_{t=1}^T\abs{\zeta_{it}}^2\nonumber \\
=&\,O_p\left(\max\left(\frac1 { N},\frac {\log N}{T}\right)\right)+O_p\left(\max\left(\frac 1{ T},\frac {\log N}{N^{6-\gamma}T^2},\frac {\log N}{ m^2 N^{4-\gamma}},\frac{\log N}{mN^{2-\gamma}T}\right)\right)\nonumber\\
=&\, O_p\left(\max\left(\frac1 { N},\frac {\log N}{T}\right)\right),
\end{align}
because we assumed $\sqrt{ T}/ (mN^{2-\gamma}) \to 0$.
Indeed, the first term in the second last line of (ref) is due to (ref) and (ref). As for the second term in the second last line of (ref), we have
\[
\max_{i=1,\ldots, N}\frac 2T\sum_{t=1}^T\abs{\delta_{it}}^2\le \max_{i=1,\ldots, N}\frac 2T\sum_{t=1}^T \norm{X_{i\cdot t}}^2 \norm{\bar J\left(\widehat{\theta}^{\text{\tiny OLS}}-\bar J\theta\right)}^2 = O_p\left(\frac 1T\right) O_p\left(1+\sqrt{\frac{\log N}{T}}\right),
\]
by Theorem (ref) and since, by Assumption (ref)(ref) and Bonferroni inequality, for any $s>0$ and all $j=1,\ldots, r+2$,
\[
\text{P}\left(\abs{\max_{i=1,\ldots, N}\frac 2T\sum_{t=1}^T X_{ij t}^2- \mathbb E[X_{ijt}^2]}>s \right)\le 2N \exp(-cTs^2),
\]
for some finite $c$ independent of $j$, $N$, and $T$. Similarly, by Assumptions (ref) and (ref)(ref) we have
\[
\max_{i=1,\ldots, N}\frac 2T\sum_{t=1}^T\abs{\eta_{it}}^2\le \max_{i=1,\ldots, N}\frac 2T\sum_{t=1}^T \norm{\widehat X_{i\cdot t}-X_{i\cdot t}}^2 \norm{\theta}^2 = O_p\left(\max\left(\frac {\log N}{N^{6-\gamma}T^2},\frac {\log N}{ m^2 N^{4-\gamma}},\frac{\log N}{mN^{2-\gamma}T}\right)\right),
\]
since the error in estimating $X_t$ is function of
\begin{equation}
\frac 1{mN^2T}\sum_{t=1}^T\sum_{j=1}^m u_j \mathcal E_{(3)tji} \left(y_{t-1}\otimes\left\{y_{t-1}\otimes I_N\right\}' \frac{\mathcal E_{(3)t-1}'U}{mN}\right)_{i\cdot},
\end{equation}
see (ref) and (ref) in the proof of Proposition (ref).
Moreover, (see also chenetal2020cnar):
\begin{align}
\max_{i=1,\ldots, N}\frac 1T\sum_{t=1}^T \abs{\widehat{\epsilon}_{it}-\epsilon_{it}}\abs{\epsilon_{it}}\le&\, \max_{i=1,\ldots, N}\left(\frac 1T\sum_{t=1}^T \abs{\widehat{\epsilon}_{it}-\epsilon_{it}}^2\left\{\frac 1T\sum_{t=1}^T\epsilon_{it}^2-\mathbb E[\epsilon_{it}^2]\right\}\right)^{1/2}\nonumber\\
&+ \max_{i=1,\ldots, N}\left(\frac 1T\sum_{t=1}^T \abs{\widehat{\epsilon}_{it}-\epsilon_{it}}^2\mathbb E[\epsilon_{it}^2]\right)^{1/2}\nonumber\\
=&\, O_p\left(\max\left(\frac1 { \sqrt N},\sqrt{\frac {\log N}{T}}\right)\right)O_p\left(\frac 1{\sqrt T}\right)
+O_p\left(\max\left(\frac1 { \sqrt N},\sqrt{\frac {\log N}{T}}\right)\right),
\end{align}
by Assumptions (ref)(ref), (ref)(ref), and (ref)(ref) and following the same reasoning as in (ref).
Therefore, from (ref) and (ref), and Assumption (ref)(ref),
\begin{align}
\norm{\widehat S-S}&\le \max_{i=1,\ldots, N}\frac 1T\sum_{t=1}^T \abs{\widehat{\epsilon}_{it}-\epsilon_{it}}^2+2\max_{i=1,\ldots, N}\frac 1T\sum_{t=1}^T \abs{\widehat{\epsilon}_{it}-\epsilon_{it}}\abs{\epsilon_{it}}+\max_{i=1,\ldots, N}\abs{\frac 1T\sum_{t=1}^T \left\{\epsilon_{it}^2-\mathbb E[\epsilon_{it}^2]\right\}}\nonumber\\
&= O_p\left(\max\left(\frac1 { N},\frac {\log N}{T}\right)\right)+O_p\left(\max\left(\frac1 { \sqrt N},\sqrt{\frac {\log N}{T}}\right)\right)+O_p\left(\frac 1{\sqrt T}\right)\nonumber\\
&= O_p\left(\max\left(\frac1 { \sqrt N},\sqrt{\frac {\log N}{T}}\right)\right).
\end{align}
And from (ref) and Assumption (ref)(ref)
\begin{align}
\norm{\widehat S^{-1}-S^{-1}}&\le \norm{S^{-1}} \norm{\widehat S-S}\norm {\widehat S^{-1}}\nonumber\\
&= \frac 1{\min_{i=1,\ldots, N} \mathbb E[\epsilon_{it}^2]} O_p\left(\max\left(\frac1 { \sqrt N},\sqrt{\frac {\log N}{T}}\right)\right)
\frac 1{\min_{i=1,\ldots, N} \frac 1T\sum_{t=1}^T\epsilon_{it}^2}\nonumber\\
&\le \frac 1{\underline M_\epsilon}O_p\left(\max\left(\frac1 { \sqrt N},\sqrt{\frac {\log N}{T}}\right)\right) \frac 1{\underline M_\epsilon+O_p\left(\max\left(\frac1 { \sqrt N},\sqrt{\frac {\log N}{T}}\right)\right)}\nonumber\\
& = O_p\left(\max\left(\frac1 { \sqrt N},\sqrt{\frac {\log N}{T}}\right)\right).
\end{align}
Thus, from (ref) and (ref)
\begin{align}
\norm{\frac{\widehat \Lambda'\widehat S^{-1}\widehat \Lambda-\Lambda'S^{-1}\Lambda}{N}} = O_p\left(\max\left(\frac1 { \sqrt N},\sqrt{\frac {\log N}{T}}\right)\right).
\end{align}
Furthermore, recalling the definition of $\widehat V^{-1}$ in (ref)
\begin{equation}
\widehat V^{-1}=\widehat S^{-1}-\widehat S^{-1} \frac{\widehat\Lambda}{\sqrt N}\left(\frac {I_N}N+ \frac{\widehat\Lambda'\widehat S^{-1}\widehat \Lambda}N\right)^{-1} \frac{\widehat\Lambda'}{\sqrt N}\widehat S^{-1},\nonumber
\end{equation}
and since we can always write
\begin{equation}
V^{-1}= S^{-1}- S^{-1} \frac{\Lambda}{\sqrt N}\left(\frac {I_N}N+ \frac{\Lambda' S^{-1} \Lambda}N\right)^{-1} \frac{\Lambda'}{\sqrt N} S^{-1},\nonumber
\end{equation}
we have
\begin{align}
\norm{\widehat V^{-1}-V^{-1}} = O_p\left(\max\left(\frac1 { \sqrt N},\sqrt{\frac {\log N}{T}}\right)\right),
\end{align}
because of (ref), (ref), and (ref). This completes the proof.
lemmaUnder Assumptions (ref)-(ref), as $N,T\to\infty$, if $\sqrt N/T\to 0$,
\begin{align}
\frac 1{\sqrt{NT}}\sum_{t=1}^T \widehat W_t^{*'}\epsilon_t=&\,\frac 1{\sqrt{NT}}\sum_{t=1}^T W_t'\epsilon_t+\sqrt{\frac NT}\xi^*_{NT}\nonumber\\
&+\sqrt NO_p(\Vert\widehat{\theta}^*-\bar J\theta \Vert^2)+O_p(\Vert\widehat{\theta}^*-\bar J\theta \Vert)+\sqrt NO_p\left(\max\left(\frac 1N, \frac 1T\right)\right),\nonumber
\end{align}
with
\[
\xi^*_{NT}:=-\frac 1{NT}\sum_{t=1}^T\bar J\left(X_t-\frac 1T\sum_{s=1}^T (G_t'G_s)X_s\right)'\Lambda\sum_{u=1}^T G_u\left(\frac 1N\sum_{i=1}^N\epsilon_{iu}\epsilon_{it} \right).
\]
proofFirst note that $M_{{\Lambda}}-M_{\widehat{\Lambda}^*}=P_{{\Lambda}}-P_{\widehat{\Lambda}^*}$, where
$P_{{\Lambda}}={\Lambda}({\Lambda}'{\Lambda})^{-1}{\Lambda}'$
and
$P_{\widehat{\Lambda}^*}=\widehat{\Lambda}^{*}(\widehat{\Lambda}^{*'}\widehat{\Lambda}^*)^{-1}\widehat{\Lambda}^{*'}$, then
\begin{align}
\frac 1{\sqrt {NT}}\sum_{t=1}^TX_t'(M_{{\Lambda}}-M_{\widehat{\Lambda}^*})\epsilon_t=&\,\frac 1{\sqrt {NT}}\sum_{t=1}^TX_t'\widehat{\Lambda}^{*}(\widehat{\Lambda}^{*'}\widehat{\Lambda}^*)^{-1}\widehat{\Lambda}^{*'}\epsilon_t-\frac 1{\sqrt {NT}}\sum_{t=1}^TX_t'{\Lambda}({\Lambda}'{\Lambda})^{-1}{\Lambda}'\epsilon_t\nonumber\\
=&\, \frac 1{\sqrt {NT}}\sum_{t=1}^TX_t'\left(\widehat{\Lambda}^{*}-\Lambda J\right)\left(J\Lambda'\Lambda J\right)^{-1}J\Lambda'\epsilon_t\nonumber\\
&+ \frac 1{\sqrt {NT}}\sum_{t=1}^TX_t'\left(\widehat{\Lambda}^{*}-\Lambda J\right)\left(J\Lambda'\Lambda J\right)^{-1}
\left(\widehat{\Lambda}^{*}-\Lambda J\right)'\epsilon_t\nonumber\\
&+\frac 1{\sqrt {NT}}\sum_{t=1}^TX_t'\Lambda J\left(J\Lambda'\Lambda J\right)^{-1}
\left(\widehat{\Lambda}^{*}-\Lambda J\right)'\epsilon_t\nonumber\\
&+ \frac 1{\sqrt {NT}}\sum_{t=1}^TX_t'\left(\widehat{\Lambda}^{*}-\Lambda J\right)\left(\widehat{\Lambda}^{*'}\widehat{\Lambda}^*-J\Lambda'\Lambda J\right)^{-1}
J\Lambda'\epsilon_t\nonumber\\
&+\frac 1{\sqrt {NT}}\sum_{t=1}^TX_t'\left(\widehat{\Lambda}^{*}-\Lambda J\right)\left(\widehat{\Lambda}^{*'}\widehat{\Lambda}^*-J\Lambda'\Lambda J\right)^{-1}
\left(\widehat{\Lambda}^{*}-\Lambda J\right)'\epsilon_t\nonumber\\
&+ \frac 1{\sqrt {NT}}\sum_{t=1}^TX_t'\Lambda J\left(\widehat{\Lambda}^{*'}\widehat{\Lambda}^*-J\Lambda'\Lambda J\right)^{-1}
\left(\widehat{\Lambda}^{*}-\Lambda J\right)'\epsilon_t\nonumber\\
&+ \frac 1{\sqrt {NT}}\sum_{t=1}^TX_t'\Lambda J\left(\widehat{\Lambda}^{*'}\widehat{\Lambda}^*-J\Lambda'\Lambda J\right)^{-1}
J\Lambda'\epsilon_t\nonumber\\
=:&\, a+b+c+d+e+f+g.
\end{align}
Then, since $\left(\widehat{\lambda}_i^{*'}-\lambda_i' J\right)\left({J\Lambda'\Lambda J}\right)^{-1}J\lambda_j$ is a scalar,
\begin{align}
a&=\frac 1{\sqrt {NT}}\sum_{t=1}^T\sum_{i=1}^N X_{it}\left(\widehat{\lambda}_i^{*'}-\lambda_i' J\right)\left(\frac{J\Lambda'\Lambda J}N\right)^{-1}J\frac 1N\sum_{j=1}^N \lambda_j\epsilon_{jt}\nonumber\\
&= \frac 1N \sum_{i=1}^N \left(\widehat{\lambda}_i^{*'}-\lambda_i' J\right)\left(\frac{J\Lambda'\Lambda J}N\right)^{-1}J\left(\frac 1{\sqrt {NT}}\sum_{t=1}^T\sum_{j=1}^N \lambda_jX_{it}\epsilon_{jt}\right),\nonumber
\end{align}
and, therefore,
\begin{align}
\Vert a\Vert \le&\, \left[\frac 1N\sum_{i=1}^N \left\Vert \widehat{\lambda}_i^{*'}-\lambda_i' J \right\Vert^2 \right]^{1/2}\, \left\Vert \left(\frac{J\Lambda'\Lambda J}N\right)^{-1}\right\Vert\,\left\Vert J\right\Vert\left[\frac 1N\sum_{i=1}^N \left\Vert
\left(\frac 1{\sqrt {NT}}\sum_{t=1}^T\sum_{j=1}^N \lambda_jX_{it}\epsilon_{jt}
\right)
\right\Vert^2
\right]^{1/2}\nonumber\\
&=\left[O_p\left(\Vert\widehat{\theta}^*-\bar J\theta\Vert\right)+ O_p\left(\max\left(\frac 1{\sqrt N},\frac 1{\sqrt T}\right)\right)\right] O_p(1),
\end{align}
because of (ref), Assumption (ref)(ref), and barigozzi2022.
Similarly,
\begin{align}
b&= \frac 1{\sqrt {NT}}\sum_{t=1}^T\sum_{i=1}^N X_{it}\left(\widehat{\lambda}_i^{*'}-\lambda_i' J\right)\left(\frac{J\Lambda'\Lambda J}N\right)^{-1}\frac 1N
\sum_{j=1}^N\left(\widehat{\lambda}_j^{*}-J\lambda_j\right)\epsilon_{jt}\nonumber\\
&=\sqrt N\frac 1{N^2}\sum_{i=1}^N\sum_{j=1}^N
\left(\widehat{\lambda}_i^{*'}-\lambda_i' J\right)\left(\frac{J\Lambda'\Lambda J}N\right)^{-1}
\left(\widehat{\lambda}_j^{*}-J\lambda_j\right)\left(\frac 1{\sqrt {T}} \sum_{t=1}^T X_{it}\epsilon_{jt}\right).\nonumber
\end{align}
Therefore,
\begin{align}
\Vert b\Vert \le&\, \sqrt N \left(\frac 1N\sum_{i=1}^N \left\Vert\widehat{\lambda}_i^{*'}-\lambda_i' J\right\Vert^2\right) \left\Vert \left(\frac{J\Lambda'\Lambda J}N\right)^{-1}\right\Vert
\left(\frac 1{N^2}\sum_{i=1}^N\sum_{j=1}^N\left\Vert \frac 1{\sqrt {T}} \sum_{t=1}^T X_{it}\epsilon_{jt}\right\Vert^2\right)^{1/2}\nonumber\\
&=\sqrt N\left[O_p\left(\Vert\widehat{\theta}^*-\bar J\theta\Vert^2\right)+ O_p\left(\max\left(\frac 1{ N},\frac 1{ T}\right)\right)\right]O_p(1),
\end{align}
because of (ref), Assumption (ref)(ref), and barigozzi2022.
Then, consider
\begin{align}
c=&\,\frac 1{\sqrt {NT}}\sum_{t=1}^T X_{t}'\Lambda J\left(\frac{J\Lambda'\Lambda J}N\right)^{-1}
\frac 1N \left(\widehat{\Lambda}^{*}-\Lambda J\right)'\epsilon_{t}\nonumber\\
=&\, \frac{\sqrt {NT}} T\left\{\frac 1{T}\sum_{t=1}^T\sum_{s=1}^T \frac{X_t'\Lambda J}{N}\left(\frac{J\Lambda'\Lambda J}{N}\right)^{-1}G_s \left(\frac 1{N} \sum_{i=1}^N \epsilon_{it}\epsilon_{is}\right)\right\}\nonumber\\
& + O_p\left(\Vert{\widehat{\theta}^*-\bar J\theta}\Vert\right)+\sqrt{\frac{N}{T}}O_p\left(\Vert{\widehat{\theta}^*-\bar J\theta}\Vert\right)+{\sqrt N}O_p\left(\max\left(\frac 1N,\frac 1T\right)\right)\nonumber\\
=&\, \sqrt {\frac NT}\psi_{NT} + O_p\left(\Vert{\widehat{\theta}^*-\bar J\theta}\Vert\right)+\sqrt{\frac{N}{T}}O_p\left(\Vert{\widehat{\theta}^*-\bar J\theta}\Vert\right)+{\sqrt N}O_p\left(\max\left(\frac 1N,\frac 1T\right)\right).\nonumber
\end{align}
And it easily seen that $\Vert \psi_{NT}\Vert=O_p(1)$ so $\sqrt{\frac NT}\psi_{NT} $ dominates $\sqrt{\frac{N}{T}}O_p\left(\Vert{\widehat{\theta}^*-\bar J\theta}\Vert\right)$, since $\Vert{\widehat{\theta}^*-\bar J\theta}\Vert=o_p(1)$ as shown in (ref) in the proof of Lemma (ref)(ref). Finally, $d$, $e$, and $f$ are dominated by $a$, $b$, and $c$, respectively, and $g$ behaves as $a$.
Therefore, from (ref) and Lemma (ref)(ref) we have
\begin{align}
\frac 1{\sqrt {NT}}&\,\sum_{t=1}^TX_t'(M_{{\Lambda}}-M_{\widehat{\Lambda}^*})\epsilon_t=\frac{\sqrt {NT}} T\psi_{NT}\\
&+O_p\left(\Vert\widehat{\theta}^*-\bar J\theta\Vert\right)+ O_p\left(\max\left(\frac 1{\sqrt N},\frac 1{\sqrt T}\right)\right)\nonumber\\
&+\sqrt N\left[O_p\left(\Vert\widehat{\theta}^*-\bar J\theta\Vert^2\right)+ O_p\left(\max\left(\frac 1{ N},\frac 1{ T}\right)\right)\right]\nonumber\\
&+ O_p\left(\Vert{\widehat{\theta}^*-\bar J\theta}\Vert\right)
+{\sqrt N}O_p\left(\max\left(\frac 1N,\frac 1T\right)\right)\nonumber\\
=&\, \frac{\sqrt {NT}} T\psi_{NT}+O_p\left(\Vert\widehat{\theta}^*-\bar J\theta\Vert\right)+\sqrt N O_p\left(\Vert\widehat{\theta}^*-\bar J\theta\Vert^2\right)+
\sqrt NO_p\left(\max\left(\frac 1{ N},\frac 1{ T}\right)\right),\nonumber
\end{align}
where
\[
\psi_{NT} := \frac 1{T}\sum_{t=1}^T\sum_{s=1}^T \frac{X_t'\Lambda J}{N}\left(\frac{J\Lambda'\Lambda J}{N}\right)^{-1}G_s \left(\frac 1{N} \sum_{i=1}^N \epsilon_{it}\epsilon_{is}\right).
\]
Let now $V_t:=\frac 1T\sum_{t=1}^T(G_s'G_t) X_t$. Then, replacing $X_t$ with $V_t$ in (ref) we have
\begin{align}
\frac 1{\sqrt {NT}}\sum_{t=1}^T&V_t'(M_{{\Lambda}}-M_{\widehat{\Lambda}^*})\epsilon_t\\
=&\, \frac{\sqrt {NT}} T\psi_{NT}^*+O_p\left(\Vert\widehat{\theta}^*-\bar J\theta\Vert\right)+\sqrt N O_p\left(\Vert\widehat{\theta}^*-\bar J\theta\Vert^2\right)+
\sqrt NO_p\left(\max\left(\frac 1{ N},\frac 1{ T}\right)\right),\nonumber
\end{align}
where
\[
\psi_{NT}^* :=- \frac 1{T}\sum_{t=1}^T\sum_{s=1}^T \frac{V_t'\Lambda J}{N}\left(\frac{J\Lambda'\Lambda J}{N}\right)^{-1}G_s \left(\frac 1{N} \sum_{i=1}^N \epsilon_{it}\epsilon_{is}\right).
\]
By combining (ref) and (ref) we complete the proof.
lemmaLet $D(\widehat{\Lambda}^*):=\frac 1{NT} \sum_{t=1}^T \bar J\widehat{W}_{t}^{*'} \widehat{W}_{t}^{*}\bar J$ and
$D({\Lambda}):=\frac 1{NT} \sum_{t=1}^T \bar J{W}_{t}^{'} {W}_{t}\bar J$.
Under Assumptions (ref)-(ref), as $N,T\to\infty$,
\begin{enumerate}[label=(\roman*)]
• $\norm{D(\widehat{\Lambda}^*)^{-1}-D({\Lambda})^{-1}}=o_p(1)$.
• $\sqrt{\frac TN}\norm{D(\widehat{\Lambda}^*)^{-1}-D({\Lambda})^{-1}}=o_p(1)$ if $\sqrt T/N\to 0$.
• $\sqrt{\frac NT}\norm{D(\widehat{\Lambda}^*)^{-1}-D({\Lambda})^{-1}}=o_p(1)$ if $\sqrt N/T\to 0$.
\end{enumerate}
proofFor part (ref), notice that
\begin{align}
D(\widehat{\Lambda}^*)-D({\Lambda})=&\, \frac 1{NT}\sum_{t=1}^TX_t'(M_{\widehat{\Lambda}^*}-M_{\Lambda})X_t\nonumber\\
&-\frac 1N\left[\frac 1{T^2}\sum_{t=1}^T\sum_{s=1}^TX_t'(M_{\widehat{\Lambda}^*}-M_{\Lambda})X_s (G_t'G_s)\right].\nonumber
\end{align}
Then the proof follows from the fact that by (ref) and barigozzi2022
\begin{align}
\norm{M_{\widehat{\Lambda}^*}-M_{\Lambda}}^2=2 tr\left(I_q-\frac{\widehat{\Lambda}^{*'} \Lambda(\Lambda'\Lambda)^{-1}\Lambda'\widehat{\Lambda}^{*'}}{N}\right) = O_p\left(\Vert{\widehat{\theta}^*-\bar J\theta}\Vert\right) + O_p\left(\max\left(\frac 1{N},\frac 1T\right)\right).
\end{align}
which follows from
\begin{align}
\frac 1N\Lambda'(\widehat{\Lambda}^*-\Lambda J)&=\frac{\Lambda'\widehat{\Lambda}^*}{N}-\frac{\Lambda'{\Lambda}J}{N}= O_p\left(\Vert{\widehat{\theta}^*-\bar J\theta}\Vert\right) + O_p\left(\max\left(\frac 1{N},\frac 1T\right)\right),\nonumber\\
\frac 1N\widehat{\Lambda}^{*'}(\widehat{\Lambda}^*-\Lambda J)&=\frac{\widehat{\Lambda}^{*'}\widehat{\Lambda}^*}{N}-\frac{\widehat{\Lambda}^{*'}\widehat{\Lambda}^{*}}{N}= O_p\left(\Vert{\widehat{\theta}^*-\bar J\theta}\Vert\right) + O_p\left(\max\left(\frac 1{N},\frac 1T\right)\right).\nonumber
\end{align}
Therefore,
\begin{align}
\norm{D(\widehat{\Lambda}^*)-D({\Lambda})}=O_p\left(\Vert{\widehat{\theta}^*-\bar J\theta}\Vert\right) + O_p\left(\max\left(\frac 1{N},\frac 1T\right)\right).
\end{align}
This proves part (ref).
For part (ref), from (ref) and the second step (ref) (which does not use this lemma), we have
\begin{align}
\sqrt{NT}\left(\widehat\theta^*-\bar J\theta\right)&=\left(D(\widehat{\Lambda}^*)\right)^{-1} \left\{\frac 1{\sqrt{NT}}\sum_{t=1}^T \bar J{W}_t^{'}\epsilon_t\right.\nonumber\\
&-\sqrt{\frac TN}\left[\frac {1}{NT}\sum_{t=1}^T \bar JX_t'M_{\widehat{\Lambda}^*}\left(\frac 1T\sum_{s=1}^T
\mathbb E[\epsilon_s\epsilon_s']\right)\widehat{\Lambda}^*\left(\frac{\Lambda'\widehat{\Lambda}^*}{N}\right)^{-1}G_t\right]\nonumber\\
&\left.-\sqrt{\frac NT} \left[\frac 1{NT}\sum_{t=1}^T\bar J\left(X_t-\frac 1T\sum_{s=1}^T (G_t'G_s)X_s\right)'\Lambda\left(\frac{\Lambda'\Lambda}{N}\right)^{-1}\sum_{u=1}^T G_u\left(\frac 1N\sum_{i=1}^N\epsilon_{iu}\epsilon_{it} \right)
\right]
\right\}+o_p(1)\nonumber\\
&= \left(D(\widehat{\Lambda}^*)\right)^{-1}\frac 1{\sqrt{NT}}\sum_{t=1}^T \bar J{W}_t^{'}\epsilon_t+\sqrt{\frac TN} \zeta_{NT} +\sqrt {\frac NT} \xi_{NT}+o_p(1),
\end{align}
where
\begin{align}
\zeta_{NT}:=-\left(D(\widehat{\Lambda}^*)\right)^{-1}\frac {1}{NT}\sum_{t=1}^T \bar JX_t'M_{\widehat{\Lambda}^*}\left(\frac 1T\sum_{s=1}^T
\mathbb E[\epsilon_s\epsilon_s']\right)\widehat{\Lambda}^*\left(\frac{\Lambda'\widehat{\Lambda}^*}{N}\right)^{-1}G_t,
\end{align}
and
\begin{align}
\xi_{NT}&:=-\left(D(\widehat{\Lambda}^*)\right)^{-1}
\frac 1{NT}\sum_{t=1}^T\bar J\left(X_t-\frac 1T\sum_{s=1}^T (G_t'G_s)X_s\right)'\Lambda\left(\frac{\Lambda'\Lambda}{N}\right)^{-1}\sum_{u=1}^T G_u\left(\frac 1N\sum_{i=1}^N\epsilon_{iu}\epsilon_{it}\right).
\end{align}
It is easily seen that $\Vert \xi_{NT}\Vert=O_p(1)$ and $\Vert \zeta_{NT}\Vert=O_p(1)$. Moreover, by Assumptions (ref)(ref) and (ref)(ref), we also have $\left(D(\widehat{\Lambda}^*)\right)^{-1}\frac 1{\sqrt{NT}}\sum_{t=1}^T \bar J{W}_t^{'}\epsilon_t=O_p(1)$.
Therefore, from (ref)
\[
\sqrt{NT}\left(\widehat\theta^*-\bar J\theta\right)=O_p\left(\sqrt{\frac TN}\right)+O_p\left(\sqrt{\frac NT}\right).
\]
Hence,
\begin{align}
\left(\widehat\theta^*-\bar J\theta\right)=O_p\left(\max\left(\frac 1N,\frac 1T\right)\right).
\end{align}
By (ref) and (ref) we have
\[
\norm{D(\widehat{\Lambda}^*)-D({\Lambda})} = O_p\left(\max\left(\frac 1{\sqrt N},\frac 1{\sqrt T}\right)\right).
\]
The proof of parts (ref) and (ref) follows immediately. This completes the proof.
lemmaUnder Assumptions (ref)-(ref), as $N,T\to\infty$,
\begin{enumerate}[label=(\roman*)]
• if $\sqrt T/N\to 0$,
\[
\sqrt {\frac TN}\norm{\zeta_{NT}-\zeta_{NT}^0}= o_p(1),
\]
where
\[
\zeta_{NT}^0:=- \left(D({\Lambda})\right)^{-1}\frac {1}{NT}\sum_{t=1}^T \bar JX_t'M_{{\Lambda}}\left(\frac 1T\sum_{s=1}^T
\mathbb E[\epsilon_s\epsilon_s']\right){\Lambda}\left(\frac{\Lambda'{\Lambda}}{N}\right)^{-1}G_t,
\]
and
$\zeta_{NT}$ is defined in (ref) in the proof of Lemma (ref).
• if $\sqrt N/T\to 0$,
\[
\sqrt {\frac NT}\norm{\xi_{NT}-\xi_{NT}^0}= o_p(1),
\]
where
\[
\xi_{NT}^0 := - \left(D({\Lambda})\right)^{-1}\frac 1{NT}\sum_{t=1}^T\bar J\left(X_t-\frac 1T\sum_{s=1}^T (G_t'G_s)X_s\right)'\Lambda\left(\frac{\Lambda'\Lambda}{N}\right)^{-1}\sum_{u=1}^T G_u\left(\frac 1N\sum_{i=1}^N\mathbb E[\epsilon_{iu}\epsilon_{it}]\right),
\]
and $\xi_{NT}$ is defined in (ref) in the proof of Lemma (ref).
\end{enumerate}
proofFor part (ref), we have
\begin{align}
\left\Vert \frac {1}{NT}\sum_{t=1}^T \bar JX_t'M_{\widehat{\Lambda}^*}\left(\frac 1T\sum_{s=1}^T
\mathbb E[\epsilon_s\epsilon_s']\right)\widehat{\Lambda}^*\left(\frac{\Lambda'\widehat{\Lambda}^*}{N}\right)^{-1}G_t\right.&-\left.
\frac {1}{NT}\sum_{t=1}^T \bar JX_t'M_{{\Lambda}}\left(\frac 1T\sum_{s=1}^T
\mathbb E[\epsilon_s\epsilon_s']\right){\Lambda}\left(\frac{\Lambda'{\Lambda}}{N}\right)^{-1}G_t\right\Vert\nonumber\\
&= O_p\left(\max\left(\frac 1{\sqrt N},\frac 1{\sqrt T}\right)\right),\nonumber
\end{align}
by (ref) and (ref) in the proof of Lemma (ref). Thus, since $\sqrt {\frac TN}O_p\left(\max\left(\frac 1{\sqrt N},\frac 1{\sqrt T}\right)\right)= o_p(1)$ if $\sqrt T/N\to 0$, by using Lemma (ref)(ref), we prove part (ref).
For part (ref), by Lemma (ref)(ref), if $\sqrt N/T\to 0$, we have
\begin{align}
\sqrt {\frac NT}\left({\xi_{NT}-\xi_{NT}^0}\right) = \sqrt {\frac NT}\left(\left(D(\widehat{\Lambda}^*)\right)^{-1}\xi_{NT}^*-\xi_{NT}^0\right)= \sqrt {\frac NT}\left(\left(D({\Lambda})\right)^{-1}\xi_{NT}^*-\xi_{NT}^0\right)+o_p(1),
\end{align}
where $\xi_{NT}^*$ is defined in Lemma (ref), and since $\Vert \xi_{NT}^*\Vert=O_p(1)$. Moreover, let
\[
A_{tu}:= G_u'\otimes \left[\left(X_t-\frac 1T\sum_{s=1}^T (G_t'G_s)X_s\right)'\frac{\Lambda}N\right],
\]
which is such that $\norm{A_{tu}}=O_p(1)$. Then, since $\norm{\left(\frac{\Lambda'\Lambda}{N}\right)^{-1}}=O_p(1)$ by Assumption (ref)(ref), we have
\begin{align}
\left(D({\Lambda})\right)^{-1}\xi_{NT}^*-\xi_{NT}^0&=-\left(D({\Lambda})\right)^{-1}\frac 1T \sum_{t=1}^T\sum_{u=1}^T A_{tu}\left\{\frac 1N\sum_{i=1}^N \epsilon_{iu}\epsilon_{it}-\mathbb E[\epsilon_{iu}\epsilon_{it}]\right\}vec\left(\left(\frac{\Lambda'\Lambda}{N}\right)^{-1}\right)\nonumber\\
&=O_p\left(\frac 1{\sqrt N}\right).
\end{align}
by Assumption (ref)(ref). Thus, by substituing (ref) into (ref), we have
$\sqrt {\frac NT}\left({\xi_{NT}-\xi_{NT}^0}\right)= O_p\left(\frac 1{\sqrt T}\right)$. This completes the proof.
lemmaUnder Assumptions (ref)-(ref), as $N,T\to\infty$,
\begin{enumerate}[label=(\roman*)]
• $\norm{\frac 1N (\widehat{\Lambda}^*-\Lambda J)\epsilon_t}= \frac 1{\sqrt N}O_p(\Vert{\widehat{\theta}^*-\bar J\theta}\Vert)+O_p\left(\max\left(\frac 1N,\frac 1T\right)\right)$.
• $\norm{\frac 1{N\sqrt T}\sum_{t=1}^T (\widehat{\Lambda}^*-\Lambda J)\epsilon_t}= \frac 1{\sqrt N}O_p(\Vert{\widehat{\theta}^*-\bar J\theta}\Vert)
+\frac 1{\sqrt T}O_p(\Vert{\widehat{\theta}^*-\bar J\theta}\Vert)+O_p\left(\frac 1{\sqrt T}\right)
+O_p\left(\max\left(\frac 1N,\frac 1T\right)\right)$.
• $\norm{\frac 1{N T}\sum_{t=1}^TG_t' (\widehat{\Lambda}^*-\Lambda J)\epsilon_t}= \frac 1{\sqrt {NT}}O_p(\Vert{\widehat{\theta}^*-\bar J\theta}\Vert)
+\frac 1{T}O_p(\Vert{\widehat{\theta}^*-\bar J\theta}\Vert)
+O_p\left(\frac 1{T}\right)
+\frac 1{\sqrt T}O_p\left(\max\left(\frac 1N,\frac 1T\right)\right)$.
• \begin{align}
&\norm{
\frac 1{NT}\sum_{t=1}^T \frac{X_t'\Lambda J}{N}\left(\frac{J\Lambda'\Lambda J}{N}\right)^{-1}(\widehat{\Lambda}^*-\Lambda J)\epsilon_t-
\frac 1{T^2}\sum_{t=1}^T\sum_{s=1}^T \frac{X_t'\Lambda J}{N}\left(\frac{J\Lambda'\Lambda J}{N}\right)^{-1}G_s \left(\frac 1{N} \sum_{i=1}^N \epsilon_{it}\epsilon_{is}\right)
}\nonumber\\
&=\frac 1{\sqrt {NT}}O_p(\Vert{\widehat{\theta}^*-\bar J\theta}\Vert)
+\frac 1{T}O_p(\Vert{\widehat{\theta}^*-\bar J\theta}\Vert)
+\frac 1{\sqrt T}O_p\left(\max\left(\frac 1N,\frac 1T\right)\right).\nonumber
\end{align}
\end{enumerate}
proofPart (ref) follows from (ref) and barigozzi2022.
For part (ref), using (ref) and (ref), we have
\begin{align}
\frac 1{N \sqrt T}\sum_{t=1}^TG_t' (\widehat{\Lambda}^*-\Lambda J)\epsilon_t&= \frac 1{N\sqrt T}\sum_{t=1}^T \left(\frac{\widehat{\Lambda}^{*'}\Lambda}{N}\right)^{-1}
\left\{\sum_{j=1}^8 I_j\right\}' \epsilon_t + o_p\left(\frac 1{\sqrt {NT}}\right)\nonumber\\
&=:A_1+A_2+A_3+A_4+A_5+A_6+A_7+A_8+ o_p\left(\frac 1{\sqrt {NT}}\right).
\end{align}
For $A_1$ in (ref), we have
\begin{align}
\norm{A_1}&\le \frac 1{\sqrt N} \norm{ \left(\frac{\widehat{\Lambda}^{*'}\Lambda}{N}\right)^{-1}}
\left(
\frac 1T\sum_{t=1}^T \norm{\sum_{s=1}^T \frac{\widehat{\Lambda}^{*'}X_s}{N}}\bar J(\bar J\theta-\widehat{\theta}^*)(\bar J\theta-\widehat{\theta}^*)'\bar J\frac{X_s'\epsilon_t}{\sqrt{NT}}
\right)\nonumber\\
&\le \frac 1{\sqrt N} \norm{ \left(\frac{\widehat{\Lambda}^{*'}\Lambda}{N}\right)^{-1}} \frac 1T\sum_{t=1}^T \norm{\frac{\widehat{\Lambda}^*}{\sqrt N}}\, \norm{\frac 1{\sqrt{NT}}\sum_{s=1}^T \bar JX_s'\epsilon_t}\,\frac{\norm{X_s}}{\sqrt N} \norm{\bar J\theta-\widehat{\theta}^*}^2\nonumber\\
&= \frac 1{\sqrt N} O_p\left(\norm{\bar J\theta-\widehat{\theta}^*}\right)O_p(1),
\end{align}
by (ref), (ref), and Assumption (ref)(ref). Similarly, for $A_2$ in (ref), we have
\begin{align}
A_2&=\frac 1{N T}\frac 1{\sqrt T} \sum_{t=1}^T \sum_{s=1}^T G_s (\bar J\theta-\widehat{\theta}^*)'\bar JX_s'\epsilon_t.\nonumber
\end{align}
Thus,
\begin{align}
\norm{A_2}=\frac 1{\sqrt N} O_p\left(\norm{\bar J\theta-\widehat{\theta}^*}\right).
\end{align}
And by the same arguments, for $A_3$ in (ref), we have
\begin{align}
\norm{A_3}&\le\frac 1{\sqrt N} \norm{ \left(\frac{\widehat{\Lambda}^{*'}\Lambda}{N}\right)^{-1}}\left(\frac 1T\sum_{t=1}^T \norm{
\sum_{s=1}^T \frac{\widehat{\Lambda}^{*'}\epsilon_s}{N} (\bar J\theta-\widehat{\theta}^*)'\bar J\frac{X_s'\epsilon_t}{\sqrt{NT}}
}\right)\nonumber\\
&\le \frac 1{\sqrt N} \norm{ \left(\frac{\widehat{\Lambda}^{*'}\Lambda}{N}\right)^{-1}}\frac 1T \sum_{t=1}^T \norm{\frac{\widehat{\Lambda}^*}{\sqrt N}}\,
\norm{\frac 1{\sqrt{NT}}\sum_{s=1}^T \bar J X_s'\epsilon_t}\, \frac{\norm{\epsilon_s}}{\sqrt N} \norm{\bar J\theta-\widehat{\theta}^*}\nonumber\\
&= \frac 1{\sqrt N} O_p\left(\norm{\bar J\theta-\widehat{\theta}^*}\right)O_p(1),
\end{align}
by barigozzi2022.
For $A_4$ in (ref), we have
\begin{align}
A_4&= \frac 1{N\sqrt T} \frac 1{NT} \sum_{t=1}^T\sum_{s=1}^T\left(\frac{\widehat{\Lambda}^{*'}\Lambda}{N}\right)^{-1} \widehat{\Lambda}^{*'} X_s \bar J(\bar J\theta-\widehat{\theta}^*)G_s'\Lambda'\epsilon_t\nonumber\\
&= \frac 1{\sqrt N} \left(\frac{\widehat{\Lambda}^{*'}\Lambda}{N}\right)^{-1}\left(\frac1 T \sum_{s=1}^T \frac{\widehat{\Lambda}^{*'} X_s\bar J}{N}\right)
(\bar J\theta-\widehat{\theta}^*)\left(\frac 1{\sqrt {NT}}\sum_{t=1}^T G_s'\Lambda'\epsilon_t\right).
\nonumber
\end{align}
Thus, by Assumption (ref)(ref), (ref) and barigozzi2022,
\begin{align}
\norm{A_4}&=\frac 1{\sqrt N} O_p\left(\norm{\bar J\theta-\widehat{\theta}^*}\right).
\end{align}
For $A_5$ in (ref), we have
\begin{align}
A_5&=\frac 1{N\sqrt T} \sum_{t=1}^T\left(\frac{\widehat{\Lambda}^{*'}\Lambda}{N}\right)^{-1} \frac 1{NT}\sum_{s=1}^T \widehat{\Lambda}^{*'} X_s\bar J (\bar J\theta-\widehat{\theta}^*)\epsilon_s'\epsilon_t\nonumber\\
&= \frac 1{\sqrt T} \left(\frac{\widehat{\Lambda}^{*'}\Lambda}{N}\right)^{-1}\left\{\frac 1{\sqrt T} \sum_{s=1}^T\frac{\widehat{\Lambda}^{*'} X_s\bar J}{N} (\bar J\theta-\widehat{\theta}^*)\frac 1N\sum_{k=1}^N\epsilon_{ks}\right\} \frac 1{\sqrt T}\sum_{t=1}^T \epsilon_{kt}.\nonumber
\end{align}
Thus, by Assumptions (ref)(ref) and (ref)(ref), and (ref)
\begin{align}
\norm{A_5}&=\frac 1{\sqrt T} O_p\left(\norm{\bar J\theta-\widehat{\theta}^*}\right).
\end{align}
For $A_6$ in (ref), we have
\begin{align}
A_6=&\, \frac 1{N\sqrt T}\sum_{t=1}^T \left(\frac{\widehat{\Lambda}^{*'}\Lambda}{N}\right)^{-1}\frac 1{NT} \sum_{s=1}^T \widehat{\Lambda}^{*'}\epsilon_s G_s'\Lambda'\epsilon_t\nonumber\\
=&\,\frac 1{N^2T}\frac 1{\sqrt T} \sum_{s=1}^T \left(\frac{\widehat{\Lambda}^{*'}\Lambda}{N}\right)^{-1} J \Lambda'\epsilon_sG_s'
\sum_{t=1}^T \Lambda'\epsilon_t+ \frac 1{N^2T}\frac 1{\sqrt T} \sum_{s=1}^T \left(\frac{\widehat{\Lambda}^{*'}\Lambda}{N}\right)^{-1} (\widehat{\Lambda}^{*'}-J \Lambda')\epsilon_sG_s'
\sum_{t=1}^T \Lambda'\epsilon_t\nonumber\\
=:&\, A_{6,1}+A_{6,2}.
\end{align}
Then, by (ref) and barigozzi2022,
\begin{align}
A_{6,1}&=\frac 1{N\sqrt T} \left(\frac{\widehat{\Lambda}^{*'}\Lambda}{N}\right)^{-1}J\left(\frac1 {\sqrt{NT}}\sum_{s=1}^T \Lambda' \epsilon_s G_s'\right) \left(\frac 1{\sqrt{NT}}\sum_{t=1}^T\Lambda'\epsilon_t \right)=O_p\left(\frac 1{N\sqrt T}\right).
\end{align}
Similarly, by (ref), (ref), and barigozzi2022,
\begin{align}
A_{6,2}&=\frac 1{N\sqrt T} \left(\frac{\widehat{\Lambda}^{*'}\Lambda}{N}\right)^{-1}\left(\frac1 {\sqrt{NT}}\sum_{s=1}^T (\widehat{\Lambda}^{*'}-J\Lambda') \epsilon_s G_s'\right) \left(\frac 1{\sqrt{NT}}\sum_{t=1}^T\Lambda'\epsilon_t \right)\nonumber\\
&=\frac 1{\sqrt N}\left\{O_p\left(\widehat{\theta}^*-\bar J\theta\right) + O_p\left(\max\left(\frac 1{\sqrt N},\frac 1{\sqrt T}\right)\right) \right\}\nonumber\\
&=
\frac 1{\sqrt N}O_p\left(\widehat{\theta}^*-\bar J\theta\right)+ O_p\left(\max\left(\frac 1{ N},\frac 1{ T}\right)\right).
\end{align}
By substituting (ref) and (ref) into (ref)
\begin{align}
\norm{A_6}&=\frac 1{\sqrt N}O_p\left(\widehat{\theta}^*-\bar J\theta\right)+ O_p\left(\max\left(\frac 1{ N},\frac 1{ T}\right)\right).
\end{align}
For $A_7$ in (ref), we have
\begin{align}
A_7&=\frac 1{N\sqrt T}\frac 1T \sum_{t=1}^T\sum_{s=1}^T G_s\epsilon_s'\epsilon_t= \frac 1{\sqrt T} \left(\frac 1{\sqrt{NT}}\sum_{s=1}^T G_s \epsilon_s'\right)
\left(\frac 1{\sqrt{NT}}\sum_{t=1}^T \epsilon_t\right).
\end{align}
Thus, by barigozzi2022,
\begin{align}
\norm{A_7}=O_p\left(\frac 1{\sqrt T}\right).
\end{align}
Finally, for $A_8$ in (ref), we have
\begin{align}
A_8=&\, \frac 1{N^2 T} \frac 1{\sqrt T} \sum_{t=1}^T\sum_{s=1}^T \left(\frac{\widehat{\Lambda}^{*'}\Lambda}{N}\right)^{-1} \widehat{\Lambda}^{*'} \epsilon_s\epsilon_s'\epsilon_t\nonumber\\
=&\,\frac 1{N^2 T} \frac 1{\sqrt T} \sum_{t=1}^T\sum_{s=1}^T \left(\frac{\widehat{\Lambda}^{*'}\Lambda}{N}\right)^{-1} J {\Lambda}'\epsilon_s\epsilon_s'\epsilon_t\nonumber\\
&+ \frac 1{N^2 T} \frac 1{\sqrt T} \sum_{t=1}^T\sum_{s=1}^T \left(\frac{\widehat{\Lambda}^{*'}\Lambda}{N}\right)^{-1} (\widehat{\Lambda}^{*'} -J\Lambda')\epsilon_s\epsilon_s'\epsilon_t\nonumber\\
=:&\, A_{8,1}+A_{8,2}.
\end{align}
Then, by (ref), barigozzi2022, and Assumptions (ref)(ref) and (ref)(ref),
\begin{align}
A_{8,1}=&\, \frac 1{NT} \sum_{s=1}^T \left(\frac{\widehat{\Lambda}^{*'}\Lambda}{N}\right)^{-1}\left\{
\left(\frac 1{\sqrt N} J\Lambda'\epsilon_s\right)
\left(\frac1{\sqrt{NT}} \sum_{t=1}^T \epsilon_s'\epsilon_t-\mathbb E[\epsilon_s'\epsilon_t]\right)
\right\}\nonumber\\
&+ \frac 1{NT} \sum_{s=1}^T \left(\frac{\widehat{\Lambda}^{*'}\Lambda}{N}\right)^{-1}\left\{
\left(\frac 1{\sqrt N} J\Lambda'\epsilon_s\right)
\left(\frac1{\sqrt{NT}} \sum_{t=1}^T \mathbb E[\epsilon_s'\epsilon_t]\right)
\right\}\nonumber\\
=\,&O_p\left(\frac 1N\right)+O_p\left(\frac 1{\sqrt{NT}}\right).
\end{align}
Similarly,
\begin{align}
A_{8,2}=&\,\frac 1{\sqrt N}\frac 1T\sum_{s=1}^T \frac 1{\sqrt{N T}}\sum_{t=1}^T \left(\frac{\widehat{\Lambda}^{*'}\Lambda}{N}\right)^{-1} \frac{(\widehat{\Lambda}^{*'} -J\Lambda' )\epsilon_s}N \left\{\epsilon_s'\epsilon_t-\mathbb E[\epsilon_s'\epsilon_t]\right\}\nonumber\\
&+\frac 1{\sqrt N}\frac 1T\sum_{s=1}^T \frac 1{\sqrt{N T}}\sum_{t=1}^T \left(\frac{\widehat{\Lambda}^{*'}\Lambda}{N}\right)^{-1} \frac{(\widehat{\Lambda}^{*'} -J\Lambda' )\epsilon_s}N \mathbb E[\epsilon_s'\epsilon_t]\nonumber\\
=:&\, A_{8,2,1}+A_{8,2,2}.
\end{align}
By (ref), (ref), and barigozzi2022,
\begin{align}
\norm{A_{8,2,1}}&\le \frac 1{\sqrt N} \norm{\left(\frac{\widehat{\Lambda}^{*'}\Lambda}{N}\right)^{-1}} \left(\frac 1T\sum_{s=1}^T\norm{\frac 1{\sqrt{NT}}\sum_{t=1}^T\left\{\epsilon_s'\epsilon_t-\mathbb E[\epsilon_s'\epsilon_t]\right\} }^2\right)^{1/2}\left(\frac 1T\sum_{s=1}^T \frac{\norm{\epsilon_s}^2}{N}\right)^{1/2}\frac{\norm{\widehat{\Lambda}^{*} -\Lambda J }}{\sqrt N}\nonumber\\
&=\frac 1{\sqrt N}\left\{O_p\left(\widehat{\theta}^*-\bar J\theta\right) + O_p\left(\max\left(\frac 1{\sqrt N},\frac 1{\sqrt T}\right)\right) \right\}\nonumber\\
&=
\frac 1{\sqrt N}O_p\left(\widehat{\theta}^*-\bar J\theta\right)+ O_p\left(\max\left(\frac 1{ N},\frac 1{ T}\right)\right).
\end{align}
And, by the same arguments, and Assumption (ref)(ref)
\begin{align}
\norm{A_{8,2,2}}&\le \frac 1{\sqrt T} \frac{\norm{\widehat{\Lambda}^{*} -\Lambda J }}{\sqrt N} \norm{\left(\frac{\widehat{\Lambda}^{*'}\Lambda}{N}\right)^{-1}}\frac 1T\sum_{t=1}^T\sum_{s=1}^T \mathbb E[\epsilon_s'\epsilon_t]\frac{\Vert \epsilon_s\Vert}{\sqrt N}\nonumber\\
&=\frac 1{\sqrt T}\left\{O_p\left(\widehat{\theta}^*-\bar J\theta\right) + O_p\left(\max\left(\frac 1{\sqrt N},\frac 1{\sqrt T}\right)\right) \right\}\nonumber\\
&=
\frac 1{\sqrt T}O_p\left(\widehat{\theta}^*-\bar J\theta\right)+ O_p\left(\max\left(\frac 1{ N},\frac 1{ T}\right)\right).
\end{align}
Therefore, from (ref) and (ref), we have
\begin{align}
\norm{A_{8,2}}= \frac 1{\sqrt N}O_p\left(\widehat{\theta}^*-\bar J\theta\right)+\frac 1{\sqrt T}O_p\left(\widehat{\theta}^*-\bar J\theta\right)+ O_p\left(\max\left(\frac 1{ N},\frac 1{ T}\right)\right).
\end{align}
By substituting (ref) and (ref) into (ref) we obtain:
\begin{align}
\norm{A_{8}}= \frac 1{\sqrt N}O_p\left(\widehat{\theta}^*-\bar J\theta\right)+\frac 1{\sqrt T}O_p\left(\widehat{\theta}^*-\bar J\theta\right)+ O_p\left(\max\left(\frac 1{ N},\frac 1{ T}\right)\right).
\end{align}
Summing up, by substituting (ref), (ref), (ref), (ref), (ref), (ref), (ref), and (ref) into (ref), we prove part (ref).
For part (ref) just divide part (ref) by $\sqrt T$ and
notice that $\Vert G_t\Vert = O_p(1)$ because of Assumption (ref)(ref).
For part (ref) just replace in part (ref) $G_t$ with $X_t'\Lambda J(J\Lambda'\Lambda J)^{-1}$ which is also $O_p(1)$ and the second term on the right-hand-side of the statement comes from the $O_p(T^{-1})$ term in part (ref) (see also the expression of $A_7$ in (ref)).
\refstepcounter{lettersection}
\oldsection{Monte Carlo simulations}
Data generating process
We assume $r=1$ and $q=1$, i.e., one network factor and one node factor.
We then generate artificial time series of $y_t$ and $\mathcal{W}_t$, for $t=1, \dots, T$, according to the three model equations:
align*[align* omitted — 196 chars of source]
where $\mathcal{F}_t$ here has size $N \times N \times 1$, $U$ is a column vector of length $m$, $G_t$ is a scalar and $\Lambda$ is the $N$-dimensional column vector of node-specific factor loadings.
The Monte Carlo (MC) simulations are implemented as follows. We consider $N \in \{10,20,50,100,200\}$, $m \in \{20,50,100\}$, and $T \in \{50,80,100\}$.
First of all, we fix the values of FNAR parameters $\beta=0.5$, $\rho=0.3$, and $\alpha=0.2$.
Also, for each value of the pair ($N ,m$), we randomly generate the entries of the loading vectors $U$ and $\Lambda$ once (and independently)
from $\mathcal{N}(1,1)$, and then we keep them fixed across MC iterations (see chenetal2020cnar).
All results are based on $S=500$ iterations. Then, at each MC iteration we simulate the model according to the following.
itemize• For all $t=1,\ldots,T$, we generate the entries of the network factor matrix $F_{1,t}$ from $\mathcal{N}(0,1)$, except for the diagonal elements, which are set to zero to ensure the interpretation of $F_{1,t}$ as a network.
• For the multi-network of idiosyncratic terms $\mathcal{E}_t$, which is $N\times N\times m$, we consider two cases.
\begin{enumerate}[label=\Roman*]
• $\mathcal{E}_t$ has zero autocorrelation and zero cross-layer correlation.
• $\mathcal{E}_t$ is serially and cross-layer correlated.
Specifically, we generate $\mathcal{E}_t$ according to the equation
$$
\mathcal{E}_{i j \cdot, t} = \rho_{\mathcal{E}} \mathcal{E}_{i j \cdot, t-1} + \mathcal{N}(0_{m}, \Sigma_{\mathcal{E}}^{(m)}), \quad i,j=1,\ldots, N.
$$
Where $\rho_{\mathcal{E}}=0.5$ and the $(i,j)$-th element of
$\Sigma_{\mathcal{E}}^{(m)}$ is given by $\tau_{\mathcal{E}}^{\vert i-j \vert }$ if $\vert i-j \vert < d $, and 0 otherwise.
We set $\tau_{\mathcal{E}}$ to 0.5 and $d$ to 5.
\end{enumerate}
In both case I and case II, we assume no cross-correlation in the first two dimensions of $\mathcal{E}_t$.
Also, we set all diagonal elements of each $N \times N$ layer $\mathcal{E}_{\cdot \cdot k, t}$ to zero, $k=1, \dots, m$ and $t = 1, \dots, T$.
• To keep the noise-to-signal ratio in check, we choose to rescale the simulated elements of $\mathcal{E}_t$ so that the (iteration-specific) sample variance across all the elements in the $N \times N \times m \times T$ tensor of idiosyncratic components $\mathcal{E}$ is equal to half the sample variance across all the elements in the $N \times N \times m \times T$ tensor of common components $\mathcal{F} \times_3 U$.
• We simulate the node factor $G_t$ according to $G_t = \rho_G G_{t-1} + \mathcal{N}(0, 1)$ where $\rho_G= 0.2$. This allows us to study the iterated estimator under serially correlated FNAR errors, i.e., when the OLS and GLS are not consistent.
• The $N$ node-specific idiosyncratic terms in vector $\epsilon_{t}$ are generated from independent $\mathcal{N}(0, 1)$.
Then, to control the noise-to-signal ratio in the FNAR errors $\nu_t$, we rescale the simulated elements of $\epsilon_{t}$ so that the sample variance of $\text{vec}((\epsilon_1 \cdots \epsilon_T))$ is half the sample variance of $\text{vec}((\Lambda G_1 \cdots \Lambda G_T))$.
• We simulate $y_t$ and $\mathcal{W}_t$ according to the previous steps.
Next, at each iteration we re-estimate
$\mathcal{F}_t$, $U$, $\beta$, $\rho$ and $\alpha$ on the simulated $y_t$ and $\mathcal{W}_t$.
When estimating $\mathcal{F}_t$ and $U$, we implement a small-$N$ adjustment to our tensor principal component estimator. In particular, recall that the $N \times N$ network matrices (slices) composing the $N \times N \times m$ weight tensor $\mathcal{W}_{t}$ have zero diagonal elements. Accordingly, the number of elements of each network matrix that are effectively used to calculate the inner product $\widehat{\Gamma}^{\mathcal{W}}$, used for PC estimation,
is not $N^2$, but $N^2 - N$.
Of course, this is not an issue for large $N$, and asymptotic results for $N \to \infty$ are derived by simply considering $N^2$ elements in each network matrix, since the term $N$ is dominated by $N^2$ and is therefore asymptotically negligible.
However, for small values of $N$, estimation improves by taking into account the presence of zeros along the diagonals.
For this reason, in the MC simulations, we introduce the following adjustment to estimated loadings and factors, replacing $N^2$ with $N^2-N$:
equation*[equation* omitted — 241 chars of source]
To estimate the FNAR parameters, we use the bai2009panel iterative estimator, given the serial correlation assumed in the node-specific factor $G_t$.
Recall that factors, loadings and the FNAR parameter $\beta$ are consistently estimated only up to a sign. To address this issue, we change the signs of $\widehat{U}$ to ensure that $\widehat{U}' U $ is positive, then $\widehat{\mathcal{F}}_t$ and $\widehat{\beta}$ are adjusted accordingly.
For different values of $T$, $N$, and $m$, we report the MC root mean squared errors of the estimates, both in absolute terms (RMSE) and relative to the true parameter values (ReRMSE).
Let us define $\mathcal{F}^{mc}$ as the $N \times N \times 1 \times T \times S$ tensor storing all simulated tensors $\mathcal{F}_t$ for all time periods $T$ and all $S$ iterations, and let us define $\widehat{\mathcal{F}}^{mc}$ analogously for factor estimates.
Also, let us define $\widehat{U}^{mc}$ as the $N \times 1 \times S$ tensor storing all estimated loadings $\widehat{U}$ for all $S$ iterations, and $U^{mc}$ the tensor of same dimension obtained by replicating the true $U$ $S$ times.
Then, in the case of network factors and loadings,
the RMSE and relative RMSE are calculated as
$ \Vert \widehat{\mathcal{F}}^{mc} - \mathcal{F}^{mc} \Vert_F $, $ \vert \vert \widehat{U}^{mc} - U^{mc} \Vert_F$,
$ \Vert \widehat{\mathcal{F}}^{mc} - \mathcal{F}^{mc} \Vert_F / \Vert \mathcal{F}^{mc} \vert \vert_F $
and $ \Vert \widehat{U}^{mc} - U^{mc} \Vert_F / \Vert U^{mc} \Vert_F $, respectively. Likewise, we compute the RMSE and relative RMSE for the elements of the estimated FNAR coefficients, averaged over all $S$ iterations and when considering the estimator $\widehat{\theta}^\dag$ (the results for the elements of $\widehat{\theta}^*$ are almost identical). Finally, we also report the MC distribution of the estimated network effect $\widehat{\beta}$.
Results for case I are below, while results for case II are in Section (ref), except those for $T=80$ which are below.
RMSEs and Histograms - case I
Tables (ref), (ref), (ref) and (ref) report the {RMSE} and {ReRMSE} of parameter estimates for $T=10$, $T=50$, $T=80$ and $T=100$, respectively, under case I.
Figures (ref), (ref) and (ref) show the histograms of the MC estimates of $\beta$, for different combinations of $N,m,T$, under case I.
table[table omitted — 2,339 chars of source]
table[table omitted — 2,323 chars of source]
table[table omitted — 2,377 chars of source]
table[table omitted — 2,328 chars of source]
landscape\begin{figure}[H]
\caption{Monte Carlo histograms of the network effect $\widehat{\beta}$ - $T=10$, case I}
\begin{subfigure}{.24\textwidth}
\caption{$N=10, m=20$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=20, m=20$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=50, m=20$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=100, m=20$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=200, m=20$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=10, m=50$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=20, m=50$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=50, m=50$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=100, m=50$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=200, m=50$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=10, m=100$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=20, m=100$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=50, m=100$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=100, m=100$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=200, m=100$}
\end{subfigure}
\subcaption*{Note: the true value is $\beta= 0.5$.}
\end{figure}
landscape\begin{figure}[H]
\caption{Monte Carlo histograms of the network effect $\widehat{\beta}$ - $T=50$, case I}
\begin{subfigure}{.24\textwidth}
\caption{$N=10, m=20$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=20, m=20$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=50, m=20$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=100, m=20$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=200, m=20$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=10, m=50$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=20, m=50$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=50, m=50$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=100, m=50$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=200, m=50$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=10, m=100$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=20, m=100$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=50, m=100$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=100, m=100$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=200, m=100$}
\end{subfigure}
\subcaption*{Note: the true value is $\beta= 0.5$.}
\end{figure}
landscape\begin{figure}[H]
\caption{Monte Carlo histograms of the network effect $\widehat{\beta}$ - $T=80$, case I}
\begin{subfigure}{.24\textwidth}
\caption{$N=10, m=20$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=20, m=20$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=50, m=20$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=100, m=20$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=200, m=20$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=10, m=50$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=20, m=50$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=50, m=50$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=100, m=50$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=200, m=50$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=10, m=100$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=20, m=100$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=50, m=100$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=100, m=100$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=200, m=100$}
\end{subfigure}
\subcaption*{Note: the true value is $\beta= 0.5$.}
\end{figure}
landscape\begin{figure}[H]
\caption{Monte Carlo histograms of the network effect $\widehat{\beta}$ - $T=100$, case I}
\begin{subfigure}{.24\textwidth}
\caption{$N=10, m=20$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=20, m=20$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=50, m=20$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=100, m=20$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=200, m=20$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=10, m=50$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=20, m=50$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=50, m=50$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=100, m=50$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=200, m=50$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=10, m=100$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=20, m=100$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=50, m=100$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=100, m=100$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=200, m=100$}
\end{subfigure}
\subcaption*{Note: the true value is $\beta= 0.5$.}
\end{figure}
RMSEs and Histograms - case II
Tables (ref) and (ref) report the {RMSE} and {ReRMSE} of parameter estimates for $T=10$ and $T=80$, under case II. The results for $T=50$ and $T=100$ are reported in Section (ref) of the paper.
Figures (ref) and (ref) show the histograms of the MC estimates of $\beta$, for different combinations of $N,m,T$, under case II.
table[table omitted — 2,339 chars of source]
table[table omitted — 2,323 chars of source]
landscape\begin{figure}[H]
\caption{Monte Carlo histograms of the network effect $\widehat{\beta}$ - $T=10$, case II}
\begin{subfigure}{.24\textwidth}
\caption{$N=10, m=20$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=20, m=20$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=50, m=20$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=100, m=20$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=200, m=20$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=10, m=50$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=20, m=50$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=50, m=50$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=100, m=50$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=200, m=50$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=10, m=100$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=20, m=100$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=50, m=100$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=100, m=100$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=200, m=100$}
\end{subfigure}
\subcaption*{Note: the true value is $\beta= 0.5$. }
\end{figure}
landscape\begin{figure}[H]
\caption{Monte Carlo histograms of the network effect $\widehat{\beta}$ - $T=50$, case II}
\begin{subfigure}{.24\textwidth}
\caption{$N=10, m=20$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=20, m=20$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=50, m=20$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=100, m=20$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=200, m=20$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=10, m=50$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=20, m=50$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=50, m=50$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=100, m=50$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=200, m=50$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=10, m=100$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=20, m=100$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=50, m=100$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=100, m=100$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=200, m=100$}
\end{subfigure}
\subcaption*{Note: the true value is $\beta= 0.5$. }
\end{figure}
landscape\begin{figure}[H]
\caption{Monte Carlo histograms of the network effect $\widehat{\beta}$ - $T=80$, case II}
\begin{subfigure}{.24\textwidth}
\caption{$N=10, m=20$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=20, m=20$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=50, m=20$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=100, m=20$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=200, m=20$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=10, m=50$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=20, m=50$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=50, m=50$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=100, m=50$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=200, m=50$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=10, m=100$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=20, m=100$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=50, m=100$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=100, m=100$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=200, m=100$}
\end{subfigure}
\subcaption*{Note: the true value is $\beta= 0.5$.}
\end{figure}
landscape\begin{figure}[H]
\caption{Monte Carlo histograms of the network effect $\widehat{\beta}$ - $T=100$, case II}
\begin{subfigure}{.24\textwidth}
\caption{$N=10, m=20$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=20, m=20$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=50, m=20$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=100, m=20$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=200, m=20$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=10, m=50$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=20, m=50$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=50, m=50$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=100, m=50$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=200, m=50$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=10, m=100$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=20, m=100$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=50, m=100$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=100, m=100$}
\end{subfigure}
\begin{subfigure}{.24\textwidth}
\caption{$N=200, m=100$}
\end{subfigure}
\subcaption*{Note: the true value is $\beta= 0.5$.}
\end{figure}
Comparison of asymptotic standard errors
In this section, we compare the standard errors of the iterative estimators $\widehat {\theta}^{\dag}$ and $\widehat{\theta}^*$
(see (ref)). We compute these in two ways.
First, we use their analytical expressions. Let $\widehat {\beta}^{\dag}$ and $\widehat{\beta}^*$ be the first elements of $\widehat {\theta}^{\dag}$ and $\widehat{\theta}^*$, respectively, and let
$\widehat{\Sigma}(\theta^\dag)= \widehat{\operatorname*{Avar}}\left[ \sqrt{NT} (\widehat{\theta}^{\dag} -\theta) \right]$ and
$\widehat{\Sigma}(\theta^*)= \widehat{\operatorname*{Avar}}\left[ \sqrt{NT} (\widehat{\theta}^{*} -\theta) \right]$
be the estimated asymptotic covariance matrices as defined in (ref) (see Section (ref) above on how to fix the sign indeterminacy). Then, at each iteration of the MC exercise we compute:
$(NT)^{-1/2}\sqrt{[\widehat{\Sigma}(\theta^\dag)]_{11}}$ and $(NT)^{-1/2}\sqrt{[\widehat{\Sigma}(\theta^*)]_{11}}$, which are the estimators of the asymptotic standard errors of $\widehat {\beta}^{\dag}$ and $\widehat{\beta}^*$. Averages of these quantities over the $S$ iterations are in Table (ref) (second and third column) for selected values of $m,N,$ and $T$.
Second, we compute the MC standard errors of $\widehat {\beta}^{\dag}$ and $\widehat{\beta}^*$. Results are in Table (ref) (fourth and fifth column) for selected values of $m,N,$ and $T$.
table[table omitted — 848 chars of source]
Comparison with TOPUP and TIPUP
In this section, we compare the Monte Carlo performance of our tensor PC estimator with that of the two estimators proposed by chenyangzhang22, TOPUP and TIPUP. These are designed for the case of idiosyncratic tensors with no serial autocorrelation. They are implemented considering different values of $h_0$, i.e., the maximum lag used in the computation of inner and outer products of the tensor data. In particular, we report results for $h_0 = 1$ and $h_0 = 10$.
In Table (ref), we report the relative RMSE of the estimates of $U$ provided by TOPUP, TIPUP, and our PC-based approach, as well as the RMSE of the TOPUP and TIPUP estimators with respect to ours.
As expected, the TOPUP and PCA estimators behave very similarly when there is no autocorrelation in $\mathcal{E}$ (columns (a) and (b)), while the performance of TOPUP keeps deteriorating as the autocorrelation parameter $\rho_{\mathcal{E}}$ increases to 0.5 (columns (c) and (d)) and up to 0.9 (columns (e) and (f)). Our estimator has instead a similar performance in all considered cases, since by construction it is not affected by serial idiosyncratic autocorrelations, as long as the autocovariances are summable as prescribed by Assumption (ref)(ref). The TIPUP estimator seems to work poorly in all considered settings.
table[table omitted — 2,348 chars of source]
\refstepcounter{lettersection}
\oldsection{Empirical application: data and additional results}
In this section we provide additional information about the data used in the empirical application considered in the paper, as well as additional results.
In Section (ref) we provide details about how the network layers are built and about GDP data.
Additional empirical results are provided in Section (ref).
Data
\noindentCountries.
The $N=24$ countries we consider are further divided into: advanced economies, which are:
Australia (AUS), Belgium (BEL), Canada (CAN), France (FRA), Germany (DEU), Italy (ITA), Japan (JAP), South Korea (KOR), Netherlands (NLD), Norway (NOR), Spain (ESP), Sweden (SWE), Switzerland (CHE), United Kingdom (GBR), and United States (USA); and emerging economies, which are: Brazil (BRA), China (CHN), Hong Kong (HKG), India (IND), Indonesia (IDN), Mexico (MEX), Saudi Arabia (SAU), South Africa (ZAF), and Turkey (TUR),
\noindentReal GDP. For real GDP growth rates, we use the dataset compiled by mohaddesraissi20 to which we refer for details on the primary data sources and data adjustments. The dataset contains time series of log real GDP indices from 1979Q2 to 2019Q4, we take the first differences to calculate quarterly growth rates.
For Hong Kong, which is not included in the dataset of mohaddesraissi20, we retrieve data from the IMF's International Financial Statistics.
\noindentNetwork layers.
The full list of the $m=25$ network layers is reported in Table (ref). Here we summarize the way we built our data.
table[table omitted — 2,523 chars of source]
\noindentTrade in goods and services.
We calculate trade weights for goods using trade data from the UN Comtrade Database and classifying products into 10 categories, based on the 1-digit Standard International Trade Classification (SITC, revisions 3-4).
To facilitate interpretation of the results, we exclude unclassified commodities (SITC code 9).
Trade weights for services are computed using data from the OECD-WTO Balanced Trade in Services Database (BPM6 edition). Services are classified into 10 categories, based on the Balance of Payments Services Classification (EBOPS, 2010 edition for the period 2005-2019 and 2002 edition for the period 2001-2004). Two categories of EBOPS 2010, namely “SA" (manufacturing services on physical inputs owned by others) and “SB" (maintenance and repair services n.i.e.), are excluded because they were absent in EBOPS 2002 and have a large number of missing values in EBOPS 2010.
Given a product/service type representing the $k$-th layer of the network, the weight of country $j$ for country $i$ in layer $k$ at time $t$ (i.e., the element $ijk$ of the weight tensor $\mathcal{W}_t$) is calculated as:
equation[equation omitted — 147 chars of source]
where $imports_{ij,k,t}$ is the dollar amount of commodities of type $k$ imported by country $i$ from country $j$ at time $t$ and $exports_{ij,k,t}$ is the amount of exports of $k$ from country $i$ to country $j$.
\noindentFinancial assets and liabilities.
To construct the financial layers of the network, we first use data on outstanding bilateral assets/liabilities, classified into three categories of financial claims: equity, short-term debt and long-term debt. The data source is the IMF's Coordinated Portfolio Investment Survey (CPIS) database. The liabilities data that we use are derived from the assets data reported by counterparties. As explained by the IMF, while many countries report liabilities data directly, “more reliable detailed cross border positions data can usually be collected on an economy's holdings of portfolio investment because the holder (creditor) will usually know what securities it holds. On the liabilities side, the issuer of a security (debtor) may not know the residency of the holder because the securities may be held by foreign custodians or other intermediaries. Using the assets data reported by CPIS participating economies, the IMF derives liabilities data for all economies (CPIS reporters as well as nonreporters); these data are termed {\it derived liabilities}” (source: https://datahelp.imf.org/knowledgebase/articles/500647-why-are-coordinated-portfolio-investment-cpis-da).
Letting $k$ identify a category of financial claims, the weight of country $j$ for country $i$ in layer $k$ is calculated as:
equation[equation omitted — 154 chars of source]
where $assets_{ij,k,t}$ is the stock of assets held by country $i$ and issued by country $j$, and $liabilities_{ij,k,t}$ is the stock of liabilities of country $i$ towards country $j$.
\noindentFlows of banks' assets and liabilities.
We use data on flows of banks' assets and liabilities, as measured in the Locational Banking Statistics by the Bank for International Settlements (BIS). We cannot calculate bilateral capital flows by simply taking the first differences of the IMF CPIS assets and liabilities. The reason is that financial positions are reported at their market values in the CPIS, so first differences also reflect changes in valuation and exchange rate movements. In its International Banking Statistics, the BIS reports changes in amounts outstanding adjusted for exchange-rate movements and breaks in reporting methodologies, in order to approximate flows. In the case of BIS banking data, we generally use data provided directly by country $i$. Five countries (China, India, Indonesia, Norway and Saudi Arabia) do not report data. For these countries, we use derived flows, i.e., data reported by their counterparties. As for bilateral flows among non-reporting countries, we assume them to be zero.
The weight of country $j$ for country $i$ in layer $k$ is calculated as:
equation[equation omitted — 215 chars of source]
where $\Delta bank\_assets_{ij,t}$ is the annual change (at year-end) in outstanding assets held by banks located in country $i$ and issued by counterparties of any type located in country $j$, and $ \allowbreak \Delta bank\_liabilities_{ij,t}$ is the annual change in liabilities of banks located in country $i$ and held by counterparties of any type located in country $j$.
\noindentMergers and acquisitions.
We collect all cross-border mergers and acquisitions (M&A) deals involving companies located in two different countries and classify them in two macro-sectors of economic activity, “goods sectors" and “services sectors", using the Standard Industrial Classification (SIC). Data are from Bureau Van Dijk's Zephyr database (160,713 completed deals between the 24 countries considered over the period 2001-2019, corresponding to an average of around 15 deals for each pair of countries per year). We consider not only operations that are strictly speaking mergers or acquisitions, i.e., involving more than 50% of the target firms' capital, but also capital increases and acquisitions of minority stakes. We classify the M&A deals using the primary economic sector of the target companies. The rationale for this choice is that acquirors tend to be larger companies operating in a larger number of sectors than targets, so using the sectors of targets should provide a more accurate classification.
Finally, given the greater volatility of M&A weights compared to the other types of weights (M&A operations are rarer than transactions in goods, services and portfolio instruments), we smooth out the M&A weights by taking 3-year moving averages.
If $k$ identifies M&A deals in a given sector, the weight of country $j$ for country $i$ is given by:
equation[equation omitted — 105 chars of source]
where $m\&a_{ij,k,t}$ denotes the dollar value of all mergers and acquisitions in sector $k$ at time $t$ involving a company located in country $i$ and a company located in country $j$. This includes both the deals in which country $i$'s companies are acquirors and the deals in which they are targets.
Additional results
Cosine similarity
The main premise of our approach is that the different layers of the network are driven by common factor networks. As a preliminary analysis, it is therefore useful to assess the degree of similarity between the observed layers. As suggested by bargiglietal15, we consider cosine similarity as a measure of layer similarity for weighted networks.
The cosine similarity ($C$) between two vectors $x$ and $y$ is their inner product normalized by the product of their norms, i.e.:
$C(x,y) = \frac{x'y}{\norm{x}\,\norm{y}}. $
To calculate the similarity between any two layers $h$ and $k$, we vectorize the matrices $\mathcal{W}_{\cdot\cdot,h,t}$ and $\mathcal{W}_{\cdot\cdot,k,t}$ for each $t$, then stack the resulting vectors for $t=1,\dots,T$ to obtain two column vectors of length $N^2T \times 1$ each. Results are presented in Table (ref).
landscape\begin{table}[t!]
\caption{Cosine similarity between layers.}
\scriptsize
{
\begin{tabular}{l|lllllllllllllllllllllllll}
\hline \hline
layer & 1 & 2 & 3 & 4 & 5 & 6 & 7 & 8 & 9 & 10 & 11 & 12 & 13 & 14 & 15 & 16 & 17 & 18 & 19 & 20 & 21 &22 &23 &24 &25 \\
\hline
1 & & 0.75 & 0.79 & 0.74 & 0.70 & 0.87 & 0.87 & 0.85 & 0.80 & 0.80 & 0.83 & 0.74 & 0.71 & 0.71 & 0.67 & 0.78 & 0.77 & 0.74 & 0.64 & 0.62 & 0.71 & 0.54 & 0.56 & 0.54 & 0.58 \\
2 & 0.75 & & 0.63 & 0.58 & 0.57 & 0.77 & 0.72 & 0.74 & 0.67 & 0.74 & 0.78 & 0.69 & 0.72 & 0.70 & 0.63 & 0.74 & 0.75 & 0.73 & 0.60 & 0.64 & 0.68 & 0.59 & 0.59 & 0.58 & 0.61 \\
3 & 0.79 & 0.63 & & 0.68 & 0.65 & 0.76 & 0.82 & 0.80 & 0.78 & 0.68 & 0.71 & 0.61 & 0.57 & 0.54 & 0.51 & 0.63 & 0.62 & 0.61 & 0.50 & 0.48 & 0.53 & 0.40 & 0.45 & 0.48 & 0.49 \\
4 & 0.74 & 0.58 & 0.68 & & 0.63 & 0.72 & 0.73 & 0.69 & 0.60 & 0.63 & 0.68 & 0.65 & 0.57 & 0.58 & 0.50 & 0.64 & 0.61 & 0.61 & 0.47 & 0.46 & 0.50 & 0.42 & 0.46 & 0.45 & 0.47 \\
5 & 0.70 & 0.57 & 0.65 & 0.63 & & 0.64 & 0.65 & 0.61 & 0.57 & 0.56 & 0.60 & 0.57 & 0.49 & 0.48 & 0.44 & 0.55 & 0.53 & 0.53 & 0.43 & 0.40 & 0.47 & 0.36 & 0.37 & 0.40 & 0.43 \\
6 & 0.87 & 0.77 & 0.76 & 0.72 & 0.64 & & 0.93 & 0.94 & 0.87 & 0.84 & 0.85 & 0.77 & 0.75 & 0.72 & 0.72 & 0.83 & 0.83 & 0.78 & 0.68 & 0.66 & 0.70 & 0.57 & 0.61 & 0.61 & 0.62 \\
7 & 0.87 & 0.72 & 0.82 & 0.73 & 0.65 & 0.93 & & 0.93 & 0.89 & 0.80 & 0.84 & 0.74 & 0.70 & 0.67 & 0.62 & 0.77 & 0.75 & 0.71 & 0.61 & 0.57 & 0.65 & 0.53 & 0.56 & 0.55 & 0.57 \\
8 & 0.85 & 0.74 & 0.80 & 0.69 & 0.61 & 0.94 & 0.93 & & 0.89 & 0.83 & 0.85 & 0.75 & 0.74 & 0.70 & 0.68 & 0.80 & 0.80 & 0.76 & 0.67 & 0.63 & 0.69 & 0.56 & 0.58 & 0.59 & 0.62 \\
9 & 0.80 & 0.67 & 0.78 & 0.60 & 0.57 & 0.87 & 0.89 & 0.89 & & 0.78 & 0.80 & 0.69 & 0.68 & 0.66 & 0.65 & 0.74 & 0.75 & 0.70 & 0.64 & 0.62 & 0.65 & 0.52 & 0.56 & 0.54 & 0.57 \\
10 & 0.80 & 0.74 & 0.68 & 0.63 & 0.56 & 0.84 & 0.80 & 0.83 & 0.78 & & 0.84 & 0.80 & 0.82 & 0.77 & 0.79 & 0.87 & 0.86 & 0.81 & 0.74 & 0.74 & 0.76 & 0.61 & 0.65 & 0.63 & 0.66 \\
11 & 0.83 & 0.78 & 0.71 & 0.68 & 0.60 & 0.85 & 0.84 & 0.85 & 0.80 & 0.84 & & 0.80 & 0.78 & 0.78 & 0.70 & 0.83 & 0.83 & 0.83 & 0.68 & 0.70 & 0.75 & 0.65 & 0.65 & 0.61 & 0.66 \\
12 & 0.74 & 0.69 & 0.61 & 0.65 & 0.57 & 0.77 & 0.74 & 0.75 & 0.69 & 0.80 & 0.80 & & 0.72 & 0.74 & 0.67 & 0.78 & 0.77 & 0.74 & 0.61 & 0.64 & 0.69 & 0.62 & 0.59 & 0.57 & 0.61 \\
13 & 0.71 & 0.72 & 0.57 & 0.57 & 0.49 & 0.75 & 0.70 & 0.74 & 0.68 & 0.82 & 0.78 & 0.72 & & 0.84 & 0.76 & 0.85 & 0.85 & 0.82 & 0.72 & 0.75 & 0.77 & 0.65 & 0.70 & 0.63 & 0.67 \\
14 & 0.71 & 0.70 & 0.54 & 0.58 & 0.48 & 0.72 & 0.67 & 0.70 & 0.66 & 0.77 & 0.78 & 0.74 & 0.84 & & 0.78 & 0.85 & 0.87 & 0.85 & 0.74 & 0.79 & 0.81 & 0.70 & 0.71 & 0.64 & 0.69 \\
15 & 0.67 & 0.63 & 0.51 & 0.50 & 0.44 & 0.72 & 0.62 & 0.68 & 0.65 & 0.79 & 0.70 & 0.67 & 0.76 & 0.78 & & 0.84 & 0.88 & 0.80 & 0.81 & 0.83 & 0.76 & 0.59 & 0.62 & 0.63 & 0.65 \\
16 & 0.78 & 0.74 & 0.63 & 0.64 & 0.55 & 0.83 & 0.77 & 0.80 & 0.74 & 0.87 & 0.83 & 0.78 & 0.85 & 0.85 & 0.84 & & 0.91 & 0.88 & 0.80 & 0.80 & 0.79 & 0.66 & 0.70 & 0.68 & 0.71 \\
17 & 0.77 & 0.75 & 0.62 & 0.61 & 0.53 & 0.83 & 0.75 & 0.80 & 0.75 & 0.86 & 0.83 & 0.77 & 0.85 & 0.87 & 0.88 & 0.91 & & 0.89 & 0.82 & 0.84 & 0.83 & 0.69 & 0.70 & 0.68 & 0.71 \\
18 & 0.74 & 0.73 & 0.61 & 0.61 & 0.53 & 0.78 & 0.71 & 0.76 & 0.70 & 0.81 & 0.83 & 0.74 & 0.82 & 0.85 & 0.80 & 0.88 & 0.89 & & 0.75 & 0.79 & 0.77 & 0.67 & 0.69 & 0.67 & 0.71 \\
19 & 0.64 & 0.60 & 0.50 & 0.47 & 0.43 & 0.68 & 0.61 & 0.67 & 0.64 & 0.74 & 0.68 & 0.61 & 0.72 & 0.74 & 0.81 & 0.80 & 0.82 & 0.75 & & 0.79 & 0.72 & 0.53 & 0.58 & 0.59 & 0.61 \\
20 & 0.62 & 0.64 & 0.48 & 0.46 & 0.40 & 0.66 & 0.57 & 0.63 & 0.62 & 0.74 & 0.70 & 0.64 & 0.75 & 0.79 & 0.83 & 0.80 & 0.84 & 0.79 & 0.79 & & 0.80 & 0.67 & 0.64 & 0.66 & 0.68 \\
21 & 0.71 & 0.68 & 0.53 & 0.50 & 0.47 & 0.70 & 0.65 & 0.69 & 0.65 & 0.76 & 0.75 & 0.69 & 0.77 & 0.81 & 0.76 & 0.79 & 0.83 & 0.77 & 0.72 & 0.80 & & 0.75 & 0.67 & 0.60 & 0.64 \\
22 & 0.54 & 0.59 & 0.40 & 0.42 & 0.36 & 0.57 & 0.53 & 0.56 & 0.52 & 0.61 & 0.65 & 0.62 & 0.65 & 0.70 & 0.59 & 0.66 & 0.69 & 0.67 & 0.53 & 0.67 & 0.75 & & 0.61 & 0.56 & 0.58 \\
23 & 0.56 & 0.59 & 0.45 & 0.46 & 0.37 & 0.61 & 0.56 & 0.58 & 0.56 & 0.65 & 0.65 & 0.59 & 0.70 & 0.71 & 0.62 & 0.70 & 0.70 & 0.69 & 0.58 & 0.64 & 0.67 & 0.61 & & 0.59 & 0.60 \\
24 & 0.54 & 0.58 & 0.48 & 0.45 & 0.40 & 0.61 & 0.55 & 0.59 & 0.54 & 0.63 & 0.61 & 0.57 & 0.63 & 0.64 & 0.63 & 0.68 & 0.68 & 0.67 & 0.59 & 0.66 & 0.60 & 0.56 & 0.59 & & 0.61 \\
25 & 0.58 & 0.61 & 0.49 & 0.47 & 0.43 & 0.62 & 0.57 & 0.62 & 0.57 & 0.66 & 0.66 & 0.61 & 0.67 & 0.69 & 0.65 & 0.71 & 0.71 & 0.71 & 0.61 & 0.68 & 0.64 & 0.58 & 0.60 & 0.61 & \\
\hline \hline
\end{tabular}
}
\subcaption*{ The table reports the cosine similarity coefficients between the 25 layers of the network. Please refer to Table (ref) for the list of layers.}
\end{table}
Network factors
Figure (ref) displays the average values of the factors over the period 2001-2019 using color scales. Green cells identify positive weights and red cells negative weights, with darker shades of color indicating larger weights in absolute value. Figure (ref) plots the factor loadings for different layers of the network.
figure[figure omitted — 1,009 chars of source]
figure[figure omitted — 958 chars of source]
Figure (ref) reports the time-varying weights of different countries in the network of US. For each network factor, the countries reported are those that have on average over the sample 2001-2019 the largest weights in absolute values. Similarly, Figure (ref) reports the time-varying weights of different countries in the network of China. Again, for each network factor, the countries reported are those that have on average over the sample 2001-2019 the largest weights in absolute values. The estimated network factors are divided by $N$.
figure[figure omitted — 548 chars of source]
figure[figure omitted — 553 chars of source]
Explained variance
Consider the order-4 tensor $\mathcal{W}$ of size $ N \times N \times m \times T$ such that $\text{mat}_{(4)}(\mathcal W)$ is $T\times N^2m$ with $t$-th row $\text{vec}(\text{mat}_{(1)} \mathcal{W}_{t})$,
and, letting $\widehat{\mathcal{W}}_t^{(k)} := \widehat{\mathcal{F}}_{\cdot\cdot,k,t} \times_3 \widehat{U}_{\cdot k}$, $t=1,\ldots, T$, $k=1,\dots,r$, we define
the tensor $\widehat{\mathcal{W}}^{(k)}$ of size $N \times N \times m \times T$ in the same way.
The fraction of variance in $\mathcal{W}$ explained by the $k$-th factor, denoted as $v^{(k)}$, is then:
\[
v^{(k)} = \frac{ \norm{ \widehat{\mathcal{W}}^{(k)}}_F^2}{ \norm{ \mathcal{W} }_F^2}, \quad k=1,\ldots, r.
\]
And we have $v^{(1)}=0.68$, $v^{(2)}=0.07$, $v^{(3)}=0.03$, $v^{(4)}=0.03$, $v^{(5)}=0.02$, and $v^{(6)}=0.02$. Thus, overall the 6 factors explain about 85% of the total variance of $\mathcal W$. However, the importance of different factors varies greatly across countries. Let $v_{ij}^{(k)}$ be the fraction of variance in the weight of country $j$ for country $i$ explained by the $k$-th factor. Then, we have that:
\[
v_{ij}^{(k)} = \frac{ \norm{ \widehat{\mathcal{W}}_{ij,\cdot\cdot}^{(k)} }_F^2}{ \norm{ \mathcal{W}_{ij,\cdot\cdot} }_F^2}, \quad i,j=1,\ldots, N, \quad k=1,\ldots, r.
\]
For each factor, Table (ref) reports the ten network links for which the factor explains the largest share of variance.
table[table omitted — 4,156 chars of source]
Details on the forecasting application
Forecasts are computed according to the following recursive forecasting scheme. We initially estimate the model on a sample window ending in 2001Q4, then we expand the sample by one-quarter increments up to 2019Q3, and we use the estimates obtained in each window to produce 1-quarter-ahead forecasts of GDP growth, from 2002Q1 to 2019Q4. The smallest window ends in 2001 as this is the first year in which data on financial networks are available, as explained in Section (ref) of the paper. To make the exercise feasible, the start date of the estimation sample is moved back in time, specifically to 1980Q1, and we assume that the network tensor until 2000 is constant at its first available value (2001).
We consider the following competing estimators.
enumerate• The tensor-based estimators developed by wangetal21 to estimate coefficients of high-dimensional VARs. Namely, we consider the multilinear low-rank (MLR) estimator and the sparse higher-order reduced-rank (SHORR) estimator. These estimators apply a Tucker decomposition to the order-3 tensor of unknown VAR coefficients, where the first two dimensions of the tensor are given by the number of variables in the VAR and the third dimension is given by the number of lags. To allow for a proper order-3 tensor, we consider a VAR with 2 lags in this case. In each estimation window, the multilinear ranks (i.e., the dimensions of the core tensor in the Tucker decomposition) are selected using the ridge-type ratio described in wangetal21.
• Two multilayer NARs (ref) and (ref), based on networks extracted from two alternative full Tucker tensor decompositions, as detailed in Appendix (ref).
• A multilayer NAR that includes all 25 layers composing the original weight tensor $\mathcal{W}_t$, estimated by means of the LASSO or Ridge estimators.
• A simple VAR(1) model that does not utilize any information on the underlying networks.
\refstepcounter{lettersection}
\oldsection{NAR estimation based on full Tucker decompositions}
Consider the $N\times N\times m$ tensor of observed data $\mathcal W_t$ having as slices the $m$ layers $W_{k,t}$.
A full Tucker decomposition of $\mathcal W_t$ reads:
align[align omitted — 140 chars of source]
where the tensor of factors is $\mathcal G_t$ of size $p\times q\times r$ and the loadings are $V_1$ of size $N\times p$, $V_2$ of size $N\times q$, and $V_3$ of size $m\times r$.
We can get estimated loadings $\widehat V_j$ in (ref), using the TOPUP estimators of chenyangzhang22. The estimated factors are then obtained by linear projection as
$\widehat{\mathcal G}_t= \mathcal W_t \times_1 (\widehat V_1'\widehat V_1)^{-1} \widehat V_1'\times_2 (\widehat V_2'\widehat V_2)^{-1} \widehat V_2'\times_3 (\widehat V_3'\widehat V_3)^{-1} \widehat V_3'$. We cannot directly use $\widehat{\mathcal G}_t$ for estimating the multilayer NAR as now these do not have the right dimensions and cannot be interpreted as networks.
We then consider two possibilities.
The first one is to use the multilayer common network
align[align omitted — 136 chars of source]
Denote as $\widehat S_{t,k}$, $k=1,\ldots, m$, the $m$ layers of $\widehat {\mathcal S}_t$, each of size $N\times N$.
Thus, we could think of estimating a new multilayer NAR of the form
align[align omitted — 198 chars of source]
However, this model has by construction collinear regressors. Indeed, write (ref) as $y_t=\widehat X_t\theta+\omega_t$ with $\widehat X_t := (\widehat{\mathsf S}_{1,t-1,\cdot}',\cdots, \widehat{\mathsf S}_{m,t-1,\cdot}', y_{t-1}, \iota_N)$ where $\widehat{\mathsf S}_{k,t-1,\cdot}$ is the $t$-th row of $\widehat{\mathsf S}_k:=\left(\widehat S_{k,0} y_{0},\cdots, \widehat S_{k,T-1} y_{T-1}\right)^\prime$. Then,
$\sum_{t=1}^T \widehat X_t' \widehat X_t$ is $m+2\times m+2$ but it is not invertible, since it has rank $r+2$. We therefore applied either LASSO or Ridge to estimate (ref).
The second possibility is to calculate the $N \times N \times r$ tensor
align[align omitted — 134 chars of source]
and then to use the matrices $\widehat{F}_{k,t}^*$, with $k=1,\ldots, r$, defined as the slices of $\widehat{\mathcal{F}}_t^*$ of size $N \times N$, in place of our factors $\widehat{F}_{k,t}$, $k=1,\ldots, r$, in the FNAR, i.e.,
align[align omitted — 210 chars of source]
We refer to Table (ref) in Section (ref) of the paper for an empirical comparison between our approach and these two alternative full Tucker approaches. The results show that our approach delivers in general better predictions.
We conclude with a comment on consistency rates. First, note that our estimator of the loadings $\widehat U$ (in the paper) has the same rate as $\widehat V_3$, however, both $\widehat{\mathcal S}_t$ and $\widehat{\mathcal F}_t^*$ depend also on the estimated loadings $\widehat V_1$ and $\widehat V_2$, while our network factors $\widehat{\mathcal F}_t$ used in the FNAR do not. Therefore the consistency rates of $\widehat{\mathcal S}_t$ and $\widehat{\mathcal F}_t^*$ might differ from those we derived for $\widehat{\mathcal F}_t$. In particular, $\widehat{\mathcal S}_t$ and $\widehat{\mathcal F}_t^*$ depend on the estimated loadings both directly in their definition and indirectly through the estimated core tensor $\widehat{\mathcal G}_t$.
Tables (ref) and (ref) summarize this comparison.
table[table omitted — 1,035 chars of source]
table[table omitted — 1,304 chars of source]
Consider estimation of the loadings in Table (ref), then the rates depend on two terms:
(A) a term due to the fact that the factors are unknown, and (B) a classical term which would be obtained by linear projection of the data onto known factors.
Similarly, considering estimation of the factors in Table (ref), the rates also depend on two terms:
(A) a term due to the fact that the loadings are unknown, and (B) a classical term which would be obtained by linear projection of the data onto known loadings.
The results in this table follow directly by noticing that the rates of the TOPUP estimator of the loadings by chenyangzhang22 coincide with those derived by helitrapani2022, and from the results in the latter paper the rates for $\widehat{\mathcal G}_t$ can be obtained. Those for $\widehat{\mathcal S}_t$ and $\widehat{\mathcal F}_t^*$ are then simply the worst between those of $\widehat{\mathcal G}_t$ and the estimated loadings.
It is then clear that the alternative approaches based on a full-Tucker decomposition have a term in the consistency rates which is always faster than ours (column B) and have a term which is faster than ours if $N/m\to 0$ (column A). Nevertheless, as we mentioned, these alternative approaches, although interesting, do not deliver networks that are economically interpretable nor are able to outperform our approach in the considered empirical application.
\let\section\oldsection