The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
182,443 characters
Robust Nearly-Efficient Estimation of Large Panels with Factor Structures
\maketitle
\begin{abstract}
This paper studies estimation of linear panel regression models with heterogeneous coefficients, when both the regressors and the residual contain a possibly common, latent, factor structure. Our theory is (nearly) efficient, because based on the GLS principle, and also robust to the specification of such factor structure, because it does not require any information on the number of factors nor estimation of the factor structure itself.
We first show how the unfeasible GLS estimator not only affords an efficiency improvement but, more importantly, provides a bias-adjusted estimator with the
conventional limiting distribution, for situations where the OLS is affected by a first-order bias. The technical challenge resolved in the paper is to show how these properties are preserved for a class of feasible GLS estimators in a double-asymptotics setting. Our theory is illustrated by means of Monte Carlo exercises and, then, with an empirical application using individual asset returns and firms' characteristics data.
\end{abstract}
\textbf{Keywords:} GLS estimation; panel; factor structure; robustness; bias-correction.
\section{Introduction}
This paper considers (nearly) {\em efficient} estimation of linear panel regression models with heterogeneous coefficients, when both the regressors and the residual contain a common, latent, factor structure. At the same time, our estimation procedure does not require any knowledge of such latent factor structure, not even the maximum possible number of latent factors, let alone the latent factors themselves. This qualifies our procedure as {\em robust}.
Factor models represent one of the most popular and successful way to capture cross-sectional and temporal dependence, especially when facing a large number of units ($N$) and time periods ($T$), although in our context factors and their loadings represent nuisance parameters.
However, the possibility of a common factor structure in both regressors and residuals, which
would typically arise when omitting relevant regressors, leads to endogeneity, making estimation by ordinary least squares (OLS) invalid such that all its asymptotic properties are not holding any longer.
We first consider an unfeasible generalized least squares (UGLS) estimator for the regression coefficients,
based on the presumption that the covariance matrix of the residuals, evaluated conditional on the latent factors, is known. It turns out that, regardless of the possibility of endogeneity (that is when regressors and residuals are correlated), the UGLS is $\sqrt{T}$-consistent and asymptotically normal
distributed {\em without } requiring any information on the factor structure, such as the number of factors or the factors themselves and their loadings.
This contrasts with the asymptotic bias plaguing the OLS estimator, under the same circumstances.
In other words, the UGLS does not only represent a more efficient estimator but provides an {\em automatic } biased-adjusted estimator with desirable asymptotic properties. This result is due to an important insight, namely the existence of a form of {\em asymptotic orthogonality } between the common factors, that affect the residuals, and the {\em inverse } of the residuals' covariance matrix. Most importantly, such asymptotic orthogonality is manifested at a very fast rate, namely the squared norm of the product between the covariance matrix and the factors is $O(T^{-1})$.
The challenge arises when considering a
feasible version of the UGLS estimator. A natural approach, here followed, is to
make use of the panel dimension, considering the sample (across $N$) covariance
matrix of the OLS residuals, which in turn have been obtained by a (time-series) regression with $T$ observations.
Unlike the OLS and UGLS cases, the asymptotic theory for the GLS requires both $N$ and $T$ to diverge. Lack of consistency of the OLS estimator for the regression coefficients unavoidably implies that such $T \times T $ sample covariance matrix is not consistent, element by element, for the true residuals' covariance matrix but will converge (element by element) instead to a pseudo-true covariance matrix.
The surprising, crucial, result here established is that such pseudo-true covariance matrix is also asymptotically orthogonal
to the latent factors, and at the same rate of convergence $O(T^{-1})$. Indeed, there is an entire class of
matrices, rather than a unique matrix, that is asymptotically
orthogonal to the factors. This is the most intriguing aspect of our theory. The technical achievement of this paper is to show that the {\em feasible }
GLS (henceforth GLS)
estimator for the regression coefficients, is $\sqrt{T}$-consistent and asymptotically normal, as {\em both } $N,T$ diverge to
infinity. Again, this holds even when the OLS remains an invalid estimator.
At the same time, since the pseudo-true value differs in general from the true covariance matrix, the GLS might not be as efficient as
the UGLS. However, evaluation of the GLS iteratively, as explained below, permits to make it close to the UGLS estimator.
In summary, the GLS estimator exhibits four main desirable, compelling, properties.
First, it permits to carry out inference on the regression coefficients based on conventional asymptotic distributions. In particular, the GLS estimator of the regression coefficients has a mixed-normal asymptotic distribution, implying the possibility of inference by means of chi-squared criteria.
Second, as in classical estimation theory, it delivers (nearly) efficient estimation.
Third, the GLS estimator does not require any knowledge of the exact number of latent
factors, or even an upper bound of such number. In particular, the number of factors can be either smaller, equal or larger than the number of regressors.
Fourth, the GLS is computationally easy to handle since it simply requires to perform ($N$) linear
regressions, without invoking any nonlinear numerical optimizations.
Our approach can be also applied to the dual case of cross-sectional regressions with time-varying coefficients.
This paper belongs to, and extends, two different strands of literature.
First, it has been demonstrated, in various contexts, that efficient estimation techniques not only lead to an improvement of precision but, most importantly, resurrect the required asymptotic properties, in terms of bias, rate of convergence and distribution, in situations where these are not warranted by non-efficient approaches.
In the context of cointegrated systems, \cite{P91b} and \cite{P91a}
show that use of the efficient, full system, maximum likelihood (ML) goes beyond an efficiency improvement: it solves the well-known issues of specification and inference in cointegrated systems, that plagues unrestricted VAR estimation such as the presence of asymptotic biases and non-standard asymptotic distributions (i.e. the Dickey-Fuller distribution).\footnote{Indeed,
especially \cite{P91b} provides a detailed explanation of these properties, namely removing second order bias, dealing with endogeneity, absence of nuisance parameters and, obviously, achievement of full efficiency.}
Note that ML is asymptotically equivalent to GLS in that set-up. Although our theoretical framework is not one of cointegrated systems, strong analogies emerge with \cite{P91b} and \cite{P91a}: in both cases, a (local)
mixed-normal distribution arises and efficient estimation mitigates the lack of strong exogeneity. Moreover, such deficiency (i.e. lack of exogeneity) is manifested through the form of the residuals' covariance matrix: non-block diagonality for (triangular) cointegrated systems of \cite{P91b}, \cite{P91a} and a factor structure such as ours, which also rules out block-diagonality, for our framework. Second, \cite{P91b} demonstrates how these remarkable properties of efficient estimation are warranted by full system regressions but not by single-equation regressions. Likewise, our method requires the full information arising from the panel, namely one needs both $T$ and $N$ to diverge.
\cite{RH97} study estimation of time series regression models, when both the regressors and the residual exhibit long-memory, and in fact spectral singularities can arise at any frequency. Under these circumstances, in particular when the spectral singularities of the regressors and residuals arise at the same frequency with sufficient intensity, the OLS estimator is no longer $\sqrt{T}$-consistent and asymptotically normal. However, under the same circumstances, \cite{RH97} show that a class of weighted least squares estimates, which includes GLS as a special case, has standard asymptotic properties.
An important difference between our approach and \cite{P91b}, \cite{P91a} and \cite{RH97} is that their estimation procedure is affected by a second-order bias, that is their estimators are consistent (although with non-standard rate of convergence and asymptotic distribution), whereas in our context a first-order bias arises, leading for instance to inconsistency of the OLS estimator. Therefore, our GLS adjustment appears compelling in our framework.
Second, inference of panel data model with a latent factor structure in the residuals and {\em heteroreneous} regression coefficients has been studied, initially, from a purely econometric perspective and, more recently, from an empirical finance angle.
In a linear cross-sectional
regression \cite{A05} shows that, when residuals and regressors share a factor
structure, $\sqrt{N}$-consistency of the OLS estimator is preserved only with uncorrelated factor loadings.\footnote{Although not spelled out, \cite{A05} can be readily applied to time-varying coefficients.}
Within a linear time regression,
\cite{P06} shows that
heterogeneous regression coefficients can be $\sqrt{T}$-consistently
estimated OLS by augmenting the regressors with cross-sectional averages
of the dependent variable and individual-specific regressors. \cite{AB15}
consider a panel model with (sparse) heterogenous coefficients and establish the asymptotics of a penalized OLS estimator.
Maintaining the assumption of a common latent factor structure in the residuals of a panel data model with heterogeneous coefficients, \cite{vel17} allow for the possibility that
the idiosyncratic innovation is non-stationary, in particular exhibiting long memory.
Motivated by empirical asset pricing, new methods to conduct robust inference on panel data models with a latent factor structure have been recently developed.
\cite{GX18} derive the asymptotics for a procedure to estimate the risk-premium of an observed factor, robust to the omission of the set of relevant (i.e. priced) factors.
Like us, they adopt a double-asymptotics approach. However, \cite{GX18} differ from us because they focus on estimation of the parameters of the second-pass regression,
that is when the asset-pricing restriction is imposed, whereas we ignore any asset-pricing content (i.e., from the point of view of the two-pass methodology, we focus on the parameters of the first-pass regression). Moreover, their procedure relies on estimating the complete space spanned by the latent factors driving the model whereas our method can avoid this aspect altogether. Gagliardini, Ossola and Scaillet (2018)
study the properties of
a diagnostic criterion to detect an approximate factor structure in the residuals of large, unbalanced, panel data models. Like us, they consider a double-asymptotic setting and ignore any
asset-pricing restrictions on the parameters of the panel data model. Moreover, \cite{GOS18} method is robust, in the sense that it does not need to explicitly estimate the latent factor structure embedded in the residuals, just like us. However their focus is specifically to check whether the unobserved residuals have a factor structure whereas our method focuses on estimation of the regression coefficients to the observed, possibly heterogenous, regressors.
Unlike the previous papers, the large majority of contributions to this literature focused on the case of constant regression coefficients.
\cite{P06} shows that a faster rate of convergence is achieved with constant regression coefficients.
\cite{B09} considers joint estimation of the constant regression coefficients and of the residuals' factor structure components through an iterative OLS procedure.
The same estimator has been studied by \cite{MW09} under weaker conditions on the observed regressors.
\cite{MW13} show that \cite{B09} and \cite{MW09} results hold, with no loss of efficiency, when the exact number of latent factors $M$ is unknown and only an upper bound is specified. \cite{BAIGLS} show that GLS estimation leads to an efficiency improvement over the \cite{B09} and \cite{MW09} OLS-type estimator.
\cite{GHS12} establish $\sqrt{NT}$-asymptotics for the OLS estimator by augmenting the regressors with the principal component estimator of the common factors extracted from the observable data.
\footnote{Several generalizations of the aforementioned results have been considered.
\cite{PT11} and \cite{CP13} confirm the same asymptotic results of \cite{P06} when spatial-dependence in the idiosyncratic component of the innovation's factor structure as well as dynamic panel, respectively, are allowed for. \cite{KUW15} show that Pesaran's estimator retains its asymptotic properties under weaker conditions,
allowing for either correlated loadings or for the number of latent factor $m$ to be larger than the number of observables, whereas \cite{WU15} discuss some limitations. \cite{S13} extends \cite{B09} to the case of non-constant regression coefficients establishing $\sqrt{T}$-asymptotics when $T/N^2 \rightarrow 0 $. Dynamic panel are permitted.}
Other contributions to this literature include Holtz-Eakin, Newey \& Rosen (1988), Ahn, Hoon Lee \& Schmidt (2001),
\cite{BN04}, \cite{PS03}, \cite{MP04} and \cite{PS07}.
None of these papers address the issue of efficient estimation, except for \cite{BAIGLS}\footnote{\cite{BAIGLS} focus on the {\em homogeneous} parameter case, unlike us, and considers joint estimation of the latent factors and parameters, generalizing \cite{B09} and \cite{MW09}.
More importantly, the motivation of \cite{BAIGLS} differs drastically from ours because they focus on the GLS approach for an efficiency improvement of an estimator that already exhibits the conventional asymptotic properties under \cite{BAIGLS} assumptions, in particular iid-ness across time. In our case, our GLS approach mitigates the {\em first-order} bias affecting the OLS estimator, where we allow for both serial and cross-sectional correlation as well as heteroskedasticity of the residuals.
}, but rather focus on various, ingenious, ways to mitigate the bias induced by the correlation between regressors and innovations.
In contrast, our GLS approach allows to tackle both issues, at the same time, without requiring {\em any } knowledge of the factor structures affecting the regressors and innovations.\footnote{In particular, given that we can afford to be completely agnostic about the need to conduct inference on the latent factor structure affecting the model, our work differ, both in terms of focus and in terms of the techniques developed, from the multitude of papers developing inference methods on latent factor structures
(on estimating the number of latent factors see \cite{BN02}, \cite{HL07}, \cite{AW07}, \cite{O09}, \cite{O10}, \cite{AH13} and on estimating latent factor structures see \cite{FHLR00}, \cite{SW02}, \cite{BN02}, \cite{B09} among others.}
Our asymptotic distribution theory requires $ T^3/N^2 \rightarrow 0 $ whereas the milder $ T/N \rightarrow 0 $ ensures consistency. This relative speed spells out a neat dichotomy in terms of the role of $T$ and $N$: the faster rate of divergence for $N$ is asked for to estimate accurately the (inverse of the) sample-covariance matrix required by the GLS formula,
which in turn mitigates the asymptotic bias.
Instead, the slower divergence of $T$ controls the asymptotic variance of the GLS estimator, dictating ultimately the estimator's rate of convergence. Noticeably, the relative speed requested by our estimator differs from the relative speed requested by the alternative procedures described above,
suggesting that our result can also be viewed as complementary to the others, for example more suitable to short panels where $N$ is much larger than $T$.\footnote{For example, \cite{P06} requires $T/N^2 \rightarrow 0$ for asymptotic normality but the weaker condition $T/N \rightarrow 0 $ is required for homogeneous regression coefficients, where the faster $\sqrt{NT}$ -rate of convergence is achieved.
Moreover, one needs the number of heterogeneous regressors to be greater than number of latent factors. \cite{B09} shows that the regression coefficients' estimator is also $ \sqrt{NT}$-consistent, when $T/N \rightarrow \kappa > 0 $ for some constant $ \kappa $. \cite{B09} estimator is asymptotically biased, in general, but an asymptotically valid bias-correction is established under slightly stronger conditions.
\cite{MW09} establish $ \sqrt{NT} $-asymptotics, again when $N$ and $T$ diverge at the same speed (i.e. $N/T \rightarrow \kappa > 0 $).}
This paper proceeds as follows.
Section~\ref{def} illustrates the general model and
the assumptions required for
estimation of regressions with unit-specific parameters. The asymptotic results for the OLS, UGLS and GLS estimators are presented in Section~\ref{res}.
Section~\ref{common} describes estimation and inference of the coefficients to common regressors.
The technical contributions of the paper are discussed and highlighted in Section~\ref{LEMMI}
Section~\ref{various} discusses various issues related to the GLS estimator. In particular,
we first explore the case when the regressors and the residuals do not depend on the same set of factors.
Second, we discuss the conditions under which the feasible GLS will still work in the context of dynamic panels.
Third, despite the inefficiency of the feasible GLS, we explain how substantial efficiency gains can be achieved by a multi-step version of the GLS estimator. Fourth, we describe how consistent estimation of the asymptotic covariance matrix can be obtained.
Fifth, we describe how to implement our estimator to cross-sectional regressions with time-varying coefficients.
Our theoretical results are corroborated by a set of Monte Carlo
experiments described in Section~\ref{MC}. An empirical application, which investigates whether firms' characteristics are relevant to individual stock returns, is presented in Section~\ref{EMP}.
Section~\ref{CONCL}
concludes. The proofs of our theorems are reported in Appendix \ref{proofmain}, relying on three technical results, enunciated in Appendix \ref{centralemmas}. Appendix \ref{auxiliary} defines some quantities of interest for the construction of the GLS estimator, in particular regarding the (inverse of the) covariance matrix of the residuals. The Supplement contains appendices \ref{someresults}-\ref{app3} with the proofs of additional material, that serve out main results.
Hereafter we use the following notation. Let $\bm{A}(R \times C)$ denotes a generic real $R\times C$ matrix with entries $a_{nm}$; in short $\bm{A}=[a_{nm}]_{n,m=1}^{R,C}$, or simply $\bm{A}=[a_{nm}]$ when the matrix's dimension is clear. Similarly, $\bm{a}$ denotes a generic column vectors of length $R$ with element $a_n$; in short $\bm{a}=[a_n]_{n=1}^R$. The transpose of a $\bm{A}$ is denoted by $\bm{A}'$. If $R=C$, $\lambda_1(\bm{A})$ and $\lambda_N(\bm{A})$ denote the minimum and the maximum eigenvalue of $\bf{A}$, respectively.
With $\bm{A} > (\geq) 0 $ we mean that $\bm{A}$ is positive definite (positive semi positive).
Let $\left\|\bm{A}\right\|_{sp}=\sqrt{\lambda_N\left(\bm{A}'\bm{A}\right)}$
denotes the spectral norm of $\bm{A}$, and $\|\bm{A}\|=\sqrt{{\rm tr }\left(\bm{A}'\bm{A}\right)}$, where ${\rm tr }(\cdot)$ denotes the trace, is the Frobenius norm. When $R=C$ we define the column and row norm of $\bm{A} $ as $\| \bm{A} \|_{col} = \max_{1\leq n \leq C} \sum_{n'=1}^C |a_{nn'}|$ and
$ \| \bm{A} \|_{row} = \max_{1\leq n' \leq C} \sum_{n=1}^C |a_{nn'}|$, respectively.
Furthermore, for $C>R$, we use $\mathscr{P}_{\bm{A}}=\bm{A(A'A)^{+}A'}$, where $\bm{A}^+$ denotes the Moore-Penrose generalized inverse of $\bm{A}$ and $\mathscr{M}_A=\bm{I}_C-\mathscr{P}_A$, where $\bm{I}_C$ is the identity matrix of dimension $C \times C $. If $\bm{A}$ has full column rank, $\bm{A}_{\bot}$ denotes the $C\times (C-R)$ matrix satisfying $\bm{A}_{\bot}'\bm{A}=\bm{0}$, where $\bm{0}$ is a matrix of zeros, and $\bm{A}'_{\bot}\bm{A}_{\bot}=\bm{I}_{C-R}$.
We use $ ``{\xrightarrow{p}}" $, $ ``{\xrightarrow{d}}"$ to denote
convergence in probability and convergence in
distribution, respectively, and $ \mathcal{N}\left(\bm{a},\bm{B}\right)$ denote a random vector normally distributed with mean and covariance matrix equal to $\bm{a},\bm{B} $, respectively.
For $\bm{A}(C\times R),\bm{B}(C\times C ) $ and $\bm{C}(C \times P)$ being three random matrices, then $ \bm{\Sigma}_{A'BC}$ denotes the probability limit (when finite) of $ C^{-1} \bm{ A'BC} $ as $C \rightarrow \infty $. For the $C\times R$ random matrices $\bm{A}_N,\bm{B}_N$ that are functions of $N$ we write $\bm{A}_N \approx\bm{B}_N$ if $\left\|\bm{A}_N-\bm{B}_N\right\|\xrightarrow{p}0$ when $N\to\infty$.
$\operatorname{\mathscr{F}}(\bm{A})$ denotes the sigma-algebra generated by the random matrix $\bm{A}$, $\mathbb{P}(\cdot)$ and $\mathbb{E}(\cdot)$ indicate the probability of an event and the expectation of a random variable, respectively. In the sequel, $\kappa$ denotes a generic, positive constant, which need not to be the same every time we use it.
\section{Model: definitions and assumptions} \label{def}
Assume that the observed variables obey a linear
regression model with $S$ common observed regressors ${\bf d }_t =(d_{t1},\dots,d_{tS})'$ and $K$ heterogeneous regressors ${\bf x }_{it}=(\mathrm{x}_{it1} \hdots \mathrm{x}_{itK})'$.
Following the convenient specification put forward by \cite{P06}, the
model for the $i$th unit can be expressed, in matrix form, as
\begin{equation} \label{eq:hetero}
\bm{\mathrm{y}}_i= \bm{\mathrm{D}} \bm{ \alpha}_{i} + \bm{\mathrm{X}_i} \bm{\beta _{i}} + {\bf u_i},
\end{equation}
for an observed $T \times 1 $ vector $ {\bf y }_i = ( \mathrm{y}_{i1}, \dots, \mathrm{y}_{iT})' $, an observed $T \times S $ matrix ${\bf D } = ( {\bf d }_{1} \hdots {\bf d }_{T})' $ of common
regressors, an observed $T \times K $ matrix $ {\bf
X }_i = ( {\bf x }_{i1} \hdots {\bf x }_{iT})' $ of unit-specific regressors, and
an unobserved $T \times 1 $ vector $ {\bf u_i} = ( \mathrm{u}_{i1} \hdots \mathrm{u}_{iT})' $. In turn, the innovation vector satisfies the factor structure:
\begin{equation}\label{eq:uu}
\bm{\mathrm{u}_i} =
{\bf F } {\bf b }_i + \bm{\varepsilon }_i,\qquad\textrm{with}\quad \bm{\Xi_{i}}:=\mathbb{E}\bm{\varepsilon_{i}
\varepsilon'_{i}},
\end{equation} for an unobserved $ M \times 1 $ vector of factor loadings $ {\bf b }_{i} $, an
unobserved $T \times M $ matrix of common factors $ {\bf F } = (
{\bf f}_{1},...,{\bf f }_{T})' $ and an unobserved $ T \times 1 $
vector of idiosyncratic innovations $ {\bm \varepsilon }_i=(
\varepsilon_{i1} \hdots \varepsilon_{iT})' $. The unit specific regressors satisfy:
\begin{equation} \label{eq:ref_def}
{\bf X }_i = {\bf D } {\bf \Delta }_i +
{\bf F } {\bf \Gamma }_i + {\bf V }_i, \end{equation} for an unobserved $
S \times K $ matrix of factor loadings $ {\bf \Delta }_{i} = ( \boldsymbol{ \delta }_{i1} \hdots \boldsymbol{ \delta }_{iS} )'$ with $\boldsymbol{\delta }_{il}=(\delta_{il1} \hdots \delta_{ilK} )' $,
an unobserved $ M \times K $ matrix of factor loadings $ {\bf \Gamma }_{i} = ( \boldsymbol{ \gamma }_{i1} \hdots \boldsymbol{ \gamma }_{iM} )' $ with $ \boldsymbol{ \gamma }_{il} = ( \gamma_{il1} \hdots \gamma_{ilK} )' $, and
an unobserved $T \times K $ matrix of idiosyncratic innovations $ {\bf V }_i=(
\bm{\mathrm{v}}_{i1} \hdots \bm{\mathrm{v}}_{iT})' $ with $ \bm{\mathrm{v}}_{it} = (v_{it1}, \dots v_{itK} )' $.
The maintained
assumption here is that $K$, $S$ and $M$ {\em do not vary } with $T$ and
$N$. Moreover, we do not need to impose {\em any } relationship between them so that, in particular, $M$ can be either {\em smaller, equal or bigger} than $K$. Although model (\ref{eq:hetero}) is written as a single regression
across time for a given $i$, we assume that in fact a panel of observations $\{{\bf y}_1 \hdots {\bf y}_N, {\bf X
}_1 \hdots {\bf X }_N \}$ is available and fully used within our methodology.
As explained below, throughout our analysis we always {\em de-mean } the data by $\mathscr{M}_{\bm{\mathrm{D}}} $. This allows to avoid making any assumptions on $\bm{ \Delta_i} $.
We now present our assumptions which, thank to the detailed specification of model (\ref{eq:hetero})-(\ref{eq:ref_def}), appear relatively primitive.
\medskip
\begin{assumption}[idiosyncratic innovation $ \varepsilon_{it} $]\label{ass eps} The $N\times 1$ vector $\bm{\varepsilon}_{t}= ( {\varepsilon}_{1t}...{\varepsilon}_{Nt})'$ satisfies the following equation
\begin{equation} \label{eq:11CHL}
\bm{\varepsilon}_{t}= \bm{R } \bm{ a }_{t},\qquad\qquad \textrm{for}\quad t=1,\dots T,
\end{equation}
where the $N\times N$ matrix of constants $\bm{R} = [ r_{ij} ] $ satisfies $ \parallel \!\! \bm{R} \!\! \parallel_{row} +\parallel \!\! \bm{R} \!\! \parallel_{col} \, < \infty $, $\inf_i\sum_{j=1}^N |r_{ij}|>\kappa$ for some $\kappa>0$,
and the elements of the $N \times 1 $ vector $
\bm{ a_{t}} = \left(a_{1t},a_{2t},\dots,a_{Nt}\right)' $ follow a linear process:
\begin{equation} \label{eq:11}
a_{it} =\sum_{s=0}^{\infty} \phi_{is}\eta_{i,t-s}, \quad \sup_i\sum_{s=0}^{\infty}s^2|\phi_{is}|<\infty,\quad \textrm{with}\quad \phi_{i0}=1,
\end{equation}
where the sequence $\{\eta_{it}\}$ is independent and identically distributed across $i$ and $t$ with $\mathbb{E} \eta_{it}=0 $ and $ \mathbb{E} | \eta_{it} |^6< \infty $. Moreover,
for every complex number $z\in\mathbb{C}$,
\begin{equation}\label{eq:11ap}
\inf_i|\phi_i(z)|>\kappa, \quad |z|\leq 1, \qquad \textrm{where}\quad\phi_i(z)=\sum_{s=0}^{\infty} \phi_{is}z^s.
\end{equation}
\end{assumption}
\medskip
\begin{remark}
Assumption \ref{ass eps} is similar to Assumptions 1 and 2 in \cite{PT11} and, with same variations,
this form of cross-sectional and time dependence has been adopted also by \cite{MW09}, \cite{MW13} and \cite{Onatski2015388}.
The above assumption turns out to be extremely convenient for establishing the asymptotic distribution of the feasible and unfeasible GLS estimators along the lines of
Theorem 1 in \cite{RH97}.
\end{remark}
\begin{remark}\label{remcum} Assumption \ref{ass eps} implies that, for every $ 2 \le h,\ell\le 6$:
$$
\sup_{i_1} \sup_{t_1} \sum_{i_2 \cdots i_{\ell}=1}^N \sum_{t_2 \cdots t_h =1}^T | {\rm cum }_h (\varepsilon_{i_1 t_1}, \varepsilon_{i_2 t_2} \cdots , \varepsilon_{i_{\ell} t_h})| < \infty ,
$$
where the summands are the cumulants of order $h$ of $\varepsilon_{i_1 t_1}, \varepsilon_{i_2 t_2} \cdots , \varepsilon_{i_{\ell} t_h}$.
\end{remark}
\medskip
\begin{remark}\label{unif lower bound} By \cite{BD91}, Proposition 4.5.3, (\ref{eq:11ap}) implies that the eigenvalues of the covariance matrices of $\bm{a_j}=(a_{j1},\dots a_{jT})'$ are bounded, and greater than $\kappa$ for every $i$. Easy calculations give
$\bm{\Xi_{i}}= \sum_{j=1}^N r_{ij}^2 \mathbb{E} \bm{a}_j \bm{a}_j'$,
implying that $ \inf_i \lambda_1(\bm{\Xi_{i}})>\kappa $ and $ \sup_i \lambda_T(\bm{\Xi_{i}})<\infty$.
\end{remark}
\medskip
\begin{assumption}[regressor innovation ${\bf{V}}_{\bm{i}} $]\label{ass V}
The sequence $\{{ v}_{itk}\} $ have zero mean, $ \sup_i \sup_k \sup_t \mathbb{E} | \! v_{itk} \! |^{12} < \infty $ and they satisfy, for every $2 \le h,\ell,s\le 14$ and $2\le j\le h$:
$$
\sup_{k_1 \cdots k_s } \sup_{i_1} \sup_{t_1} \sum_{i_2 \cdots i_{\ell} =1}^N \sum_{t_2 \cdots t_h =1}^T (1 + t_j^2) | {\rm cum }_h (v_{i_1 t_1 k_1}, \cdots , v_{i_{\ell} t_h,k_s})| \le\infty .
$$
Moreover, $\inf_i\lambda_1\left(\mathbb{E}\bm{\mathrm{v}}_{it}\bm{\mathrm{v}}'_{it}\right)>\kappa$, where $\bm{\mathrm{v}'_{it}}=(v_{1it},\dots,v_{Kit})$.
\end{assumption}
\begin{remark}\label{inverseV}
Assumption \ref{ass V} implies that $T^{-1}{\bf{V}}_{\bm{i}}'{\bf{V}}_{\bm{i}}\xrightarrow{p}T^{-1}\sum_{t=1}^T\mathbb{E}\bm{\mathrm{v}}_{it}\bm{\mathrm{v}}'_{it}=: \bm{\Sigma}_{{\bf{V}}_{\bm{i}}'{\bf{V}}_{\bm{i}}}$, with $\inf_i \lambda_1\left(\bm{\Sigma}_{{\bf{V}}_{\bm{i}}'{\bf{V}}_{\bm{i}}}\right)>\kappa$. It follows that $\sup_i\left\|\left({\bf{V}}_{\bm{i}}'{\bf{V}}_{\bm{i}}/T\right)^{-1}\right\|=O_p(1)$.
\end{remark}
\begin{remark}
The $ v_{itk} $ can be interpreted as the high-rank components of the regressors $\mathrm{x}_{itk}$, adopting \cite{MW13} terminology, as opposed to the $ \bm{\mathrm{D}} $ which represent the low-rank components. For instance, if for each $k$ the $ v_{itk} $ are generated as $\varepsilon_{it}$ in Assumption \ref{ass eps}, one obtains $\bm{V_k}=\left[v_{kti}\right]_{t,i=1}^{T,N}= O_p( \sqrt{ \max ( N,T) } ) $
for every $k$ (see the discussion in \cite{MW13}, Appendix 1 and \cite{Onatski2015388}). In contrast, $ ( \sum_{i=1}^N \parallel \!\! \bm{\mathrm{D}} \!\! \parallel^2 )^{1 \over 2} =
( N \parallel \!\! \bm{\mathrm{D}} \!\! \parallel^2 )^{1 \over 2} = O_p(\sqrt{(NT} ) $.
\end{remark}
\medskip
\begin{assumption}[latent and observed factors]\label{ass factors}
Set $\bm{\mathrm{Z}} = ( \bm{\mathrm{D}} , \bm{\mathrm{F}} ) = [ \mathrm{z}_{ti} ]$
for $ 1 \le t \le T $ and $ 1 \le j \le M+S < \infty $.
Then,
\begin{equation}\label{eq:maggio1}
\frac{\bm{\mathrm{Z}}'\bm{\mathrm{Z}}}{T}\xrightarrow{p}\bm{\Sigma_{\bm{\mathrm{Z}}'\bm{\mathrm{Z}}}},\quad\textrm{with}\quad \bm{\Sigma_{\bm{\mathrm{Z}}'\bm{\mathrm{Z}}}}:=\left[
\begin{array}{cc}
\bm{\Sigma}_{\bm{\mathrm{D}}'\bm{\mathrm{D}}} & \bm{\Sigma}_{\bm{\mathrm{D}}'\bm{\mathrm{D}}}\\
\bm{\Sigma}_{\bm{\mathrm{F}}'\bm{\mathrm{D}}} & \bm{\Sigma}_{\bm{\mathrm{F}}'\bm{\mathrm{F}}}
\end{array}
\right]>0,
\end{equation}
and $ \bm{\Sigma}_{\bm{\mathrm{D}}' \bm{\mathrm{D}}}>0 $, $ \bm{\Sigma}_{\bm{\mathrm{F}}' \bm{\mathrm{F}}}>0 $. Moreover, we assume
$
\mathbb{E}\left\|\bm{\mathrm{z}_{t}}\right\|^4<\infty.
$
where $\bm{\mathrm{z}_{t}}=(\mathrm{z}_{t,1},\dots,\mathrm{z}_{t,M+S})'$
\end{assumption}
\medskip
\begin{remark}\label{fpd}
Equation (\ref{eq:maggio1}) implies that $\bm{\mathrm{F}}'\mathscr{M}_{\bm{\mathrm{D}}}\bm{\mathrm{F}}>0$ (see \cite{ltk}, Result (4), Section 9.11.2).
\end{remark}
\medskip
\begin{remark}
Although not strictly necessary, we are ruling out trending behaviours in $ \bm{\mathrm{D}} $ and $ \bm{\mathrm{F}} $. However, $\bm{\mathrm{D}}$ and $\bm{\mathrm{F}} $ are allowed to be cross-correlated as well as serially correlated. although not perfectly collinear. For instance, the joint dynamics of $\bm{\mathrm{Z}} $ could be described by a multivariate stationary ARMA.
\end{remark}
\medskip
\begin{assumption}[regressors]\label{ass X} For every $i$, the matrix of unit specific regressors $\bm{\mathrm{X}_i} $ and the matrix of common regressors $\bm{\mathrm{D}} $ have full row rank. Moreover, setting $ \bf{Z}_i := [ \bm{\mathrm{D}} , \bm{\mathrm{X}_i} ] $,
$ N^{-1}\sum_{i=1}^N \mathscr{M}_{\bf{Z}_i} \bm{\mathrm{u}_i} \bm{\mathrm{u}_i}' \mathscr{M}_{\bf{Z}_i}$ has always rank $T-S$ for sufficiently large $N$ and $T$.
\end{assumption}
\begin{remark}\label{remark ass X}
Assumption \ref{ass X} requires enough cross section heterogeneity of the $\bm{\mathrm{X}_i}$'s across individuals. Simple manipulations show that
$$
\bm{\mathrm{D}}_{\bot}'\left(\frac{1}{N}\sum_{i=1}^N \mathscr{M}_{\bf{Z}_i} \bm{\mathrm{u}_i} \bm{\mathrm{u}_i}' \mathscr{M}_{\bf{Z}_i}\right)\bm{\mathrm{D}}_{\bot}=\frac{1}{N}\sum_{i=1}^N \bm{M_{(\bm{\mathrm{D}}_{\bot}'\bm{\mathrm{X}_i})}} \bm{\mathrm{D}}_{\bot}'\bm{\mathrm{u}_i} \bm{\mathrm{u}_i}'\bm{\mathrm{D}}_{\bot} \bm{M_{(\bm{\mathrm{D}}_{\bot}'\bm{\mathrm{X}_i})}}>0 ,
$$
implying that the empirical covariance matrix $\bm{\hat{\mathcal{S}}_N}$ defined in (\ref{eq:charlie2}) is invertible.
\end{remark}
\medskip
\begin{assumption}[loadings $\bm{\mathrm{b}_i}$ and $ {\bf \Gamma }_i $]\label{ass loading}
$\bm{\Gamma_{i}}$ and $\bm{\mathrm{b}_i}$ are non-random such that $\left\| \bm{\Gamma_{i}} \right\|<\infty$ and $\left\|\bm{\mathrm{b}_i}\right\|<\infty$ and, for $N> M$,
\begin{equation}
\bm{ \mathrm {B}_N}:={1 \over N } \sum_{i=1}^N \bm{\mathrm{b}_i} \bm{\mathrm{b}_i}' >0 . \label{eq:seconmomb}
\end{equation}
and
\begin{eqnarray}
\bm{A_N}&:=&\frac{1}{N}\sum_{i=1}^N\left(\bm{I}_M-\bm{\Gamma_{i}}\bm{\Psi_i}^{-1}\bm{\Gamma_{i}}'\frac{\bm{\mathrm{F}}\mathscr{M}_{\bm{\mathrm{D}}}\bm{\mathrm{F}}}{T}\right)
\bm{\mathrm{b}_i}\bm{\mathrm{b}_i}'
\left(\bm{I}_M-\frac{\bm{\mathrm{F}}'\mathscr{M}_{\bm{\mathrm{D}}}\bm{\mathrm{F}}}{T}\bm{\Gamma_{i}}\bm{\Psi_i}^{-1}\bm{\Gamma_{i}}'\right), \label{eq:seconmomb2}
\end{eqnarray}
is positive definite with
\begin{equation}\label{eq:trenta}
\bm{\Psi}_i:=\bm{\Gamma_{i}}'\frac{\bm{\mathrm{F}}'\mathscr{M}_{\bm{\mathrm{D}}}\bm{\mathrm{F}}}{T}\bm{\Gamma_{i}}+\bm{\Sigma}_{{\bf{V}}_{\bm{i}}'{\bf{V}}_{\bm{i}}}.
\end{equation}
\end{assumption}
\begin{remark}
Condition (\ref{eq:seconmomb}) implies that the factor structure (\ref{eq:uu}) is strong, as defined in \cite{PT11}. This is commonly assumed in the literature.
The technical condition (\ref{eq:seconmomb2}) is used in the proof of Theorem \ref{Theorem_GLS}. As shown in Section \ref{lemmata1} in the Supp. Material, Lemma \ref{lemmata1}\ref{zerone}, the matrices in brackets are of full rank. Hence, (\ref{eq:seconmomb2}) will be satisfied when there is enough cross-sectional heterogeneity in the sample. Finally, our results will not change if random loadings are assumed (and cross-sectionally independent from other parameters).
\end{remark}
\begin{assumption}[independence]\label{independence} The $f_{mt}, v_{ksi},\varepsilon_{uj} $ are mutually independent for every $i,j$ and $t,s,u$ and $m,k$.
\end{assumption}
\begin{remark}
We are not allowing for any correlation between any entries of $ \bm{\varepsilon_j}$ and $ \bm{\mathrm{X}_i} $. This rules out the possibility that $ \bm{\mathrm{X}_i} $ contains a weakly exogenous component, and in this respect we are similar to \cite{P06} and \cite{B09}. The implications from generalizing this assumption, in particular when considering dynamic panels where one element of $\bm{\mathrm{X}_i}$ represents the lagged dependent variable, are discussed in Section~\ref{dynamic}.
\end{remark}
\begin{remark}\label{mreg}
Assumptions \ref{ass V}, \ref{ass factors} and \ref{independence} and Remark \ref{fpd} imply that $T^{-1}\bm{\mathrm{X}_i}'\bm{\mathrm{X}_i}\xrightarrow{p}\bm{\Sigma}_{\bm{\mathrm{X}_i}'\bm{\mathrm{X}_i}}>0$ and $T^{-1}\bm{\mathrm{X}_i}'\mathscr{M}_{\bm{\mathrm{D}}}\bm{\mathrm{X}_i}\xrightarrow{p}\bm{\Sigma}_{\bm{\mathrm{X}_i}'\mathscr{M}_{\bm{\mathrm{D}}}\bm{\mathrm{X}_i}}>0$, for every $i$. Hence $\left\|(T^{-1}\bm{\mathrm{X}_i}'\bm{\mathrm{X}_i})^{-1}\right\|=O_p(1)$ and $\left\|(T^{-1}\bm{\mathrm{X}_i}'\mathscr{M}_{\bm{\mathrm{D}}}\bm{\mathrm{X}_i})^{-1}\right\|=O_p(1)$ for $T$ large enough.
\end{remark}
\section{Estimators: definitions and asymptotics} \label{res}
Our main objective is to estimate the heterogeneous slope coefficients $ \bm{ \beta }_{i}$ of (\ref{eq:hetero}). However, estimation of the coefficients $ {\bm \alpha }_{i} $ of the common regressors is also discussed in Section \ref{common}. Hence, without loss of generality, we premultiply both sides of
(\ref{eq:hetero}) by the projection matrix $\mathscr{M}_{\bm{\mathrm{D}}}$, obtaining
\begin{eqnarray}
\mathscr{M}_{\bm{\mathrm{D}}} \bm{\mathrm{y}}_i= \mathscr{M}_{\bm{\mathrm{D}}} \bm{\mathrm{X}_i}\bm{\beta_i}+\mathscr{M}_{\bm{\mathrm{D}}} \bm{\mathrm{u}_i} . \label{eq:mm11}
\end{eqnarray}
We consider three different estimators for the parameters $ {\bm \beta }_{i}$, namely the OLS, the unfeasible and feasible GLS estimators.
Regarding the OLS estimator for $ \bm{ \beta _{i} }$:
\begin{equation}\label{eq:charlie1}
\bm{\hat{ \beta }_i^{OLS}} := ( \bm{\mathrm{X}_i}' \mathscr{M}_{\bm{\mathrm{D}}} \bm{\mathrm{X}_i} )^{-1} \bm{\mathrm{X}_i} \mathscr{M}_{\bm{\mathrm{D}}} \bm{\mathrm{y}}_i.
\end{equation}
We now consider GLS estimation.
Define the cross-sectional averages of the individual covariance matrices of the $\mathscr{M}_{\bm{\mathrm{D}}} \bm{\mathrm{u}_i}$, conditional on sigma algebra generated by $ \bm{Z}$, defined in Assumption \ref{ass factors}:
\begin{equation}\label{eq:deffeb}
\mathscr{M}_{\bm{\mathrm{D}}}\bm{\mathrm{S}}_N\mathscr{M}_{\bm{\mathrm{D}}},\; \mbox{ setting }\;\bm{\mathrm{S}}_N:= \bm{\mathrm{F}}\bm{ \mathrm {B}_N} \bm{\mathrm{F}}'+ \bm{\Xi_{N}},\quad\mbox{with}\quad
\bm{\Xi_{N}}:=\frac{1}{N}\sum_{i=1}^N\bm{\Xi_{i}}.
\end{equation}
We assume without loss of generality that $\bf{d_t}$ includes an element equal to one, i.e. we allow for an intercept term, leading to $ \mathbb{E} \bm{\mathrm{u}_i} =0 $.
The presence of $\mathscr{M}_{\bm{\mathrm{D}}} $ could cause some complications in the definition of the GLS estimator since the $\mathscr{M}_{\bm{\mathrm{D}}} \bm{\mathrm{u}_i} $ have a singular covariance matrix. We show how to solve this issue and obtain a model with a non-singular residual covariance matrix that can be used to construct the GLS estimator.
Proceeding along the lines of \cite{MN88}, Section 11 in Chapter 13, one gets the UGLS estimator when the residual covariance matrix to model (\ref{eq:mm11}) is singular:
\begin{equation}\label{eq:11giugno}
\bm{\hat{\beta}}_i^{UGLS}:=\left(\bm{\mathrm{X}_i}'\mathscr{M}_{\bm{\mathrm{D}}} \left(\mathscr{M}_{\bm{\mathrm{D}}} \bm{\mathrm{S}}_N \mathscr{M}_{\bm{\mathrm{D}}}\right)^+ \mathscr{M}_{\bm{\mathrm{D}}} \bm{\mathrm{X}_i}\right)^{-1}\bm{\mathrm{X}_i}'
\mathscr{M}_{\bm{\mathrm{D}}} (\mathscr{M}_{\bm{\mathrm{D}}}\bm{\mathrm{S}}_N\mathscr{M}_{\bm{\mathrm{D}}})^+ \mathscr{M}_{\bm{\mathrm{D}}} \bm{\mathrm{y}}_i.
\end{equation}
By Lemma \ref{fact6410} in the Supp. Material
$$ (\mathscr{M}_{\bm{\mathrm{D}}}\bm{\mathrm{S}}^{-1}_N \mathscr{M}_{\bm{\mathrm{D}}})^+ =\bm{\mathrm{D}}_{\bot}\left(\bm{\mathrm{D}}_{\bot}'\bm{\mathrm{S}}_N \bm{\mathrm{D}}_{\bot}\right)^{-1}\bm{\mathrm{D}}_{\bot}' , $$ where $\bm{\mathrm{D}}_{\bot}$ is the $T\times (T\!-\!S)$ full rank matrix such that
$ \mathscr{M}_{\bm{\mathrm{D}}} = \bm{\mathrm{D}}_{\bot} \bm{\mathrm{D}}_{\bot}' $ where $\bm{\mathrm{D}}_{\bot}'\bm{\mathrm{D}}_{\bot} =\bm{I}_{T-S} $. Assumption \ref{ass X} and display (\ref{eq:deffeb}) imply that the inverse in (\ref{eq:11giugno}) is well defined for any $T$. By substitution, setting for simplicity
\begin{equation} \label{eq:storti}
\bm{\mathcal{y}_i}=\bm{\mathrm{D}}_{\bot}' \bm{\mathrm{y}}_i,\quad \bm{\mathcal{X}_i}= \bm{\mathrm{D}}_{\bot}'\bm{\mathrm{X}_i},\quad
\bm{\epsilon_i}=\bm{\mathrm{D}}_{\bot}' \bm{\varepsilon_i},\quad \bm{\mathcal{F}}=\bm{\mathrm{D}}_{\bot}' \bm{\mathrm{F}},\quad \bm{\mathcal{u}_i}= \bm{\mathrm{D}}_{\bot}' \bm{\mathrm{u}_i},
\end{equation}
one obtains
\begin{eqnarray}
\bm{\hat{\beta}}_i^{UGLS} &&
= \left(\bm{\mathrm{X}_i}'\bm{\mathrm{D}}_{\bot}
\left(\bm{\mathrm{D}}_{\bot}'\bm{\mathrm{S}}^{-1}_N\bm{\mathrm{D}}_{\bot}\right)^{-1}
\bm{\mathrm{D}}_{\bot}'\bm{\mathrm{X}_i}'\right)^{-1}
\bm{\mathrm{X}_i}'\bm{\mathrm{D}}_{\bot}\left(\bm{\mathrm{D}}_{\bot}'\bm{\mathrm{S}}^{-1}_N \bm{\mathrm{D}}_{\bot}\right)^{-1}\bm{\mathrm{D}}_{\bot}\bm{\mathrm{y}}_i \nonumber\\
&& = \left(\bm{\mathcal{X}_i}' \bm{\mathcal{S}}^{-1}_N \bm{\mathcal{X}_i} \right)^{-1}
\bm{\mathcal{X}_i}' \bm{\mathcal{S}}^{-1}_N \bm{\mathcal{y}_i},\label{eq:UGLSdef}
\end{eqnarray}
where we set
$
\bm{\mathcal{S}}_N= \bm{\mathrm{D}}_{\bot}'\bm{\mathrm{S}}_N\bm{\mathrm{D}}_{\bot}.
$
This means that the UGLS has now the more conventional expression of the generalized least squares for the model
\begin{eqnarray}\label{eq:modstorto}
\bm{\mathcal{y}_i}=\bm{\mathcal{X}_i}\bm{\beta}_i+\bm{\mathcal{u}_i},\qquad\textrm{with}\qquad \bm{\mathcal{u}_i}=\bm{\mathcal{F}}\bm{\mathrm{b}_i}+\bm{\mathcal{u}_i},
\end{eqnarray}
without involving Moore-Penrose matrices.
Pre-multiplying the data by $ \bm{\mathrm{D}}_{\bot}' $ reduces the sample size by $S$ units since now the $ \bm{\mathcal{y}_i} $ and the $ \bm{\mathcal{X}_i} $ have $T-S$ rows. Likewise, considering again model (\ref{eq:modstorto}), an equivalent representation of (\ref{eq:charlie1}) is $ \bm{\hat{\beta }_i^{OLS}} = (\bm{\mathcal{X}_i} ' \bm{\mathcal{X}_i} )^{-1} \bm{\mathcal{X}_i} ' \bm{\mathcal{y}_i} $.
Along the same lines, our proposed {\em feasible } GLS estimator is given by
\begin{equation}\label{eq:bgls}
\bm{\hat{ \beta }_i^{GLS}} := \left(\bm{\mathcal{X}_i}' \bm{\hat{\mathcal{S}}_N}^{-1} \bm{\mathcal{X}_i} \right)^{-1}
\bm{\mathcal{X}_i}' \bm{\hat{\mathcal{S}}_N}^{-1} \bm{\mathcal{y}_i},
\end{equation}
where
\begin{equation}\label{eq:charlie2}
\bm{\hat{\mathcal{S}}_N}:=
N^{-1} \sum_{i=1}^N \bm{\hat{\mathcal{u}}_i}\bm{\hat{\mathcal{u}}_i}' , \,\,\, \mbox{ with } \bm{\hat{\mathcal{u}}_i}:=
\bm{\mathcal{y}_i}-\bm{\mathcal{X}_i} \bm{\hat{\beta}}_i^{OLS} = \mathscr{M}_{\bm{\mathcal{X}_i}} \bm{\mathcal{u}_i},
\end{equation}
for $N$ and $T$ large enough, by Assumption~\ref{ass X} and Remark~\ref{remark ass X}, $ \bm{\hat{\mathcal{S}}_N} $
has full rank.
The following two theorems enunciates the asymptotic distribution of the OLS, UGLS and GLS estimators, respectively. The proofs are given in Appendixes \ref{thuno} and \ref{proofT2}, respectively. Further details are provided in the Supp. Material.
\begin{theorem} \label{Theorem_OLS}
When Assumptions \ref{ass eps}, \ref{ass V}, \ref{ass factors}, \ref{ass X}, \ref{ass loading} and \ref{independence} hold, for any $N$ and as $T\to\infty$
\noindent
(i) (OLS estimator)
$$
T^{1 \over 2 } ( \bm{\hat{ \beta }}_i^{OLS} - \bm{ \beta }_{i} - {\bf
\tau }_i^{ OLS } ) \xrightarrow{d}
\mathcal{N}\left(\bm{0},\bm{\Sigma_i}\right),\quad
$$
where
\begin{equation}\label{eq:bias17}
{\bf \tau }_i^{OLS} := \bm{\Sigma_{\bm{\mathcal{X}_i} ' \bm{\mathcal{X}_i}}}^{-1} \bm{\Sigma_{\bm{\mathcal{X}_i}'\bm{\mathcal{F}}}} \bm{\mathrm{b}_i} ,
\end{equation}
is the bias term, and the asymptotic covariance matrix equals
\begin{equation}\label{eq:asycm17}
\bm{\Sigma_i}
:= \bm{\Sigma_{\bm{\mathcal{X}_i} ' \bm{\mathcal{X}_i}}}^{-1} \bm{\Sigma_{ \bm{\mathcal{X}_i} ' \bm{\mathrm{D}}_{\bot}'\bm{\Xi_{i}}\bm{\mathrm{D}}_{\bot} \bm{\mathcal{X}_i} }} \bm{\Sigma_{\bm{\mathcal{X}_i} ' \bm{\mathcal{X}_i}}}^{-1} ,
\end{equation}
setting
\begin{eqnarray}
\bm{ \Sigma }_{\bm{\mathcal{X}_i}' \bm{\mathcal{F}} } &:=& \bm{\Gamma_{i}}' \Big( \bm{\Sigma }_{\bm{\mathrm{F}}'\bm{\mathrm{F}}} - \bm{\Sigma }_{\bm{\mathrm{F}}'\bm{\mathrm{D}}} \bm{\Sigma }_{\bm{\mathrm{D}}'\bm{\mathrm{D}}}^{-1} \bm{\Sigma }_{\bm{\mathrm{D}}'\bm{\mathrm{F}}} \Big) , \label{eq:sigmaXF} \\
\bm{ \Sigma }_{\bm{\mathcal{X}_i}' \bm{\mathcal{X}_i} } &:=& {\bf \Gamma}_i' \Big( \bm{\Sigma }_{\bm{\mathrm{F}}'\bm{\mathrm{F}}} - \bm{\Sigma }_{\bm{\mathrm{F}}'\bm{\mathrm{D}}} \bm{\Sigma }_{\bm{\mathrm{D}}'\bm{\mathrm{D}}}^{-1} \bm{\Sigma }_{\bm{\mathrm{D}}'\bm{\mathrm{F}}} \Big) {\bf \Gamma}_i + {\bf \Sigma }_{ {\bf{V}}_{\bm{i}}'{\bf{V}}_{\bm{i}}} , \label{eq:sigmaXX} \\
\bm{ \Sigma }_{\bm{\mathcal{X}_i}' \bm{\mathrm{D}}_{\bot}'\bm{\Xi_{i}}\bm{\mathrm{D}}_{\bot} \bm{\mathcal{X}_i} } &:=& \bm{\Gamma_{i}} ' ( - \bm{ \Sigma }_{\bm{\mathrm{F}}'\bm{\mathrm{D}}} \bm{ \Sigma }_{\bm{\mathrm{D}}'\bm{\mathrm{D}}}^{-1} , \bm{ I }_m ) \bm{ \Sigma }_{\bm{\mathrm{Z}}' \bm{\Xi_{i}} \bm{\mathrm{Z}}} ( - \bm{ \Sigma }_{\bm{\mathrm{F}}'\bm{\mathrm{D}}} \bm{ \Sigma }_{\bm{\mathrm{D}}'\bm{\mathrm{D}}}^{-1} , \bm{ I }_m )' \bm{\Gamma_{i}} \label{eq:sigmaXHX}\\
&&+ \bm{ \Sigma }_{{\bf{V}}_{\bm{i}}' \bm{\Xi_{i}} {\bf{V}}_{\bm{i}}} . \nonumber
\end{eqnarray}
\medskip
\noindent (ii) (UGLS estimator)
\begin{equation}\label{eq:theorem1part2}
T^{1 \over 2 } \left( \bm{\hat{ \beta }_i^{UGLS}} - \bm{ \beta }_{i} \right) \xrightarrow{d}
\mathcal{N}\left(\bm{0},\bm{\Sigma}^{\star}_N\right),
\end{equation}
with $\bm{\Sigma}^{\star}_N:=\bm{\Sigma}_{{\bf{V}}_{\bm{i}}'\bm{\Xi^{-1}_{N}}{\bf{V}}_{\bm{i}}}^{-1}\bm{\Sigma}_{{\bf{V}}_{\bm{i}}'\bm{\Xi^{-1}_{N}}\bm{\Xi_{i}}\bm{\Xi^{-1}_{N}}{\bf{V}}_{\bm{i}}}\bm{\Sigma}_{{\bf{V}}_{\bm{i}}'\bm{\Xi^{-1}_{N}}{\bf{V}}_{\bm{i}}}^{-1}$.
\end{theorem}
\medskip
\begin{remark}
The OLS estimator is affected by a {\em first-order bias}. It will be asymptotically unbiased if either $ \bm{\mathrm{b}_i} = \bm{0} $ or $ \bm{\Gamma_{i}} = \bm{0} $ or, alternatively, for diagonal $\bm{ \Sigma }_{\bm{\mathcal{F}}' \bm{\mathcal{F}} } $ as well as with $ \bm{\Gamma_{i}} $ and $ \bm{\mathrm{b}_i} $ satisfying $ \gamma_{il} b_{il} = 0 $ for every $l$ and $i$. Essentially, this means that the entries of $ \bm{\Gamma_{i}} $ are non zero whenever the corresponding entries of $ \bm{\mathrm{b}_i} $ are zero, for the same row $l$, and viceversa. More in general,
no bias arises if ${\bf b}_i $ belongs to the null space of $ \bm{ \Sigma }_{\bm{\mathcal{F}}' \bm{\mathcal{F}} } \bm{\Gamma_{i}} $, assuming $M>K$.
\end{remark}
\begin{remark}
One can assume without loss of generality that the same latent factors $\bm{\mathrm{F}} $ enter into $ \bm{X}_i $ and $\bm{\mathrm{u}_i} $. In fact, assume $ \bm{\mathrm{u}_i} = \bm{\mathrm{G}} \bm{b}_i + \bm{\varepsilon }_i $ with the rows of $ \bm{\mathrm{G}} $ correlated, but not identical to the rows of $ \bm{\mathrm{F}} $. Then the bias takes the form
\[
{\bf \tau }_i^{OLS} = \bm{\Sigma_{\bm{\mathcal{X}_i} ' \bm{\mathcal{X}_i}}}^{-1} \bm{\Sigma }_{\bm{\mathrm{F}}' \mathscr{P}_F \bm{\mathrm{G}}} \bm{\mathrm{b}_i} ,
\]
exploiting the decomposition $ \bm{\mathrm{G}} = \mathscr{P}_F\bm{\mathrm{G}}+ \mathscr{M}_F\bm{\mathrm{G}}$. Hence the bias will only be non-zero due to the portion of $\bm{\mathrm{G}}$ correlated with $\bm{\mathrm{F}}$. The same consideration applies to the GLS estimator. In Section~\ref{diff-fac} we explore more in details the implications of having different, yet correlated, factor structures for regressors and innovations.
\end{remark}
\bigskip
\begin{remark}
The UGLS estimator is asymptotically unbiased, consistent and asymptotically normal as $T \rightarrow \infty $. Moreover, the UGLS estimator can be efficient in the GLS sense. In particular, when
the $ \bm{\varepsilon }_i $ are not (unconditionally) heteroskedastic, namely $ \bm{\Xi_{i}} = {\bm \Xi } $, then the UGLS asymptotic covariance matrix does not have the sandwich form, unlike for the OLS estimator. One can define the UGLS differently, for instance replacing $\bm{\mathrm{S}}_N $ with $\bm{\mathrm{F}}\bm{ \mathrm {B}_N} \bm{\mathrm{F}}'+ \bm{\Xi_{i}} $ in (\ref{eq:11giugno}). However, our definition of the UGLS estimator
makes it closer to the population counterpart to the class of feasible GLS estimators here studied.
\end{remark}
We now present the main result of the paper.
\begin{theorem} \label{Theorem_GLS} When Assumptions \ref{ass eps}, \ref{ass V}, \ref{ass factors}, \ref{ass X}, \ref{ass loading} and \ref{independence} hold,
as $ 1 / T + T / N \rightarrow 0 $,
\begin{equation}\label{eq:theorem2A}
\bm{\hat{ \beta }_i^{GLS}} \xrightarrow{p}\bm{\beta _{i}},\qquad
\end{equation}
and, as $(1/T)+(T^3/N^2)\to 0$, then
\begin{equation}\label{eq:theorem2B}
\left({\bf{V}}_{\bm{i}}'\bm{C^{-1}_N}\bm{\Xi_{i}}\bm{C^{-1}_N}{\bf{V}}_{\bm{i}}\right)^{-\frac{1}{2}}
\left({\bf{V}}_{\bm{i}}'\bm{C^{-1}_N}{\bf{V}}_{\bm{i}}\right)
( \bm{\hat{ \beta }_i^{GLS}} - \bm{ \beta _{i}} ) \xrightarrow{d}
\mathcal{N}\left(\bm{0},\bm{I}_{\bm{K}}\right) ,
\end{equation}
where
\begin{equation}\label{eq:MMM2}
\bm{\mathrm{C}_N}:= \frac{1}{N}\sum_{i=1}^N \left(\bm{\Xi_{i}}+\bm{\Theta_i}\right) \mbox{with}\;
\bm{\Theta_i}:=\mathbb{E}\left[{\bf{V}}_{\bm{i}}\bm{\Sigma_{\bm{\mathcal{X}_i} ' \bm{\mathcal{X}_i}}}^{-1} \bm{\Gamma_{i}}\bm{\Sigma_{\bm{\mathcal{F}}'\bm{\mathcal{F}}}}\bm{\mathrm{b}_i}\bm{\mathrm{b}_i}'\bm{\Sigma_{\bm{\mathcal{F}}'\bm{\mathcal{F}}}}\bm{\Gamma_{i}}'\bm{\Sigma_{\bm{\mathcal{X}_i} ' \bm{\mathcal{X}_i}}}^{-1} {\bf{V}}_{\bm{i}}'\right],
\end{equation}
with $\bm{\Sigma_{\bm{\mathcal{X}_i} ' \bm{\mathcal{X}_i} }}$ defined in (\ref{eq:sigmaXX}).
\end{theorem}
\begin{remark}
The GLS estimator is asymptotically unbiased, consistent and asymptotically normal as both $N,T \rightarrow \infty $ such that $ T^3/N^2 \rightarrow 0 $. The feasible GLS estimator is not efficient in general. A multi-step generalization achieves substantial efficiency gains, see Section~\ref{various}.
\end{remark}
\section{Common regressors} \label{common}
We now consider estimation of the coefficients $ \bm{ \alpha }_{i} $ to the common regressors $ \bm{\mathrm{D}} $ in model (\ref{eq:hetero}). A natural generalization of the GLS estimator would be
\[
\Big( \begin{array}{c} \tilde{\bm \alpha }_i^{GLS} \\ \tilde{\bm \beta }_i^{GLS} \end{array} \Big) := \left(\bf{Z}_i' \tilde{\bm{S}}^{+}_N\bf{Z}_i \right)^{+}
\bf{Z}_i ' \tilde{\bm{S}}^{+}_N \bm{\mathrm{y}}_i,
\]
where $ \bf{Z}_i$ has been defined in Assumption \ref{ass X} and
\begin{equation}\label{eq:gidio}
\tilde{\bm{S}}_N = N^{-1} \sum_{i=1}^N \bm{\hat{\mathrm{u}}_i}\bm{\hat{\mathrm{u}}_i}',\qquad
\bm{\hat{\mathrm{u}}_i}:=
\bm{\mathrm{y}}_i - \bm{\mathrm{D}} \hat{\bm{ \alpha }}_i^{OLS} - \bm{\mathrm{X}_i} \hat{\bm{ \beta }}_i^{OLS} ,
\end{equation}
that is $\bm{\hat{\mathrm{u}}_i}$ are the OLS residuals, for $ ( \hat{\bm{ \alpha }}_i^{OLS'} , \hat{\bm{ \beta }}_i^{OLS'} )':= (\bf{Z}_i' \bf{Z}_i )^{-1} \bf{Z}_i' \bm{\mathrm{y}}_i$.
However, we show in Theorem \ref{TI1} (Supp. Material, Appendix~\ref{app2}) that $ \tilde{\bm{ \alpha }}_i^{GLS} = {\bm 0}_M $ and $ \tilde{\bm{ \beta }}_i^{GLS} = \hat{\bm{ \beta }}_i^{GLS} $ due to a cancellation that occurs as a consequence of $ \bm{\mathrm{D}} $ being common across units.
If the joint distribution for the estimators of $ ( {\bm{ \alpha }}_{i}' , {\bm{ \beta }}_{i}' )' $ is not required, one can estimate the $ \bm{ \alpha }_{i} $
as the projection of
$ \bm{\mathrm{y}}_i- \bm{\mathrm{X}_i} \hat{\bm{ \beta }}_i^{GLS} $ on $ \bm{\mathrm{D}} $ yielding
\begin{equation}\label{eq:gidio2}
\widetilde{\bm \alpha }_i := (\bm{\mathrm{D}}' \bm{\mathrm{D}} )^{-1} \bm{\mathrm{D}}' ( \bm{\mathrm{y}}_i - \bm{\mathrm{X}_i} \hat{\bm{ \beta }}_i^{GLS} ) .
\end{equation}
Using our theory, its asymptotic distribution follows (see Theorem \ref{TI2} ins the Spp. Material for further details).
Note that the additional assumption $ \bm{\mathrm{F}} ' \bm{\mathrm{D}} = {\bm 0 } $ is required. For example, if we are interested in a model with an intercept term, heterogenous across units,
such as $ \bm{\mathrm{D}} {\bm \alpha }_{i} = \bm{ \iota }_T \alpha_{i1} + \bm{\mathrm{D}}_2 \bm{ \alpha }_{i2} $, with $ \bm{ \alpha }_{i} = (\alpha_{i1} ,\bm{ \alpha }_{i2} ')' , \bm{\mathrm{D}} = ( \bm{ \iota }_T , \bm{\mathrm{D}}_2 ) $,
then one of the restrictions $\bm{\mathrm{F}}' \bm{\mathrm{D}}= {\bm 0 } $ is simply $ \sum_{t=1}^T \bm{f}_t = \bm{0} $. If, moreover, a grand-mean is also allowed for, such as $ \bm{\mathrm{D}} \bm{ \alpha }_{i} = \bm{ \iota }_T \alpha_{3} + \bm{ \iota}_T \alpha_{i1} + \bm{\mathrm{D}}_2 \bm{ \alpha }_{i2} $, then the additional
restriction $ \sum_{i=1}^N \alpha_{i1} = 0 $ is needed.
Similar identification conditions are discussed in \cite{B09} and \cite{MW09}.\footnote{Most of the papers on estimation of panel regressions with so-called interactive fixed effects, such as ours, focus exclusively on the coefficients to the heterogeneous time-varying regressors. Among the few exceptions, is \cite{B09} who shows that, without further identification
assumption, estimation of the coefficient to common regressors is possible only for constant parameters. In contrast, for the case of non-constant coefficients further identification assumptions similar to ours are needed. \cite{MW09} study the same estimator of \cite{B09} under weaker conditions on the regressors, allowing for instance for pre-determinatedness. Our identification condition for the coefficients to common regressors implies their weaker corresponding assumption. They focus exclusively on the case of constant regression coefficients.}
If instead the joint distribution for estimators of $ \bm{\alpha }_{i} $ and $ \bm{ \beta }_{i} $ is required, this can be achieved by a slight modification of our GLS estimator, namely
\begin{equation}
\left( \begin{array}{c} \breve{\bm \alpha }_i^{GLS } \\ \breve{\bm \beta }_i^{GLS} \end{array} \right) := \left(\bf{Z}_i' \breve{\bm{S}}_N^{-1} \bf{Z}_i \right)^{-1}
\bf{Z}_i' \breve{\bm{S}}_N^{-1} \bm{\mathrm{y}}_i, \label{dagger}
\end{equation}
for the non-singular matrix
\begin{equation}\label{eq:gidio3}
\breve{\bm{S}}_N := \tilde{\bm{S}}_N + \left( { {\rm tr }( \tilde{\bm{S}}_N ) \over N }\right) \mathscr{P}_{\bm{\mathrm{D}}} .
\end{equation}
Non-singularity of $\breve{\bm{S}}_N $ follows by augmenting the matrix $\breve{\bm{S}}_N $, of rank $T-S$, with the projection matrix $ \mathscr{P}_{\bm{\mathrm{D}}} $ of rank $S$. Scaling by
$N^{-1}{\rm tr }\left( \tilde{\bm{S}}_N\right) $ in not required by the asymptotic theory but could be relevant in finite-samples to ensure the same order of magnitude of the two terms in $\breve{\bm{S}}_N $. It turns out that the same identification condition $ \bm{\mathrm{F}} ' \bm{\mathrm{D}} = {\bm 0 } $, discussed above, is required. Monte Carlo experiments are reported in Section~\ref{MC} to assess the small-sample properties of these estimators.
\section{Technical contributions} \label{LEMMI}
The asymptotics for the GLS estimator requires four key auxiliary results, enunciated in Appendix \ref{centralemmas}, which could be useful in a broader set of statistical problems.
The main reason for this complexity is that, unlike most of the existing theoretical results on GLS estimation, we are not restricting the number of free elements of the weighting matrix to be finite. Indeed, in our case the number of free elements of the weighting matrix is $O(T^2)$ and hence rapidly increasing with $T$. To tackle the curse-of-dimensionality issue, we exploit the approximate factor structure of the weighted matrix, that we write as $\bm{E=FAF'+C}$, for (possibly random) $ M_1 \times M_1 $ matrix $ \bm{A}>0 $ and $ T \times T $ matrix $ \bm{C} > 0 $ for every finite $T$. The inverse has a convenient form thanks to the Sherman-Morrison formula (see Appendix \ref{smw}).
Lemma \ref{PZ} establishes the asymptotic orthogonality between the inverse of the matrix $\bm{E}$, and the factor $\bm{\mathrm{F}}$. More precisely, when $ \bm{A} $ and $ \bm{C} $ satisfy a set of mild regularity conditions
\begin{equation}
\parallel \bm{E}^{-1} \bm{F} \parallel^2 = O_p (T^{-1}), \label{asyorth}
\end{equation}
This is a remarkably fast rate given that
$\bm{E}^{-1} \bm{\mathrm{F}} $ is $T \times M_1 $ dimensional, with $M_1$ fixed, hence with its number of rows increasing with $T$. It implies that for a large class of $ T \times M_2 $ matrices $\bm{P}$ (satisfying the mild regularity conditions of Lemma \ref{PZ}) possibly unrelated to both $\bm{E}$ and $\bm{F} $, then $ \bm{P}' \bm{E}^{-1} \bm{F} = O_p(1) $ and, when the entries of $\bm{P}$ have zero mean and are stochastically independent of $\bm{E} $ and $\bm{F} $, then $ \bm{P}' \bm{E}^{-1} \bm{F} = O_p(T^{-\nicefrac{1}{2}}) $. These rates are very different from the usual case, arising when $ \bm{E} $ and $ \bm{F} $ are unrelated. For example, when $ \bm{A}=\bm{0} $, under the same assumptions on $\bm{P}$, one gets that $ \bm{P}' \bm{E}^{-1} \bm{F} $ is of order $ O_p(T^{\nicefrac{1}{2}} ) $ or $ O_p(T) $, depending on whether $ \bm{P}$ has zero or non-zero mean, respectively, assuming that $ \bm{F} $, $\bm{A} $ and $\bm{P}$ are mutually independent.
The asymptotic orthogonality (\ref{asyorth}) plays a crucial role in establishing the asymptotics for the GLS (and UGLS) estimator. To better understand this, consider the following decomposition of the GLS estimator:
\begin{eqnarray}
\hat{\bm \beta }_i^{GLS} - {\bm \beta }_{i} && = \left(\bm{\mathcal{X}_i}' \bm{\hat{\mathcal{S}}_N}^{-1} \bm{\mathcal{X}_i} \right)^{-1}\nonumber
\bm{\mathcal{X}_i}' \bm{\hat{\mathcal{S}}_N}^{-1} \bm{\mathcal{u}_i} \\ && = \left(\bm{\mathcal{X}_i}' \bm{\hat{\mathcal{S}}_N}^{-1} \bm{\mathcal{X}_i} \right)^{-1}
\bm{\mathcal{X}_i}' \bm{\hat{\mathcal{S}}_N}^{-1} \bm{\mathcal{F}} \bm{\mathrm{b}_i} + \left(\bm{\mathcal{X}_i}' \bm{\hat{\mathcal{S}}_N}^{-1} \bm{\mathcal{X}_i} \right)^{-1}\label{eq:franzferd}
\bm{\mathcal{X}_i}' \bm{\hat{\mathcal{S}}_N}^{-1} \bm{\epsilon_i}.
\end{eqnarray}
Since $ \bm{\hat{\mathcal{S}}_N}
= \bm{\mathcal{F}} \bm{\hat{A}_N} \bm{\mathcal{F}}' + \bm{\hat{\bm{\mathcal{C}}}_N} $, for some random matrices $\bm{\hat{A}_N}, \bm{\hat{ \bm{\mathcal{C}}}_N} $ (specified in Appendix~\ref{auxiliary}) function of both $T$ and $N$, we show that Lemma \ref{PZ} applies to the bias term, namely the first term in (\ref{eq:franzferd}). In particular, by (\ref{asyorth}) term $ \bm{\mathcal{X}_i}' \bm{\hat{\mathcal{S}}_N}^{-1} \bm{\mathcal{F}} $
is of a smaller order of magnitude (and vanishes asymptotically) than $\bm{\mathcal{X}_i}' \bm{\hat{\mathcal{S}}_N}^{-1} \bm{\epsilon_i} $ as long as $N$ is diverging faster than $T$.
Note that both the dimension of $ \bm{\hat{ \bm{\mathcal{C}}}}_N $, as well as its elements, are changing with $T$. The faster rate for $N$ is demanded for by the need to have $\bm{\hat{A}_N} $ and $ \bm{\hat{\bm{\mathcal{C}}}}_N $ with the desired limiting properties.
To save notation we rename $\bm{\hat{\beta}_i^{GLS}}$ in (\ref{eq:bgls}) as $\bm{\hat{\beta}_i(\bm{\hat{\mathcal{S}}_N})}$ and define $\bm{\hat{\beta}_i(\bm{\mathcal{H}_N})}$ the estimator obtained replacing the weighting matrix $\bm{\hat{\mathcal{S}}_N}$ with $\bm{\mathcal{H}_N}$ in (\ref{eq:bgls}).
The matrix $ \bm{\mathcal{H}_N} =\bm{\mathcal{F}} \bm{A_N} \bm{\mathcal{F}}' + \bm{\mathcal{C}_N}$, defined in Appendix~\ref{proofT2} (display (\ref{eq:MMM1})),
is non-stochastic, if $\bm{\mathrm{Z}}$ is fixed, and, most importantly, satisfies the assumption of Lemma \ref{PZ}. Existing conditions for the asymptotic equivalence of $\bm{\hat{\beta}_i(\bm{\hat{\mathcal{S}}_N})}$ and $\bm{\hat{\beta}_i(\bm{\mathcal{H}_N})}$ cannot be applied here. Although the inverse operator of a matrix is an analytic function, one cannot rely on the delta method due to the curse of dimensionality, namely the fact that the elements as well as the size of $\bm{\hat{\mathcal{S}}_N} $ are varying with $T$ and $N$. For similar reasons, element-wise convergence of $\bm{\hat{\mathcal{S}}_N}$ cannot be combined with the Slutsky's Theorem, as discussed in \cite{mf94}. Considering the absolute convergence $\left\|\bm{\hat{\mathcal{S}}_N}^{-1}-\bm{\mathcal{H}^{-1}_N}\right\|_{sp}$ and use random matrix theory is not a viable option. Even if we were able to obtain the optimal convergence rate $\left\|\bm{\hat{\mathcal{S}}_N}-\bm{\mathcal{H}_N}\right\|_{sp}=O_p(\sqrt{T/N})$ established by for i.i.d. data, it would be not sufficient to obtain $\sqrt{T}\left(\bm{\hat{\beta}_i(\bm{\hat{\mathcal{S}}_N})}-\bm{\hat{\beta}_i(\bm{\mathcal{H}_N})}\right)=o_p(1)$ without strengthening the Assumptions of Theorem \ref{Theorem_GLS}. The latter convergence requires indeed proving that $T^{-\nicefrac{1}{2}}\bm{\mathcal{X}_i}'\left(\bm{\hat{\mathcal{S}}_N}^{-1}-\bm{\mathcal{H}^{-1}_N}\right)\bm{\mathcal{u}_i}=o_p(1)$. We would have
\begin{equation}\label{eq:casa}
\left\| T^{-\nicefrac{1}{2}}\bm{\mathcal{X}_i}'\left(\bm{\hat{\mathcal{S}}_N}^{-1}-\bm{\mathcal{H}^{-1}_N}\right)\bm{\mathcal{u}_i}\right\| \leq \left\| T^{-\nicefrac{1}{2}}\bm{\mathcal{X}_i}\right\|
\left\| \bm{\hat{\mathcal{S}}_N}^{-1}-\bm{\mathcal{H}^{-1}_N} \right\|_{sp} \left\| \bm{\mathcal{u}_i} \right\| =O_p\left(\frac{T}{\sqrt{N}}\right).
\end{equation}
To prove our results we found convenient proceeding in two steps: firstly we prove that
$T^{-\nicefrac{1}{2}}\bm{\mathcal{X}_i}'\bm{\hat{\mathcal{S}}_N}^{-1}\bm{\mathcal{u}_i}$
is asymptotically equivalent
$T^{-\nicefrac{1}{2}}\bm{\mathcal{X}_i}'\bm{\Omega_N}^{-1}\bm{\mathcal{u}_i}$, with $\bm{\Omega_N}$ as defined in Appendix~\ref{auxiliary}, equation (\ref{eq:silver}) ,and secondly that the latter term is in turn asymptotically equivalent to
$T^{-\nicefrac{1}{2}}\bm{\mathcal{X}_i}'\bm{\mathcal{H}^{-1}_N}\bm{\mathcal{u}_i}$.
Lemma \ref{approx inv}, a simple extension of a well known result in matrix algebra, entails that
\begin{eqnarray}\label{eq:formula}
\bm{\hat{\mathcal{S}}_N}^{-1}
= \bm{\Omega_N}^{-1} - \bm{\Omega_N}^{-1} ( \bm{\hat{\mathcal{S}}_N} - \bm{\Omega_N}) \bm{\Omega_N}^{-1}+
\bm{\Omega_N}^{-1} ( \bm{\hat{\mathcal{S}}_N} - \bm{\Omega_N} ) \bm{\hat{\mathcal{S}}_N}^{-1} ( \bm{\hat{\mathcal{S}}_N} - \bm{\Omega_N} ) \bm{\Omega_N}^{-1},
\end{eqnarray}
implying that $
\bm{\mathcal{X}_i}' \bm{\hat{\mathcal{S}}_N}^{-1} \bm{\mathcal{u}_i} $ can be re-written as the (algebric) sum of $ \bm{\mathcal{X}_i}' \bm{\Omega_N}^{-1} \bm{\mathcal{u}_i} $, $ \bm{\mathcal{X}_i}'
\bm{\Omega_N}^{-1} ( \bm{\hat{\mathcal{S}}_N} - \bm{\Omega_N} ) \bm{\Omega_N}^{-1} \bm{\mathcal{u}_i} $ and
$ \bm{\mathcal{X}_i}'
\bm{\Omega_N}^{-1} ( \bm{\hat{\mathcal{S}}_N} - \bm{\Omega_N} ) \bm{\Omega_N}^{-1} ( \bm{\hat{\mathcal{S}}_N} - \bm{\Omega_N} )
\bm{\Omega_N}^{-1} \bm{\mathcal{u}_i} $. In the proof we show that the second last term is of order $O_p(T/N^{\nicefrac{1}{2}}) $ and that the last term is of order $O_p(T^2/N) $, respectively. Instead, the first leading term will exhibit the usual $O_p(T^{\nicefrac{1}{2}}) $ rate of convergence. For the second and third terms to be asymptotically negligible, in terms of asymptotic distribution, one requires that $T^{3 \over 2}/N $ goes to zero as $T$ increases. Unfortunately our approach requires lengthy calculations involving high order cumulants that are bounded using the diagram formula (see Appendix \ref{auxiliary1} in the Supp. Material).
A further step necessary to derive the convergence in distribution of the estimator is the derivation of the asymptotic distribution of $T^{-\nicefrac{1}{2}}\bm{\mathcal{X}_i}'\bm{\mathcal{H}^{-1}_N}\bm{\mathcal{u}_i}$.
We first show that the latter term is equivalent to $T^{-\nicefrac{1}{2}}{\bf{V}}_{\bm{i}}'\bm{\mathrm{C}_N}^{-1}\bm{\varepsilon_i}$. In order to exploit the results in \cite{RH97}, the absolute row/column summability of $\bm{\mathrm{C}_N}^{-1}$ needs to be shown. To accomplish this task we first approximate $\bm{\mathrm{C}_N}$ with a circulant symmetric matrix, as further discussed in the Supp. Material, Appendix \ref{invcov}. By Lemma \ref{paolo}, that extends a result in \cite{Z93}, we prove that the inverse of the latter matrix has indeed bounded row norm.
\section{Discussion and generalizations} \label{various}
In this section we describe various generalizations to our framework. In particular, we explain the consequences of allowing for different, yet related, factor structures in the regressors and
residuals, respectively. Then, we show to derive a consistent estimator for the asymptotic covariance matrix of the GLS estimator. We also discuss how to achieve efficiency improvements
of the GLS by an iterative procedure. Finally, we explain how our results apply to cross-sectional regressions with time-varying coefficients.
\subsection{Different factor structures} \label{diff-fac}
So far we have assumed that the unit-specific regressors $\bm{\mathrm{X}_i}$ and the true residuals $\bm{\mathrm{u}_i} $ of (\ref{eq:hetero}) share the same common, latent, factors. We now explore the implications of allowing that possibly different, yet correlated, set of factors affect the regressors and the residuals, respectively. Let us here illustrate the UGLS case and then provide more details in the Supp. Material (Appendix~\ref{app3}) for the (feasible) GLS.
To simplify the exposition we assume that $\bm{\mathrm{D}}=\bm{0}$, and
\begin{equation}\label{eq:diffact}
\bm{\mathrm{y}}_i=\bm{\mathrm{X}_i}\bm{\beta}_{i}+\bm{\mathrm{u}_i}, \quad \bm{\mathrm{X}_i}=\bm{\mathrm{F}}_1\bm{\Gamma_{i}}+{\bf{V}}_{\bm{i}}, \quad \bm{\mathrm{u}_i}=\bm{\mathrm{F}}_2\bm{\mathrm{b}_i}+\bm{\varepsilon_i} ,
\end{equation}
where $\bm{\mathrm{F}}_1(T\times M_1)$ and $\bm{\mathrm{F}}_2(T\times M_2)$ satisfy Assumption \ref{ass factors}. By Remark \ref{remarkPZ},
\begin{equation}\label{eq:tikka}
T^{-\nicefrac{1}{2}}\bm{\mathrm{F}}_1'\bm{\mathrm{S}}^{-1}_N\bm{\mathrm{u}_i} \approx
T^{-\nicefrac{1}{2}}\bm{\mathrm{F}}_1'\left[\bm{I}_T-\bm{\Xi^{-1}_{N}}\bm{\mathrm{F}}_2\left(\bm{\mathrm{F}}_2'\bm{\Xi^{-1}_{N}}\bm{\mathrm{F}}_2\right)^{-1}\bm{\mathrm{F}}_2'\right]\bm{\Xi^{-1}_{N}}\bm{\mathrm{u}_i} .
\end{equation}
Let $\bm{W}:=\bm{\Xi^{-1}_{N}}\bm{\mathrm{F}}$, then
\begin{eqnarray*}
&&\bm{\mathrm{F}}_1'\left[\bm{I}_T-\bm{\Xi^{-1}_{N}}\bm{\mathrm{F}}_2\left(\bm{\mathrm{F}}_2'\bm{\Xi^{-1}_{N}}\bm{\mathrm{F}}_2\right)^{-1}\bm{\mathrm{F}}_2'\right]\bm{\Xi^{-1}_{N}}\bm{\mathrm{u}_i}=
\bm{\mathrm{F}}_1'\left[\bm{I}_T-\bm{W}\left(\bm{\mathrm{F}}_2'\bm{W}\right)^{-1}\bm{\mathrm{F}}_2'\right]\bm{\Xi^{-1}_{N}}\bm{\mathrm{u}_i}\\
&=& \bm{\mathrm{F}}_1'\bm{\mathrm{F}}_{2\bot}\left(\bm{W}'_{\bot}\bm{\mathrm{F}}_{2\bot}\right)^{-1}\bm{W}'_{\bot}\bm{\Xi^{-1}_{N}}\bm{\mathrm{u}_i}
=\bm{\mathrm{F}}_1'\bm{\mathrm{F}}_{2\bot}\left(\bm{W}'_{\bot}\bm{\mathrm{F}}_{2\bot}\right)^{-1}\bm{W}'_{\bot}\bm{\Xi^{-1}_{N}}\bm{\varepsilon_i},
\end{eqnarray*}
where $\bm{\mathrm{F}}'_{2\bot}\bm{\mathrm{F}}_2=\bm{0}$.
It follows that even if $\bm{\mathrm{F}}_1\notin\textrm{sp}(\bm{\mathrm{F}}_2)$, where $\textrm{sp}(\bm{\mathrm{F}}_2)$ denote the space spanned by $\bm{\mathrm{F}}_2$, the UGLS estimator is still consistent but the term $\breve{\bm{\mathrm{F}}}_1'\bm{\Xi_i}^{-1}\bm{\varepsilon_i}$, with $\breve{\bm{\mathrm{F}}}_1:=\bm{W}_{\bot}\left(\bm{\mathrm{F}}_{2\bot}'\bm{W}_{\bot}\right)^{-1}\bm{\mathrm{F}}_2'\bm{\mathrm{F}}_1\in\textrm{sp}(\bm{W}_{\bot})$, will contribute to the asymptotic distribution of the estimator.\footnote{We conjecture that one needs to assume on the $\breve{\mathrm{g}}_t$ the same conditions assumed by Robinson and Hidalgo (1997) about their $ x_t$ (see their Condition 7) .}
In Appendix~\ref{app3} (Supp. Material) we show, heuristically, that the FGLS estimator of (\ref{eq:diffact}) enjoys the same asymptotic properties stated in Theorem \ref{Theorem_GLS}.
\subsection{Dynamic models} \label{dynamic}
Although our set-up allows for dynamics, through the dynamic autocorrelation of either the factors and the idiosyncratic error,
our results extend to the case of dynamic panel with factor structure such as
\begin{equation}\label{eq:dynamic1}
\bm{\mathrm{y}}_i= \bm{\mathrm{X}_i}\bm{\beta}_{i}+ \bm{y}_{-1,i} {\rho}_{i}+ \bm{\mathrm{u}_i},
\end{equation}
where $ \bm{\mathrm{X}_i} $ and $ \bm{\mathrm{u}_i} $ satisfy (\ref{eq:ref_def}) and (\ref{eq:uu}), respectively, and we set $ \bm{y}_{-1,i} = ( \mathrm{y}_{i0}, \dots, \mathrm{y}_{iT-1})' $, with first-order autoregressive coefficients satisfying $ -1 < \rho_i < 1 $ for every $i$. We set $ \bm{D} = \bm{0} $ to simplify the exposition.
Obviously one can re-write (\ref{eq:dynamic1}) as
\begin{equation} \label{wrongdyn}
\bm{\mathrm{y}}_i=\bm{\mathrm{X}_i}^* \bm{\beta}_{i}^* + \bm{\mathrm{u}_i} \mbox{ setting } \bm{\mathrm{X}_i}^* := ( \bm{\mathrm{X}_i} , \bm{y}_{-1,i} ) \mbox{ and } \bm{\beta}_{i}^* := ( \bm{\beta}_{i} ' , {\rho}_{i} )'.
\end{equation}
It turns out that applying our GLS estimator (\ref{eq:bgls}) to specification (\ref{wrongdyn}) will still work when
further conditions are assumed on the idiosyncratic part of the residuals $ \bm{\mathrm{u}_i} = {\bf F } {\bf b }_i + \bm{\varepsilon }_i $, namely that
the $ \varepsilon_{it} $ are i.i.d. across time but have some degree of cross-correlation across $i$. A similar assumption is made by \citet{CP13}, Assumption 1, also in the context of dynamic panel data models. Notice that our result is rather strong because we are {\em not } ruling out that the model residuals $u_{it}$ are dependent across time (and across $i$), through the factors ${\bf f }_t $. Moreover, again thanks to the factors ${\bf f }_t $, regressors and residuals are correlated and thus we are violating the classical strong-exogeneity assumption typically advocated in a GLS framework.
A technical proof goes beyond the scope, and the page limit, of the present paper but details,
corroborated by Monte Carlo simulations, are available upon request.
\subsection{Efficiency improvements}
The form of the asymptotic covariance matrix of the GLS, indicated in Theorem~\ref{Theorem_GLS}, denotes lack of efficiency, unlike for the UGLS estimator case (in the special sense discussed). This arises because although $ \bm{\hat{\mathcal{S}}_N}$ is {\em approximated } by the matrix $ \bm{\mathcal{H}_N}$ defined in Appendix~\ref{proofT2}, (in the sense that $\bm{\mathcal{X}_i}'\bm{\hat{\mathcal{S}}_N}\bm{\mathcal{u}_i} = \bm{\mathcal{X}_i}'\bm{\mathcal{H}^{-1}_N}\bm{\mathcal{u}_i} + o_p(\sqrt{T})$), the latter does not coincide with the true covariance matrix $\bm{\mathcal{S}}_N $. This is to be expected since $\bm{\mathcal{H}_N}$ is constructed based on the OLS residuals $ \bm{\hat{\mathcal{u}}_i} = \bm{\mathcal{y}_i} - \bm{\mathcal{X}_i} \bm{\hat{ \beta }}_i^{OLS} $ where $\bm{\hat{ \beta }}_i^{OLS} $ is non consistent for $ \bm{\beta }_{i} $. However, a multi-step procedure can be envisaged that could achieve (near) asymptotic efficiency, or more precisely an estimator with an asymptotic distribution arbitrarily close to the UGLS estimator. We shall call the outcome of this procedure the iterated-GLS estimator. The first step would be to construct the GLS estimator as explained in the previous sections, which we now denominate as $\hat{\bm \beta }_{i}^{(1)} $. We then construct the associated residuals $ \bm{\hat{\mathcal{u}}_i}^{(1)} = \bm{\mathcal{y}_i} - \bm{\mathcal{X}_i} \bm{\hat{ \beta }}_i^{(1)} $. Notice that now $\bm{\hat{ \beta }}_i^{(1)} $ is a consistent estimator for $\bm{\beta }_{i} $. The second step entails constructing
$\bm{\hat{\mathcal{S}}_N}^{(1)} = N^{-1} \sum_{i=1}^N \bm{\mathrm{D}}_{\bot}' \bm{\hat{\mathcal{u}}_i}^{(1)} \bm{\hat{\mathcal{u}}_i}^{(1)'} \bm{\mathrm{D}}_{\bot} $ and using it to obtain
$
\bm{\hat{ \beta }}_i^{(2)} =
\left(
\bm{\mathcal{X}_i}' \left(\bm{\hat{\mathcal{S}}_N}^{(1)}\right)^{-1} \bm{\mathcal{X}_i} \right)^{-1} \bm{\mathcal{X}_i}'
\left(\bm{\hat{\mathcal{S}}_N}^{(1)} \right)^{-1} \bm{\mathcal{y}_i}.
$
In general the $h$th step entails constructing $ \bm{\hat{ \beta }}_i^{(h)} = \left(\bm{\mathcal{X}_i}' \left( \bm{\hat{\mathcal{S}}_N}^{(h-1)} \right)^{-1} \bm{\mathcal{X}_i} \right)^{-1} \bm{\mathcal{X}_i}' \left(\bm{\hat{\mathcal{S}}_N}^{(h-1)} \right)^{-1} \bm{\mathcal{y}_i}$,
where $
\bm{\hat{\mathcal{S}}_N}^{(h-1)} $ is obtained based on $ \bm{\hat{ \beta }}_i^{(h-1)} $. We conjecture that as $h$ increases, the asymptotic distribution of $\bm{\hat{ \beta }}_i^{(h)} $
is getting arbitrarily close to the one of the UGLS. This is confirmed by the Monte Carlo experiments presented in Section~\ref{MC}. Although the theoretical analysis of this iterated-GLS is not developed here, techniques along the lines of the ones developed in the current paper would allow to establish the asymptotics. Indeed, since $\bm{\hat{ \beta }}_i^{(h)} $ is consistent for $ \bm{ \beta }_{i} $ for any $h \ge 2 $, the asymptotics should follow more easily than for the GLS estimator.
\subsection{ Estimation of asymptotic covariance matrix} \label{SE}
Consistent estimation of the GLS asymptotic covariance matrix can be obtained in different ways, depending on the type of heteroskedasticity and correlation assumed for the $ \bm{\varepsilon_i} $. For instance, using the results of
\cite{NW87}, Theorem 2, one obtains the covariance matrix estimator for $\hat{\bm \beta }_i^{GLS} $
\begin{equation}
\left( { \bm{\mathcal{X}_i}' \bm{\hat{\mathcal{S}}_N}^{-1} \bm{\mathcal{X}_i} \over T } \right)^{-1}
\left( \hat{\bm A }_{i0} + \sum_{h=1}^n {\scriptstyle (1 - {h \over (n+1) })} \left( \hat{\bm A }_{ih} + \hat{\bm A }_{ih}' \right) \right)
\left( { \bm{\mathcal{X}_i}' \bm{\hat{\mathcal{S}}_N}^{-1} \bm{\mathcal{X}_i} \over T } \right)^{-1}, \label{con_est}
\end{equation}
setting $
\hat{\bm A }_{ih}:= T^{-1} \sum_{t=h+1}^T \hat{u}_{it}^{GLS} \hat{u}_{it-h}^{GLS} \hat{\mathcal{x} }_{it} \hat{\bf \mathcal{x} }_{it-h}', \,\,\, h=0,1,...T-1$,
where $ \hat{ \bm{\mathcal{X}_i} }:= ( \hat{\mathcal{x} }_{i1} ... \hat{\mathcal{x} }_{iT} )' = \bm{\hat{\mathcal{S}}_N}^{-1} \bm{\mathcal{X}_i} $ and $ \hat{ \bm{\mathcal{u}_i} }^{GLS}:= ( \hat{u}_{1i}^{GLS} ... \hat{u }_{iT}^{GLS} )' = \bm{\mathcal{y}_i} - \bm{\mathcal{X}_i} \hat{\bm \beta }_i^{GLS} $ and the bandwidth $n=n(T,N)$ grows slowly with $N$ and $T$.
The same approach has been used in \cite{P06}, eq. (51) and (52). \cite{B09}, Section 7, provides estimators of the asymptotic covariance matrix when correlation and heteroskedasticity of either series or cross-section form is allowed for, using \cite{NW87} and a partial-samplig approach, respectively. Note that these approaches cannot be applied to our case since require constant regression coefficients, involving averaging across both $N$ and $T$. Similar approaches have been used by \cite{MW09} and \cite{MW13} under more restrictive dependence assumptions.
Notice that, although $\hat{\bm{\mathcal{u}_i}}^{GLS} $ contains a factor structure, as it is evident by from its population counterpart $ \bm{\mathcal{u}_i} = \bm{\mathcal{F}} \bm{\mathrm{b}_i} + \bm{\epsilon_i} $, the contribution of
$ \bm{\hat{\mathcal{S}}_N}^{-1} \bm{\mathcal{F}} \bm{\mathrm{b}_i} $ to $ \bm{\hat{\mathcal{S}}_N}^{-1} \hat{\bm{\mathcal{u}_i} }^{GLS} $ is (asymptotically) negligible with respect to the contribution of $\bm{\hat{\mathcal{S}}_N}^{-1} \bm{\epsilon_i} $, whose asymptotic variance is require for consistent
estimation of the asymptotic covariance matrix of $\bm{\hat{ \beta }}_i^{GLS} $.
An alternative approach consists of estimating the idiosyncratic component of $\hat{\bm{\mathcal{u}_i}}^{GLS} $ directly, for example by principal components, yielding $ \hat{\bm{\epsilon_i}}^{GLS}:=( \hat{\epsilon }_{i1}^{GLS} \cdots \hat{\epsilon }_{iT}^{GLS} )' $ and then replacing $ \hat{\bm A }_{ih} $ by $
T^{-1} \sum_{t=h+1}^T \hat{\epsilon }_{it}^{GLS} \hat{\epsilon }_{it-h}^{GLS} \hat{\bf x }_{it} \hat{\bf x }_{it-h}' $ into (\ref{con_est}). Preliminary testing for the number of factors $M$ is required in this case, making it less appealing.
Consistent estimation of the asymptotic covariance matrix for
$ ( \breve{\bm \alpha }_i^{GLS \prime} , \breve{\bm \beta }_i^{GLS \prime } )' $
of (\ref{dagger}) follows along the same lines, leading to:
\[
\left({ \bf{Z}_i' \breve{\bm{S}}_N^{-1} \bf{Z}_i \over T } \right)^{-1}
\left( \breve{\bm A }_{i0} + \sum_{h=1}^n {\scriptstyle (1 - {h \over (n+1) })} \left( \breve{\bm A }_{ih} + \breve{\bm A }_{ih}' \right) \right)
\left({ \bf{Z}_i' \breve{\bm{S}}_N^{-1} \bf{Z}_i \over T } \right)^{-1},
\]
where $
\breve{\bm A }_{ih} = {1 \over T } \sum_{t=h+1}^T \breve{u}_{it}^{GLS} \breve{u}_{it-h}^{GLS} \breve{\bf z }_{it} \breve{\bf z }_{it-h}', \,\,\, h=0,1,...T-1$,
where $ \breve{ \bm{\mathrm{Z}} }_i = ( \breve{\bf z }_{i1} ... \breve{\bf z }_{iT} )' = \breve{\bm{S}}_N^{-1} \bm{\mathrm{Z}}_i $ and $ \breve{ \bm{\mathcal{u}_i} }^{GLS} = ( \breve{u}_{1i}^{GLS} ... \breve{u }_{iT}^{GLS} )' =
\bm{\mathrm{y}}_i - \bm{\mathrm{D}} \breve{\bm \alpha }_i^{GLS } - \bm{\mathrm{X}_i} \breve{\bm \beta }_i^{GLS } $.
\subsection{ Cross-sectional regressions}
As an example of a cross-sectional regression with factor structure, consider \cite{A05} model:
\begin{equation} y_{it} = {\bm \vartheta }_{t} ' (1 \, {\bf x }_{it}'
)' + u_{it}, \label{fac0} \end{equation} where $(y_{it}, {\bf x }_{it} )$ are
assumed $i.i.d.$ across units conditional on $ {\bf c }_{1t}, \,
{\bf C }_{2t}$ by \cite{A05}, Assumption 1, with \begin{eqnarray} u_{it}
&=& {\bf c }_{1t}' {\bf u }_i^* + \varepsilon_{it}, \label{fac1} \\
{\bf x }_{it} &=& {\bf C }_{2t} {\bf x }_i^* + {\bf v }_{it} ,
\label{fac2}
\end{eqnarray} with ${\bf c
}_{1t}, \, {\bf u }_i^* $ are $ d_1 \times 1 $ random vectors and $
{\bf C }_{2t}, \, {\bf x }_i^* $ respectively a random matrix of
dimension $k \times d_2 $, with $ d_2 \ge k $, and a random vector
of dimension $ d_2 \times 1 $ and $ \varepsilon_{it} $ and $ {\bf v
}_{it} $ are $i.i.d.$ innovations across $i$ and $t$, respectively
scalar and $k \times 1 $, with zero mean and variances $ \xi_{i,t} $
and $ {\bf \Sigma }_{V_t'V_t} $, respectively. We focus here on \cite{A05}'s {\em
standard factor} structure, spelled out in his Assumption SF1, here
slightly extended to allow for an
idiosyncratic component in both the regression error $u_{it}$ and
the regressors ${\bf x }_{it}$ as well as time-variation in
parameters, common factors and covariance matrices. The first
extension is unavoidable for us since when $ \varepsilon_{it} = 0 \,\,\,
a.s.$ our theory does not apply. Model (\ref{fac0})-(\ref{fac1})-(\ref{fac2}) can be rewritten as
\begin{equation} \label{time}
{\bf y}_t = {\bf D } {\bf \alpha }_{t} + {\bf X }_t {\bf \beta }_{t} + \bm{\mathrm{u}_t},
\end{equation}
where we set $ {\bf y }_t = ( y_{1t} \hdots y_{Nt})' $ and $ {\bf D } = {\bm \iota }_N , {\bf X }_t = {\bf X }^* {\bf C }_{2t}' $,
with $ {\bf X }^* = ( {\bf x }_1^* \hdots {\bf x }_N^* )' $ and parameters $ {\bm \vartheta }_{t} = ( {\bm \alpha }_{t}' , {\bm \beta }_{t}' )' $, and the $ {\bm u}_t = ( u_{1t} \hdots u_{Nt})' $ satisfy the factor structure (\ref{eq:uu})
\[
{\bm u}_t =
{\bf B } {\bf f }_t + {\bm \varepsilon }_t,
\]
with ${\bf f}_t = {\bf c }_{1t} $, $ {\bf B } = ( {\bf u }_1^* \hdots {\bf u }_N^* )' $ and
$ \bm{\varepsilon }_t=( \varepsilon_{1t} \hdots \varepsilon_{Nt} )' $.
In analogy with Section~\ref{def},
our proposed {\em feasible } GLS estimator is
\[
\hat{\bm \beta }_t^{GLS} = \left(\bm{\mathcal{X}_t}' \hat{\bm{\mathcal{S}}}_T^{-1} \bm{\mathcal{X}_t} \right)^{-1}
\bm{\mathcal{X}_t}' \hat{\bm{\mathcal{S}}}_T^{-1}\bm{\mathcal{y}_t},
\]
setting $ \bm{\mathcal{y}_t} = \bm{D}_{\bot}' \bm{Y}_t, \bm{\mathcal{X}_t} = \bm{D}_{\bot}' \bm{X}_t $,
assuming large enough $N$ and $T$ to ensure invertibility of $ \hat{\bm{\mathcal{S}_T}} $, given by
\begin{equation}\label{eq:charlie2T}
\bm{\hat{\mathcal{S}}_T} =
T^{-1} \sum_{t=1}^T \bm{\hat{\mathcal{u}}_t} \bm{\hat{\mathcal{u}}_t}' , \,\,\, \mbox{ with } \bm{\hat{\mathcal{u}}_t}=
\bm{\mathcal{y}_t}-\bm{\mathcal{X}_t} \bm{\hat{\beta}}_t^{OLS} = \mathscr{M}_{\bm{\mathcal{X}_t}} \bm{\hat{\mathrm{u}}_t} ,
\end{equation}
where $ \bm{\hat{\beta}}_t^{OLS} = \left(\bm{\mathcal{X}_t}' \bm{\mathcal{X}_t} \right)^{-1}
\bm{\mathcal{X}_t}' \bm{\mathcal{y}_t} $ and ${\bm{\hat{\mathrm{u}}_t}}$ are the OLS estimator and the OLS regression residuals, respectively, of regression (\ref{time}).
Notice that now $\mathscr{M}_{\bm{\mathrm{D}}} = \bm{I}_N-\bm{\mathrm{D}}(\bm{\mathrm{D}}'\bm{\mathrm{D}})^{-1}\bm{\mathrm{D}}' = \bm{D}_{\bot} \bm{D}_{\bot}' $ is a $N \times N $ matrix and $ \bm{\mathrm{D}}_{\bot} $ is a $N-1 \times N $ matrix.
Given the duality between $\hat{\bm \beta }_t^{GLS} $ and $ \hat{\bm \beta }_i^{GLS} $, we conjecture that under a set of regularity conditions analogous to Assumptions 2.1-2.6 one obtains consistency of $\hat{\bm \beta }_t^{GLS} $ for $ 1/N + N/T \rightarrow 0 $ and asymptotic normality of $ \sqrt{N}( \hat{\bm \beta }_t^{GLS} - {\bm \beta }_{t} ) $ for $ 1/N + N^3/T^2 \rightarrow 0 $. Extension to a more general form of common observed regressors, other than $ {\bf D } = {\bm \iota }_N $, can be obtained along the lines of Section~\ref{common}.
\section{Monte Carlo analysis} \label{MC}
We conduct a set of Monte Carlo experiments to appreciate the relevance of our asymptotic results for the GLS estimator in finite samples.
\subsection{Design}
The data generating process is
\begin{eqnarray} \label{unit}
y_{it} && = \alpha_{i0} + \beta_{i0} x_{it} + b_{i10} f_{1t} + b_{i20} f_{2t} +\varepsilon_{it},
\end{eqnarray}
where the single regressor satisfies
\begin{eqnarray} \label{reg}
x_{it} = 0.5 + \delta_{i10} f_{1t} + \delta_{i30} f_{3t} + v_{it}.
\end{eqnarray}
Note that the model implies an observed common factor equal to $1$ for all observations. The single regressor is allowed to be contemporaneously correlated with the innovation through one of the latent common factors (whenever $ b_{i10} \delta_{i10} \neq 0 $). The factor loadings are normally distributed random variables, $i.i.d.$ across unit:
\begin{eqnarray}
&& \left( \begin{array}{c} b_{i10} \\ b_{i20} \end{array} \right) \sim NID \left( \left( \begin{array}{c} 1 \\ 0 \end{array} \right), \left( \begin{array}{cc} 0.2 & 0 \\ 0 & 0.2 \end{array} \right) \right), \\
&& \left( \begin{array}{c} \delta_{i10} \\ \delta_{i30} \end{array} \right) \sim NID \left( \left( \begin{array}{c} 0.5 \\ 0 \end{array} \right), \left( \begin{array}{cc} 0.5 & 0 \\ 0 & 0.5 \end{array} \right) \right),
\label{rank}
\end{eqnarray}
and the latent common factors and the idiosyncratic components are stationary stochastic processes, mutually independent to each other, satisfying
\[
{ f }_{j,t} = 0.5 { f }_{j,t-1} + \sqrt{0.5} { \eta }_{jf,t}, \, j=1,2,3,
\]
where each $ { \eta }_{jf,t} \sim NID(0,1) $, mutually independent for $j=1,2,3$, and
\begin{eqnarray*}
&& \varepsilon_{it} = \rho_{i\varepsilon } { \varepsilon }_{it-1} + { \eta }_{i\varepsilon,t}, \,\,\,
{ \eta }_{i\varepsilon,t} \sim NID( 0, \sigma_i^2 ( 1 - \rho_{i\varepsilon }^2) ), i=1,...,N, \\
&& v_{it} = \rho_{iv }
{ v }_{it-1} + { \eta }_{iv,t}, \,\,\,
{ \eta }_{iv ,t} \sim NID( 0, ( 1 - \rho_{i v }^2) ), i=1,...,N,
\end{eqnarray*}
with $ \rho_{i \varepsilon } \sim UID(0.05,0.95), \, \rho_{i v } \sim UID(0.05,0.95), \sigma_{i \varepsilon }^2 \sim UID(0.5,1.5) $ where $NID, UID$ means $iid$ normally and uniformly distributed respectively.
Finally, the parameters of interest are constant across replications and equal to $ \alpha_{i0} = 1, \gamma_{i0} = 0.5 $ and, assuming $N$ even,
\[
\beta_{i0} = \left\{ \begin{array}{ll} 1 & \mbox{ for } i=1,...,{N \over 2} , \\ 3 & \mbox{ for } i={N \over 2}+1,...,N. \end{array} \right.
\]
This Monte Carlo design is a simplified version of \cite{P06}, designed in such a way that (through (\ref{rank})) the rank condition in \cite{P06}, eq. (21), is not satisfied. \cite{P06} shows that
under this circumstance his individual specific estimator for $ \beta_{i0} $ is invalid whereas his pooled estimators for $\beta_{0} = \mathbb{E} \beta_{i0} $ remains consistent.
We consider $2000$ Monte Carlo replications with sample sizes $(N,T) \in \{ 60,200,600 \} \times \{ 30,100,300 \}$, where $N>T$.
The results are summarized in Tables 4,5, and 6, where we report the sample mean and the root mean square error
for the estimates of the parameter ${\bf \alpha }_{i0}, {\bf \beta }_{i0} $,
averaged across the Monte Carlo iterations. We consider four estimators which corresponds to four panels of each table: the GLS,
the multi-step GLS (described in Section~\ref{various}) where the iteration is carried out $J=4$ times, the OLS
and the UGLS
estimators.
In particular, for each of these four estimators, we report the average across all $N$ units of the sample mean (denoted by \texttt{mean})
$ MM^{-1} \sum_{m=1}^{MM} \hat{\alpha}_i^m $ and of the root mean square error (denoted by (denoted by \texttt{rmse}) $ \left( MM^{-1} \sum_{m=1}^{MM} ( \hat{\alpha}_i^m - 1 )^2 \right)^{1 \over 2} $
and the average across the units $i=N/2+1,...,N$ of
$ MM^{-1} \sum_{m=1}^{MM} \hat{\beta}_i^m $ and $ \left(MM^{-1} \sum_{m=1}^{MM} ( \hat{\beta}_i^m - 3 )^2 \right)^{1 \over 2} $ with $MM=2,000$. Recall that we assumed that the true intercept coefficients are constant across units whereas the regression coefficients take two different values for the first half and second half of the $N$ units.
Here
$\hat{\alpha}_i^m $ and $ \hat{\beta}_i^m $ denote, respectively, the estimates of the intercept and regression coefficients corresponding to the $m$th Monte Carlo iteration for a generic estimator.
\subsection{Results}
Our comments below apply to each table, with minor differences. Since the GLS and multi-step GLS estimators requires $N \ge T $, each panel is made by a lower triangular matrix. Obviously, the OLS and the UGLS estimator do not require this constraint since they can be also
evaluated when $N<T$ but we did not report the results for this case.
The upper left panel describes the GLS results. One can see how the bias diminishes as both $N,T$ grow or when $N$ increases for a given $T$. This is because the
inverse of the pseudo-covariance matrix is better estimated in these circumstances. In contrast, although still negligible in absolute terms, the bias, if any, tends to increase when $T$ grows for a given $N$.
Instead, as expected, the \texttt{rmse} always diminishes when $T$ increases for a given $N$ or when they both increase.
In general these results suggest that the bias of the estimates varies mainly with $N$ and their variance varies with $T$.
The same pattern is observed with respect to the multi-step GLS results, reported in the upper right panel.
The only difference is that now the bias and the \texttt{rmse} are always much smaller than the GLS case. The lower right panel reports the results for the UGLS which is unfeasible in practice since it involves the
true covariance matrix ${\bf S }_N $. As a consequence, the results do not depend on $N$ but only on $T$. The bias is negligible even for small samples and, for larger sample sizes,
it is remarkably comparable to the iterated GLS although the latter exhibit a slightly larger \texttt{rmse}. Finally, the lower left panel reports the OLS results which also do not depend on $N$, as expected.
Under our design, the OLS estimator is non-consistent obtaining a bias which is much larger than for any other estimators and, more importantly, only marginally varying as $N$ or $T$ increases.
The \texttt{rmse} diminishes suggesting that the variance of the OLS estimator is converging to zero with the squared bias converging to the squared of $ \tau_i^{OLS} $.
\section{Empirical Application: Firms' Characteristics and Expected Returns} \label{EMP}
We present an empirical application of our methodology, inspired by asset-pricing theory. According to so-called beta-pricing models, asset returns follow a factor model:
\begin{equation}
R_{i,t} = \tilde{\alpha}_{i} + \tilde{\boldsymbol {\gamma }}_{i}' \bm{\mathrm{D}}_t + u_{i,t} , \label{eq:facmod}
\end{equation}
where $ R_{i,t}$ defines the rate of return for asset $i$, in excess of the risk-free rate, and $ \bm{\mathrm{D}}_t $ is a vector of observed factors, with coefficients $ \tilde{\boldsymbol \gamma }_{i} $.
Important, special, cases of model (\ref{eq:facmod}) are the Capital Asset Pricing Model (CAPM) of \cite{S64} and \cite{L65}, when $\bm{\mathrm{D}}_t$ is the (scalar) excess market return with $ \tilde{\alpha}_{i} =0 $ for every $i$, and the Arbitrage Pricing Theory (APT) of \cite{R76}, when $\bm{\mathrm{D}}_t$ is a vector of possibly non-traded factors.\footnote{Focusing on the special case when $\bm{\mathrm{D}}_t$ are the excess returns of traded assets, the APT holds when the $ \tilde{ \alpha}_{i} $, although not zero, satisfy the condition $ \tilde{\boldsymbol \alpha }' ( var( \bm{\mathrm{u}_t} ) )^{-1} \tilde{\boldsymbol \alpha} < \infty $, setting $ \tilde{\boldsymbol \alpha } = ( \tilde{\alpha}_{1} , \cdots , \tilde{ \alpha}_{N} )' $.}
Model (\ref{eq:facmod}), together with some form of no-arbitrage and some constraints of the covariance matrix of the $u_{i,t}$, implies that expected excess returns $\mathbb{E}(R_{it}) $ are linear in the coefficients $ \boldsymbol{ \alpha }_i $ only, namely that the $\bm{\mathrm{D}}_t$ are the only source of risk (see Corollary 1, \cite{C83}). However, this fundamental paradigm has been challenged empirically. For instance, \cite{DT97}
and \cite{DFF} provide strong evidence according to which stocks characteristics, such as market capitalization (size), valuation (book-to-market) and other characteristics do influence expected returns well beyond the betas. One can extend model (\ref{eq:facmod}) to allow for characteristics by specifying:
\begin{equation}
R_{i,t} = \tilde{\alpha}_{i} + \tilde{\boldsymbol {\gamma}}_{i}' {\bf D}_t + \boldsymbol{\beta}_i' {\bf X}_{i,t} + u_{i,t} = \boldsymbol{ \alpha}_i' (1, {\bf D}_t')' + \boldsymbol{ \beta}_i' {\bf X}_{i,t} + u_{i,t}, \label{eq:facmod2}
\end{equation}
where now $X_{i,t}$ defines a vector of characteristics associated with the $i$th stock, setting $ \boldsymbol{\alpha}_i = ( \tilde{{\alpha}}_{i}, \tilde{\boldsymbol{ \gamma}}_{i}' )' $. Model (\ref{eq:facmod2}) can be interpreted
as, and in fact is equivalent to, our basic model (\ref{eq:hetero}). Moreover, it is conceivable that the error term has a factor structure, such as (\ref{eq:uu}), possibly correlated with both the $\bm{\mathrm{D}}_t$ and the $ {\bm X}_t$. For instance, this is arises whenever one suspects the possibility of missing, pervasive, factors.
Our asymptotic distribution theory can be used to assess whether the $\boldsymbol{ \alpha }_i $ or the $ \boldsymbol{ \beta }_i $ or both are significant or not.
We use a data set of monthly observations, from January 1966 to December 1994, of individual asset returns extracted from CRSP and of firms' characteristics extracted from COMPUSTAT.\footnote{See \cite{BCS98} for details}. In particular, the eight characteristics that we consider are SIZE (the natural logarithm of the market value of the equity of the firm as of
the end of the second to last month), BM (the natural logarithm of the ratio of the book value of equity plus deferred taxes to the market value of equity, using the end of the previous year market and book values)\footnote{As in Fama and French (1993), the value of BM for July of year t to June of year t+1 was computed using accounting data at the end of year t-1.}, DVOL (the natural logarithm of the dollar volume of trading in the security in the second to last month), PRICE (the natural logarithm of the reciprocal of the share price as reported at the end of the second to last month), YLD (the dividend yield as measured by the sum of all dividends paid over the previous 12 months, divided by the share price at the end of the second to last month), RET2-3 (the natural logarithm of the cumulative return over the two months ending at the beginning of the previous month), RET4-6 (the natural logarithm of the cumulative return over the three months ending three months previously),
RET7-12 (the natural logarithm of the cumulative return over the 6 months ending 6 months previously).\footnote{Lagged return variables were constructed to exclude the return during the immediate prior month in order to avoid any spurious association between the prior month return and the current month return caused by thin trading or bid-ask spread effects.}
We report the results in Table 1,2 and 3. In particular, we consider three different factor models, depending on the set of common factors. Table 1 refers to the CAPM model augmented with the eight characteristics.
We report the average, across the $N=356 $ assets, of the GLS estimates $ (\breve{\boldsymbol{ \alpha }}_i^{GLS \prime } , \breve{\boldsymbol \beta }_i^{GLS \prime} )' $ in (\ref{dagger}) for each regression parameter, together with their $10$th and $90$th percentiles, out of the $N$ assets.
Similarly, we report the average, across the $N=356 $ assets, of the t-$ratio$s for each regression parameter, together with their $10$th and $90$th percentiles, out of the $N$ assets.
Finally, we report the F test statistics corresponding to three different joint hypotheses, namely for all $ \boldsymbol{ \alpha}_i = 0 $, or all $ \boldsymbol{\beta}_i=0 $ or both. Again, we report
the average across the $N$ assets of the F test statistics, and their $10$th and $90$th percentile. Table 2 refers to the
three-factor model of \cite{FF3}, augmented with the eight characteristics, whereby the elements of $\bm{\mathrm{D}}_t$ are the market, the small-minus-large (SML) and the high-minus-low (HML) portfolio returns, respectively. Finally, Table 3 refers to the five-factor model of \cite{FF5}, augmented with the eight characteristics, whereby the elements of $\bm{\mathrm{D}}_t$, with respect to the three-factor model, are augmented by the profitability (RMW) and investment (CMA) portfolio returns.
Across all the three asset-pricing models, the results strongly indicate that characteristics influences excess returns, and highly significantly so. This emerges both by considering individual t-$ratios$ as well as the F test for the joint hypothesis that the coefficients to the characteristics (i.e. the $ \boldsymbol{ \beta}_i $) are all zero. Noticeably, the effects of the common factors, for example the market return
for the CAPM, are also strongly significant, across the three asset-pricing models. Indeed, their effects appear unambiguously stronger than for the characteristics, although
they are both highly significant.
\section{Concluding remarks} \label{CONCL}
This paper proposes a feasible GLS estimator for linear panel with common factor structure in both the regressors and the innovation.
We establish our results for time regressions with unit-specific coefficients, and present several generalisations such as dynamic panels, cross-section regressions with time varying coefficients and different factor structures for regressors and residuals.
The GLS estimator is consistent and asymptotically normal, when both the cross-section $N$ and time series $T$ dimensions diverge to infinity where, under the same
circumstances, the OLS is first-order biased.
In summary, the GLS estimator exhibits four main properties:
first, it permits to carry out inference on the regression coefficients based on conventional distributions;
second, as in classical estimation theory, it delivers (almost) efficient estimation;
third, it does not require any knowledge of the exact number of latent
factors, or even an upper bound of such number; and
fourth, the GLS is computationally easy to handle without invoking any nonlinear numerical optimizations.
Our results are corroborated by a set of Monte Carlo experiments and illustrated by an asset-pricing empirical application.
\vspace{-2in}
\begin{table} \label{capmtable}
\centering
{\footnotesize\bf $\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,$ Table 1: \\ $\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,$ CAPM: \\ Testing the effect of characteristics }
\footnotesize
\begin{center}
\begin{tabular}{ |c|ccc|ccc| }
\hline
\multicolumn{7}{|c|}{Panel A} \\ \hline \\
\multicolumn{1}{|c|}{} & \multicolumn{3}{|c|}{t-ratios} & \multicolumn{3}{|c|}{GLS estimates} \\
parameter: & 10-th quantile & average & 90-th quantile & 10-th quantile & average & 90-th quantile \\
intercept &-182.9 & -32.6 & 125.2 & -0.123 & -0.024 & 0.063 \\
Mkt & 1507.2 & 2488.4 & 3568.2 & 0.061 & 0.010 & 0.014 \\
SIZE & -239.5 & -83.6 & 77.1 & -0.217 & -0.069 & 0.073 \\
BM & -118.4 & 14.7 & 153.5 & -0.063 & 0.006 & 0.076 \\
DVOL & -12.1 & 118.9 & 272.4 & -0.01 & 0.038 & 0.085 \\
PRICE & -266.2 & -126.7 & 4.6 & -0.088 & -0.044 & 0.001 \\
YLD & -301.9 & -103.6 & 59.9 & -0.150 & -0.051 & 0.010 \\
RET23 & -289.1 & -143.4 & -26.4 & -0.021 & -0.012 & -0.002 \\
RET46 & -264.9 & -139.9 & -14.2 & -0.022 & -0.012 & -0.001 \\
RET712 & -234.4 & -121.4 & -18.7 & -0.019 & -0.011 & -0.001 \\
\hline
\multicolumn{7}{|c|}{Panel B} \\ \hline
\multicolumn{1}{|c|}{test statistic:} & \multicolumn{2}{|c}{10-th quantile} & \multicolumn{2}{c}{average} & \multicolumn{2}{c|}{90-th quantile} \\
\multicolumn{1}{|c|}{ $F_\gamma$ } & \multicolumn{2}{|c}{1444899.20} & \multicolumn{2}{c}{ 5135622.34} & \multicolumn{2}{c|}{9905946.98 } \\
\multicolumn{1}{|c|}{$ F_\beta $} & \multicolumn{2}{|c}{82958.69} & \multicolumn{2}{c}{ 296084.32} & \multicolumn{2}{c|}{ 612749.73 } \\
\multicolumn{1}{|c|}{$ F_{ \beta , \gamma }$ } & \multicolumn{2}{|c}{7690568.38} & \multicolumn{2}{c}{20110261.61} & \multicolumn{2}{c|}{ 37936804.13 } \\
\hline
\end{tabular}
\end{center}
\end{table}
\noindent \footnotesize{{\bf Note to Table~1}: Panel A reports t-$ratios$ and parameter estimates corresponding to the CAPM model, augmented with characteristics SIZE,
BM, DVOL, PRICE, YLD, RET23, RET46 and RET712:
\[
R_{it} = \tilde{\alpha }_{i} + \tilde{\gamma }_{i} R_{Mkt,t} + \boldsymbol{ \beta}_{i}' {\bf X}_{it} + u_{it}, \,\,\, t=1, \cdots , T, \,\,\, i=1, \cdots , N ,
\]
where $R_{it}$ defines the excess return on asset $i$, $R_{Mkt,t} $ is the $S\& P 500 $ excess return and $X_{it}$ the $8 \times 1 $ vector of characteristics.
Panel B reports the $F$ test statistics corresponding to the null hypotheses $H_0: \tilde{ \gamma }_{i} = 0 $,
$H_0: \boldsymbol{ \beta}_{i} = 0 $ and $H_0: \boldsymbol{ \beta}_i = 0 , \tilde{ \gamma }= 0 $, given by $F_\gamma , F_\beta $ and $ F_{\beta , \gamma } $ respectively.
The data are monthly and makes a panel of monthly observations with $T=348, N=356 $. The characteristics have been cross-sectionally standartized.
Column 2 to 4 of Panel A report the $10$-th decile, the average and the $90$th decile of the t-$ratios$ across the $ N$ assets. Columns 5 to 7 of Panel A report the same quantities with respect to the
parameter estimates, using the GLS estimator $ (\breve{ \boldsymbol \alpha }_i^{GLS \prime} , \breve{ \boldsymbol \beta }_i^{GLS \prime} )' $ in (\ref{dagger}). Their covariance matrix
is estimated using the approach described in Section~\ref{SE}. Column 2 to 4 of Panel B report the $10$-th decile, the average and the $90$th decile of the three F test statistics
across the $ N$ assets.
\noindent
}
\newpage
\begin{table} \label{ff3table}
\centering
{\footnotesize\bf $\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,$ Table 2: \\ $\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,$ Fama French (1993) 3-factor model: \\ $\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,$ Testing the effect of characteristics }
\footnotesize
\begin{center}
\begin{tabular}{ |c|ccc|ccc| }
\hline
\multicolumn{7}{|c|}{Panel A} \\ \hline \\
\multicolumn{1}{|c|}{} & \multicolumn{3}{|c|}{t-ratios} & \multicolumn{3}{|c|}{GLS estimates} \\
parameter: & 10-th quantile & average & 90-th quantile & 10-th quantile & average & 90-th quantile \\
intercept & -210.6 & -36.6 & 133.5 & -0.116 & -0.021 & 0.069 \\
Mkt & 1473.1 & 2325 & 3349.8 & 0.006 & 0.009 & 0.0128 \\
SMB & -541.1 & 256.6 & 1072.5 & -0.003 & 0.003 & 0.010 \\
HML & -340.2 & 348.3 & 1068.7 & -0.002 & 0.002 & 0.006 \\
SIZE & -252.3 & -91.7 & 81.6 & -0.205 & -0.068 & 0.065 \\
BM & -138.3 & 14.6 & 169.3 & -0.06 & 0.004 & 0.076 \\
DVOL & -12.8 & 125.1 & 278.6 & -0.003 & 0.036 & 0.081 \\
PRICE & -303.3 & -140.4 & -2.1 & -0.089 & -0.042 & -0.001 \\
YLD & -327.1 & -113.7 & 72.7 & -0.148 & -0.050 & 0.009 \\
RET23 & -305.8 & -155.2 & -30.8 & -0.021 & -0.011 & -0.002 \\
RET46 & -316.6 & -160.3 & -25.8 & -0.021 & -0.011 & -0.002 \\
RET712 & -272.5 & -141.3 & -28.5 & -0.018 & -0.010 & -0.002 \\
\hline
\multicolumn{7}{|c|}{Panel B} \\ \hline
\multicolumn{1}{|c|}{test statistic:} & \multicolumn{2}{|c}{10-th quantile} & \multicolumn{2}{c}{average} & \multicolumn{2}{c|}{90-th quantile} \\
\multicolumn{1}{|c|}{ $F_\gamma$ } & \multicolumn{2}{|c}{3148154.05} & \multicolumn{2}{c}{8766878.36} & \multicolumn{2}{c|}{16270455.79} \\
\multicolumn{1}{|c|}{$ F_\beta $} & \multicolumn{2}{|c}{87651.83} & \multicolumn{2}{c}{ 315163.52} & \multicolumn{2}{c|}{ 634901.58} \\
\multicolumn{1}{|c|}{$ F_{ \beta , \gamma }$ } & \multicolumn{2}{|c}{ 12894492.39} & \multicolumn{2}{c}{ 29647189.57} & \multicolumn{2}{c|}{ 55497464.04} \\
\hline
\end{tabular}
\end{center}
\end{table}
\noindent \footnotesize{{\bf Note to Table~2}: Panel A reports t-$ratios$ and parameter estimates corresponding to the \cite{FF3} 3-factor model, augmented with characteristics SIZE,
BM, DVOL, PRICE, YLD, RET23, RET46 and RET712:
\[
R_{it} = \tilde{ \alpha}_{i} + \tilde{\gamma }_{i1} R_{Mkt,t} + \tilde{ \gamma }_{i2} R_{SMB,t} + \tilde{\gamma }_{i3} R_{HML,t} + \boldsymbol{ \beta}_{i}' X_{it} + u_{it}, \,\,\, t=1, \cdots , T, \,\,\, i=1, \cdots , N ,
\]
where $R_{it}$ defines the excess return on asset $i$, $R_{Mkt,t} $ is the $S\& P 500 $ excess return, $R_{SMB,t}$ is the size factor, $ R_{HML,t} $ is
the value factor and ${\bf X}_{it}$ the the $8 \times 1 $ vector of characteristics.
Panel B reports the $F$ test statistics corresponding to the null hypotheses $H_0: \tilde{ \boldsymbol \gamma}_{i} = 0 $, $H_0: \boldsymbol{ \beta}_{i} = 0 $ and
$H_0: \boldsymbol{ \beta}_i = 0 , \tilde{ \boldsymbol \gamma }_i = 0 $, given by $F_\gamma , F_\beta $ and $ F_{ \beta , \gamma } $ respectively, setting $ \tilde{ \boldsymbol \gamma }_i = ( \tilde{ \gamma }_{i1},
\tilde{ \gamma }_{i2}, \tilde{ \gamma }_{i3})' $. For details refer to the notes to Table~1.
}
\newpage
\begin{table} \label{ff5table}
\centering
{\footnotesize\bf $\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,$ Table 3: \\ $\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,$ Fama French (2105) 5-factor model: \\ $\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,$ Testing the effect of characteristics }
\footnotesize
\begin{center}
\begin{tabular}{ |c|ccc|ccc| }
\hline
\multicolumn{7}{|c|}{Panel A} \\ \hline \\
\multicolumn{1}{|c|}{} & \multicolumn{3}{|c|}{t-ratios} & \multicolumn{3}{|c|}{GLS estimates} \\
parameter: & 10-th quantile & average & 90-th quantile & 10-th quantile & average & 90-th quantile \\
intercept & -208.7 & -36.6 & 129.1 & -0.120 & -0.021 & 0.075 \\
Mkt & 1303 & 2172.2 & 3207.8 & 0.006 & 0.009 & 0.012 \\
SMB & -605.5 & 259.9 & 1143.2 & -0.003 & 0.002 & 0.010 \\
HML & -326.7 & 202.5 & 863.8 & -0.003 & 0.001 & 0.007 \\
RMW & -440.7 & -57.1 & 344.4 & -0.006 & -0.001 & 0.006 \\
CMA & -488.1 & -0.661 & 484.4 & -0.007 & 0.0001 & 0.007 \\
SIZE & -255.8 & -90.3 & 78.5 & -0.206 & -0.067 & 0.062 \\
BM & -141.7 & 16.4 & 163.8 & -0.063 & 0.005 & 0.075 \\
DVOL & -11.1 & 124.7 & 266.5 & -0.002 & 0.036 & 0.081 \\
PRICE & -300.1 & -138.9 & 0.813 & -0.087 & -0.042 & -0.001 \\
YLD & -319.1 & -113.1 & 71.9 & -0.151 & -0.050 & 0.009 \\
RET23 & -304.6 & -154.5 & -26.7 & -0.021 & -0.011 & -0.002 \\
RET46 & -318.8 & -158.8 & -21.7 & -0.021 & -0.011 & -0.001 \\
RET712 & -265.6 & -141.1 & -31 & -0.018 & -0.010 & -0.002 \\
\hline
\multicolumn{7}{|c|}{Panel B} \\ \hline
\multicolumn{1}{|c|}{test statistic:} & \multicolumn{2}{|c}{10-th quantile} & \multicolumn{2}{c}{average} & \multicolumn{2}{c|}{90-th quantile} \\
\multicolumn{1}{|c|}{ $F_\gamma$ } & \multicolumn{2}{|c}{ 4930084.6} & \multicolumn{2}{c}{11721601.5} & \multicolumn{2}{c|}{20418779.1} \\
\multicolumn{1}{|c|}{$ F_\beta $} & \multicolumn{2}{|c}{83573.5} & \multicolumn{2}{c}{303480.1} & \multicolumn{2}{c|}{ 569989.4} \\
\multicolumn{1}{|c|}{$ F_{\beta , \gamma }$ } & \multicolumn{2}{|c}{14961844.2} & \multicolumn{2}{c}{34707178.6} & \multicolumn{2}{c|}{60856307.5 } \\
\hline
\end{tabular}
\end{center}
\end{table}
\noindent \footnotesize{{\bf Note to Table~3}: Panel A reports t-$ratios$ and parameter estimates corresponding to the Fama and French \cite{FF5} 5-factor model, augmented with characteristics SIZE,
BM, DVOL, PRICE, YLD, RET23, RET46 and RET712:
\[
R_{it} = \tilde{\alpha}_{i} + \tilde{\gamma}_{i1} R_{Mkt,t} + \tilde{\gamma}_{i2} R_{SMB,t} + \tilde{\gamma}_{i3} R_{HML,t} + \tilde{\gamma}_{i4} R_{RMW,t}
+ \tilde{\gamma}_{i5} R_{CMA,t} + \boldsymbol{ \beta }_{i}' \boldsymbol{ X}_{it} + u_{it}, \,\,\, t=1, \cdots , T, \,\,\, i=1, \cdots , N ,
\]
where $R_{it}$ defines the excess return on asset $i$, $R_{Mkt,t} $ is the $S\& P 500 $ excess return, $R_{SMB,t}$ is the size factor, $ R_{HML,t} $ is the value factor,
$R_{RMW,t}$ is the size factor, $ R_{CMA,t} $ is the value factor and ${\bf X}_{it}$ the the $8 \times 1 $ vector of characteristics.
Panel B reports the $F$ test statistics corresponding to the null hypotheses $H_0: \tilde{ \boldsymbol \gamma}_i = 0 $, $H_0: \boldsymbol{ \beta}_{i} = 0 $ and $H_0: \boldsymbol{ \beta}_i = 0 , \tilde{ \boldsymbol \gamma}_i = 0 $,
given by $F_\gamma, F_\beta $ and $ F_{ \beta , \gamma } $ respectively, setting $\tilde{ \boldsymbol \gamma}_i = ( \tilde{\gamma}_{i1}, \tilde{\gamma}_{i2}, \tilde{\gamma}_{i3},
\tilde{\gamma}_{i4}, \tilde{\gamma}_{i5})' $. For details refer to the notes to Table~1.
\noindent
}
\newpage
\begin{table} \label{MC1}
\centering
{\footnotesize\bf $\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,$ Table 4: \\ time regression with unit-specific coefficients \\ intercept term $ \alpha_{i0} = 1, \,\,\, i=1,...,N.$ }
\footnotesize
\begin{tabular}{|lcccccccccccc|}
\hline
\multicolumn{1}{|c|}{ } &
\multicolumn{6}{|c|}{ } & \multicolumn{6}{|c|}{ }
\\
\multicolumn{1}{|c|}{ } & \multicolumn{6}{|c|}{ \begin{centering}
$GLS$ \end{centering}
} &
\multicolumn{6}{|c|}{ \begin{centering} $GLS$ (multi-step) \end{centering} }
\\ \hline
\multicolumn{1}{|c|}{ } & \multicolumn{3}{|c|}{ \begin{centering}
\texttt{mean}
\end{centering}
} & \multicolumn{3}{|c|}{ \begin{centering}
\texttt{rmse} \end{centering}
} & \multicolumn{3}{|c|}{ \begin{centering} \texttt{mean}
\end{centering}
} & \multicolumn{3}{|c|}{ \begin{centering}
\texttt{rmse} \end{centering}
} \\ \hline
\multicolumn{1}{|c|}{ $(N,T)$ } & \multicolumn{1}{|r}{
\hspace{0.6pt} $30$ } & \multicolumn{1}{r}{\hspace{0.6pt} $100$ } &
\multicolumn{1}{c|}{ $300$ } & \multicolumn{1}{|r}{ \hspace{0.6pt}
$30$ } & \multicolumn{1}{r}{\hspace{0.6pt} $100$ } &
\multicolumn{1}{c|}{ $300$ } & \multicolumn{1}{|r}{ \hspace{0.6pt}
$30$ } & \multicolumn{1}{r}{\hspace{0.6pt} $100$ } &
\multicolumn{1}{c|}{ $300$ } & \multicolumn{1}{|r}{ \hspace{0.6pt}
$30$ } & \multicolumn{1}{r}{\hspace{0.6pt} $100$ } &
\multicolumn{1}{c|}{ $300$
} \\
\multicolumn{1}{|c|}{ } & \multicolumn{3}{|c|}{ \begin{centering} \end{centering} } & \multicolumn{3}{|c|}{ \begin{centering} \end{centering} } & \multicolumn{3}{|c|}{ \begin{centering} \end{centering} } & \multicolumn{3}{|c|}{ \begin{centering} \end{centering} } \\
\multicolumn{1}{|l|}{ $60$ } & \multicolumn{3}{|l|}{ \begin{centering} $0.944$ \hspace{0.1in} $-$ \hspace{0.16in} $-$ \end{centering} } & \multicolumn{3}{|l|}{ \begin{centering} $0.523$ \hspace{0.1in} $-$ \hspace{0.16in} $-$ \end{centering} } & \multicolumn{3}{|l|}{ \begin{centering} $0.976$ \hspace{0.1in} $-$ \hspace{0.16in} $-$ \end{centering} } & \multicolumn{3}{|l|}{ \begin{centering} $0.531$ \hspace{0.1in} $-$ \hspace{0.16in} $-$ \end{centering} } \\
\multicolumn{1}{|l|}{ $200$ } & \multicolumn{3}{|l}{ \begin{centering} $0.967$ $ 0.951 $ \hspace{0.1in} $-$ \end{centering} } & \multicolumn{3}{|l|}{ \begin{centering} $0.518$ $ 0.315 $ \hspace{0.1in} $-$ \end{centering} } & \multicolumn{3}{|l|}{ \begin{centering} $0.986$ $0.987 $ \hspace{0.1in} $-$ \end{centering} } & \multicolumn{3}{|l|}{ \begin{centering} $0.527$ $ 0.309$ \hspace{0.1in} $-$ \end{centering} } \\
\multicolumn{1}{|l|}{ $600$ } & \multicolumn{3}{|l|}{
\begin{centering} $0.981$ $ 0.982$ $0.955$ \end{centering} } & \multicolumn{3}{|l|}{ \begin{centering} $0.524$ $0.308$ $0.200$ \end{centering} } & \multicolumn{3}{|l|}{ \begin{centering} $0.994$ $0.998$ $0.991$ \end{centering} } & \multicolumn{3}{|l|}{ \begin{centering} $0.531$ $0.310$ $0.184$
\end{centering} } \\ \hline
\multicolumn{1}{|c|}{ } &
\multicolumn{6}{|c|}{ } & \multicolumn{6}{|c|}{ }
\\
\multicolumn{1}{|c|}{ } & \multicolumn{6}{|c|}{ \begin{centering}
$OLS$ \end{centering}
} &
\multicolumn{6}{|c|}{ \begin{centering} $UGLS$ \end{centering} }
\\ \hline
\multicolumn{1}{|c|}{ } & \multicolumn{3}{|c|}{ \begin{centering}
\texttt{mean}
\end{centering}
} & \multicolumn{3}{|c|}{ \begin{centering}
\texttt{rmse} \end{centering}
} & \multicolumn{3}{|c|}{ \begin{centering} \texttt{mean}
\end{centering}
} & \multicolumn{3}{|c|}{ \begin{centering}
\texttt{rmse} \end{centering}
} \\ \hline
\multicolumn{1}{|c|}{ $(N,T)$ } & \multicolumn{1}{|r}{
\hspace{0.6pt} $30$ } & \multicolumn{1}{r}{\hspace{0.6pt} $100$ } &
\multicolumn{1}{c|}{ $300$ } & \multicolumn{1}{|r}{ \hspace{0.6pt}
$30$ } & \multicolumn{1}{r}{\hspace{0.6pt} $100$ } &
\multicolumn{1}{c|}{ $300$ } & \multicolumn{1}{|r}{ \hspace{0.6pt}
$30$ } & \multicolumn{1}{r}{\hspace{0.6pt} $100$ } &
\multicolumn{1}{c|}{ $300$ } & \multicolumn{1}{|r}{ \hspace{0.6pt}
$30$ } & \multicolumn{1}{r}{\hspace{0.6pt} $100$ } &
\multicolumn{1}{c|}{ $300$
} \\
\multicolumn{1}{|c|}{ } & \multicolumn{3}{|c|}{
\begin{centering}
\end{centering}
} & \multicolumn{3}{|c|}{ \begin{centering}
\end{centering}
} & \multicolumn{3}{|c|}{ \begin{centering}
\end{centering}
} & \multicolumn{3}{|c|}{ \begin{centering}
\end{centering}
} \\
\multicolumn{1}{|l|}{ $60$ } & \multicolumn{3}{|l|}{ \begin{centering} $0.897$ \hspace{0.1in} $-$ \hspace{0.16in} $-$ \end{centering} } & \multicolumn{3}{|l|}{ \begin{centering} $0.560$ \hspace{0.1in} $-$ \hspace{0.16in} $-$ \end{centering} } & \multicolumn{3}{|l|}{ \begin{centering} $0.993$ \hspace{0.1in} $-$ \hspace{0.16in} $-$ \end{centering} } & \multicolumn{3}{|l|}{ \begin{centering} $ 0.369$ \hspace{0.1in} $-$ \hspace{0.16in} $-$ \end{centering} } \\
\multicolumn{1}{|l|}{ $200$ } & \multicolumn{3}{|l|}{ \begin{centering} $0.892$ $ 0.901$ \hspace{0.1in} $-$ \end{centering} } & \multicolumn{3}{|l|}{ \begin{centering} $0.563$ $ 0.361$ \hspace{0.1in} $-$ \end{centering} } & \multicolumn{3}{|l|}{ \begin{centering} $0.993$ $0.999$ \hspace{0.1in} $-$ \end{centering} } & \multicolumn{3}{|l|}{ \begin{centering} $0.368$ $0.221 $ \hspace{0.1in} $-$ \end{centering}
} \\
\multicolumn{1}{|l|}{$600$}&\multicolumn{3}{|l|}{
\begin{centering} $0.898$ $0.902$ $0.904$ \end{centering} } & \multicolumn{3}{|l|}{
\begin{centering} $0.567$ $0.363$ $0.261$ \end{centering} } &
\multicolumn{3}{|l|}{ \begin{centering} $0.994$ $0.998$ $0.999$
\end{centering} } & \multicolumn{3}{|l|}{ \begin{centering} $0.369$
$0.222$ $0.134$ \end{centering} } \\ \hline
\end{tabular}
\end{table}
\noindent \footnotesize{{\bf Note to Table~4}: data are generated according to model
\[
y_{it} = \alpha_{i0} + \beta_{i0} x_{it} + b_{i10} f_{1t} + b_{i20} f_{2t} +\varepsilon_{it}
\]
with regressor $ x_{it} = \gamma_{i0} + \delta_{i10} f_{1t} + \delta_{i30} f_{3t} + v_{it}$.
Factor loadings are normally distributed random variables, $iid$ across units and mutually independent, satisfying
$b_{i10} \sim NID (1,0.2), $ \hspace{0.1in} $ b_{i20} \sim NID (0,0.2) ,$ $ \delta_{i10} \sim NID(0.5,0.5),$ $ \delta_{i30} \sim NID(0,0.5). $
Latent common factors are
$ { f }_{j,t} = $ $ 0.5 { f }_{j,t-1} $ $ + \sqrt{0.5} { \eta }_{jf,t}, $
with $ { \eta }_{jf,t} $ $ \sim NID(0,1) $, mutually independent for $j=1,2,3$, and
idiosyncratic innovation are
$
\varepsilon_{it} = $ $ \rho_{i\varepsilon } { \varepsilon }_{it-1}$ $ + { \eta }_{i\varepsilon,t}$ with $
{ \eta }_{i\varepsilon,t} $ $ \sim NID( 0, \sigma_i^2 ( 1 - \rho_{i\varepsilon }^2) ),$ $
v_{it} = $ $ \rho_{iv } { v }_{it-1} + { \eta }_{iv,t},$ with $
{ \eta }_{iv ,t} $ $ \sim NID( 0, ( 1 - \rho_{i v }^2) ),$
with $ \rho_{i \varepsilon } $ $ \sim UID(0.05,0.95),$ $ \rho_{i v } $ $ \sim UID(0.05,0.95),$ $ \sigma_{i \varepsilon }^2 $ $ \sim UID(0.5,1.5)$, $iid$ across $i=1,...,N$ and mutually independent.
\noindent Parameters of interest are constant across replications and equal to $ \alpha_{i0} = 1, \gamma_{i0} = 0.5 $ and, assuming $N$ even, $
\beta_{i0} = 1 $ for $i=1,...,N/2$ and $\beta_{i0} = 3 $ for $ i=N /2+1,...,N.$
\noindent Panels headed by \texttt{mean} and \texttt{rmse} report, respectively, $ N^{-1} \sum_{i=1}^N \left( MM^{-1} \sum_{m=1}^{MM} \hat{\alpha}_i^m \right) $ and \newline
$ N^{-1} \sum_{i=1}^N \left( MM^{-1} \sum_{m=1}^{MM} ( \hat{\alpha}_i^m - 1 )^2 \right)^{1 \over 2} $ with $MM=2,000$. Here $ \hat{ \alpha }_i^m $ denotes the estimate, based on either the GLS
$ \tilde{\bm \alpha }_i^{GLS} $ of Section~(\ref{common}) (top left panel), multi-step GLS with $J=4$ steps (top right panel), OLS $ \bm{\hat{ \alpha }_i^{OLS}} $ (bottom left panel) and UGLS $ \bm{\hat{ \alpha }_i^{UGLS}} $ (bottom right panel) of $ \alpha_{i0} $ for the $m$ Monte Carlo iteration.
}
\newpage
\begin{table} \label{MC2}
\vspace{-1.30in}
\centering {\footnotesize\bf
$\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,$ Table 5: \\
time
regression with unit-specific coefficients \\ regression coefficient
$ \beta_{i0} = 1, \,\,\, i=1,...,N/2.$ }
\footnotesize
\begin{tabular}{|lcccccccccccc|}
\hline
\multicolumn{1}{|c|}{ } &
\multicolumn{6}{|c|}{ } & \multicolumn{6}{|c|}{ }
\\
\multicolumn{1}{|c|}{ } & \multicolumn{6}{|c|}{ \begin{centering}
$GLS$ \end{centering}
} &
\multicolumn{6}{|c|}{ \begin{centering} $GLS$ (multi-step) \end{centering} }
\\ \hline
\multicolumn{1}{|c|}{ } & \multicolumn{3}{|c|}{ \begin{centering}
\texttt{mean}
\end{centering}
} & \multicolumn{3}{|c|}{ \begin{centering}
\texttt{rmse} \end{centering}
} & \multicolumn{3}{|c|}{ \begin{centering} \texttt{mean}
\end{centering}
} & \multicolumn{3}{|c|}{ \begin{centering}
\texttt{rmse} \end{centering}
} \\ \hline
\multicolumn{1}{|c|}{ $(N,T)$ } & \multicolumn{1}{|r}{
\hspace{0.6pt} $30$ } & \multicolumn{1}{r}{\hspace{0.6pt} $100$ } &
\multicolumn{1}{c|}{ $300$ } & \multicolumn{1}{|r}{ \hspace{0.6pt}
$30$ } & \multicolumn{1}{r}{\hspace{0.6pt} $100$ } &
\multicolumn{1}{c|}{ $300$ } & \multicolumn{1}{|r}{ \hspace{0.6pt}
$30$ } & \multicolumn{1}{r}{\hspace{0.6pt} $100$ } &
\multicolumn{1}{c|}{ $300$ } & \multicolumn{1}{|r}{ \hspace{0.6pt}
$30$ } & \multicolumn{1}{r}{\hspace{0.6pt} $100$ } &
\multicolumn{1}{c|}{ $300$
} \\
\multicolumn{1}{|c|}{ } & \multicolumn{3}{|c|}{ \begin{centering} \end{centering} } & \multicolumn{3}{|c|}{ \begin{centering} \end{centering} } & \multicolumn{3}{|c|}{ \begin{centering} \end{centering} } & \multicolumn{3}{|c|}{ \begin{centering} \end{centering} } \\
\multicolumn{1}{|l|}{ $60$ } & \multicolumn{3}{|l|}{ \begin{centering} $1.089$ \hspace{0.1in} $-$ \hspace{0.16in} $-$ \end{centering} } & \multicolumn{3}{|l|}{ \begin{centering} $0.266$ \hspace{0.1in} $-$ \hspace{0.16in} $-$ \end{centering} } & \multicolumn{3}{|l|}{ \begin{centering} $1.028$ \hspace{0.1in} $-$ \hspace{0.16in} $-$ \end{centering} } & \multicolumn{3}{|l|}{ \begin{centering} $0.198$ \hspace{0.1in} $-$ \hspace{0.16in} $-$ \end{centering} } \\
\multicolumn{1}{|l|}{ $200$ } & \multicolumn{3}{|l|}{ \begin{centering} $1.039$ $ 1.079 $ \hspace{0.1in} $-$ \end{centering} } & \multicolumn{3}{|l|}{ \begin{centering} $0.175$ $ 0.191 $ \hspace{0.1in} $-$ \end{centering} } & \multicolumn{3}{|l|}{ \begin{centering} $1.008$ $1.015 $ \hspace{0.1in} $-$ \end{centering} } & \multicolumn{3}{|l|}{ \begin{centering} $0.148$ $ 0.096$\hspace{0.1in} $-$ \end{centering} } \\
\multicolumn{1}{|l|}{ $600$ } & \multicolumn{3}{|l|}{
\begin{centering} $1.026$ $1.028$ $1.078$ \end{centering} } & \multicolumn{3}{|l|}{ \begin{centering} $0.155$ $0.097$ $0.167$ \end{centering} } & \multicolumn{3}{|l|}{ \begin{centering} $1.006$ $1.002$ $1.012$ \end{centering} } & \multicolumn{3}{|l|}{ \begin{centering} $0.138$ $0.069$ $0.055$
\end{centering} } \\ \hline
\multicolumn{1}{|c|}{ } &
\multicolumn{6}{|c|}{ } & \multicolumn{6}{|c|}{ }
\\
\multicolumn{1}{|c|}{ } & \multicolumn{6}{|c|}{ \begin{centering}
$OLS$ \end{centering}
} &
\multicolumn{6}{|c|}{ \begin{centering} $UGLS$ \end{centering} }
\\ \hline
\multicolumn{1}{|c|}{ } & \multicolumn{3}{|c|}{ \begin{centering}
\texttt{mean}
\end{centering}
} & \multicolumn{3}{|c|}{ \begin{centering}
\texttt{rmse} \end{centering}
} & \multicolumn{3}{|c|}{ \begin{centering} \texttt{mean}
\end{centering}
} & \multicolumn{3}{|c|}{ \begin{centering}
\texttt{rmse} \end{centering}
} \\ \hline
\multicolumn{1}{|c|}{ $(N,T)$ } & \multicolumn{1}{|r}{
\hspace{0.6pt} $30$ } & \multicolumn{1}{r}{\hspace{0.6pt} $100$ } &
\multicolumn{1}{c|}{ $300$ } & \multicolumn{1}{|r}{ \hspace{0.6pt}
$30$ } & \multicolumn{1}{r}{\hspace{0.6pt} $100$ } &
\multicolumn{1}{c|}{ $300$ } & \multicolumn{1}{|r}{ \hspace{0.6pt}
$30$ } & \multicolumn{1}{r}{\hspace{0.6pt} $100$ } &
\multicolumn{1}{c|}{ $300$ } & \multicolumn{1}{|r}{ \hspace{0.6pt}
$30$ } & \multicolumn{1}{r}{\hspace{0.6pt} $100$ } &
\multicolumn{1}{c|}{ $300$
} \\
\multicolumn{1}{|c|}{ } & \multicolumn{3}{|c|}{
\begin{centering}
\end{centering}
} & \multicolumn{3}{|c|}{ \begin{centering}
\end{centering}
} & \multicolumn{3}{|c|}{ \begin{centering}
\end{centering}
} & \multicolumn{3}{|c|}{ \begin{centering}
\end{centering}
} \\
\multicolumn{1}{|l|}{ $60$ } & \multicolumn{3}{|l|}{ \begin{centering} $1.179$ \hspace{0.1in} $-$ \hspace{0.16in} $-$ \end{centering} } & \multicolumn{3}{|l|}{ \begin{centering} $0.451$ \hspace{0.1in} $-$ \hspace{0.16in} $-$ \end{centering} } & \multicolumn{3}{|l|}{ \begin{centering} $1.013 $ \hspace{0.1in} $-$ \hspace{0.16in} $-$ \end{centering} } & \multicolumn{3}{|l|}{ \begin{centering} $ 0.176$ \hspace{0.1in} $-$ \hspace{0.16in} $-$ \end{centering} } \\
\multicolumn{1}{|l|}{ $200$ } & \multicolumn{3}{|l|}{ \begin{centering} $1.179$ $ 1.170$ \hspace{0.1in} $-$ \end{centering} } & \multicolumn{3}{|l|}{ \begin{centering} $0.452$ $ 0.374$ \hspace{0.1in} $-$ \end{centering} } & \multicolumn{3}{|l|}{ \begin{centering} $1.014$ $1.003$ \hspace{0.1in} $-$ \end{centering} } & \multicolumn{3}{|l|}{ \begin{centering} $0.177$ $0.087$ \hspace{0.1in} $-$ \end{centering}
} \\
\multicolumn{1}{|l|}{$600$}&\multicolumn{3}{|l|}{ \begin{centering}
$1.179$ $1.171$ $1.169$ \end{centering} } & \multicolumn{3}{|l|}{
\begin{centering} $0.451$ $0.373$ $0.347$ \end{centering} } &
\multicolumn{3}{|l|}{ \begin{centering} $1.014$ $1.004$ $1.001$
\end{centering} } & \multicolumn{3}{|l|}{ \begin{centering} $0.177$
$0.088$ $0.049$ \end{centering} } \\ \hline
\multicolumn{13}{l}{
\footnotesize{{\bf Note to Table~5}: data are generated according to the same model described in Table~1.}} \\
\multicolumn{13}{l}{
\footnotesize{ Panels headed by \texttt{mean} and \texttt{rmse} report, respectively,
$ (N/2)^{-1} \sum_{i=1}^{N/2} \left( MM^{-1} \sum_{m=1}^{MM} \hat{\beta}_i^m \right) $ }} \\
\multicolumn{13}{l}{ \footnotesize{ and $ (N/2)^{-1} \sum_{i=1}^{N/2} \left( MM^{-1} \sum_{m=1}^{MM} ( \hat{\beta}_i^m - 1 )^2 \right)^{1 \over 2} $ with $MM=2,000$. Here $ \hat{ \beta }_i^m $ denotes the
}} \\
\multicolumn{13}{l}{ \footnotesize{
estimate, based on either the GLS $ \bm{\hat{ \beta }_i^{GLS}} $ in equation (\ref{eq:bgls}) (top left panel), multi-step GLS $\bm{\hat{ \beta }}_i^{(J)} $ }} \\
\multicolumn{13}{l}{ \footnotesize{ with $J=4$ steps
(top right panel), OLS $ \bm{\hat{ \beta }_i^{OLS}} $ (bottom left panel) and UGLS $ \bm{\hat{ \beta }_i^{UGLS}} $ }} \\
\multicolumn{13}{l}{ \footnotesize{ (bottom right panel) of $ \beta_{i0} $
for the $m$th Monte Carlo iteration.}}
\end{tabular}
\end{table}
\begin{table} \label{MC3}
\centering {\footnotesize\bf
$\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,\,$ Table 6: \\ time
regression with unit-specific coefficients \\ regression coefficient
$ \beta_{i0} = 3, \,\,\, i=N/2+1,...,N.$ }
\footnotesize
\begin{tabular}{|lcccccccccccc|}
\hline
\multicolumn{1}{|c|}{ } &
\multicolumn{6}{|c|}{ } & \multicolumn{6}{|c|}{ }
\\
\multicolumn{1}{|c|}{ } & \multicolumn{6}{|c|}{ \begin{centering}
$GLS$ \end{centering}
} &
\multicolumn{6}{|c|}{ \begin{centering} $GLS$ (multi-step) \end{centering} }
\\ \hline
\multicolumn{1}{|c|}{ } & \multicolumn{3}{|c|}{ \begin{centering}
\texttt{mean}
\end{centering}
} & \multicolumn{3}{|c|}{ \begin{centering}
\texttt{rmse} \end{centering}
} & \multicolumn{3}{|c|}{ \begin{centering} \texttt{mean}
\end{centering}
} & \multicolumn{3}{|c|}{ \begin{centering}
\texttt{rmse} \end{centering}
} \\ \hline
\multicolumn{1}{|c|}{ $(N,T)$ } & \multicolumn{1}{|r}{
\hspace{0.6pt} $30$ } & \multicolumn{1}{r}{\hspace{0.6pt} $100$ } &
\multicolumn{1}{c|}{ $300$ } & \multicolumn{1}{|r}{ \hspace{0.6pt}
$30$ } & \multicolumn{1}{r}{\hspace{0.6pt} $100$ } &
\multicolumn{1}{c|}{ $300$ } & \multicolumn{1}{|r}{ \hspace{0.6pt}
$30$ } & \multicolumn{1}{r}{\hspace{0.6pt} $100$ } &
\multicolumn{1}{c|}{ $300$ } & \multicolumn{1}{|r}{ \hspace{0.6pt}
$30$ } & \multicolumn{1}{r}{\hspace{0.6pt} $100$ } &
\multicolumn{1}{c|}{ $300$
} \\
\multicolumn{1}{|c|}{ } & \multicolumn{3}{|c|}{ \begin{centering} \end{centering} } & \multicolumn{3}{|c|}{ \begin{centering} \end{centering} } & \multicolumn{3}{|c|}{ \begin{centering} \end{centering} } & \multicolumn{3}{|c|}{ \begin{centering} \end{centering} } \\
\multicolumn{1}{|l|}{ $60$ } & \multicolumn{3}{|l|}{ \begin{centering} $3.105$ \hspace{0.1in} $-$ \hspace{0.16in} $-$ \end{centering} } & \multicolumn{3}{|l|}{ \begin{centering} $0.314$ \hspace{0.1in} $-$ \hspace{0.16in} $-$ \end{centering} } & \multicolumn{3}{|l|}{ \begin{centering} $3.041$ \hspace{0.1in} $-$ \hspace{0.16in} $-$ \end{centering} } & \multicolumn{3}{|l|}{ \begin{centering} $0.277$ \hspace{0.1in} $-$ \hspace{0.16in} $-$ \end{centering} } \\
\multicolumn{1}{|l|}{ $200$ } & \multicolumn{3}{|l|}{ \begin{centering} $3.053$ $ 3.095 $ \hspace{0.1in} $-$ \end{centering} } & \multicolumn{3}{|l|}{ \begin{centering} $0.228$ $ 0.227 $ \hspace{0.1in} $-$ \end{centering} } & \multicolumn{3}{|l|}{ \begin{centering} $3.013$ $3.024 $ \hspace{0.1in} $-$ \end{centering} } & \multicolumn{3}{|l|}{ \begin{centering} $0.215$ $ 0.146$\hspace{0.1in} $-$ \end{centering} } \\
\multicolumn{1}{|l|}{ $600$ } & \multicolumn{3}{|l|}{
\begin{centering} $3.034$ $3.037$ $3.091$ \end{centering} } & \multicolumn{3}{|l|}{ \begin{centering} $0.209$ $0.133$ $0.198$ \end{centering} } & \multicolumn{3}{|l|}{ \begin{centering} $3.012$ $3.004$ $3.019$ \end{centering} } & \multicolumn{3}{|l|}{ \begin{centering} $0.200$ $0.111$ $0.090$
\end{centering} } \\ \hline
\multicolumn{1}{|c|}{ } &
\multicolumn{6}{|c|}{ } & \multicolumn{6}{|c|}{ }
\\
\multicolumn{1}{|c|}{ } & \multicolumn{6}{|c|}{ \begin{centering}
$OLS$ \end{centering}
} &
\multicolumn{6}{|c|}{ \begin{centering} $UGLS$ \end{centering} }
\\ \hline
\multicolumn{1}{|c|}{ } & \multicolumn{3}{|c|}{ \begin{centering}
\texttt{mean}
\end{centering}
} & \multicolumn{3}{|c|}{ \begin{centering}
\texttt{rmse} \end{centering}
} & \multicolumn{3}{|c|}{ \begin{centering} \texttt{mean}
\end{centering}
} & \multicolumn{3}{|c|}{ \begin{centering}
\texttt{rmse} \end{centering}
} \\ \hline
\multicolumn{1}{|c|}{ $(N,T)$ } & \multicolumn{1}{|r}{
\hspace{0.6pt} $30$ } & \multicolumn{1}{r}{\hspace{0.6pt} $100$ } &
\multicolumn{1}{c|}{ $300$ } & \multicolumn{1}{|r}{ \hspace{0.6pt}
$30$ } & \multicolumn{1}{r}{\hspace{0.6pt} $100$ } &
\multicolumn{1}{c|}{ $300$ } & \multicolumn{1}{|r}{ \hspace{0.6pt}
$30$ } & \multicolumn{1}{r}{\hspace{0.6pt} $100$ } &
\multicolumn{1}{c|}{ $300$ } & \multicolumn{1}{|r}{ \hspace{0.6pt}
$30$ } & \multicolumn{1}{r}{\hspace{0.6pt} $100$ } &
\multicolumn{1}{c|}{ $300$
} \\
\multicolumn{1}{|c|}{ } & \multicolumn{3}{|c|}{
\begin{centering}
\end{centering}
} & \multicolumn{3}{|c|}{ \begin{centering}
\end{centering}
} & \multicolumn{3}{|c|}{ \begin{centering}
\end{centering}
} & \multicolumn{3}{|c|}{ \begin{centering}
\end{centering}
} \\
\multicolumn{1}{|l|}{ $60$ } & \multicolumn{3}{|l|}{ \begin{centering} $3.201$ \hspace{0.1in} $-$ \hspace{0.16in} $-$ \end{centering} } & \multicolumn{3}{|l|}{ \begin{centering} $0.488$ \hspace{0.1in} $-$ \hspace{0.16in} $-$ \end{centering} } & \multicolumn{3}{|l|}{ \begin{centering} $3.013 $ \hspace{0.1in} $-$ \hspace{0.16in} $-$ \end{centering} } & \multicolumn{3}{|l|}{ \begin{centering} $ 0.181$ \hspace{0.1in} $-$ \hspace{0.16in} $-$ \end{centering} } \\
\multicolumn{1}{|l|}{ $200$ } & \multicolumn{3}{|l|}{ \begin{centering} $3.204$ $ 3.196$ \hspace{0.1in} $-$ \end{centering} } & \multicolumn{3}{|l|}{ \begin{centering} $0.490$ $ 0.415$ \hspace{0.1in} $-$ \end{centering} } & \multicolumn{3}{|l|}{ \begin{centering} $3.012$ $3.003$ \hspace{0.1in} $-$ \end{centering} } & \multicolumn{3}{|l|}{ \begin{centering} $0.180$ $0.090$ \hspace{0.1in} $-$ \end{centering}
} \\
\multicolumn{1}{|l|}{$600$}&\multicolumn{3}{|l|}{ \begin{centering}
$3.204$ $3.197$ $3.193$ \end{centering} } & \multicolumn{3}{|l|}{
\begin{centering} $0.492$ $0.417$ $0.389$ \end{centering} } &
\multicolumn{3}{|l|}{ \begin{centering} $3.012$ $3.003$ $3.001$
\end{centering} } & \multicolumn{3}{|l|}{ \begin{centering} $0.180$
$0.090$ $0.051$ \end{centering} } \\ \hline
\multicolumn{13}{l}{
\footnotesize{{\bf Note to Table~6}: data are generated according to the same model described in Table~1.}} \\
\multicolumn{13}{l}{
\footnotesize{ Panels headed by \texttt{mean} and \texttt{rmse} report, respectively,
$ (N/2)^{-1} \sum_{i=N/2+1}^N \left( MM^{-1} \sum_{m=1}^{MM} \hat{\beta}_i^m \right) $ }} \\
\multicolumn{13}{l}{ \footnotesize{ and $ (N/2)^{-1} \sum_{i=N/2+1}^N \left( MM^{-1} \sum_{m=1}^{MM} ( \hat{\beta}_i^m - 3 )^2 \right)^{1 \over 2} $ with $MM=2,000$. Here $ \hat{ \beta }_i^m $ denotes the
}} \\ \multicolumn{13}{l}{ \footnotesize{
estimate, based on either the GLS $ \bm{\hat{ \beta }_i^{GLS}} $ in equation (\ref{eq:bgls}) (top left panel), multi-step GLS $\bm{\hat{ \beta }}_i^{(J)} $ }} \\
\multicolumn{13}{l}{ \footnotesize{ with $J=4$ steps
(top right panel), OLS $ \bm{\hat{ \beta }_i^{OLS}} $ (bottom left panel) and UGLS $ \bm{\hat{ \beta }_i^{UGLS}} $ }} \\
\multicolumn{13}{l}{ \footnotesize{ (bottom right panel) of $ \beta_{i0} $
for the $m$th Monte Carlo iteration.}}
\end{tabular}
\end{table}
\par
\setcounter{section}{0}
\setcounter{subsection}{0}
\setcounter{table}{0}
\setcounter{figure}{0}
\setcounter{equation}{0}
\numberwithin{equation}{section}
\gdef\thetable{\Alph{table}}
\gdef\thefigure{\Alph{figure}}
\gdef\theequation{\Alph{section}.\arabic{equation}}
\gdef\thesection{\Alph{section}}
\normalfont
\begin{center}
\begin{Huge}
\textsc{
\textbf{Appendices}
}
\end{Huge}
\end{center}
All the proofs and technical details are reported in the three appendixes (Appendixes A,B and C) of the manuscripts and in the seven appendixes of the Supplement
(Appendixes D,E,F,G,H,I,J).
\bigskip
In particular, Appendix A states (without proof; see Appendix F) Lemmas A.1 to A.3 and Appendix B contains the proofs to Theorem 3.1 and 3.2 and states
(without proofs; see Appendix G) Propositions B.1 to B.17.
Finally, Appendix C defines the $\bm{\Omega_N} $ matrix and associated quantities that characterize the asymptotic distribution of the GLS estimator.
\bigskip
Regarding the Supplement, Appendix D contains needed results of linear matrix algebra (Lemmas D.1 to D.6), Appendix E contains some results used to construct bounds on the inverse of various covariance matrixes (Lemmas E.1 to E.3 and Corollary E.3), Appendix F contains the proofs to Lemmas A.1 to A.3, Appendix G
contains the proofs to Propositions B.1 to B.17, Appendix H contains important auxiliary results for the proof of Theorem 3.2 (Lemmas H.1 to H.19), Appendix I formalizes the
asymptotic properties of the estimators for the common observed regressors' coefficient (Theorems I.1 to I.3) and, finally, Appendix J provides some technical details for the case when the regressors and the residuals have different, yet correlated, factor structures.
\section{Central Lemmas}\label{centralemmas}
In the following $m,m_1,m_3,m_3$ denote positive constants. The proofs of the lemmas stated in this section are provided in Appendix \ref{proofAppA}.
\begin{lemma}\label{PZ}
Let ${\bf A }(m_1
\times m_1)$, ${\bf C
}(m_2 \times m_2)$, and ${\bf B }(m_1 \times m_2) $, $m_1>m_2$ be random matrices.
Set $ {\bf E:=BCB' + A }$ and assume that $\lambda_1\left(\bf{A}\right)=O_p(1)$. Assume further that $m_2<\infty$, and
\begin{enumerate}[label=(\alph*)]
\item $\left\|\frac{\bm{B'A^{-1}B}}{m_1}\right\|=O_p(1)\;$ and $\quad\left\|\left(\frac{\bm{B'A^{-1}B}}{m_1}\right)^{-1}\right\|=O_p(1)$.
\item $\left\|\left(\frac{\bm{ C^{-1} +
B'A^{-1}B}}{m_1} \right)^{-1}\right\| =O_p(1)$.
\item $\left\|\frac{\bm{B'(A^{-1})'A^{-1}B}}{m_1}\right\|=O_p(1)$.
\item $\left\| \bm{C} \right\|=O_p(1)$ and $\left\| \bm{C}^{-1} \right\|=O_p(1)$.
\end{enumerate}
Then,
\begin{enumerate}[label=(\roman*)]
\item
$
\parallel \bm{ E }^{-1} \bm{ B } \parallel_2=O_p\left( m_1^{-\nicefrac{1}{2}}\right). $
\end{enumerate}
Let $\bm{D}$ be a $m_1\times m_3$ matrix, with $m_3<\infty$.
\begin{enumerate}[label=(\roman*)]
\setcounter{enumi}{1}
\item
If $ \left\| m_1^{-1}\bm{D}' \bm {A }^{-1
} \bm{ B }\right\|_2
= O_p ( 1 ) $,
then
$ \left\| \bm{D}' \bm{ E }^{-1} \bm{ B }\right\| =O_p\left( 1 \right). $
\item If $ \left\| m_1^{-\nicefrac{1}{2}}\bm{D}' \bm{ A }^{-1}
\bm{ B }\right\|_2
= O_p (1) $, then $\left\| \bm{D}' \bm{ E }^{-1} \bm{ B } \right\| =O_p\left(m_1^{-\nicefrac{1}{ 2 }} \right)$.
\end{enumerate}
\end{lemma}
\begin{remark}\label{remarkPZ}
Let $\bm{\bar{E}}:=\bm{A}^{-1}-
\bm{A}^{-1}\bm{B}
\left(
\bm {B'A^{-1}B
}
\right)^{-1}\bm{B}'\bm{A}^{-1}$. Write $\bm{H}:=\bm{A^{-1}B}$; following \cite[p.40]{j96} we find
\begin{eqnarray*}
\bm{\bar{E}}&=&
\bm{A}^{-1}\left(\bm{I}_T-\bm{A}^{-1}\bm{B}\left(\bm {B'A^{-1}B}\right)^{-1}\bm{B}'\right)=\bm{H_{\bot}}\left(\bm{B_{\bot}'}\bm{H_{\bot}}\right)^{-1}\bm{B'_{\bot}},
\end{eqnarray*}
implying that $\bm{\bar{E}}\bm{B}=\bm{0}$. Because
$$
\left\| \bm{E}^{-1}-\bm{\bar{E}}\right\|=\left\|\bm{A^{-1}BC^{-1}B'A^{-1}}\right\|\leq \left\|\frac{\bm{A^{-1}B}}{m^{\nicefrac{1}{2}}}\right\|^2\left\|\frac{\bm{C^{-1}}}{m}\right\|=O_p\left(\frac{1}{m^{\nicefrac{1}{2}}}\right),
$$
it follows that the inverse of the matrix $\bm{E}$ approximates a matrix orthogonal to $\bm{B}$.
\end{remark}
\begin{lemma}\label{approx inv}
Let $\bm{A},\bm{B}$ nonsingular matrices. Then
\begin{eqnarray*}
(i)\quad\bm{A}^{-1}&=&\bm{B}^{-1}-\bm{B}^{-1}(\bm{A}-\bm{B})\bm{B}^{-1}+\bm{B}^{-1}(\bm{A}-\bm{B})\bm{A}^{-1}(\bm{A}-\bm{B})\bm{B}^{-1}\\
(ii)\quad\bm{A}^{-1}&=& \bm{B}\sum_{j=0}^J\left[(-1)\left(\bm{A}-\bm{B}\right)\bm{B}\right]^{j}+\bm{B}(\bm{A}-\bm{B})\bm{A}^{-1}\left[(-1)\bm{B}\left(\bm{A}-\bm{B}\right)\right]^{j}
\end{eqnarray*}
for $J=1,2,\dots.$
\end{lemma}
\begin{lemma}\label{paolo}
Let $g( \omega) $ be a periodic (mod $2 \pi $) symmetric function defined over $ - \pi \le \omega \le \pi $ with bounded $r$-th order derivative $g^{(r)}(\cdot)$, some $ r \ge 1 $. Then the Fourier coefficients $ \varsigma_h = \int_{-\pi }^\pi g( \omega ) \cos ( h \omega ) d \omega $ satisfy:
\[
| \varsigma_h | = O \left( {1 \over h^r } \right)\quad \mbox{ as } \qquad h \rightarrow \infty .
\]
\end{lemma}
\section{Proof of the main theorems}\label{proofmain}
The proofs of the theorems rely on several propositions, the proofs of which are relegated in Appendix \ref{propth1}
\subsection{Proof of Theorem~\ref{Theorem_OLS} }\label{thuno}
\subsubsection*{Proof of part (i)}
Rewrite equation (\ref{eq:charlie1}) as
\begin{equation}
\bm{\hat{ \beta }_i^{OLS}} - \bm{{ \beta }_i} - ( \bm{\mathcal{X}_i}'\bm{\mathcal{X}_i} )^{-1} \bm{\mathcal{X}_i}'\bm{\mathcal{F}} \bm{\mathrm{b}_i} = ( \bm{\mathcal{X}_i}'\bm{\mathcal{X}_i} )^{-1} \bm{\mathcal{X}_i} '\bm{\epsilon_i}. \label{eq:ols1}
\end{equation}
Proposition \ref{propo1} implies that
\begin{equation}\label{eq:radio106}
{ \bm{\mathcal{X}_i}'\bm{\mathcal{X}_i} \over T } = \bm{\Gamma_{i}}' { \bm{\mathcal{F}} ' \bm{\mathcal{F}} \over T } \bm{\Gamma_{i}} + { \bm{\mathcal{V}}_{\bm{i}}' \bm{\mathcal{V}}_{\bm{i}} \over T } +
{ \bm{\mathcal{V}}_{\bm{i}}' \bm{\mathcal{F}} \over T } \bm{\Gamma_{i}} + \bm{\Gamma_{i}}' { \bm{\mathcal{F}}'\bm{\mathcal{V}}_{\bm{i}} \over T } \xrightarrow{p} \bm{\Sigma_{\bm{\mathcal{X}_i} ' \bm{\mathcal{X}_i} }}>0,
\end{equation}
and
\begin{equation}\label{eq:radio107}
\frac{\bm{\mathcal{X}_i}'\bm{\mathcal{F}}}{T}=\bm{\Gamma_{i}}\frac{\bm{\mathcal{F}}' \bm{\mathcal{F}}}{T}+\frac{\bm{\mathcal{V}}_{\bm{i}}' \bm{\mathcal{F}}}{T}.
\xrightarrow{p} \bm{\Sigma_{\bm{\mathcal{X}_i}'\bm{\mathcal{F}}}} \bm{\mathrm{b}_i}.
\end{equation}
Hence $( \bm{\mathcal{X}_i}'\bm{\mathcal{X}_i} )^{-1} \bm{\mathcal{X}_i}'\bm{\mathcal{F}} \bm{\mathrm{b}_i}\bm{\mathrm{b}_i}\xrightarrow{p}\bm{\tau_i^{OLS}}$, proving $(\ref{eq:bias17})$.
Next, we derive the asymptotic distribution of
\begin{equation}\label{eq:twoterms}
T^{-{\nicefrac{1}{2}}} \bm{\mathcal{X}_i}'\bm{\epsilon_i} = T^{-\nicefrac{1}{2}} \bm{\Gamma_{i}} ' \bm{\mathrm{F}} ' \mathscr{M}_{\bm{\mathrm{D}}} \bm{\varepsilon_i} + T^{-{\nicefrac{1}{2}}} {\bf{V}}_{\bm{i}}' \mathscr{M}_{\bm{\mathrm{D}}}\bm{\varepsilon_i} .
\end{equation}
We first show that the first term in (\ref{eq:twoterms}) satisfies
\begin{equation}\label{eq:200117a}
T^{-\nicefrac{1}{2}}{ \bm{\Gamma_{i}} '\bm{\mathrm{F}}' \mathscr{M}_{\bm{\mathrm{D}}} \bm{\varepsilon_i} }
\xrightarrow{d}
\mathcal{N}\left(0, \bm{\Gamma_{i}} ' ( - \bm{ \Sigma} _{\bm{\mathrm{D}}'\bm{\mathrm{D}}}^{-1} , \bm{I}_M ) \bm{ \Sigma }_{\bm{\mathrm{Z}}' \bm{\Xi_{i}} \bm{\mathrm{Z}}} ( - \bm{ \Sigma} _{\bm{\mathrm{F}}'\bm{\mathrm{D}}} \bm{ \Sigma}_{\bm{\mathrm{D}}'\bm{\mathrm{D}}}^{-1} , \bm{I}_M )' \bm{\Gamma_{i}} \right) ,
\end{equation}
Note that
$\bm{\Gamma_{i}} '\bm{\mathrm{F}}' \mathscr{M}_{\bm{\mathrm{D}}} \bm{\varepsilon_i}=\bm{\Gamma_{i}} ' ( - \bm{\mathrm{F}} '\bm{\mathrm{D}}( \bm{\mathrm{D}}' \bm{\mathrm{D}})^{-1}, \bm{I}_M ){ \bm{\mathrm{Z}}' \bm{\varepsilon_i} },$
with $\bm{\mathrm{Z}}$ defined in Assumption \ref{ass factors}. Because of (\ref{eq:maggio1}), to prove (\ref{eq:200117a}) it suffices to prove that
\begin{equation}\label{eq:radio110}
T^{-\frac{1}{2}}\bm{\mathrm{Z}}\bm{'\varepsilon_i}\xrightarrow{d}\mathcal{N}\left(0,\bm{ \Sigma }_{\bm{\mathrm{Z}}' \bm{\Xi_{i}} \bm{\mathrm{Z}}}\right).
\end{equation}
Adapting \cite[Proof of Theorem 1]{RH97}, and using equations (\ref{eq:11CHL}) and (\ref{eq:11}) we have
\begin{equation}\label{eq:cdieci}
T^{-\nicefrac{1}{2}}\bm{\mathrm{Z}}' \bm{\varepsilon_i} = \sum_{j=1}^N r_{ij} \left( T^{-\nicefrac{1}{2}} \sum_{t=1}^T \bm{\mathrm{z}}_{t} a_{jt} \right) =
\sum_{j=1}^N r_{ij}\bm{w_{j}},
\end{equation}
where $\bm{w_{j}}:=T^{-\nicefrac{1}{2}} \sum_{t=1}^T \bm{\mathrm{z}}_{t} a_{jt} $, and $\bm{\mathrm{z}}'_t=(z_{t1},\dots,z_{t(M+S)})'$ is the $t-$th row of $\bm{\mathrm{Z}}$. For $\tau_0=\tau_0(T)$ yet to be chosen, define
\begin{equation}\label{eq:cdieci3}
\bm{w}_{j0}:= T^{-{1 \over 2} } \sum_{u=- \tau_0 }^T \bm{s}_{ju} \eta_{ju},
\qquad \textrm{with} \qquad
\bm{s}_{ju}:= \sum_{t=1}^T \bm{\mathrm{z}}_{t} \phi_{jt-u},\;
\phi_{jh} = 0, \mbox{ for } h < 0 ,
\end{equation}
and let $\bm{w}_{j1}:=\bm{w}_{j}-\bm{w}_{j0}$, $\bm{W}_{j1}:=\bm{W}_{j}-\bm{W}_{j0}$,
where
\begin{equation}\label{eq:cidieci3FEB}
\bm{W}_{j}:= \mathbb{E} \left( \bm{w}_{j} \bm{w}_{j}' | \operatorname{\mathscr{F}}(\bm{\mathrm{Z}}) \right),\qquad \mbox{and}
\qquad
\bm{W}_{j0} := \mathbb{E} ( \bm{w}_{j0} \bm{w}_{j0}' | \operatorname{\mathscr{F}}(\bm{\mathrm{Z}}) ).
\end{equation}
Write
$$
\bm{W_{j}}^{-1/2}\bm{w_{j}}=\left(\bm{I_{(M+S)}}+\bm{W_{j1}}\bm{W_{j0}}^{-1}\right)^{-1/2}\bm{W_{j0}}^{-1/2}\bm{w_{j0}}+\bm{W_{j}}^{-1/2}\bm{w_{j1}}.
$$
Noting that $\mathbb{E}\left\|\bm{W_{j1}}\right\|\leq \mathbb{E}\left\|\bm{w_{j1}}\right\|^2$, Propositions \ref{prop190117a} and \ref{prop190117c} implies that, for $\tau_0$ increasing suitably with $T$, $\left\| \bm{W_{j}}^{-1/2}\bm{w_{j1}}\right\|=o_p(1)$ and $\left\|\bm{W_{j1}}\bm{W_{j0}}^{-1}\right\|=o_p(1)$. Hence, $ \bm{W_j}^{-1/2}\bm{w_j}\approx\bm{W_{j0}}^{-1/2}\bm{w_{j0}}$ as $T\to\infty$. Therefore, Bernstein's Lemma (see \cite{H70}, p. 242), Propositions \ref{prop190117b} and \ref{prop190117c} imply that, for every given $j$, $\bm{w_{j}}\xrightarrow{d}\mathcal{N}(\bm{0},\bm{\mathcal{W}_{j}})$
with $\bm{\mathcal{W}_j}$ defined in Equation (\ref{eq:190117a}).
However, the terms $\bm{w_j}$ (and their limits) are uncorrelated across $j$'s implying that
$
T^{-{1 \over 2}}\bm{\mathrm{Z}}' \bm{\varepsilon_i}=\sum_{j=1}^N r_{ij}\bm{w_j} \xrightarrow{d} \mathcal{N}(\bm{0}, \sum_{j=1}^N r_{ij}^2 \bm{\mathcal{W}_{j}} ),
$
where notice that by easy calculations $ \sum_{j=1}^N r_{ij}^2 {\bm{\mathcal{W} }}_{j} = \bm{ \Sigma }_{\bm{\mathrm{Z}}' \bm{\Xi_{i}} \bm{\mathrm{Z}} }$, proving (\ref{eq:radio110}).
It is worth noting that as $N$ increases, the latter distribution
can be made arbitrarily close to $ \mathcal{N}\left(\bm{0}, \sum_{j=1}^\infty r_{ij}^2 \bm{ \mathcal{W}_{j}} \right) $ because of the absolute summability of the $r_{ij}$.
\bigskip
The second term in (\ref{eq:twoterms}) satisfies $T^{-\nicefrac{1}{2}}{\bf{V}}_{\bm{i}}'\mathscr{M}_{\bm{\mathrm{D}}}\bm{\varepsilon_i}\approx T^{-\nicefrac{1}{2}}{\bf{V}}_{\bm{i}}'\bm{\varepsilon_i}$ (see Proposition \ref{propo1}\ref{propo1iv}). The proof of the weak convergence,
\begin{equation}\label{eq:radio105}
T^{-\nicefrac{1}{2}}{\bf{V}}_{\bm{i}}'\bm{\varepsilon_i}\xrightarrow{d}\mathcal{N}\left(0,\bm{\Sigma}_{{\bf{V}}_{\bm{i}}'\bm{\Xi_{i}}{\bf{V}}_{\bm{i}}}\right),
\end{equation}
is very similar to the proof of (\ref{eq:radio110}), and hence omitted.
By assumption (\ref{independence}), the two terms on the RHS of (\ref{eq:twoterms}) are uncorrelated, so by (\ref{eq:200117a}) and (\ref{eq:radio105}) its LHS converges weakly to a random variable with mean zero, and variance given in Equation (\ref{eq:sigmaXHX}).
\begin{prop}\label{propo1}
\begin{enumerate}[label=(\roman*)]
\item[]
\item $\frac{\bm{\mathcal{F}}'\bm{\mathcal{F}}}{T}={ \bm{\mathrm{F}} ' \mathscr{M}_{\bm{\mathrm{D}}} \bm{\mathrm{F}} \over T } \xrightarrow{p} \bm{ \Sigma }_{\bm{\mathrm{F}}'\mathscr{M}_{\bm{\mathrm{D}}} \bm{\mathrm{F}}} = \bm{ \Sigma }_{\bm{\mathrm{F}}' \bm{\mathrm{F}}} - \bm{ \Sigma }_{\bm{\mathrm{F}}'\bm{\mathrm{D}}} \bm{ \Sigma }_{\bm{\mathrm{D}}'\bm{\mathrm{D}}}^{-1} \bm{ \Sigma }_{\bm{\mathrm{D}}'\bm{\mathrm{F}}}$ .\label{propo1D1}
\item $\frac{\bm{\mathcal{V}}_{\bm{i}}'\bm{\mathcal{V}}_{\bm{i}}}{T}={ {\bf{V}}_{\bm{i}} '\mathscr{M}_{\bm{\mathrm{D}}}{\bf{V}}_{\bm{i}} \over T } \xrightarrow{p} \bm{ \Sigma}_{{\bf{V}}_{\bm{i}}'{\bf{V}}_{\bm{i}}}.
$\label{propo1D2}
\item $ \left\|\frac{\bm{\mathcal{V}}_{\bm{i}}'\bm{\mathcal{F}}}{T}\right\|=\left\|{ {{\bf{V}}_{\bm{i}} ' \mathscr{M}_{\bm{\mathrm{D}}} \bm{\mathrm{F}}} \over T }\right\| = O_p(T^{-\nicefrac{1}{2}} )$.\label{propo1D3}
\item $ \left\|{{{\bf{V}}_{\bm{i}} ' \bm{\varepsilon_i}} \over \sqrt{T}}-{{{\bf{V}}_{\bm{i}} ' \mathscr{M}_{\bm{\mathrm{D}}} \bm{\varepsilon_i}} \over \sqrt{T} }\right\| = O_p(T^{-\nicefrac{1}{2}} )$.\label{propo1iv}
\item $\bm{\Sigma}_{\bm{\mathcal{X}_i}'\bm{\mathcal{X}_i}}>0.$\label{propo15M}
\end{enumerate}
\end{prop}
\begin{prop}\label{prop190117a}
For $\tau_0$ increasing suitably with $T$
$$
\lim_{T\to\infty}\mathbb{E}\left\|\bm{w}_{j1}\right\|^2=0.
$$
\end{prop}
\begin{prop}\label{prop190117b}
As $T\to\infty$
\begin{equation}\label{eq:scott0}
\bm{W}_{j0} ^{-{1 \over 2}} \bm{w}_{j0} \xrightarrow{d} \mathcal{N}\left(\bm{0}, \bm{I_{(M+S)}} \right).
\end{equation}
\end{prop}
\begin{prop}\label{prop190117c}
As $T\to\infty$
\begin{equation}\label{eq:190117a}
\bm{W}_j \xrightarrow{p} \mathbb{E} (\eta_{j0}^2 ) \bm{\Sigma }_{\bm{\mathrm{Z}}' \bm{\Phi }_j \bm{\mathrm{Z}} }=:\bm{\mathcal{ W}_j} > 0,
\end{equation} and
$\bm{\Phi }_j $ is the $T\times T$ matrix with $(t,s)-$th element equal to $ \sum_{u=- \infty }^{\min(t,s)} \phi_{j,t-u} \phi_{j,s-u} = \sum_{v=0 }^{\infty} \phi_{j,v}
\phi_{j,v+|t-s|}$.
\end{prop}
\subsubsection*{Proof of part (ii)}
All the limits below hold as $T \rightarrow \infty $. By simple manipulation of equation (\ref{eq:UGLSdef})
\begin{equation}
\sqrt{T}\left(\bm{\hat{\beta }_i^{UGLS} - { \beta }_{i}} -
\left(\bm{\mathcal{X}_i}' \bm{\mathcal{S}}^{-1}_N \bm{\mathcal{X}_i} \right)^{-1}
\bm{\mathcal{X}_i}' \bm{\mathcal{S}}^{-1}_N \bm{\mathcal{F}} \bm{\mathrm{b}_i}\right) =
\left(\frac{\bm{\mathcal{X}_i}' \bm{\mathcal{S}}^{-1}_N \bm{\mathcal{X}_i}}{T} \right)^{-1}
\left(\frac{\bm{\mathcal{X}_i}' \bm{\mathcal{S}}^{-1}_N \bm{\epsilon_i}}{\sqrt{T}}\right) . \label{eq:ugls11}
\end{equation}
We first show that estimator is asymptotically unbiased.
Proposition \ref{lemma160117a} implies
\begin{equation}\label{eq:radio111D}
\frac{\bm{\mathcal{X}_i}' \bm{\mathcal{S}}^{-1}_N \bm{\mathcal{X}_i}}{ T } = \bm{\Gamma_{i}}' \frac{ \bm{\mathcal{F}} ' \bm{\mathcal{S}}^{-1}_N \bm{\mathcal{F}}}{T } \bm{\Gamma_{i}} + \frac{ \bm{\mathcal{V}}_{\bm{i}} ' \bm{\mathcal{S}}^{-1}_N \bm{\mathcal{V}}_{\bm{i}}}{T } + \frac{ \bm{\mathcal{V}}_{\bm{i}} ' \bm{\mathcal{S}}^{-1}_N \bm{\mathcal{F}}}{ T } \bm{\Gamma_{i}} + \bm{\Gamma_{i}}' \frac{ \bm{\mathcal{F}} ' \bm{\mathcal{S}}^{-1}_N \bm{\mathcal{V}}_{\bm{i}}}{ T}\xrightarrow{p}\bm{ \Sigma }_{{\bf{V}}_{\bm{i}}' \bm{\Xi^{-1}_{N}} {\bf{V}}_{\bm{i}} } ,
\end{equation}
and
$$
\left\| \bm{\mathcal{X}_i} ' \bm{\mathcal{S}}^{-1}_N\bm{\mathcal{F}} \bm{\mathrm{b}_i}\right\| \leq \left\|\bm{\Gamma_{i}}\right\|\left\| \bm{\mathcal{F}} ' \bm{\mathcal{S}}^{-1}_N\bm{\mathcal{F}} \right\| \left\|\bm{\mathrm{b}_i}\right\| +\left\| \bm{\mathcal{V}}_{\bm{i}} ' \bm{\mathcal{S}}^{-1}_N\bm{\mathcal{F}} \bm{\mathrm{b}_i}\right\|=O_p(1),
$$
implying that the the bias term is $O_p\left(T^{-\nicefrac{1}{2}}\right)$.
To complete the proof we need to derive the limiting distribution of the latter term in (\ref{eq:ugls11}). By Proposition \ref{lemma170117a}, $T^{-\nicefrac{1}{2}}\bm{\mathcal{X}_i} ' \bm{\mathcal{S}}^{-1}_N\bm{\epsilon_i}\approx T^{-\nicefrac{1}{2}}{\bf{V}}_{\bm{i}} ' \bm{\Xi^{-1}_{N}} \bm{\varepsilon_i}$.
Proposition \ref{prop020217a}$(i)$ shows that, $T^{-\nicefrac{1}{2}}{\bf{V}}_{\bm{i}} ' \bm{\Xi^{-1}_{N}} \bm{\varepsilon_i}\approx$ $ T^{-\nicefrac{1}{2}}$ ${\bf{V}}_{\bm{i}} ' \bm{\bar{\Xi}_{N}} \bm{\varepsilon_i}$, where $\bm{\bar{\Xi}_{N}}$, defined in the same proposition, is circulant and symmetric. Hence,
as stated by the second part of Proposition \ref{prop020217a},
$\|\bm{\bar{\Xi}_{N}}\|_{row}<\infty$, allowing us to exploit again \cite[Propositons 1 and 2]{RH97}.
Similarly to (\ref{eq:cdieci})-(\ref{eq:cidieci3FEB}), we define
\begin{equation}\label{eq:radio112}
T^{-\nicefrac{1}{2}} {{\bf{V}}_{\bm{i}}'\bm{\bar{\Xi}_{N}} \bm{\varepsilon_i}} =
\sum_{j=1}^N r_{ij} \left(T^{-{\nicefrac{1}{2}}} \sum_{t=1}^T\sum_{s=1}^T {\bm{\mathrm{v}}}_{it}\bar{\xi}_N(t-s) \mathnormal{a_{js}} \right)
=\sum_{j=1}^N r_{ij}\bm{w^{(\bar{\xi})}_{ij}} ,
\end{equation}
where $\bar{\xi}_N(t-s)$, denoting the $(t,s)-$element of $\bm{\bar{\Xi}_{N}}$, satisfies $\bar{\xi}_N(h)=\bar{\xi}_N(T-h)$, $h=0,1,\dots,T-1$. Write
\begin{equation}\label{eq:LEO}
\bm{w}^{(\bar{\xi})}_{ij0} := T^{-{\nicefrac{1}{2}} } \sum_{u=- \tau_0 }^T \bm{s}^{(\bar{\xi})}_{iju} \eta_{ju},\quad
\bm{s}^{(\bar{\xi})}_{iju}:= \sum_{s=1}^T \bm{\ell}^{(\bar{\xi})}_{is} \phi_{js-u}, \quad\bm{\ell}^{(\bar{\xi})}_{is}:=\sum_{t=1}^{T} \bm{\mathrm{v}}_{it}\bar{\xi}_N(t-s),
\end{equation}
and define $ \bm{w}^{(\bar{\xi})}_{ij1}:= \bm{w}^{(\bar{\xi})}_{ij}- \bm{w}^{(\bar{\xi})}_{ij0}$, and $ \bm{W}^{(\bar{\xi})}_{ij1}:= \bm{W}^{(\bar{\xi})}_{ij}- \bm{W}^{(\bar{\xi})}_{ij0}$, where
\begin{equation}\label{eq:radio112b}
\bm{W}^{(\bar{\xi})}_{ij} := \mathbb{E} \left( \bm{w}^{(\bar{\xi})}_{ij} \bm{w}^{(\bar{\xi})'}_{ij}| \operatorname{\mathscr{F}}({\bf{V}}_{\bm{i}})\right),\quad \mbox{and}
\quad
\bm{W}^{(\bar{\xi})}_{ij0} := \mathbb{E} \left( \bm{w}^{(\bar{\xi})}_{ij0} \bm{w}^{(\bar{\xi})'}_{ij0}| \operatorname{\mathscr{F}}({\bf{V}}_{\bm{i}})\right).
\end{equation}
Proceeding as in the proof of the first part of the theorem, Propositions \ref{propCHLOE1}-\ref{propCHLOE3} allow us to establish that, for any $i,j$
$\bm{w^{(\bar{\xi})}_{ij}}\xrightarrow{d}\mathcal{N}(\bm{0},\bm{\mathcal{W}^{(\bar{\xi})}_{ij}})$, where $\bm{\mathcal{W}^{(\bar{\xi})}_{ij}}$ defined in (\ref{eq:06021717a})
below.
It follows that, by (\ref{eq:radio112})
\begin{equation}\label{eq:radio115}
T^{-{\nicefrac{1}{2}}}{{\bf{V}}_{\bm{i}}'\bm{\bar{\Xi}_{N}} \bm{\varepsilon_i}} \xrightarrow{d} \mathcal{N}\left(\bm{0}, \sum_{j=1}^N r_{ij}^2 \bm{{\cal {\bm W}}^{(\bar{\xi})}_{ij}} \right),
\end{equation}
and $ \sum_{j=1}^N r_{ij}^2 \bm{\mathcal{W}^{(\bar{\xi})}_{ij}}=\bm{ \Sigma }_{{\bf{V}}_{\bm{i}}' \bm{\Xi^{-1}_{N}}\bm{\Xi_{i}}\bm{\Xi^{-1}_{N}} {\bf{V}}_{\bm{i}} }$, as required. The result in (\ref{eq:theorem1part2}) is proved by (\ref{eq:radio111D}) and (\ref{eq:radio115}).
\begin{prop}\label{lemma160117a}
\begin{enumerate}[label=(\roman*)]
\item[]
\item $\left\|\bm{\mathcal{F}} ' \bm{\mathcal{S}}^{-1}_N \bm{\mathcal{F}}\right\| = O_p(1).$
\item $\left\|\bm{\mathcal{V}}_{\bm{i}} ' \bm{\mathcal{S}}^{-1}_N \bm{\mathcal{F}}\right\| = O_p\left(T^{-\nicefrac{1}{2}} \right).$
\item $T^{-1}\left( \bm{\mathcal{V}}_{\bm{i}} ' \bm{\mathcal{S}}^{-1}_N\bm{\mathcal{V}}_{\bm{i}}\right) \xrightarrow{d} \bm{ \Sigma }_{{\bf{V}}_{\bm{i}}' \bm{\Xi^{-1}_{N}}{\bf{V}}_{\bm{i}} } > 0.$ \label{eq:lemma160117aTRE}
\end{enumerate}
\end{prop}
\begin{prop}\label{lemma170117a}
\begin{enumerate}[label=(\roman*)]
\item[]
\item $\left\|\bm{\mathcal{F}}' \bm{\mathcal{S}}^{-1}_N \bm{\epsilon_i}\right\| = O_p\left(T^{-\nicefrac{1}{2}}\right).$
\item $\left\| {\bm{\mathcal{V}}_{\bm{i}}'} {\bm{\mathcal{S}}^{-1}_N} {\bm{\epsilon_i}} -{\bf{V}}_{\bm{i}}'{\bm{\Xi^{-1}_{N}}}{\bm{\varepsilon_i}}\right\|=O_p(1).$
\end{enumerate}
\end{prop}
\begin{prop}\label{prop020217a}
Let
$\bm{\bar{\Xi}_{N}}:=2\pi\bm{P}\bm{G}_{\xi_N}^{-1}\bm{P}'$, with $\bm{P}$ defined in (\ref{eq:peigenvectors}) and $\bm{G}_{\xi_N}={\rm diag }({\bm g}(\xi_N,\omega))$ is defined in (\ref{eq:gt}) with $g_{\xi_N}(\omega)=\sum_{h=-\infty}^\infty \xi_{N}(h)\cos(h\omega)$, where $\xi_N(h)=\xi_{N,ts}$ is the $(t,s)-$entry of the matrix $\bm{\Xi_{N}}$ defined in (\ref{eq:deffeb}). Then,
\begin{enumerate}[label=(\roman*)]
\item $\left\| {\bf{V}}_{\bm{i}} '\bm{\Xi^{-1}_{N}} \bm{\varepsilon_i} -{\bf{V}}_{\bm{i}}'\bm{\bar{\Xi}_{N}}\bm{\varepsilon_i}\right\|=O_p(1)$.
\item $\left\| \bm{\bar{\Xi}_{N}}\right\|_{row}<\infty$.
\end{enumerate}
\end{prop}
\begin{prop}\label{propCHLOE1}
For $\tau_0$ increasing suitably with $T$
$$
\lim_{T\to\infty}\mathbb{E}\left\| \bm{w}^{(\bar{\xi})}_{ij1}\right\|^2=0.
$$
\end{prop}
\begin{prop}\label{propCHLOE2}
As $T\to\infty$
\begin{equation}\label{eq:scott0B}
\left(\bm{W}^{(\bar{\xi})}_{ij0}\right) ^{-{1 \over 2}} \bm{w}^{(\bar{\xi})}_{ij0} \xrightarrow{d} \mathcal{N}(\bm{0}, \bm{I_{m+s}} ).
\end{equation}
\end{prop}
\begin{prop}\label{propCHLOE3} Let $\bm{\mathcal{ W}^{(\bar{\xi})}_{ij}}$ as defined in (\ref{eq:06021717a}) below. Then,
\begin{equation}\label{eq:06021717a}
\bm{W}^{(\bar{\xi})}_{ij} \xrightarrow{p} \mathbb{E} (\eta_{j0}^2 ) \bm{\Sigma }_{{\bf{V}}_{\bm{i}}'\bm{\Xi^{-1}_{N}} \bm{\Phi }_j \bm{\Xi^{-1}_{N}}{\bf{V}}_{\bm{i}} } =: \bm{\mathcal{W}^{(\bar{\xi})}_{ij}} > 0 ,
\end{equation}
where the matrix $\bm{\Phi_j}$ has been defined in (\ref{eq:190117a})
\end{prop}
\subsection{Proof Theorem~\ref{Theorem_GLS}}\label{proofT2}
By Proposition \ref{23feb18} below
\begin{equation}\label{eq:ugls11N}
\sqrt{T}\left(\bm{\hat{\beta}_i^{FGLS}}-\bm{\hat{\beta}_i}\right)\approx
\left(T^{-1}\bm{\mathcal{X}_i}'\bm{\mathcal{H}^{-1}_N}\bm{\mathcal{X}_i}\right)^{-1} \left(T^{-\nicefrac{1}{2}}\bm{\mathcal{X}_i}'\bm{\mathcal{H}^{-1}_N}\bm{\epsilon_i}\right),
\end{equation}
where
\begin{equation}\label{eq:MMM1}
\bm{\mathcal{H}_N}=\bm{\mathrm{D}}_{\bot}'\left(\bm{\mathrm{F}} \bm{A_N}\bm{\mathrm{F}}'+\bm{C_N}\right)\bm{\mathrm{D}}_{\bot},
\end{equation}
and $\bm{A_N}$ and $\bm{\mathrm{C}_N}$ are defined in equations (\ref{eq:seconmomb2}) and (\ref{eq:MMM2}), respectively.
The matrix $\bm{\mathcal{H}_N}$ can be seen as the FGLS counterpart of $\bm{\mathcal{S}}^{-1}_N$. The proof follows closely that of Theorem~\ref{Theorem_OLS} part (ii). However, here we consider the joint asymptotics, for (N,T) diverging simultaneously, as can be appreciated from the inspection of Proposition \ref{23feb18}.
For the first term on the LHS of (\ref{eq:ugls11N}),
proposition \ref{lemma160117aT2} implies that
\begin{eqnarray}\label{eq:radio111}
{ \bm{\mathcal{X}_i} ' \bm{\mathcal{H}^{-1}_N} \bm{\mathcal{X}_i} \over T } &=& \bm{\Gamma_{i}}' { \bm{\mathcal{F}} ' \bm{\mathcal{H}^{-1}_N} \bm{\mathcal{F}} \over T } \bm{\Gamma_{i}} + { \bm{\mathcal{V}}_{\bm{i}} ' \bm{\mathcal{H}^{-1}_N} \bm{\mathcal{V}}_{\bm{i}} \over T } + { \bm{\mathcal{V}}_{\bm{i}} ' \bm{\mathcal{H}^{-1}_N} \bm{\mathcal{F}} \over T } \bm{\Gamma_{i}} + \bm{\Gamma_{i}}' { \bm{\mathcal{F}} ' \bm{\mathcal{H}^{-1}_N} \bm{\mathcal{V}}_{\bm{i}} \over T}\nonumber\\
&\approx & \frac{{\bf{V}}_{\bm{i}}'\bm{\mathrm{C}_N}^{-1}{\bf{V}}_{\bm{i}}}{T}
\xrightarrow{p}\bm{ \Sigma }_{{\bf{V}}_{\bm{i}}' \bm{\mathrm{C}_N}^{-1} {\bf{V}}_{\bm{i}} }
\end{eqnarray}
The matrix $\bm{ \Sigma }_{{\bf{V}}_{\bm{i}}' \bm{\mathrm{C}_N}^{-1} {\bf{V}}_{\bm{i}} }=\mathbb{E}\left({\bf{V}}_{\bm{i}}'\bm{\mathrm{C}_N}^{-1}{\bf{V}}_{\bm{i}}\right)$ is non-stochastic, but does depend on $N$, in general.
To complete the proof we need to derive the limiting distribution of the latter term in (\ref{eq:ugls11N}). By and
Proposition \ref{lemma170117aT2},
\[
\left\| { \bm{\mathcal{X}_i} ' \bm{\mathcal{H}^{-1}_N} \bm{\epsilon_i}- {\bf{V}}_{\bm{i}} ' \bm{\mathrm{C}_N}^{-1} \bm{\varepsilon_i} \over \sqrt{T }} \right\| \leq
\left\|{ \bm{\Gamma_{i}} ' \bm{\mathcal{F}} ' \bm{\mathcal{H}^{-1}_N} \bm{\epsilon_i} \over \sqrt{T}}\right\|+
\left\|{ \bm{\mathcal{V}}_{\bm{i}} ' \bm{\mathcal{H}^{-1}_N} \bm{\epsilon_i} - {\bf{V}}_{\bm{i}} ' \bm{C^{-1}_N} \bm{\varepsilon_i} \over \sqrt{T} }\right\|=O_p\left(\frac{1}{\sqrt{T}}\right),
\]
that is, $T^{-\nicefrac{1}{2}}\bm{\mathcal{X}_i} ' \bm{\mathcal{H}^{-1}_N}\bm{\epsilon_i}\approx T^{-\nicefrac{1}{2}}{\bf{V}}_{\bm{i}} ' \bm{C^{-1}_N} \bm{\varepsilon_i}$.
Proposition \ref{mamma}.(i) show that, in turn, $T^{-1/2}{\bf{V}}_{\bm{i}} ' \bm{C^{-1}_N} \bm{\varepsilon_i}\approx T^{-1/2}{\bf{V}}_{\bm{i}} ' \bm{\bar{C}_N} \bm{\varepsilon_i}$, where $\bm{\bar{C}_N}$ is a circulant matrix defined in the same proposition. Hence, the LHS of (\ref{eq:ugls11N}) can be further approximated as
\begin{equation}\label{eq:2luglio}
\sqrt{T}\left(\bm{\hat{\beta}_i^{FGLS}}-\bm{\hat{\beta}_i}\right)
\approx \left(\frac{{\bf{V}}_{\bm{i}}'\bm{\mathrm{C}_N}^{-1}{\bf{V}}_{\bm{i}}}{T}\right)^{-1}
\frac{{\bf{V}}_{\bm{i}}'\bm{\bar{C}_N}\bm{\varepsilon_i}}{\sqrt{T}}.
\end{equation}
The second part of Proposition \ref{mamma} states that the rows of the matrix $\bm{\bar{C}_N}$ are absolutely summable, allowing us to exploit again \cite[Propositons 1 and 2]{RH97}.
Similarly to (\ref{eq:radio112})-(\ref{eq:radio112b}), we write
\begin{equation}\label{eq:radio112N}
T^{-{1 \over 2}}{\bf{V}}_{\bm{i}}'\bm{\bar{C}_N} \bm{\varepsilon_i} = \sum_{j=1}^N r_{ij} \left( T^{-{1 \over 2}} \sum_{t=1}^T \bm{\mathrm{v}}_{it}\bar{c}_N(t-s) a_{js} \right)=\sum_{j=1}^N r_{ij}\bm{w^{(\bar{c})}_{ij}}\nonumber,
\end{equation}
where $\bar{c}_N(t-s)=\bar{c}_{N,ts}$ is the $(t,s)$-entry of the $(T\times T)$ matrix $\bm{\bar{C}_N}$ and
$$
\bm{w}^{(\bar{c})}_{ij} := \left( T^{-\nicefrac{1}{2} } \sum_{u=- \infty }^T \bm{s}^{(\bar{c})}_{iju} \eta_{ju} \right),\quad
\bm{s}^{(\bar{c})}_{iju}:= \sum_{t=1}^T \bm{\ell}^{(\bar{c})}_{it} \phi_{jt-u}
,\quad
\bm{\ell}^{(\bar{c})}_{it}:=\sum_{s=1}^{T} \bm{\mathrm{v}}_{it}\bar{c}_N(t-s),
$$
with $\phi_{jh} = 0$ for $h < 0$.
By Assumption \ref{ass eps} it also follows that, for any $N$
\begin{equation}\label{eq:27feb}
T^{-1}\mathbb{E}\left({\bf{V}}_{\bm{i}}'\bm{\bar{C}_N} \bm{\varepsilon_i}\bm{\varepsilon_i}'\bm{\bar{C}_N}{\bf{V}}_{\bm{i}}\left|\right.\left\{ \operatorname{\mathscr{F}}({\bf{V}}_{\bm{i}}) \right\}\right)=\sum_{i=1}^N r^2_{ij}\bm{W}^{(\bar{c})}_{ij},
\end{equation}
with $\bm{W}^{(\bar{c})}_{ij}:= \mathbb{E} \left( \bm{w}^{(\bar{c})}_{ij} \bm{w}^{(\bar{c})'}_{ij}| \{ \operatorname{\mathscr{F}}({\bf{V}}_{\bm{i}}) \} \right).$
Proceeding along the lines of Theorem 1, Part (ii) using Propositions \ref{propCHLOE1N}-\ref{propCHLOE2N} we establish that,
\begin{equation}\label{eq:radio115N}
\left(\sum_{i=1}^N r^2_{ij}\bm{W}^{(\bar{c})}_{ij}\right)^{-1/2} T^{-{1 \over 2}}{\bf{V}}_{\bm{i}}' \bm{\bar{C}_N}\bm{\varepsilon_i} \xrightarrow{d} \mathcal{N} \left(\bm{0}, \bm{I}_{K} \right).
\end{equation}
The result in display (\ref{eq:theorem2B}) follows from display (\ref{eq:2luglio}) and Proposition \ref{propMarch1}.
\begin{prop}\label{23feb18}
For $1/T+T^3/N^2\to 0$:
\begin{enumerate}[label=(\roman*)]
\item $\left\| T^{-1}\bm{\mathcal{X}_i}'(\bm{\hat{\mathcal{S}}_N}^{-1}-\bm{\mathcal{H}^{-1}_N})\bm{\mathcal{X}_i}\right\|=o_p(1)$.
\item $\left\| T^{-\nicefrac{1}{2}}\bm{\mathcal{X}_i}'(\bm{\hat{\mathcal{S}}_N}^{-1}-\bm{\mathcal{H}^{-1}_N})\bm{\epsilon_i}\right\|=o_p(1)$.
\end{enumerate}
\end{prop}
\begin{prop}\label{lemma160117aT2}
For any $N$:
\begin{enumerate}[label=(\roman*)]
\item $\left\|\bm{\mathcal{F}} ' \bm{\mathcal{H}^{-1}_N} \bm{\mathcal{F}}\right\| = O_p(1).$
\item $\left\|\bm{\mathcal{V}}_{\bm{i}} ' \bm{\mathcal{H}^{-1}_N} \bm{\mathcal{F}}\right\| = O_p\left(T^{-\nicefrac{1}{2}} \right).$
\item $\left\| T^{-1}\left( \bm{\mathcal{V}}_{\bm{i}} ' \bm{\mathcal{H}^{-1}_N} \bm{\mathcal{V}}_{\bm{i}}-{\bf{V}}_{\bm{i}}'\bm{\mathrm{C}_N}^{-1}{\bf{V}}_{\bm{i}}\right)\right\|=o_p(1).$
\item $T^{-1}\left( \bm{\mathcal{V}}_{\bm{i}} ' \bm{\mathcal{H}^{-1}_N} \bm{\mathcal{V}}_{\bm{i}}\right) \xrightarrow{p} \bm{ \Sigma }_{{\bf{V}}_{\bm{i}}' \bm{\mathrm{C}_N}^{-1} {\bf{V}}_{\bm{i}} } > 0.$
\end{enumerate}
\end{prop}
\begin{prop}\label{lemma170117aT2}
For any $N$
\begin{enumerate}[label=(\roman*)]
\item $\left\|\bm{\Gamma_{i}}'\bm{\mathcal{F}}' \bm{\mathcal{H}^{-1}_N} \bm{\epsilon_i}\right\| = O_p\left(T^{-\nicefrac{1}{2}}\right).$
\item $\left\| \bm{\mathcal{V}}_{\bm{i}} ' \bm{\mathcal{H}^{-1}_N}\bm{\epsilon_i} -{\bf{V}}_{\bm{i}}'\bm{\mathrm{C}_N}^{-1}\bm{\varepsilon_i}\right\|=O_p(1).$
\end{enumerate}
\end{prop}
\begin{prop}\label{mamma}
Let
$\bm{\bar{C}_N}:=2\pi\bm{P}\bm{G}_{c_N}^{-1}\bm{P}'$, with $\bm{P}$ defined in (\ref{eq:peigenvectors}) and $\bm{G}_{c_N}={\rm diag }({\bm g}(c_N,\omega))$ is defined as in display (\ref{eq:gt}) with $g_{c_N}(\omega)=\sum_{h=-\infty}^\infty c_{N}(h)\cos(h\omega)$. The scalar $c_N(h)=c_{N,ts}$ is the $(t,s)-$entry of the matrix $\bm{\mathrm{C}_N}$ defined in (\ref{eq:MMM2}). Then, $\forall N$
\begin{enumerate}[label=(\roman*)]
\item $\left\| {\bf{V}}_{\bm{i}} '\bm{\mathrm{C}_N}^{-1} \bm{\varepsilon_i} -{\bf{V}}_{\bm{i}}'\bm{\bar{C}_N}\bm{\varepsilon_i}\right\|=O_p(1)$.
\item $\left\| \bm{\bar{C}_N}\right\|_{row}<\infty$.
\end{enumerate}
\end{prop}
\begin{prop}\label{propCHLOE1N}
For any $N$ and $T_0$ increasing suitably with $T$
$$
\lim_{T\to\infty}\mathbb{E}\left\| \sqrt{T}\bm{w}^{(\bar{c}_N)}_{ij1}\right\|^2=0.
$$
\end{prop}
\begin{prop}\label{propCHLOE2N}
For any $N$, as $T\to\infty$
$$
\bm{W}_{ij0} ^{-{1 \over 2}} \bm{w}_{ij0} \xrightarrow{d} \mathcal{N}(\bm{0}, \bm{I}_K).
$$
\end{prop}
\begin{prop}\label{propMarch1} For any $N$
$$
\left\| \sum_{i=1}^N r^2_{ij}\bm{W}_{ij}^{(\bar{c})}-\frac{1}{T}{\bf{V}}_{\bm{i}}'\bm{\mathrm{C}_N}^{-1}\bm{\Xi_{i}}\bm{\mathrm{C}_N}^{-1}{\bf{V}}_{\bm{i}}\right\|=O_p\left(\frac{1}{\sqrt{T}}\right).
$$
\end{prop}
\section{The matrix $\bm{\Omega_N}$} \label{auxiliary}
By Equation (\ref{eq:ref_def}), (\ref{eq:storti}) and (\ref{eq:charlie2})
\begin{eqnarray*}
\bm{\hat{\mathcal{S}}_N} &=&\frac{1}{N}\sum_{i=1}^N\mathscr{M}_{\bm{\mathcal{X}_i}}\bm{\mathcal{u}_i} \bm{\mathcal{u}_i} ' \mathscr{M}_{\bm{\mathcal{X}_i}}\\
&=&\frac{1}{N}\sum_{i=1}^N\left(\bm{I}-\bm{\mathcal{F}}\bm{\Gamma_{i}}\bm{\mathcal{X}_i}^+-\bm{\mathcal{V}}_{\bm{i}}\bm{\mathcal{X}_i}^+\right)\left(\bm{\mathcal{F}}\bm{\mathrm{b}_i}+\bm{\epsilon_i} \right)\left(\bm{\mathcal{F}}\bm{\mathrm{b}_i}+\bm{\epsilon_i} \right)'
\left(\bm{I}-\bm{\mathcal{F}}\bm{\Gamma_{i}}\bm{\mathcal{X}_i}^+-\bm{\mathcal{V}}_{\bm{i}}\bm{\mathcal{X}_i}^+\right)'\\
&=:&
\bm{\mathcal{F}} \bm{\hat{A}_N} \bm{\mathcal{F}}' + \bm{\hat{\bm{\mathcal{C}}}_N},
\end{eqnarray*}
with
$$
\bm{\hat{A}_N}:=\frac{1}{N}\sum_{i=1}^N \bm{\hat{A}_i},
\qquad
\bm{\hat{\bm{\mathcal{C}}}_N}:=\frac{1}{N}\sum_{i=1}^N\left(\bm{\hat{\bm{\mathcal{C}}}_{1i}}+\bm{\hat{\bm{\mathcal{C}}}_{2i}}+\bm{\hat{\bm{\mathcal{C}}}_{3i}}
+\bm{\hat{\bm{\mathcal{C}}}_{4i}}
\right).
$$
To define $\bm{\hat{A}}_i$, note that
\begin{eqnarray*}
&&\left(\bm{I}-\bm{\mathcal{F}}\bm{\Gamma_{i}}\bm{\mathcal{X}_i}^+\right)\bm{\mathcal{F}}\bm{\mathrm{b}_i}\bm{\mathrm{b}_i}'\bm{\mathcal{F}}'
\left(\bm{I}-\bm{\mathcal{F}}\bm{\Gamma_{i}}\bm{\mathcal{X}_i}^+\right)'+\left(\bm{\mathcal{F}}\bm{\Gamma_{i}}\bm{\mathcal{X}_i}^+\right)\bm{\epsilon_i}\bm{\epsilon_i}'
\left(\bm{\mathcal{F}}\bm{\Gamma_{i}}\bm{\mathcal{X}_i}^+\right)'\\
&=&
\bm{\mathcal{F}}\left(\bm{I}-\bm{\Gamma_{i}}\bm{\mathcal{X}_i}^+\bm{\mathcal{F}}\right)\bm{\mathrm{b}_i}\bm{\mathrm{b}_i}'
\left(\bm{I}-\bm{\Gamma_{i}}\bm{\mathcal{X}_i}^+\bm{\mathcal{F}}\right)'\bm{\mathcal{F}}'+\bm{\mathcal{F}}\left(\bm{\Gamma_{i}}\bm{\mathcal{X}_i}^+\right)\bm{\epsilon_i}\bm{\epsilon_i}'
\left(\bm{\Gamma_{i}}\bm{\mathcal{X}_i}^+\right)'\bm{\mathcal{F}}'.
\end{eqnarray*}
Hence,
$
\bm{\hat{A}_i}=\left(\bm{I}-\bm{\Gamma_{i}}\bm{\mathcal{X}_i}^+\bm{\mathcal{F}}\right)\bm{\mathrm{b}_i}\bm{\mathrm{b}_i}'
\left(\bm{I}-\bm{\Gamma_{i}}\bm{\mathcal{X}_i}^+\bm{\mathcal{F}}\right)'+\left(\bm{\Gamma_{i}}\bm{\mathcal{X}_i}^+\right)\bm{\epsilon_i}\bm{\epsilon_i}'
\left(\bm{\Gamma_{i}}\bm{\mathcal{X}_i}^+\right)'.
$
Likewise,
$
\bm{\hat{\bm{\mathcal{C}}}_{1i}}= \left(\bm{I}-\bm{\mathcal{V}}_{\bm{i}}\bm{\mathcal{X}_i}^+\right)\bm{\epsilon_i}\bm{\epsilon_i}'
\left(\bm{I}-\bm{\mathcal{V}}_{\bm{i}}\bm{\mathcal{X}_i}^+\right)',$ and $
\bm{\hat{\bm{\mathcal{C}}}_{2i}}=\left(\bm{\mathcal{V}}_{\bm{i}}\bm{\mathcal{X}_i}^+\right)\bm{\mathcal{F}}\bm{\mathrm{b}_i}\bm{\mathrm{b}_i}'\bm{\mathcal{F}}'
\left(\bm{\mathcal{V}}_{\bm{i}}\bm{\mathcal{X}_i}^+\right)'.
$
The term $\bm{\hat{\bm{\mathcal{C}}}_{3i}}$ is defined as
$
\bm{\hat{\bm{\mathcal{C}}}_{3i}}=\sum_{j=1}^{13} \left(
\bm{\hat{\bm{\mathcal{C}}}_{3i,j}}+\bm{\hat{\bm{\mathcal{C}}}'_{3i,j}}\right),
$
where
$$
\begin{array}{lcl}
\bm{\hat{\bm{\mathcal{C}}}_{3i,1}}= -\bm{\mathcal{V}}_{\bm{i}}\bm{\mathcal{X}_i}^+\bm{\mathcal{F}}\bm{\mathrm{b}_i}\bm{\mathrm{b}_i}'\bm{\mathcal{F}}', & \phantom{ccccc}&
\bm{\hat{\bm{\mathcal{C}}}_{3i,2}}= \bm{\mathcal{V}}_{\bm{i}}\bm{\mathcal{X}_i}^+\bm{\mathcal{F}}\bm{\mathrm{b}_i}\bm{\mathrm{b}_i}'\bm{\mathcal{F}}'
\left(\bm{\mathcal{F}}\bm{\Gamma_{i}}\bm{\mathcal{X}_i}^+\right)' ,\\
\bm{\hat{\bm{\mathcal{C}}}_{3i,3}}=\bm{\mathcal{F}}\bm{\mathrm{b}_i}\bm{\epsilon_i}', & &
\bm{\hat{\bm{\mathcal{C}}}_{3i,4}}=\bm{\mathcal{V}}_{\bm{i}}\bm{\mathcal{X}_i}^+\bm{\mathcal{F}}\bm{\mathrm{b}_i}\bm{\epsilon_i}'\left(\bm{\mathcal{V}}_{\bm{i}}\bm{\mathcal{X}_i}^+\right)' ,\\
\bm{\hat{\bm{\mathcal{C}}}_{3i,5}}=-\bm{\mathcal{V}}_{\bm{i}}\bm{\mathcal{X}_i}^+\bm{\mathcal{F}}\bm{\mathrm{b}_i}\bm{\epsilon_i}', & &
\bm{\hat{\bm{\mathcal{C}}}_{3i,6}}=-\bm{\mathcal{F}}\bm{\mathrm{b}_i}\bm{\epsilon_i}'\left(\bm{\mathcal{V}}_{\bm{i}}\bm{\mathcal{X}_i}^+\right)' ,\\
\bm{\hat{\bm{\mathcal{C}}}_{3i,7}}=-\bm{\mathcal{F}}\bm{\Gamma_{i}}\bm{\mathcal{X}_i}^+\bm{\mathcal{F}}\bm{\mathrm{b}_i}\bm{\epsilon_i}', &&
\bm{\hat{\bm{\mathcal{C}}}_{3i,8}}=\bm{\mathcal{V}}_{\bm{i}}\bm{\mathcal{X}_i}^+\bm{\mathcal{F}}\bm{\mathrm{b}_i}\bm{\epsilon_i}'\left(\bm{\mathcal{F}}\bm{\Gamma_{i}}\bm{\mathcal{X}_i}^+\right)',\\
\bm{\hat{\bm{\mathcal{C}}}_{3i,9}}=\bm{\mathcal{F}}\bm{\Gamma_{i}}\bm{\mathcal{X}_i}^+\bm{\mathcal{F}}'\bm{\mathrm{b}_i}\bm{\epsilon_i}'\left(\bm{\mathcal{V}}_{\bm{i}}\bm{\mathcal{X}_i}^+\right)', &&
\bm{\hat{\bm{\mathcal{C}}}_{3i,10}}=-\bm{\mathcal{F}}\bm{\Gamma_{i}}\bm{\mathcal{X}_i}^+\bm{\epsilon_i}\bm{\epsilon_i}',\\
\bm{\hat{\bm{\mathcal{C}}}_{3i,11}}=\bm{\mathcal{F}}\bm{\Gamma_{i}}\bm{\mathcal{X}_i}^+\bm{\epsilon_i}\bm{\epsilon_i}'\left(\bm{\mathcal{V}}_{\bm{i}}\bm{\mathcal{X}_i}^+\right)', & &
\bm{\hat{\bm{\mathcal{C}}}_{3i,12}}=\bm{\mathcal{F}}\bm{\Gamma_{i}}\bm{\mathcal{X}_i}^+\bm{\mathcal{F}}\bm{\mathrm{b}_i}\bm{\epsilon_i}'\bm{\mathcal{X}_i}^+\bm{\Gamma_{i}}'\bm{\mathcal{F}}',
\\
\bm{\hat{\bm{\mathcal{C}}}_{3i,13}}=
-\bm{\mathcal{F}}\bm{\mathrm{b}_i}\bm{\epsilon_i}'\left(\bm{\mathcal{F}}\bm{\Gamma_{i}}\bm{\mathcal{X}_i}^+\right)' .&&
\end{array}
$$
Next, define the matrices
\begin{equation}\label{eq:silver}
\bm{\Omega_N}:=\bm{\mathcal{F}}'\bm{A_N}\bm{\mathcal{F}}'+\bm{\breve{\bm{\mathcal{C}}}_{N}},\qquad\bm{\breve{\bm{\mathcal{C}}}_{N}}:=\bm{\mathrm{D}}_{\bot}'\bm{\breve{\mathrm{C}}_N}\bm{\mathrm{D}}_{\bot},\qquad
\bm{\breve{\mathrm{C}}_N}:= \frac{1}{N}\sum_{i=1}^N \left(\bm{\Xi_{i}}+\bm{\breve{\Theta}_i}\right),\quad
\end{equation}
where $\bm{\Xi_{i}}$ and $\bm{A_N}$ are defined in Equations (\ref{eq:uu}) and (\ref{eq:seconmomb2}),
and
\begin{equation}\label{eq:defpsi}
\bm{\breve{\Theta}_i}:=\mathbb{E}\left[{\bf{V}}_{\bm{i}}\bm{\Psi_i}^{-1} \bm{\Gamma_{i}}\frac{\bm{\mathcal{F}}'\bm{\mathcal{F}}}{T}\bm{\mathrm{b}_i}\bm{\mathrm{b}_i}'\frac{\bm{\mathcal{F}}'\bm{\mathcal{F}}}{T}\bm{\Gamma_{i}}'\bm{\Psi_i}^{-1}{\bf{V}}_{\bm{i}}'\left|\operatorname{\mathscr{F}}(\bm{\mathrm{Z}})\right.\right],
\end{equation}
with $\bm{\Psi}_i$ defined as in (\ref{eq:trenta})
Some properties of the matrices defined above are established Lemmas \ref{boundinverse} and \ref{expansion} in Appendix \ref{lemmata1}. In particular, Lemma \ref{boundinverse} verifies that the matrix $\bm{\Omega_N}$ satisfies the assumptions in Lemma \ref{PZ}, on which rely the proof of most of the results in Lemma \ref{cor}.
\newpage
\title{
\begin{center}
\Huge
Supplementary Material to\\
``{Robust Nearly-Efficient Estimation of Large Panels with Factor Structures}"
\end{center}
}
\maketitle
\newpage
\par
\setcounter{section}{0}
\setcounter{subsection}{0}
\setcounter{table}{0}
\setcounter{figure}{0}
\setcounter{equation}{0}
\numberwithin{equation}{section}
\gdef\thetable{\Alph{table}}
\gdef\thefigure{\Alph{figure}}
\gdef\theequation{\Alph{section}.\arabic{equation}}
\gdef\thesection{\Alph{section}}
\normalfont
\begin{center}
\begin{Huge}
\textsc{
\textbf{Appendices}
}
\end{Huge}
\end{center}
\setcounter{section}{3}
This Supplement is made by seven appendixes: Appendix D contains some results of linear matrix algebra (Lemmas D.1 to D.6), Appendix E contains some results used to construct bounds on the inverse of various covariance matrices (Lemmas E.1 to E.3 and Corollary E.3), Appendix F contains the proofs to Lemmas A.1 to A.3 (stated in Appendix A of the manuscript), Appendix G
contains the proofs to Propositions B.1 to B.17 (stated in Appendix B of the manuscript), Appendix H contains auxiliary results for the proof of Theorem 3.2 (Lemmas H.1 to H.19), Appendix I formalises the asymptotic properties of the estimators for the common observed regressors' coefficient (Theorems I.1 to I.3) and, finally, Appendix J provides some technical details for the case when the regressors and the residuals have different, yet correlated, factor structures.
\section{Some results on matrix algebra}\label{someresults}
In this section, we report for reference some auxiliary results on matrix algebra. The first two Lemmas are reported without proof. In the following, all the matrices have real entries.
\begin{lemma}[Sherman-Morrison-Woodbury formula]\label{smw}
For every matrices $\bm{A},\bm{B}$ and $\bm{C}$ of suitable dimension:
\begin{equation}\label{eq:smw}
({\bm{BCB}' + \bm{A} })^{-1} = {\bm{A}^{-1} - \bm{A}^{-1} \bm{B} ( \bm{C}^{-1} +
\bm{B}'\bm{A}^{-1}\bm{B} )^{-1} \bm{B}' \bm{A}^{-1} },
\end{equation}
if all the involved inverses exist.
\end{lemma}
\bigskip
\begin{lemma}[\cite{bern}, Fact 6.4.10]\label{fact6410}
Let $\bm{A}(n\times m)$ a matrix of rank $m$ and $\bm{B}(m\times m)$ a positive definite matrix. Then,
\begin{equation}\label{eq:fact6410}
(\bm{ABA}')^+=\bm{A}\left(\bm{A}'\bm{A}\right)^{-1}\bm{B}^{-1}\left(\bm{A}'\bm{A}\right)^{-1}\bm{A}'.
\end{equation}
\end{lemma}
\begin{lemma}\label{MN}
Let $\bf{A}$ be defined as in Lemma \ref{fact6410}. Assume further that $\bf{A}'\bf{A}=\bm{I_m}$. Let $\bm{C}(n\times n)$ a positive definite matrix, and define $\bm{E}:=\mathscr{P}_{\bm{A}}\bm{C}\mathscr{P}_{\bm{A}}$. Then,
$$
\bm{C^{-1}}-\bm{E^+}=\bm{Q}\left(\bm{I_n-\bm{L}}\right)\bm{Q'}\geq 0,
$$
where $\bm{Q}(n\times n)$ satisfies $\bm{Q}'\bm{Q}=\bm{C^{-1}}$ and the diagonal matrix $\bm{L}$ is obtained setting equal to zero the last $n-m$ diagonal elements of the matrix $\bm{I}_n$.
\end{lemma}
\begin{proof}
By simple manipulation and Lemma \ref{fact6410}
\begin{eqnarray*}
\bm{C^{-1}}-\bm{E^+}&=&
\bm{C^{-\frac{1}{2}}}\left(
\bm{I_n}-\bm{C^{\frac{1}{2}}}\bm{E^+}\bm{C^{\frac{1}{2}}}
\right)\bm{C^{-\frac{1}{2}}}\\
&=&
\bm{C^{-\frac{1}{2}}}\left[
\bm{I_n}-\bm{C^{\frac{1}{2}}}\bf{A}\left(\bf{A}'\bm{C}\bf{A}\right)^{-1}\bm{C^{\frac{1}{2}}}
\right]
\bm{C^{-\frac{1}{2}}}=\bm{C^{-\frac{1}{2}}}\mathscr{M}_{\mathcal{A}}\bm{C^{-\frac{1}{2}}},
\end{eqnarray*}
with $\bm{\mathcal{A}}=\bm{C^{\frac{1}{2}}}\bf{A}$. Rewriting $\mathscr{M}_{\mathcal{A}}=\bm{PLP'}$, with $\bm{P'P}=\bm{I}_n$, the proof is completed setting $\bm{Q=C^{\frac{1}{2}}P}$.
\end{proof}
\begin{lemma}\label{emme}
\begin{enumerate}[label=(\roman*)]
\item[]
\item\label{emme1} Let $\bm{A}(m\times m)$ positive semidefinite, and $\bm{B}(n\times m)$. Then
$$
\lambda_n(\bm{BAB'})\leq \lambda_m(\bm{A})\lambda_n(\bm{BB'}),\quad {\rm tr }(\bm{BAB'})\leq {\rm tr }(\bm{A})\lambda_n\left(\bm{B}\bm{B}'\right).
$$
\item\label{emme2}Let $\mathscr{P}_{B}$ a projection matrix, with $\bm{B}(n\times m)$, $n>m$, and $\bm{A}(m\times m)$. Then,
$$
\lambda_n\left(\mathscr{P}_{B}\bm{A}\mathscr{P}_{B}\right)\leq \lambda_n\left(\bm{A}\right).
$$
\item\label{emme3}For any symmetric matrix $\bm{A}(n\times n)$ and positive semidefinite matrix $\bm{B}(n\times n)$
$$
\lambda_1(\bm{A}+\bm{B})\geq \lambda_1(\bm{A}),\quad
\lambda_n(\bm{A}+\bm{B})\leq \lambda_n(\bm{A})+\lambda_n(\bm{B}),
$$
and
$$
\lambda_1\left(\bm{A}\right){\rm tr }\left(\bm{B}\right)
\leq {\rm tr }\left(\bm{AB}\right)
\leq
\lambda_n\left(\bm{A}\right){\rm tr }\left(\bm{B}\right).
$$
\item\label{emme4}Let $\bm{A}(n\times n)$ be a positive semidefinite matrix, then
$$
\max_{i,j}|a_{ij}|\leq \max_ia_{ii}\leq \lambda_n(\bm{A}),\quad\textrm{for}\quad i,j=1,\dots,n.
$$ .
\end{enumerate}
\end{lemma}
\begin{proof}
The first inequality of part \ref{emme1} can be found in \cite[p. 237]{MN88}. About the second inequality, first note that ${\rm tr }(\bm{B'}\bm{A}\bm{B})={\rm tr }(\bm{B}\bm{B}'\bm{A})$. The proof is concluded using the quasilinear representation of the extremal eigenvalues \cite[p. 204]{MN88}. For part \ref{emme2}, by Reayleigh quotient we have
$$
\max_{\bm{x}}\frac{\bm{x}'\mathscr{P}_{B}\bm{A}\mathscr{P}_{B}\bm{x}}{\bm{x}'\bm{x}}\leq \max_{\bm{x}}\frac{\bm{x}'\mathscr{P}_{B}\bm{A}\mathscr{P}_{B}\bm{x}}{\bm{x}'\mathscr{P}_{B}\bm{x}}\leq \max_{\bm{x}} \frac{\bm{x}'\bm{A}\bm{x}}{\bm{x}'\bm{x}}=\lambda_n\left(\bm{A}\right) .
$$
For part \ref{emme3} see \cite[p. 204]{MN88}. Note that for the second inequality to hold, only requires $\bm{B}$ to be symmetric. Part \ref{emme4} follows from \cite[Exercise 8.7]{AM05} and \cite[Section 9.13.4(4)]{ltk}.
\end{proof}
\begin{lemma}\label{matrix}
Let $\bm{A}(m\times m)$ a matrix satisfying $\|\bm{A}\|_{col}<\infty$ and $\|\bm{A}\|_{row}<\infty$, and $\bm{p}$ and $\bm{q}$ two conformable vectors. Then,
$$
\sup_{\bm{p},\bm{q}}\left|\sum_{i,j=1}^m p_{i}a_{ij}q_{j}\right|\leq
\frac{\|\bm{A}\|_{row}\|\bm{A}\|_{col}
\|\bm{p}\|\|\bm{q}\|}{\left\|\bm{A}\right\|_{sp}}.
$$
\end{lemma}
\begin{proof}
Follows from (20) and (14), on pages 110 and 112 in \cite{ltk}, respectively.
\end{proof}
\begin{lemma}\label{matrix2}
Let $\bm{A}(m\times m)$ positive semi-definite matrix with $\lambda_m(\bm{A})<\infty$, and denote with $\bm{a_i}$ its $i$-th row. Then, $\|\bm{a_i}\|<\lambda_m(\bm{A})$.
\end{lemma}
\begin{proof}
Let $\bm{B}=\bm{A}\bm{A}=[b_{ij}]_{i,j=1}^m$. By the second inequality in Lemma \ref{emme}\ref{emme4}, $\|\bm{a_i}\|^2=b_{ii}\leq \lambda_{m}(\bm{B})=\lambda_m^2(\bm{A})$.
\end{proof}
\section{The inverse of a covariance matrix}\label{invcov}
\setcounter{equation}{0}
Lemma \ref{circulant} reports a result well known in time series analysis. We prefer to report it anyway with proof, both for completeness and to introduce some further notation.
Approximating the matrix $\bm{\mathfrak{C}}$ with the circulant symmetric matrix $\bm{\mathfrak{C}}^{(s)}$, is convenient for deriving the properties of the inverse of $\bm{\mathfrak{C}}$. Diagonalizing $\bm{\mathfrak{C}}^{(s)}$ we obtain the matrix $\Lambda$ of easy interpretation (see \ref{eq:lambda17}). Lemma \ref{020217b} highlights the importance of Lemma \ref{paolo}. The latter implies that the inverse of the matrix $\bm{\mathfrak{C}}^{(s)}$ has bounded row-norm. The same property is inherited by the matrix $\bm{P'G^{-1}_{\varsigma}\bm{P}}$, that is more convenient to work with. The main advantage is that the nonzero entries of the diagonal matrix $\bm{G^{-1}_{\varsigma}}$ do not depend on $T$. Corollary \ref{circulantN} adapt the result derived for time series analysis to a panel framework.
\begin{lemma}\label{circulant}
Let $\varsigma(\cdot)$ a real function defined on the integers. Assume that $\varsigma(\cdot)$ is symmetric, non-negative definite and $\sum_{h=1}^{\infty}h^{\delta}\left|\varsigma(h)\right|<\infty$, with $\delta\geq 1$.
Let
\begin{equation}\label{eq:gto}
g_{\varsigma}(\omega):=\frac{1}{2\pi}\sum_{h=-\infty}^{\infty}\varsigma(h)\cos(h\omega),\qquad -\pi\leq \omega\leq\pi
\end{equation}
and ${\bm{G}}_{\varsigma}:={\rm diag }\left({\bm{g}}(\varsigma,\omega)\right)$ be the $T\times T$ matrix, with
\begin{equation}\label{eq:gt}
{\bm{g}}(\varsigma,\omega):=\left\{
\begin{array}{l}
\left(g_{\varsigma}(0),g_{\varsigma}(\omega_1),g_{\varsigma}(\omega_1),\dots,g_{\varsigma}\left(\omega_{[T/2]}\right),g_{\varsigma}\left(\omega_{[T/2]}\right)\right)' \qquad\qquad\qquad\textrm{if T is odd,}\\
\left(g_{\varsigma}(0),g_{\varsigma}(\omega_1),g_{\varsigma}(\omega_1),\dots,g_{\varsigma}\left(\omega_{(T-2)/2}\right),g_{\varsigma}\left(\omega_{(T-2)/2}\right),g_{\varsigma}\left(\omega_{T/2}\right)\right)' \;\textrm{otherwise,}
\end{array}
\right.
\end{equation}
where $\omega_j=2\pi j/T$, $j=0,\dots,[T/2]$ and $[x]$ denotes the integer part of $x$.
Define the real orthogonal $T \times T$ matrix ${\bm P } $ by
\begin{equation}\label{eq:peigenvectors}
\begin{array}{lll}
{\bm P } &:= \left[ {\bm q }_0,\; {\bm q }_1,\; {\bm s }_1,\; \cdots, \;{\bm q }_{[T/2]},\; {\bm s }_{[T/2]} \right]' &\qquad \textrm{if $T$ is odd},\\
{\bm P } &:= \left[ {\bm q }_0,\; {\bm q }_1,\;{\bm s }_1,\; \cdots, \; 2^{-1/2}{\bm q }_{T/2} \right]'&\qquad \textrm{otherwise},
\end{array}
\end{equation}
where
\begin{eqnarray*}
\bm{q_j}&=&\sqrt{2/T}\left[
1\quad \cos\omega_j\quad \cos2\omega_j \quad \cdots \quad \cos(T-1)\omega_j
\right]',\\
\bm{s_j}&=&\sqrt{2/T}\left[
0\quad \sin\omega_j\quad \sin 2\omega_j \quad \cdots \quad \sin(T-1)\omega_j
\right]',
\end{eqnarray*}
and the matrix $\bm{\mathfrak{C}}=\left[\varsigma(i-j)\right]_{i,j=1}^T$. Denote by $\Delta_{pq}^{(T)} $ is the $p,q$ component of the matrix
$
{\bm P } \bm{\mathfrak{C}} {\bm P }' - 2\pi\bm{G}_{\varsigma},
$
then
\begin{equation}\label{eq:300117a}
\sup_{1 \le i,j \le T } | \Delta_{ij}^{(T)}|=O\left(\frac{1}{T}\right) \qquad\textrm{for}\quad i,j=1,\dots,T.
\end{equation}
\end{lemma}
\begin{proof} The proof follows closely \cite[Proposition 4.5.2]{BD91}. Define the $T\times T$ circulant symmetric matrix
\begin{equation}\label{eq:CS}
\bm{\mathfrak{C}}^{(s)}=\left[
\begin{array}{cccccc}
\varsigma(0) & \varsigma(1) & \varsigma(2) & \cdots & \varsigma(2) & \varsigma(1)\\
\varsigma(1) & \varsigma(0) & \varsigma(1) & \cdots & \varsigma(3) & \varsigma(2)\\
\varsigma(2) & c(1) & \varsigma(0) & \cdots & \varsigma(4) & \varsigma(3)\\
\vdots & \vdots & \vdots & & \vdots & \vdots\\
\varsigma(1) & \varsigma(2) & \varsigma(3) & \cdots & \varsigma(1) & \varsigma(0)
\end{array}
\right].
\end{equation}
For brevity's sake we only consider the case of $T$ odd\footnote{For the case when $T$ is even the reader is referred to \cite{BD91}, p. 135}. Following \cite[Section 4.5]{BD91}, the above matrix can be diagonalized as $\bm{P}\bm{\mathfrak{C}}^{(s)}\bm{P}'=\bm{\Lambda}$,
\begin{equation}\label{eq:lambda17}
\bm{\Lambda}=\textrm{diag}\left(\lambda_0,\;\lambda_1,\;\lambda_1,\;\dots,\;\lambda_{[T/2]},\;\lambda_{[T/2]}\right),
\end{equation}
with
\begin{equation}\label{eq:settembre}
\lambda_0=\sum_{|h|\leq [T/2]}\varsigma(h),\qquad
\lambda_j=
\sum_{|h|\leq [T/2]}\varsigma(h)\exp(ih\omega_j),\quad j=1,2,\dots [T/2].
\end{equation}
Let
$
\bm{p_i}=\left[p_{i1},p_{i2},\dots,p_{iT}\right]
$
and $\bm{e_i}$
denote the $i^{th}$ row of the matrices $\bm{P}$, $\bm{I_T}$, respectively. Hence we have to show that
\begin{equation}\label{eq:310117a}
\left|\bm{p_i} \bm{\mathfrak{C}}\bm{p_j}'-2\pi\bm{e_i}\bm{G}_{\varsigma}\bm{e_j}'\right|\leq
\left|\bm{p_i} \bm{\mathfrak{C}}^{(s)}\bm{p_j}'-2\pi\bm{e_i}\bm{G}_{\varsigma}\bm{e_j}'\right|+
\left|\bm{p_i} \bm{\mathfrak{C}}^{(s)}\bm{p_j}'-\bm{p_i} \bm{\mathfrak{C}}\bm{p_j}'\right|
\end{equation}
The first term on the right hand side of the above inequality
is bounded in absolute value by
\begin{equation}\label{eq:310117b}
\sum_{|h|>[T/2]}|\varsigma(h)|\leq \frac{2}{T^{\delta}}\sum_{|h|>[T/2]}h^{\delta}|\varsigma(h)|=O\left(\frac{1}{T^{\delta}}\right).
\end{equation}
For the second term we have
\begin{equation}\label{eq:310117c}
\left|
\bm{p_i\left(\bm{\mathfrak{C}}^{(s)}-\bm{\mathfrak{C}}\right)p_j'}\right|\leq \frac{4}{T}\left(2\sum_{h=1}^{[(T-1)/2]}h|\varsigma(h)|
+2\sum_{h=1}^{[(T-1)/2]}h|\varsigma(T-h)|\right)=O\left(\frac{1}{T}\right).
\end{equation}
Since both (\ref{eq:310117b}),(\ref{eq:310117c}) terms are independent of $i$ and $j$, the proof is completed.
\end{proof}
\begin{lemma}\label{020217b}
Let $\bm{P}$, $\bm{G_{\varsigma}}$ and $\bm{\Lambda}$ be as in (\ref{eq:gt}), (\ref{eq:peigenvectors}) and (\ref{eq:lambda17}), respectively.
Let $\lambda(\omega)=\sum_{|h|\leq [T/2]}\varsigma(h)\exp(ih\omega)$, with $\lambda(\omega_j)=\lambda_j$. Suppose that $\inf_\omega \lambda(\omega)>0$, $\inf_\omega g_{\varsigma}(\omega)>0$, and the conditions
assumptions of Lemma \ref{circulant} are satisfied with $\delta\geq2$.
Then, $\left\|\bm{P}'\bm{G_{\varsigma}}^{-1}\bm{P}\right\|_{row}<\infty$,
\end{lemma}
\begin{proof}
For sake of brevity, we consider only the case of $T$ odd, as in the proof of Lemma \ref{circulant}. Write
$$
\left\|\bm{P}'\bm{G_{\varsigma}}^{-1}\bm{P}\right\|_{row}\leq \left\|\left(\bm{\mathfrak{C}}^{(s)}\right)^{-1}\right\|_{row}+\left\|\left(\bm{\mathfrak{C}}^{(s)}\right)^{-1}-\bm{P}'\bm{G^{-1}_{\varsigma}}\bm{P}\right\|_{row},
$$
where $\bm{\mathfrak{C}}^{(s)}$ has been defined in (\ref{eq:CS}).
The inverse of $\bm{\mathfrak{C}}^{(s)}$ is circulant and symmetric, with eigenvalues $1/\lambda_j$, $j=0, \dots, \left[T/2\right]$, where $\lambda_j$ is defined as (\ref{eq:settembre}).
Because
\begin{equation}\label{eq:derivative}
\frac{d^2}{d\omega^2}\frac{1}{\lambda(\omega)}=-\frac{1}{\pi \lambda(\omega)^3}\sum_{|h|\leq [T/2]}h\varsigma(h)\sin(h\omega)-
\frac{1}{2\pi \lambda(\omega)^2}\sum_{|h|\leq [T/2]}h^2\varsigma(h)\cos(h\omega)
\end{equation}
is continuous in $\omega$, Lemma \ref{paolo} implies that $\varsigma(h)$ is absolutely summable, and $\left\|\left(\bm{\mathfrak{C}}^{(s)}\right)^{-1}\right\|_{row}<\infty$.
By the norm inequalities (6) and (12) in \cite[8.5.12, (12)]{ltk}
$$
\left\|\left(\bm{\mathfrak{C}}^{(s)}\right)^{-1}-\bm{P}'\bm{G^{-1}_{\varsigma}}\bm{P}\right\|_{row}
=
\left\|\bm{P}\left(\bm{\Lambda}^{-1}-\bm{G^{-1}_{\varsigma}}\right)\bm{P}'\right\|_{row}
\leq \sqrt{T}\left\|\bm{\Lambda}^{-1}-\bm{G^{-1}_{\varsigma}}\right\| =\left(\frac{1}{T^{\delta-\nicefrac{1}{2}}}\right),
$$
where the latter equality follows from(\ref{eq:310117b}).
\end{proof}
\begin{corollary}\label{circulantN}
Let $\varsigma_i(\cdot)$, $i=1,\dots,N$ a set of real functions defined on the integers. Assume that,
\begin{enumerate}[label=(\alph*)]
\item $\varsigma_i(\cdot)$ is symmetric, non-negative definite and $\sup_i\sum_{h=1}^{\infty}h^{\delta}\left|\varsigma_i(h)\right|<\infty$, with $\delta\geq 2$ , $\forall \;i$;
\item $m_g:=\inf_i\inf_{\omega}g_i(\omega)>0$, and $M_g:=\sup_i\sup_{\omega}g_i(\omega)<\infty$;
\end{enumerate}
where $g_{\varsigma_i}(\omega)=(2\pi)^{-1}\sum_{h=-\infty}^{\infty}\varsigma_i(h)\cos(h\omega)$, $-\pi\leq \omega\leq\pi$ and $m_g,M_g$ are positive constants.
Let $\bm{\mathcal{C}_{N}}=\left[\varsigma_N(i-j)\right]_{\ell,j=1}^T$, with $\varsigma_N(h)=N^{-1}\sum_{i=1}^N\varsigma_i(h)$. Define the function
\begin{equation}\label{eq:gtoN}
g_{\varsigma_N}(\omega)=\frac{1}{2\pi}\sum_{h=-\infty}^{\infty}\varsigma_N(h)\cos(h\omega),\qquad -\pi\leq \omega\leq\pi,
\end{equation}
and $T\times T$ diagonal matrix $G_{\varsigma_N}$ as in display(\ref{eq:gt}).
\begin{comment}
\begin{equation}\label{eq:gtN}
\bm{G_{NT}}=\left\{
\begin{array}{l}
\textrm{diag}\left\{g_N(0),g_N(\omega_1),g_N(\omega_1),\dots,g_N\left(\omega_{[T/2]}\right),g_N\left(\omega_{[T/2]}\right)\right\} \\
\textrm{if}\; T\;\textrm{is odd,}
\\
\\
\textrm{diag}\left\{g_N(0),g_N(\omega_1),g_N(\omega_1),\dots,g_N\left(\omega_{(T-2)/2}\right),g_N\left(\omega_{(T-2)/2}\right),g_N\left(\omega_{T/2}\right)\right\} \\\textrm{if}\; T\;\textrm{is even.}\\
\end{array}
\right.
\end{equation}
\end{comment}
Denote by $\Delta_{pq}^{(NT)} $ is the $p,q$ component of the matrix
$
{\bm P } \bm{\mathcal{C}_{N} } {\bm P }' - 2\pi\bm{G_{\varsigma_N}},
$
with $\bm{P}$ defined in (\ref{eq:peigenvectors}). Then,
\begin{enumerate}[label=(\roman)]
\item[(i)]
$
\sup_{1 \le \ell,j \le T } | \Delta_{\ell,j}^{(NT)}|=O(1/T), \quad\textrm{as}\quad N,T\to\infty , \qquad \ell,j=1,\dots,T.
$
\item[(ii)] $\left\|\bm{P'}\bm{G}^{-1}_{\varsigma_N}\bm{P}\right\|_{row}<\infty$.
\end{enumerate}
\end{corollary}
\begin{proof}
Tonelli's lemma imply that, under Assumption $(a)$, $\sum_h\varsigma_N(h)=N^{-1}\sum_i\sum_h\varsigma_i(h)$. The proof of part $(i)$ then follows very closely equations (\ref{eq:310117a})-(\ref{eq:310117c}).
Part $(ii)$ of the lemma is proved following the proof of Lemma \ref{020217b}.
A slight modification of the calculations in (\ref{eq:derivative}) show that $d^2\lambda^{-1}_{N}(\omega)/d\omega^2$ is continuous in $\omega$, with
$\lambda_N(\omega)=\sum_{|h|\leq [T/2]}\varsigma_N(h)\exp(ih\omega)$ .
\end{proof}