EconBase
← Back to paper

Interactive Effects Panel Data Models with General Factors and Regressors

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

206,230 characters · 13 sections · 104 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

\newtheorem{theorem}{Theorem}[section] \newtheorem{corollary}{Corollary}[section] \newtheorem{lemma}{Lemma}[section] \theoremstyle{definition} \newtheorem{assumption}{Assumption} \newtheorem{remark}{Remark} \newtheorem{step}{Step} \newtheorem{dgp}{DGP}

\numberwithin{corollary}{section} \numberwithin{equation}{section} \numberwithin{lemma}{section} \numberwithin{remark}{section} \numberwithin{theorem}{section}

\allowdisplaybreaks[4]

titlepage{ \begin{center} { \bf Interactive Effects Panel Data Models with General Factors and Regressors \begingroup \footnote{Previous versions of the paper were presented at seminars at Aarhus University, Lund University and Michigan State University. Westerlund would like to thank seminar participants, and in particular Richard Baillie, Nicholas Brown, David Edgerton, Yousef Kaddoura, Yana Petrova, Simon Reese, Tim Vogelsang, Jeffrey Wooldridge, and Morten {\O}rregaard Nielsen. Westerlund thanks the Knut and Alice Wallenberg Foundation for financial support through a Wallenberg Academy Fellowship. Peng and Westerlund also acknowledge the Australian Research Council Discovery Grants Program for its financial support under grant number DP210100476. Correspondence: Department of Econometrics and Business Statistics, Monash University, Australia. Email: [email removed]. } \addtocounter{footnote}{-1} \endgroup } {\sc Bin Peng$^{\star}$, Liangju Su$^\ddagger$, Joakim Westerlund$^{\sharp}$ and Yanrong Yang$^{\dag}$} $^\star$Monash University, $^\ddagger$Tsinghua University, $^\sharp$Lund University, $^{\dag}$Australian National University \today \begin{abstract} This paper considers a model with general regressors and unobservable factors. An estimator based on iterated principal components is proposed, which is shown to be not only asymptotically normal and oracle efficient, but under certain conditions also free of the otherwise so common asymptotic incidental parameters bias. Interestingly, the conditions required to achieve unbiasedness become weaker the stronger the trends in the factors, and if the trending is strong enough unbiasedness comes at no cost at all. In particular, the approach does not require any knowledge of how many factors there are, or whether they are deterministic or stochastic. The order of integration of the factors is also treated as unknown, as is the order of integration of the regressors, which means that there is no need to pre-test for unit roots, or to decide on which deterministic terms to include in the model. \end{abstract} \end{center} {\em Keywords}: Panel data, non-stationarity, principal components, interactive effects. {\em JEL classification}: C14, C23, C51 }

Introduction

The use of panel data in regression analysis has attracted considerable attention in the empirical literature in economics and elsewhere. A major reason for this is the ability to deal with the presence of unobserved heterogeneity, and the problem that this causes when said heterogeneity is correlated with the regressors. The simplest approach is to assume that the unobserved heterogeneity is made up of additive individual- and time-specific constants, or “fixed effects”. Such effects can add hugely to the fit of the model, and are typically found to be much more important than the regressors predicted by economic theory. However, in many applications fixed effects are unlikely to be enough to capture the unobserved heterogeneity (see, for example, Lemmonetal2008, and DeAngeloRoll2015), and this has in turn spurred much interest in so-called “interactive effects” models, in which the individual and time effects enter in an multiplicative way. One of the most common approaches to such interactive effects models by far is the principal components (PC) approach of Bai. In fact, it is so common that it has given rise to a separate strand of literature (see Ando, bai2014, LiQianSu, MoonWeidner2015, to mention a few).\footnote{While popular, PC is not the only approach that can be used to estimate interactive effects models. One alternative is the so-called “common correlated effects” (CCE) estimator of Pesaran2006. The problem with this approach is that it requires that the number of time effects, or factors, is bounded by the number of observables, and that the regressors load on the same factors as the dependent variable, which is not necessary in PC. There is also the generalized method of moments (GMM) approach of ALS13. However, this approach supposes that the number of time periods is “small”, and is not suitable for the type of large panel data sets that we have in mind.} The present paper aims to contribute to this strand, and it does so in at least three ways.

The first contribution of the paper is to consider a very general data generating process (DGP) that includes most of the specifications considered previously in the literature as special cases. The only requirement is that suitably normalized sample second moment matrices of the factors and regressors have positive definite limits. This is noteworthy because the existing literature is almost exclusively based on the assumption that both the factors and regressors are stationary. The only exceptions known to us are BaiKaoNg, DGP2020 and HJPS, but they assume instead that either the regressors or the factors and the regressors are pure unit root processes, which is also not realistic. Indeed, regressors and factors of different order of magnitude are likely to be the rule rather than the exception, especially in economic and financial data, due to differences in persistence over time.

The unrestricted DGP is important in itself but also because it can be accommodated without requiring any knowledge thereof. Hence, not only do we treat the factors and their number as unknown, but we also do not require any knowledge of the order of magnitude of both factors and regressors. An important implication of this is that there is no need to distinguish between deterministic and stochastic factors, or stationary and non-stationary factors. In the existing literature, deterministic factors are often treated as known, and are projected out prior to the application of PC (see, for example, MoonWeidner2015). The problem here is that there is typically great uncertainty over which deterministic terms to include, which raises the issue of model misspecification. The fact that in the present paper deterministic terms are treated as additional factors means that the problem of deciding on which terms to include does not arise. Similarly, while the regressors can be tested for unit roots, and the estimation can be made conditional on the test outcome, this raises the issue of pre-testing bias. In the present paper we do not require any knowledge about the order of integration of the regressors, which means that there is no need for any pre-testing.

Equally as important as the general model formulation and its empirical appeal is the extension of the existing econometric theory, which has not yet ventured much outside the stationary or pure unit root environments. This is our second major contribution. The main difficulty here is not the unrestricted specification of the factors and regressors per se, but rather that the order of magnitude of the factors may differ. In particular, the problem is that the nonlinearity of the PC estimator distorts the signal coming from the factors, just as it does in estimation of nonlinear regression models with mixtures of integrated regressors (see, for example, ParkPhillips2000). This is true if both the number and order of the factors are known, and the problem does not become any simpler when these quantities are treated as unknown, as they are here. An additional problem that then arises is that existing studies on the selection of the number of factors all require that the data are stationary (see, for example, BaiNg2002, and AH13), and it is not obvious how one should go about this when the order of magnitude of the factor is unknown.

Intuitively, the factors whose order is largest should dominate the PC estimator. This motivates the use of an iterative estimation procedure in which the factors and their number are estimated in order according to their magnitude with relatively larger factors being estimated first. We begin by prescribing a large number of factors, and estimate the resulting model by PC. The estimated factors only capture the most dominating factors whose order of magnitude is largest. In spite of this, we can show that the estimator is consistent, albeit at a relatively low rate of convergence. The rate is, however, high enough to ensure that the number of dominating factors can be consistently estimated using a version of the eigenvalue ratio approach of LY12, and AH13. We then apply PC conditional on the first-step factor estimates, and estimate the second most dominating factors. This procedure continues until we cannot identify any more factors, and we can show that both the estimated factors and their number are consistent. Because of the iterative fashion in which the factors are estimated, we refer to the new estimation procedure as “iterative PC” (IPC), which is shown to be asymptotically normal and “oracle efficient”.

Our third contribution is to point to a “blessing” of trending factors. The blessing occurs if the magnitude of the factors is sufficiently large, in which case the otherwise so common asymptotic bias of the PC approach can be completely eliminated without imposing any additional restrictions on the cross-sectional and time series dependencies of the regression errors. This is noteworthy, because the conclusion made in the previous literature suggests that in order to eliminate the asymptotic bias, the errors have to be independent.

The reminder of the paper is organized as follows. We begin by describing the model that we will be considering and the proposed IPC approach that we will use to estimate it. This is done in Section (ref). Section (ref) presents the formal assumptions and our main asymptotic results, whose small-sample accuracy is evaluated using Monte Carlo simulations in Section (ref). Section (ref) reports the results obtained from an empirical application to the long-run relationship between US house prices and income. For the sake of space, we present an extra empirical study on the returns to scale in the US banking industry in the supplementary appendix. Section (ref) concludes. All proofs are relegated to the online appendix.

Model and estimation procedure

Consider the panel data variable $y_{i,t}$, observable for $i=1,\ldots,N$ cross-sectional units and $t=1,\ldots,T$ time periods. The model of this variable that we will be considering is given by

eqnarray[eqnarray omitted — 125 chars of source]

where $\*x_{i,t} = (x_{1,i,t},\ldots,x_{d_x,i,t})^{\prime }$ is a $d_x\times 1$ vector of regressors, $\*f_{t}^0=(f_{1,t}^{0},\ldots,f_{d_f,t}^{0})^{\prime }$ is a $d_f\times 1$ vector of unobservable common factors with $\+\gamma_{i}^0 = (\gamma_{1,i}^0,\ldots,\gamma_{d_f,i}^0)^{\prime }$ being a conformable vector of factor loadings, and $\varepsilon_{i,t}$ is an idiosyncratic error term. Moreover, the factors are divided into groups according to their order of magnitude. There are $G$ groups of size $d_1,\ldots,d_G$, which means that $d_1+\cdots+d_G = d_f$. Because the grouping is unknown, we may without loss of generality assume that the factors are ordered, such that the first $d_1$ factors have the highest order of magnitude, the next $d_2$ factors have the second highest order, and so on. Hence, if we denote by $\*f_{g,t}^0$ and $\+\gamma_{g,i}^0$ the $d_g\times 1$ vectors of factors and loadings associated with group $g$, respectively, then $\+\gamma_{i}^{0\prime} \*f_{t}^0 = \sum_{g=1}^G \+\gamma_{g,i}^{0\prime}\*f_{g,t}^0$, where $\*f_{t}^0 =(\* f_{1,t}^{0\prime},\ldots,\*f_{G,t}^{0\prime})^{\prime}=(f_{1,t}^{0},\ldots,f_{d_f,t}^{0})^{\prime }$ and $\+\gamma_{i}^0 = (\+\gamma_{1,i}^{0\prime},\ldots,\+\gamma_{G,i}^{0\prime})^{\prime }= (\gamma_{1,i}^0,\ldots,\gamma_{d_f,i}^0)^{\prime }$. If $d_f = 0$, then $G=0$ and $\+\gamma_{i}^{0\prime} \*f_{t}^0 = 0$.

remarkThe (implicit) assumption of a common slope coefficient $\+\beta^0$ is standard in the literature (see, for example, Bai, and MoonWeidner2015). We speculate that our results can be extended to accommodate heterogeneous slope coefficients as in, for example, Castagnetti2015, Castagnetti2019, LCL2020, and HJPS.

It is useful to write (ref) on stacked vector form. Let us therefore introduce the $T\times 1$ vectors $\*y_i =(y_{i,1},\ldots, y_{i,T})^{\prime }$ and $\+\varepsilon_i = (\varepsilon_{i,1},\ldots, \varepsilon_{i,T})^{\prime }$, the $T\times d_f$ matrix $\*F^0 = (\*f_{1}^0,\ldots,\*f_{T}^0)^{\prime }$, and the $T\times d_x$ matrix $\*X_i =(\*x_{i,1},\ldots, \*x_{i,T})^{\prime }$. Analogous to $\*f_{t}^0$, $\*F^0$ is further partitioned as $\*F^0 =(\*F_{1}^{0},\ldots,\*F_{G}^{0})^{\prime }$, where the $T\times d_g$ matrix $\*F^0_g = (\*f_{g,1}^0,\ldots,\*f_{g,T}^0)^{\prime }$ contains the stacked factor observations for group $g$. In this notation, (ref) can be written as

eqnarray[eqnarray omitted — 87 chars of source]

The goal of this paper is to infer $\+\beta^0$. As mentioned in Section (ref), however, because of the generality of the model being considered, the main difficulty in the estimation process is how to control for $\*F^0$. Our proposed estimation procedure consists of three steps. We first initialize the estimation procedure by applying the PC estimator of Bai. However, because the first group of factors dominates all the other groups in terms of order of magnitude, the first-step PC factor estimator will only be estimating (the space spanned by) $\*F_{1}^0$. The second step of the procedure therefore involves iteratively applying PC conditional on previous factor estimates to estimate all subsequent groups of factors; hence, the “I” in IPC. In the third and final step, we estimate $\+\beta^0$ conditional on the second-step IPC estimator of $\*F^0$ and the first-step PC estimator of $\+\beta^0$.

step[Initial estimation]\ The (concentrated) ordinary least squares (OLS) objective function that we consider is the same as in Bai. It is given by \begin{eqnarray} \mathrm{SSR}(\+\beta,\*F)= \sum_{i=1}^N (\*y_i-\*X_i\+\beta)^{\prime }\*M_F(\*y_i-\*X_i\+\beta), \end{eqnarray} where $\*F$ is a $T\times d_{max}$ matrix satisfying $T^{-\delta}\*F^{\prime}\*F = \*I_{d_{max}}$, $d_{max} \in \mathbb{N} \ge d_f$ and $\delta\in[0,\infty)$ are user-specified numbers, and $\*M_A=\*I_T- \*A(\*A^{\prime }\*A)^{-1}\*A^{\prime }=\*I_T- \*P_A$ for any $T$-rowed full column rank matrix $\*A$. As we explain in Remark (ref), the IPC estimator of $\+\beta^0$ is invariant to the choice of $\delta$ and the need to select $d_{max}$ is standard. The initial estimator that we consider is the minimizer of $\mathrm{SSR}(\+\beta,\*F)$: \begin{eqnarray} (\widehat{\+\beta}_0, \widehat{\*F}_0) = \operatorname*{\arg\!\min}_{(\+\beta,\*F)\in \mathbb{D}}\, \mathrm{SSR}(\+\beta,\*F), \end{eqnarray} where $\mathbb{D} = \mathbb{R}^{d_x} \times \mathbb{D}_F $ and $\mathbb{D}_F= \{ \*F : T^{-\delta}\*F^{\prime }\*F = \*I_{d_{max}} \}$. It is useful to note that $\widehat{\+\beta}_0$ satisfies $\widehat{\+\beta}_0 = \widehat{\+\beta}(\widehat{\*F}_0)$, where \begin{align} \widehat{\+\beta}(\*F) = \left(\sum_{i=1}^N \*X_i^{\prime }\*M_{F}\*X_i\right)^{-1} \sum_{i=1}^N \*X_i^{\prime }\*M_{F}\*y_i. \end{align}
step[Iterative estimation of factors]\ As already pointed out, the factors are estimated in order according to magnitude. Therefore, $\widehat{\*F}_0$ is estimating $\*F_1^0$. Since $d_1 \leq d_f \leq d_{max}$, in general the dimension of $\widehat{\*F}_0$ will be larger than that of $\*F_1^0$. We therefore begin this step of the estimation procedure by estimating $d_1$, and for this purpose we employ a version of the ratio of eigenvalue-based estimator considered by, for example, LY12, and AH13, which is given by \begin{eqnarray} \widehat{d}_1 = \operatorname*{\arg\!\min}_{0\leq d\le d_{max}}\left\{ \frac{\widehat{\lambda}_{1,d+1}}{\widehat{\lambda}_{1,d}}\cdot \mathbb{I}\left( \frac{\widehat{\lambda}_{1,d}}{\widehat{\lambda}_{1,0}}\ge \tau_{N}\right) + \mathbb{I}\left( \frac{\widehat{\lambda}_{1,d}}{\widehat{\lambda}_{1,0}} <\tau_{N}\right)\right\}, \end{eqnarray} where $\mathbb{I}(A)$ is the indicator function for the event $A$ taking the value one if $A$ is true and zero otherwise, $\tau_{N} = 1/\ln\, (\max\{\widehat{\lambda}_{1,0},N\})$, $\widehat{\lambda}_{1,0} = N^{-1}\sum_{i=1}^N (\*y_i -\*X_i\widehat{\+\beta}_0)^{\prime }\*M_{\widehat{F}_0}(\*y_i -\*X_i \widehat{\+\beta}_0)$ and $\widehat{\lambda}_{1,1}\geq \cdots \geq \widehat{\lambda}_{1,d_{max}}$ are the $d_{max}$ largest eigenvalues of the following $T\times T$ matrix: \begin{eqnarray} \widehat{\+\Sigma}_1 = \frac{1}{N}\sum_{i=1}^N (\*y_i -\*X_i\widehat{\+\beta}_0)(\*y_i -\*X_i\widehat{\+\beta}_0)^{\prime }. \end{eqnarray} The threshold $\tau_{N}$, the “mock” eigenvalue $\widehat{\lambda}_{1,0}$, and the indicator function are there to ensure that the estimator is consistent. The need for these will be explained later. Given $\widehat d_1$, we update the estimate of $\*F_{1}^0$ by setting $\widehat{\*F}_1$ equal to $T^{\delta/2}$ times the eigenvectors associated with $\widehat{\lambda}_{1,1},\ldots, \widehat{\lambda}_{1,\widehat{d}_1}$, and estimate $\+\gamma_{1,i}^0$ by $\widehat{\+\gamma}_{1,i} = T^{-\delta}\widehat{\*F}_1^{\prime }(\*y_i -\*X_i\widehat{\+\beta}_0)$. The estimation of $\*F_2^0,\ldots,\*F_G^0$ is analogous to that of $\*F_1^0$. The main difference is that we have to condition on all previous estimates. Let us therefore use $\widehat{\*F}_{-g} = (\widehat{\*F}_1,\ldots,\widehat{\*F}_{g-1})$ and $\widehat{\+\gamma}_{-g,i} = (\widehat{\+\gamma}_{1,i}^{\prime },\ldots,\widehat{\+\gamma}_{g-1,i}^{\prime})^{\prime }$ to denote the matrices containing the previously estimated factors and loadings, respectively, when estimating group $g > 1$. The estimator of $d_g$ is given quite naturally by \begin{eqnarray} \widehat{d}_g = \operatorname*{\arg\!\min}_{0\leq d\le d_{max}}\left\{ \frac{\widehat{\lambda}_{g,d+1}}{\widehat{\lambda}_{g,d}}\cdot \mathbb{I}\left( \frac{\widehat{\lambda}_{g,d}}{\widehat{\lambda}_{g,0}}\ge \tau_{N}\right) + \mathbb{I}\left( \frac{\widehat{\lambda}_{g,d}}{\widehat{\lambda}_{g,0}} < \tau_{N}\right)\right\}, \end{eqnarray} where $\widehat{\lambda}_{g,0} = N^{-1}\sum_{i=1}^N (\*y_i -\*X_i\widehat{\+\beta}_0)^{\prime }\*M_{\widehat{F}_{-g}}(\*y_i -\*X_i\widehat{\+\beta}_0)$ and $\widehat{\lambda}_{g,1}\geq \cdots \geq \widehat{\lambda}_{g,d_{max}}$ are the $d_{max}$ largest eigenvalues of \begin{eqnarray} \widehat{\+\Sigma}_g = \frac{1}{N}\sum_{i=1}^N (\*y_i -\*X_i\widehat{\+\beta}_0 - \widehat{\*F}_{-g}\widehat{\+\gamma}_{-g,i})(\*y_i -\*X_i\widehat{\+\beta}_0- \widehat{\*F}_{-g}\widehat{\+\gamma}_{-g,i})^{\prime }. \end{eqnarray} The resulting estimator $\widehat{\*F}_g$ of $\*F_{g}^0$ is given by the eigenvectors associated with $\widehat{\lambda}_{g,1},\ldots, \widehat{\lambda}_{g,\widehat{d}_g}$ and $\widehat{\+\gamma}_{g,i} = T^{-\delta}\widehat{\*F}_g^{\prime }(\*y_i -\*X_i\widehat{\+\beta}_0 - \widehat{\*F}_{-g}\widehat{\+\gamma}_{-g,i})$. New groups of factors are estimated until $\widehat{d}_{g} =0$. At this point, we set $\widehat G = g-1$ and define $\widehat{\*F} = (\widehat{\*F}_1,\ldots,\widehat{\*F}_{\widehat{G}})$. This is the IPC estimator of $\*F^0$.
step[Estimation of $\+\protect\beta^0$]\ Given $\widehat{\*F}$, we compute $\widehat{\+\beta}_1 = \widehat{\+\beta}(\widehat{\*F})$ using (ref). The IPC-based estimator of $\+\beta^0$ is given by \begin{eqnarray} \widehat{\+\beta} = \widehat{\+\beta}_0+\left(\sum_{i=1}^N \widehat{\*Z}_i^{\prime }\widehat{\*Z}_i \right)^{-1}\sum_{i=1}^N \*X_i^{\prime }\*M_{\widehat{F}}\*X_i ( \widehat{\+\beta}_1-\widehat{\+\beta}_0) , \end{eqnarray} where $\widehat{\*Z}_i = \*M_{\widehat{F}}\*X_i - \sum_{j=1}^{N}\*M_{\widehat{F}}\*X_j\widehat a_{ij}$ with $\widehat a_{ij} = \widehat{\+\gamma}_{i}^{\prime }(\widehat{\+\Gamma}^{\prime }\widehat{\+\Gamma})^{-1}\widehat{\+\gamma}_{j}$, $\widehat{\+\gamma}_i =(\widehat{\+\gamma}_{1,i}^{\prime},\ldots,\widehat{\+\gamma}_{\widehat G,i}^{\prime })^{\prime }$ and $\widehat{\+\Gamma} =(\widehat{\+\gamma}_{1},\ldots,\widehat{\+\gamma}_{N})^{\prime }$.
remarkThe reason for why in Steps (ref) and (ref) $\delta$ can be set arbitrarily is that the only use of this parameter is to normalize $\widehat{\*F}$ and $T^{-\delta/2}\|\widehat{\*F}_0\| = \sqrt{d_{max}}$ for any $\delta\in [0,\infty)$. This result is nice from both applied and theoretical points of view. In the bulk of the previous literature, the appropriate value of $\delta$ to use depends on whether $\* F^0$ is stationary or unit root non-stationary (see, for example, Bai2004). The assumed knowledge of $\delta$ is therefore tantamount to assuming that the order of integration of $\*F^0$ is known, which is again not a requirement here. The need to specify a maximum $d_{max}$ for the number of factors is standard in the literature (see, for example, BaiNg2002). In Sections (ref) and (ref), we explain how $d_{max}$ is set in the Monte Carlo and empirical studies, respectively.
remarkMany papers use information criteria to select the number of factors (see, for example, BaiNg2002). We do not. The reason is that the eigenvalue ratio is self-normalizing, which makes it possible to handle factors that are of different order of magnitude. The general idea behind the eigenvalue ratio approach in Step (ref) is similar to the one used in ocular inspection of “scree plots"; one looks for the point at which the ordered eigenvalues drop substantially, and set the number of estimated factors accordingly. The most natural way to mimic this decision rule is to minimize $\widehat{\lambda }_{g,d+1}/\widehat{\lambda }_{g,d}$ over $1\leq d\leq d_{max}$. However, this raises two issues. One is that since $\widehat{\lambda }_{g,d+1}/\widehat{\lambda }_{g,d}$ is not defined for $d=0$, we cannot have $d_{f}=0$, and we want to be able to entertain the possibility that there are no factors. The use of the mock eigenvalue $\widehat{\lambda }_{g,0}$ allows us to do just that. The other problem is that under the conditions of this paper (laid out in Section (ref)) the limiting behavior of $\widehat{\lambda }_{g,d+1}/\widehat{\lambda }_{g,d}$ when $d>d_{g}$ is unknown. LY12 discuss this issue at length. They conjecture that $\widehat{\lambda }_{g,d+1}/\widehat{\lambda }_{g,d}\asymp 1$, where $a\asymp b$ means that $a=O_{P}(b)$ and $b=O_{P}(a)$, but there is no proof. The challenge when studying $\widehat{\lambda }_{g,d+1}/\widehat{\lambda }_{g,d}$ for $d>d_{g}$ is to bound this ratio from below, as both of the numerator and denominator converge to zero for a certain choice of $\delta $ at the same rate. The use of the indicator function allows us to circumvent this problem. The idea is to look at $\widehat{\lambda }_{g,d}$ only. If this eigenvalue is “small", we take it as a sign of $d>d_{g}$ and set $\widehat{\lambda }_{g,d+1}/\widehat{\lambda }_{g,d}$ to one. However, because the order of $\*f_{t}^{0}$ is assumed to be unknown, we cannot look at $\widehat{\lambda }_{g,d}$ directly but rather we look at $\widehat{\lambda }_{g,d}/\widehat{\lambda }_{g,0}$, which in contrast to $\widehat{\lambda }_{g,d}$ is self-normalizing. BT2021 also look at ratios of eigenvalues as a way to determine the number of factors when their order is unknown. However, they only consider factors that are either trending linearly, unit root non-stationary or stationary, and the implementation of their procedure is quite involved. We therefore do not consider it here.
remarkIntuition suggests to take $\widehat{\+\beta}_1$, the OLS estimator conditional on $\widehat{\*F}$, as the final estimator of $\+\beta^0$ in Step (ref). Interestingly, while consistent, because of the stepwise estimation of the factors, the asymptotic distribution of $\widehat{\+\beta}_1$ is generally not (mixed) normal and nuisance parameter-free. In fact, as we show later in Lemma (ref), even $\widehat{\+\beta}_0$ is consistent but again with non-normal asymptotic distribution. In Section (ref), we use Monte Carlo simulations to shed some light on the consistency and non-normality of $\widehat{\+\beta}_0$ and $\widehat{\+\beta}_1$.

Assumptions and asymptotic results

Before we state our assumptions and asymptotic results, we introduce some notation. Specifically, if $\*A$ is a matrix, $\lambda_{min}(\*A)$ and $\lambda_{max}(\*A)$ signify its smallest and largest eigenvalues, respectively, $\text{tr}\,\*A$ signifies its trace, and $\|\*A\| = \sqrt{\text{tr}\,\*A'\*A}$ and $\|\*A\|_2 = \sqrt{\lambda_{max}(\*A)}$ signify its Frobenius and spectral norms, respectively. We write $\*A > 0$ to signify that $\*A$ is positive definite. If $\*B$ is also a matrix, then $\mathrm{diag}(\*A, \*B)$ denotes the block-diagonal matrix that takes $\*A$ ($\*B$) as the upper left (lower right) block. The symbols $\to_D$, $\to_P$ and $MN(\cdot, \cdot)$ signify convergence in distribution, convergence in probability and a mixed normal distribution, respectively. We use $N,\,T\to\infty$ to indicate that the limit has been taken while passing both $N$ and $T$ to infinity. We use w.p.a.1 to denote with probability approaching one.

Assumption (ref) is concerned with the order of magnitude of $\*f_t^0$ and $\*x_{i,t}$. It is therefore key. The assumption is stated in terms of the required moment conditions rather than primitive assumptions on $\*f_t^0$ and $\*x_{i,t}$. The reason is that we would like to be agnostic about these variables.

assumption[Moments]\ \begin{itemize} • There exists a matrix $\+\Sigma_X$ such that $E\|(NT)^{-1}\sum_{i=1}^N \*D_T \*X_i^{\prime }\*M_{F^0}\*X_i \*D_T -\+\Sigma_X \|^2 =o(1)$, where $\*D_T = \mathrm{diag}( T^{ -\kappa_1/2} ,\ldots, T^{-\kappa_{d_x}/2 } )$ with $0\le \kappa_j< \infty$ for $j=1,\ldots,d_x$, $E\|\+\Sigma_X\|^2<\infty$, and $0<\lambda_{min}(\+\Sigma_X) \le \lambda_{max}(\+\Sigma_X) <\infty$ w.p.a.1. • $\|(NT)^{-1}\sum_{i=1}^N \*D_T\*X_i^{\prime }\+\varepsilon_i\| =O_P ( 1/\min\{\sqrt{N}, \sqrt{T}\})$ and $\|\+\varepsilon\|_{2} =O_P (\max\{\sqrt{N}, \sqrt{T}\} )$ with $\+\varepsilon = (\+\varepsilon_1,\ldots,\+\varepsilon_N)$. • There exists a matrix $\+\Sigma_{F^0}$ such that $E\|\*C_T \*F^{0\prime}\*F^0 \*C_T- \+\Sigma_{F^0} \|^2 =o(1)$, where $\*C_T = \mathrm{diag}( T^{ -\nu_1/2} \*I_{d_1} ,\ldots, T^{-\nu_G/2 } \*I_{d_G})$ with $\nu_1>\cdots >\nu_G> 1/2$, $E\| \+\Sigma_{F^0}\|^2<\infty$ and $0<\lambda_{min}(\+\Sigma_{F^0} )\le \lambda_{max}(\+\Sigma_{F^0} )<\infty$ w.p.a.1. • There exists a matrix $\+\Sigma_{\Gamma^0}$ such that $\|N^{-1}\+\Gamma^{0\prime}\+\Gamma^0- \+\Sigma_{\Gamma^0}\| =o_P(1)$ and $\max_{i\ge 1}E\| \+\gamma_{i}^0\|^4<\infty$, where $\+\Gamma^0 =(\+\gamma_{1}^0,\ldots,\+\gamma_{N}^0)^{\prime }$ is $N\times d_f$ and $ 0<\lambda_{min}(\+\Sigma_{\Gamma^0} )\le \lambda_{max}(\+\Sigma_{\Gamma^0})<\infty$. \end{itemize}

Assumption (ref) (a) is very general in that it imposes almost no restrictions on the type of trending behaviour that $\*x_{i,t}$ may have. The trending can be deterministic but it can also be stochastic, as in the presence of unit roots. Either way, the degree of the trending is not restricted in any way, provided that it is finite. The only requirement is that $0<\lambda _{min}(\+\Sigma _{X})\leq \lambda _{max}(\+\Sigma _{X})<\infty$ w.p.a.1, which means that the elements of $\*x_{i,t}$ cannot be asymptotically collinear. Note that $\+\Sigma _{X}$ is not required to be a constant matrix, as this would rule out regressors that are stochastically integrated. This is very different when compared to the bulk of the previous PC-based interactive effects literature in which $\*x_{i,t}$ is assumed to be either stationary (see, for example, Bai, and MoonWeidner2015), such that $\kappa _{1}=\cdots =\kappa _{d_{x}}=0$, or unit-root non-stationary (see BaiKaoNg, and DGP2020), such that $\kappa _{1}=\cdots =\kappa _{d_{x}}=1$. Assumption (ref) (a) is analogous to the work of, for example, DongLinton, wherein time series of different order of magnitude are considered.

The first requirement of Assumption (ref) (b) is quite mild and holds if a central limit theorem in only one of the two panel dimensions applies to the normalized sum of $\mathbf{D}_{T}\mathbf{X}_{i}^{\prime }\+\varepsilon_{i}$. The second requirement is quite common in the literature, and is expected to hold as long as $\varepsilon_{i,t}$ has zero mean, and weak serial and cross-sectional correlation (see MoonWeidner2015, for a discussion).

Assumption (ref) (c) is similar to (a) in that it leaves the trending behavior of the factors essentially unrestricted. A majority of previous PC-based works assume that $T^{-1}\*F^{0\prime }\*F^{0}$ converges to positive definite matrix (see, for example, Bai, and MoonWeidner2015). Notable exceptions include Bai2004, BaiKaoNg, and Choi2017, in which $\*f_{t}^{0}$ is assumed to follow pure unit root process, and BaiNg2004, who allow for a mix of stationary and unit root factors. The only study that comes close to ours in terms of the generality of the factors is that of Westerlund2018. However, he assumes that $\*x_{i,t}$ has a factor structure that loads on the same set of factors as $y_{i,t}$, which is not required here. Also, unlike Westerlund2018, we do not require $\nu _{G}\geq 1$, but also allow $1/2<\nu _{G}<1$, such that $T^{-1}\*F_{G}^{0\prime }\*F_{G}^{0}$ is asymptotically singular. The signal coming from $\*F_{G}^{0}$ is therefore even weaker than under stationarity, and we will therefore refer to these factors as “signal-weak", as opposed to “weak" in usual sense (see, for example, Chudiketal2011 and UY2020, for discussions). One example of such signal-weak factors is when $\*f_{G,t}^{0}$ is stationary and sparse.

Assumption (ref) (d) is standard and ensures that each factor has a nontrivial contribution to the variance of $y_{i,t}$.

assumption[Identification]\ $\inf_{\*F\in \mathbb{D}_F} \lambda_{min}(\*B(\*F))\geq c_0 >0$ for all $N$ and $T$, where \begin{align} \*B(\*F) = \frac{1}{NT}\sum_{i=1}^{N}\*D_T \*Z_i(\*F)^{\prime }\*Z_i(\*F)\*D_T \end{align} with $\*Z_i(\*F) = \*M_{F}\*X_i - \sum_{j=1}^{N}\*M_F\*X_ja_{ij}$ and $a_{ij} =\+\gamma_{i}^{0\prime}(\+\Gamma^{0\prime}\+\Gamma^0)^{-1}\+\gamma_{j}^0$.

The requirement that the smallest eigenvalue of $\*B(\*F)$ should be positive is equivalent to requiring that $\*B(\*F)$ be positive definite uniformly in $\*F$, which is a non-collinearity condition that rules out low-rank regressors (see MoonWeidner2015, for a discussion). It demands that the regressors in $\*X_{i}$ have enough variation after projecting out all variation that can be explained by $\*F^{0}$ and $\+\gamma _{i}^{0}$. Time-invariant regressors are therefore not allowed, which is similar to the conventional cross-section fixed effects-only OLS condition. Assumption (ref) is a high-level condition, just like Assumption (ref) (see, for example, Bai, and MoonWeidner2015, for similar high-level conditions). The reason for formulating it in this way is again that we want to avoid making highly specific and potentially invalid assumptions about the DGP of $\*x_{i,t}$, $\*f_{t}^{0}$ and $\+\gamma _{i}^{0}$. If all the regressors are stationary, Assumption (ref) reduces to Assumption A in Bai.

Assumptions (ref) and (ref) are enough to ensure that the initial estimator $\widehat{\+\beta}_0$ of $\+\beta^0$ is consistent.

lemma[Consistency of $\protect\widehat{\+\protect\beta}_0$]\ Under Assumptions (ref) and (ref), as $N,T\to \infty$, \begin{eqnarray} \min\{\sqrt{N}, \sqrt{T}\}\*D_T^{-1}(\widehat{\+\beta}_0-\+\beta^0) =O_P(1). \end{eqnarray}

Lemma (ref) establishes that $\widehat{\+\beta}_0$ is consistent for $\+\beta^0$ and that the rate of convergence is $\|\*D_T\|/\min\{\sqrt{N}, \sqrt{T}\} = \max\{T^{-\kappa_1/2}, \ldots, T^{-\kappa_{d_x}/2}\}/\min\{\sqrt{N}, \sqrt{T}\}$. To put this into perspective, suppose that $\*x_{i,t}$ is stationary, such that $\kappa_1 = \cdots = \kappa_{d_x} = 0$. In this case, $\*D_T = \*I_{d_x}$ and the rate of convergence is given by $1/\min\{\sqrt{N}, \sqrt{T}\}$, which is the slowest of the regular rates in pure time series and cross-section regressions. Still, the rate is fast enough for the estimation of the number of factors. This brings us to Step (ref) of the estimation procedure. In order to be able to show that $\widehat d_1,\ldots,\widehat d_{G+1}$ and $\widehat{\*F}$ are consistent, however, we need more structure.

assumption[Errors]\ \begin{itemize} • $E(\varepsilon_{i,t})=0$ and $E(\+\varepsilon_{i}\+\varepsilon_i^{\prime })=\+\Sigma_{\varepsilon,i} $. • Let $\+\varepsilon_{t} = (\varepsilon_{1,t},\ldots,\varepsilon_{N,t})^{\prime }$ in this assumption only. $\{\+\varepsilon_{t} : \ t\ge 1\}$ is strictly stationary and $\alpha$-mixing such that $\max_{i\geq 1}E|\varepsilon_{i,1}|^{4+\mu} <\infty$ for some $\mu >0$ and the mixing coefficient $\alpha(t) = \sup_{A\in \mathcal{F}_{-\infty}^0, B\in \mathcal{F}_t^\infty} |P(A)P(B) -P(AB)|$ satisfies $\sum_{t=1}^\infty \alpha(t)^{\mu/(2+\mu)}<\infty$, where $\mathcal{F}_{-\infty}^t$ and $\mathcal{F}_{t}^\infty$ are the sigma-algebras generated by $\{\+\varepsilon_{s}: s \leq 0\}$ and $\{\+\varepsilon_{s}: s \geq t\}$, respectively. • Suppose that $\sum_{i=1}^N\sum_{j=1}^N\sum_{t=1}^T\sum_{s=1}^T|E\varepsilon_{i,t}\varepsilon_{j,s}|=O(NT)$ and $\sum_{i=1}^N \sum_{j\ne i}|\sigma_{\varepsilon,ij}|=O(N)$, where $\sigma_{\varepsilon,ij}=E(\varepsilon_{i,t}\varepsilon_{j,t})$. • $\varepsilon_{i,t}$ is independent of $\+\gamma_j^0$, $\*f_s^0$ and $\*x_{j,s}$ for all $i$, $j$, $t$ and $s$. \end{itemize}

Assumption (ref) is similar to Assumptions C and D in Bai. Assumption (ref) (b) and (c) ensure that the serial and cross-sectional dependencies of $\varepsilon_{i,t}$ are at most weak. Assumption (ref) (d) requires that the regressors are exogenous, which rules out the presence of lagged dependent variables in $\*x_{i,t}$. Assumption (ref) (d) can be relaxed by instead requiring that $\varepsilon_{i,t}$ satisfies a martingale type condition similar to Assumption 3 in DGP2020. However, since this will add to the already lengthy derivations and heavy notation, in the present paper we maintain Assumption (ref).

If $\nu_G \geq 1$, Assumptions (ref)--(ref) are enough to ensure that $\widehat d_1$ is consistent. If, however, $\nu_G<1$, so that some of the factors are signal-weak, then we also need Assumption (ref).

assumption[Signal-weak factors] \ If $\nu_G<1$, then $T/N^2\to c_1\in [0,\infty)$.

As already pointed out, $\*F_{1}^0$ dominates all other factors and is thus easiest to estimate. It is therefore not surprising to find that $\widehat{d}_1$ is consistent for $d_1$. Lemma (ref) states this result formally.

lemma[Consistency of $\protect\widehat{d}_1$]\ Under Assumptions (ref)--(ref), as $N,\,T\to \infty$, \begin{eqnarray} P (\widehat{d}_1= d_1 )\to 1. \end{eqnarray}

Hence, $\widehat{d}_1$ is consistent. In order to ensure that also $\widehat d_2,\ldots,\widehat d_{G+1}$ are consistent, the groups have to be “distinct”. The next assumption formalizes this requirement. The assumption is stated in terms of $\*F_{g}^0$ and $\+\Gamma_{g}^0$, where $\*F_{g}^0$ is as before and $\+\Gamma_{g}^0 = (\+\gamma_{g,1}^{0},\ldots,\+\gamma_{g,N}^{0})^{\prime }$ is the $N\times d_g$ matrix of stacked factor loadings for group $g$.

assumption[Orthogonal groups] $\max_{ g\ne h}\|\+\Gamma_{g}^{0\prime}\+\Gamma_{h}^0\| =O_P(N^p)$ and $\max_{ g\ne h}\|\*F_{g}^{0\prime}\*F_{h}^0\| =O_P(T^q)$, where $g,h = 1,\ldots, G$, $G>1$, $p < 1$, $q<(\nu_G+\nu_{G-1})/2$ and $\nu_{G-1}\ge 1$.

The condition that $\nu _{G-1}\geq 1$ means that we only allow for one group of weak factors. This can be seen as a form of normalization and is not particularly restrictive. Let us therefore instead consider $\max_{g\neq h}\Vert \*F_{g}^{0\prime }\*F_{h}^{0}\Vert =O_{P}(T^{q})$, which is less restrictive that the exact orthogonality condition typically required in papers on grouped factor structures (see, for example, Ando). As is well known, $\+\gamma _{i}^{0\prime }\*f_{t}^{0}=\+\gamma _{i}^{0\prime }\*W^{-1}\*W\*f_{t}^{0}$ for any positive definite matrix $\*W$. Now set $\*W=(\*F^{0\prime }\*F^{0}\*C_{T}^{2})^{-1/2}$. This implies

equation[equation omitted — 107 chars of source]

which means that the renormalized factors are exactly orthogonal, and hence that Assumption (ref) is satisfied for $\mathbf{F}^{0 }\*W'$ with $q=-\infty $. This is one way in which $\max_{g\neq h}\Vert \*F_{g}^{0\prime }\*F_{h}^{0}\Vert =O_{P}(T^{q})$ can be justified. Another way is through orthogonal basis functions on $[0,1]$. For example, if $\*f_{t}^{0}=(1,t)^{\prime }$ and $\*C_{T}=\mathrm{diag}(T^{-1/2},T^{-3/2})$, then $\sqrt{T}\*C_{T}\*f_{t}^{0}\rightarrow \*f^{0}(w)=(1,w)^{\prime }$ and $\*C_{T}\*F^{0\prime }\*F^{0}\*C_{T}\rightarrow \int_{0}^{1}\*f^{0}(w)\*f^{0}(w)^{\prime }dw$ as $T\rightarrow \infty $ with $w\in \lbrack 0,1]$. Because they are square integrable, the elements of $\*f^{0}(w)$ can be represented in terms of orthogonal basis functions (see, for example, DongLinton). This is true not only for the simple example given here but for most commonly used deterministic trend terms. The loading condition, $\max_{g\neq h}\Vert \+\Gamma _{g}^{0\prime }\+\Gamma _{h}^{0}\Vert =O_{P}(N^{p})$, can be rationalized in the same way as the factor condition. Another possibility is that $\+\Gamma _{1}^{0},\ldots ,\+\Gamma _{G}^{0}$ are independent and at most one of them has non-zero mean. Independence is often assumed and we therefore do not justify it here (see, for example, Chudiketal2011, and Pesaran2006). In order to justify the zero mean assumption, suppose for simplicity that $G=2$, that $d_{1}=\+\gamma _{1,i}^{0}=1$ and that $\+\gamma_{2,i}^{0}=\+\gamma _{2}^{0}+\+\eta _{i}$ with $E(\+\eta _{i})=\mathbf{0}_{d_{2}\times 1}$. Hence, $\+\gamma _{i}^{0}=(\+\gamma _{1,i}^{0}, \+\gamma_{2,i}^{0\prime })^{\prime }=(1,\+\gamma _{2}^{0\prime }+\+\eta _{i}^{\prime})^{\prime }$ and $\*f_{t}^{0}=(\*f_{1,t}^{0},\*f_{2,t}^{0\prime })^{\prime }$, which in turn implies that

equation[equation omitted — 216 chars of source]

The common component with $\+\gamma _{i}^{0}$ as loading and $\*f_{t}^{0}$ as factor can therefore be written equivalently as one in which $(\*f_{1,t}^{0}+\+\gamma _{2}^{0\prime }\*f_{2,t}^{0},\*f_{2,t}^{0\prime})^{\prime }$ is the factor and $(1,\+\eta _{i}^{\prime })^{\prime }$ is the loading.

lemma[Consistency of $\widehat d_2,\ldots,\widehat d_{G+1}$ and $\widehat{\*F}$]\ Suppose that Assumptions (ref)--(ref) are satisfied. Then, the following results hold as $N,\,T\to \infty$: \begin{itemize} • $P (\widehat{d}_g = d_g)\to 1$ for $g=2,\ldots, G+1$, where $G> 1$ and $d_{G+1} =0$; • $\| \*P_{\widehat{F}} -\*P_{F^0}\| =o_P(1)$. \end{itemize}

Together with Lemma (ref), Lemma (ref) (a) implies that $\widehat d_1,\ldots,\widehat d_{G+1}$ are all consistent. The consistency of $\widehat d_1,\ldots,\widehat d_G$ is important for obvious reasons. The consistency of $\widehat d_{G+1}$ ensures that the stopping rule in Step (ref) of the estimation procedure is asymptotically valid, which in turn implies that $\widehat G$ is consistent. Hence, under the conditions of Lemma (ref), we have

eqnarray[eqnarray omitted — 39 chars of source]

as $N,\,T\to \infty$. The consistency of $\widehat G$ further implies that $d_f$ can be consistently estimated using $\widehat d_f = \widehat d_1 +\cdots+ \widehat d_{\widehat G}$.

As we alluded to in the discussion of Assumption (ref) above, $\*F^{0}$ and $\+\gamma _{i}^{0}$ are only identified up to a rotation matrix (see, for example, BaiNg2002). However, we cannot claim that $\widehat{\*F}$ is rotationally consistent for $\*F^{0}$, as the number of rows of both objects is growing with $T$. We therefore have to resort to alternative consistency concepts. This is where Lemma (ref) (b) comes in. It shows that the spaces spanned by $\widehat{\*F}$ and $\*F^{0}$ are asymptotically the same. This establishes that all the Step (ref) estimates are consistent. We therefore move on to investigate the asymptotic properties of the IPC estimator $\widehat{\+\beta }$ of $\+\beta ^{0}$ obtained in Step (ref).

assumption[Rates]\ \begin{itemize} • $N/T^{\nu_G}\to \rho_1 \in [0,\infty)$; • $T^{2-\nu_G}/N \to \rho_2 \in [0,\infty)$. \end{itemize}

It is important to note that if all the factors in $\*f_t^0$ are stationary such that $G=1$ and $\nu_G = \nu_1 = 1$, then Assumption (ref) (a) and (b) require that $N/T \to \rho_1 = 1/\rho_2 \in (0,\infty)$, which is the condition considered in Bai, and MoonWeidner2015. Assumption (ref) (a) and (b) are therefore more general than the condition considered in these other papers.

assumption[Asymptotic normality]\ \begin{align} \frac{1}{\sqrt{NT}}\sum_{i=1}^N\*D_T\*Z_i(\*F^0)'\+\varepsilon_i \to_D MN(\*0_{d_x\times 1},\+\Omega) \end{align} as $N,\,T\to\infty$, where \begin{align} \+\Omega = \operatorname*{p\!\lim}_{N,T\to\infty} \frac{1}{NT}\sum_{i=1}^N\sum_{j=1}^N\*D_T E[\*Z_i(\*F^0)'\+\varepsilon_i\+\varepsilon_j\*Z_j(\*F^0) | \mathcal{C}] \*D_T, \end{align} with $\mathcal{C}$ being the sigma-algebra generated by $\*F^0$.

Assumption (ref) is a central limit theorem that is analogous to Assumption E of Bai. The reason for requiring that the asymptotic distribution is mixed normal as opposed to normal is that by doing so we can accommodate stochastically integrated factors (see, for example, BaiKaoNg, and HJPS). In the absence of such integrated factors, the mixed normal becomes normal. Either way, Assumption (ref) ensures that standard normal and chi-squared inference based on $\widehat{\+\beta }$ is possible.

theorem[Asymptotic distribution of $\widehat{\+\beta}$]\ Under Assumptions (ref)--(ref), as $N,\, T\to \infty$, \begin{align} \sqrt{NT}\*D_T^{-1}(\widehat{\+\beta}-\+\beta^0) \to_D MN( \*B_0^{-1}(\sqrt{\rho_1}\*A_{1} + \sqrt{\rho_2}\*A_{2}) ,\*B_0^{-1}\+\Omega\*B_0^{-1}) \end{align} where \begin{align} \*B_0 & = \operatorname*{p\!\lim}_{N,T\to\infty} E[\*B(\*F^0) | \mathcal{C}], \\ \*A_{1} & = -\operatorname*{p\!\lim}_{N,T\to\infty} \frac{1}{T^{(1-\nu_G)/2}}\sum_{i=1}^N \*D_T E[\*X_i'\*M_{F^0} \+\Sigma_{\varepsilon} \*F_{G}^{0} (\*F_{G}^{0\prime}\*F_{G}^{0})^{-1} (\+\Gamma_G^{0\prime}\+\Gamma_G^0)^{-1} \+\gamma_{G,i}^0 | \mathcal{C}],\\ \*A_{2} & = -\operatorname*{p\!\lim}_{N,T\to\infty} \frac{1}{T^{(3-\nu_G)/2}} \sum_{i=1}^N \sum_{j=1}^N \*D_T E[\*Z_i(0)'\*F_G^0 (\*F_G^{0\prime} \*F_G^0 )^{-1} (\+\Gamma_G^{0\prime}\+\Gamma_G^0)^{-1} \+\gamma_{G,j}^{0} \+\varepsilon_j ' \+\varepsilon_i | \mathcal{C}], \end{align} with $\+\Sigma_{\varepsilon} = N^{-1}\sum_{i=1}^N \+\Sigma_{\varepsilon,i} $ and $\*Z_i(0) = \*X_i - \sum_{j=1}^{N}\*X_ja_{ij}$.

According to Theorem (ref), the asymptotic bias is driven by the factors and loadings of group $G$, which is intuitive as the factors of this group are smallest in order of magnitude. They therefore dominate the asymptotic bias. By bounding $\nu_G$ from below Assumption (ref) ensures that $\*B_0^{-1}(\sqrt{\rho_1}\*A_{1} + \sqrt{\rho_2}\*A_{2})$ is not diverging. It is important to note that $\rho_1$ and $\rho_2$ may both be zero, which will be the case if $\nu_G>1$ and $T/N$ converges to a constant. In general, the larger $\nu_{G}$ is, the less restrictive the condition on $T/N$ has to be for $\rho_1$ and $\rho_2$ to be zero. For example, if $\nu _{G}=2$, then we only require $N/T^{2}\to 0$. In this sense, trending factors are a blessing. In order to put this into perspective, suppose again that all the factors are stationary, such that $G=1$, $\nu_G = \nu_1 = 1$ and $\rho_1 = 1/\rho_2 \in (0,\infty)$, and that the regressors are stationary, too, such that $\*D_T = \*I_{d_x}$. In this case, the bias in Theorem (ref) reduces to $\*B_0^{-1}(\sqrt{\rho_1}\*A_{1} + \rho_1^{-1/2}\*A_{2})$, which is identical to the bias reported in Theorem 3 of Bai. In fact, it is not difficult to show that under these conditions, $\sqrt{NT} (\widehat{\+\beta}-\widehat{\+\beta}_0) = o_P(1)$, so that the IPC estimator is asymptotically equivalent to the first step PC estimator. The point about the bias is that while $\sqrt{\rho_1}\*A_{1}$ can be made arbitrarily small (large) by just taking $\rho_1$ to zero (infinity), this will make $\rho_1^{-1/2}\*A_{2}$ divergent (negligible). It follows that unless $\rho_1 \in (0,\infty)$ the bias will diverge, which in turn means that there is no way to make the bias disappear by just manipulating $\rho_1$, which in practical terms means restricting $T/N$. By allowing $\nu_G > 1$, we break the inverse relationship between $\rho_1$ and $\rho_2$, which means that one can be zero without for that matter forcing the other to infinity.

remarkBai shows that PC can be unbiased even if $\*f_t^0$ and $\*x_{i,t}$ are stationary. However, this requires either that $\varepsilon_{i,t}$ is independent and identically distributed across both $i$ and $t$, or that it is uncorrelated and homoskedastic in $t$ ($i$) with $T/N\to 0$ ($N/T \to 0$).
remarkNote that while biased, $\widehat{\+\beta }$ is still consistent at the best achievable rate. Again, if $\*x_{i,t}$ is stationary, $\*D_{T}=\*I_{d_{x}}$ and the rate of convergence is given by $1/\sqrt{NT}$, which is the same as in Bai. This is in contrast to Lemma (ref) and the relatively slow rate of convergence reported there. The reason for this difference is that unlike Theorem (ref), which requires that Assumptions (ref)--(ref) all hold, Lemma (ref) only requires Assumptions (ref) and (ref), and under these very relaxed conditions the Theorem (ref) rate is not attainable. If, however, the conditions of Theorem (ref) are met, then we can show that both $\widehat{\+\beta }_{0}$ and $\widehat{\+\beta}_{1}$ are consistent at the same rate as $\widehat{\+\beta }$ (see the appendix for a formal proof). However, unlike $\widehat{\+\beta }$, they are generally not mixed normal, which we verify using Monte Carlo simulations in Section (ref).
corollary[Unbiased asymptotic distribution]\ Suppose that the conditions of Theorem (ref) are met and that $\rho_1 = \rho_2 =0$. Then, as $N,\, T\to \infty$, \begin{align} \sqrt{NT}\*D_T^{-1}(\widehat{\+\beta}-\+\beta^0) \to_D MN( \*0_{d_x\times 1} ,\*B_0^{-1}\+\Omega\*B_0^{-1}). \end{align}

In the appendix, we provide some alternative conditions that ensure $\*A_{1} = \*A_{2} = \*0_{d_x\times 1}$. If $\rho_1$, $\rho_2$, $\*A_{1}$ and $\*A_{2}$ are all different from zero, one possibility is to use bias correction. In the appendix, we explain how the Jackknife approach can be used for this purpose.

Theorem (ref) imposes only minimal conditions on the correlation and heteroskedasticity of $\varepsilon_{i,t}$, and is in this sense very general. Such generality is, however, not possible if we also want to ensure consistent estimation of $\+\Omega$. Let us therefore assume for a moment that $E(\varepsilon_{i,t}\varepsilon_{j,s}) = 0$ for all $(i,t)\ne (j,s)$, so that $\varepsilon_{i,t}$ is serially and cross-sectionally uncorrelated. In this case,

eqnarray[eqnarray omitted — 214 chars of source]

A natural estimator of this matrix is given by

eqnarray[eqnarray omitted — 104 chars of source]

where $\widehat{\sigma}_{\varepsilon,i}^2 = T^{-1}\sum_{t=1}^T (\*y_i -\*X_i\widehat{\+\beta})^{\prime }\*M_{\widehat F}(\*y_i -\*X_i\widehat{\+\beta})$ and $\widehat{\*Z}_i$ is as in the definition of $\widehat{\+\beta}$. It is not difficult to show that under the conditions of Theorem (ref),

eqnarray[eqnarray omitted — 161 chars of source]

Of course, in this paper we do not assume knowledge of the order of the regressors, which in practice means that the appropriate normalization matrix $\*D_T$ to use is unknown. This is not a problem, however, as the usual Wald and $t$-test statistics are self-normalizing. As an illustration, consider testing the null hypothesis of $H_0: \*R \+\beta^0 = \*r$, where $\*R$ is a $r_0 \times d_x$ matrix of rank $r_0 \leq d_x$ and $\*r$ is a $r_0 \times 1$ vector. The Wald test statistic for testing this hypothesis is given by

eqnarray[eqnarray omitted — 378 chars of source]

which has a limiting chi-squared distribution with $r_0$ degrees of freedom under $H_0$, as is clear from

align[align omitted — 538 chars of source]

In the next section, we use Monte Carlo simulations as a means to evaluate the accuracy of this last results in small samples.

The above results are for the case when $\varepsilon _{i,t}$ is serially and cross-sectionally uncorrelated. If $\varepsilon _{i,t}$ is serially and/or cross-sectionally correlated, we recommend following Bai, who discusses the issue of consistent covariance matrix estimation at length. The same arguments can be applied without change in current context.

Monte Carlo results

This section reports the results obtained from a small-scale Monte Carlo simulation exercise. One important goal of the simulations is to examine the finite sample performance of $\widehat{\+\beta}_0$, $\widehat{\+\beta}_1$, and $\widehat{\+\beta}$, and justify the fact that $\widehat{\+\beta}$ is superior to $\widehat{\+\beta}_0$ and $\widehat{\+\beta}_1$.

We consider two DGPs, henceforth denoted “DGP (ref)” and “DGP (ref)”, which we now describe.

dgp[Artificial data]\ In this DGP, $y_{i,t}$ is generated according to a restricted version of (ref) that sets $d_x = 2$, $\+\beta^0 = \*1_{2\times 1}$, $\varepsilon_{i,t}\sim N(0,1)$ and $N,\,T \in \{40, 80, 160, 320\}$. We further set $d_f = 3$ and generate the factor loadings as $\gamma_{1,i}^0 \sim N(1,1)$, $\gamma_{2,i}^0 \sim N(0,1)$ and $\gamma_{3,i}^0 \sim N(0,1)$. The factors themselves are generated as $f_{1,t}^0=t$, $f_{2,t}^0= \mu_t$ and $f_{3,t}^0= c_t$, where $\mu_t= \mu_{t-1} + \xi_{t}$, $\mu_0 =0$, $\xi_{t} \sim N(0,1/4)$ and $c_{t} = \sin (8\pi t/T )$. Hence, in this DGP, the common component is a random walk with drift and cycle. Also, $d_1= d_2 = d_3=1$ and $(\nu_1,\nu_2,\nu_3)=(3,2,1)$. Let us denote by $x_{j,i,t}$ the $j$-th element of $\*x_{i,t}$. The following specification makes $x_{j,i,t}$ correlated with the common component of $y_{i,t}$: \begin{eqnarray} x_{j,i,t} =\frac{1}{d_x}\left(\sum_{j=1}^{d_f}| \gamma_{j,i}^0| +|\xi_t | +|c_t|\right)+ \left(\frac{t}{4}\right)^{(j-1)/4}+v_{j,i,t}, \end{eqnarray} where $v_{j,i,t}$ is the $i$-th element of the $N\times 1$ vector $\*v_{j,t} = ( v_{j,1,t}, \ldots ,v_{j,N,t} )'$, which we generate as \begin{eqnarray} \*v_{j,t} = 0.5 \*v_{j,t-1} + \+\omega_{j,t}, \end{eqnarray} where $\+\omega_{j,t} \sim N(\*0,\+\Sigma_\omega)$ and $\+\Sigma_\omega$ has $0.5^{|m-n|}$ in row $m$ and column $n$. Thus, $v_{j,i,t}$ is weakly correlated across both $i$ and $t$.
dgp[Calibrated data]\ This DGP is calibrated by using the US commercial banks data set studied in Section (ref) in which $N=466$ and $T=80$. The $d_x=5$ regressors are the same as those included in Section (ref) and $\+\beta^0 = (0.4063, 0.0199, 0.3805, 0.0912, 0.04726)'$ is calibrated using the empirical estimates taken from Table (ref). The factors are the estimated ones reported in Figure (ref). Hence, in this DGP, $d_f=2$. The loadings are generated as in DGP 1 with $\gamma_{1,i}^0 \sim N(1,1)$ and $\gamma_{2,i}^0 \sim N(0,1)$.

For each combination of $N$ and $T$, we report the correct selection frequency for $(\widehat d_1,\ldots,\widehat d_{\widehat G})$ when seen as an estimator of $(d_1,\ldots,d_{G})$ and for $\widehat d_g$ individually for each group $g = 1,\ldots,G$. Hence, while the former frequency captures the accuracy of the estimation of both $(d_1,\ldots,d_{G})$ and $G$, the latter frequency only captures the accuracy of the estimation of $d_g$ for each $g = 1,\ldots,G$. We also report the root mean squared error (RMSE) of $\widehat{\+\beta}$ and $\*P_{\widehat{F}}$, as measured by the square root of the average of $\|\widehat{\+\beta}-\+\beta^0\|^2$ and $\|\*P_{\widehat{F}} -\*P_{F^0}\|^2$, respectively, over the replications. In interest of comparison, RMSEs of $\widehat{\+\beta}$ is compared to that of $\widehat{\+\beta}_0$, $\widehat{\+\beta}_1$ and the infeasible OLS estimator of $\+\beta^0$ based on taking $\*F^0$ as known, $\widehat{\+\beta}(\*F^0)$. Some results on the 5% size of the Wald test based on $\*R =\*I_{d_x}$ and $\*r =\+\beta^0$ are also reported. In particular, we report the size of $W_{\widehat{\+\beta}}$, $W_{\widehat{\+\beta}_0}$ and $W_{\widehat{\+\beta}_1}$, which are computed in an obvious fashion by replacing $(\widehat{\+\beta},\widehat{\*F})$ with $(\widehat{\+\beta}_0,\widehat{\*F}_0)$ and $(\widehat{\+\beta}_1,\widehat{\*F})$, respectively, and $W_{\widehat{\+\beta} (\+F^0)}$, which is calculated in the same way as $W_{\widehat{\+\beta}}$ but with $\*M_F\*X_i$ in place of $\widehat{\*Z}_i$. The critical values are taken from $\chi^2(d_x)$. As pointed out in Section (ref), the IPC estimator is invariant with respect to $\delta$. The results reported here are based on $\delta=1$. In time series analysis, rules such as (the integer part of) $8 (T/100)^{1/4}$ are sometimes used for the maximum model dimension to be considered. In our theory, however, $d_{max}$ is fixed and some of the arguments used in the proofs fail if $d_{max}$ is allowed to increase with $N$ and $T$. In this section, we therefore follow the bulk of the previous literature and set $d_{max}$ to a fixed large number (see, for example, AH13, BaiNg2002, and MoonWeidner2015). We choose $d_{max}=10$, which led to the same results as some of the other values we tried. The number of replications is set to 1,000.

We begin by considering the results for DGP (ref), which are reported in Tables (ref) and (ref). In particular, while Table (ref) contains the results for the estimated common component, Table (ref) contains the results for the estimated slopes. We first look at the correct selection frequencies for all and each individual group of factors reported in Table (ref). The results for the individual groups suggest that the accuracy is generally degreasing in $g$, which is partly expected because the signal strength of the factors, as measured by $\nu_g$, decreases with increasing values of $g$. The accuracy of $(\widehat d_1,\ldots,\widehat d_{\widehat G})$ is almost identical to that of $\widehat d_{G}$, suggesting that the accuracy of $\widehat G$ is driven by the accuracy of the group whose factors has the weakest signal. Looking next at the RMSE results reported in Table (ref) for estimating $\+\beta^0$, we see that there is a clear improvement as the sample size increases. The best overall performance is generally obtained when taking $\*F^0$ as known, which is in accordance with our priori expectations. However, the improvement is not very large and it decreases with increases in $N$ and $T$. The reason for this is the accuracy of the estimated factors, which according to the RMSE of $\*P_{\widehat{F}}$ reported Table (ref) is high and increasing in $N$ and $T$. The second best performance is obtained by using $\widehat{\+\beta}$, followed by $\widehat{\+\beta}_1$, and then $\widehat{\+\beta}_0$. If our asymptotic theory is correct, while the size of $W_{\widehat{\+\beta}}$ and $W_{\widehat{\+\beta} (\+F^0)}$ should converge to 5% as the sample size increases, that of $W_{\widehat{\+\beta}_0}$ and $W_{\widehat{\+\beta}_1}$ should not, and this is exactly what we see in Table (ref). Note in particular how the size of $W_{\widehat{\+\beta}_0}$ and $W_{\widehat{\+\beta}_1}$ is not only nonconvergent but that it is in fact increasing in $N$ and $T$.

The RMSEs of $\widehat{\+\beta}_0$, $\widehat{\+\beta}_1$, $\widehat{\+\beta}$ and $\widehat{\+\beta}(\*F^0)$ in DGP (ref) are given by 0.0444, 0.0405, 0.0405 and 0.0391, respectively, and the sizes of the Wald tests associated with these estimators are given by 0.285, 0.103, 0.078 and 0.061, respectively. Hence, just as in DGP (ref), the best performance is obtained by using $\widehat{\+\beta}(\*F^0)$ with $\widehat{\+\beta}$, $\widehat{\+\beta}_1$ and $\widehat{\+\beta}_0$ ending up in second, third and fourth place, respectively.

Empirical illustration --- house prices and income

Economists have become concerned that recently house prices have grown too quickly, and that prices are now too high relative to per capita incomes. If this is correct and there is any truth to the theory on the matter, prices should stagnate or fall until they are better aligned with income, which in statistical terms mean that house prices should be cointegrated with income. The validity of this assumption has important implications for policy, because a failure could be due to a housing bubble.

In this section, we revisit the real house price data set of Hollyetal2010, which comprises data on log real house prices ($p_{i,t}$) and log real per capita income ($w_{i,t}$) for 49 US states across the 1975--2003 period. According to theory, $p_{i,t}$ and $w_{i,t}$ should be cointegrated with cointegrating vector $(1,\,-1)'$. The previous empirical evidence of this prediction has, however, been mixed and far from convincing (see, for example, Gallin2006). Hollyetal2010 argue that this lack of empirical support can be attributed in part to a failure to account for cross-sectional dependence, leading to deceptive conclusions. The authors therefore apply the “CIPS” panel unit root test of Pesaran2007, which allow for cross-section dependence in the form of a common factor. The test is applied both to $p_{i,t}$ and $w_{i,t}$ separately, and to $p_{i,t}-w_{i,t}$. According to the results, while the variables are unit root non-stationary, their difference is not. Hollyetal2010 also report CCE results suggesting that the estimated income elasticity is indeed close to one. They therefore conclude that $p_{i,t}$ and $w_{i,t}$ are cointegrated with cointegrating vector $(1,\,-1)'$, just as predicted by theory.

figure[figure omitted — 199 chars of source]

Our interest in the work of Hollyetal2010 stems from their preference to apply the CIPS test, which tests for a unit root in the defactored data. This means that if $p_{i,t}$ and $w_{i,t}$ are not cointegrated by themselves, but only when conditioning on unit root common factors, because of the way that the data are defactored prior to the testing, the unit root null hypothesis is likely to be rejected by the CIPS test. That is, the test is likely to lead to the conclusion of cointegration when in fact there is none. In this section, we use IPC as a means to investigate this possibility.

figure[figure omitted — 204 chars of source]

As in the RTS illustration, we begin by plotting the variables. This is done in Figure (ref). As expected, both variables are highly persistent and the ADF test provides no evidence against the unit root null. This corroborates the unit root test results reported by Hollyetal2010. However, we also see that the trending behaviours of $p_{i,t}$ and $w_{i,t}$ are very different, suggesting that their stochastic trends are not the same, which they should be under cointegration. We also see that the trending behaviour is very similar across states, which is suggestive of non-stationary common factors. Of course, the IPC procedure does not require cointegration and it does allow for very general types of factors. We therefore proceed with the estimation of the model. Hence, in this illustration, $y_{i,t}=p_{i,t}$ and $\*x_{i,t} = w_{i,t}$. For comparison purposes, the IPC results are presented together with the results obtained by applying the PC estimator of Bai, as well as the usual OLS estimator with time and state fixed effects. The results reported in Table (ref). The first thing to note is that the estimated slopes vary a lot depending on the estimator used. Interestingly, the point estimates are increasing in the generality of the estimator with fixed effects OLS (IPC) leading to the lowest (highest) estimate. We therefore begin by considering the IPC results. The point estimate of 2.1024 is far from the theoretically predicted value of one, which is also not included in the reported 95% confidence interval. As for the factors, we estimate $\widehat d_1 = \widehat d_2 = \widehat d_3 = 1$, implying that $\widehat d_f =3$. The estimated factors, denoted $\widehat f_{1,t}$, $\widehat f_{2,t}$ and $\widehat f_{3,t}$, are plotted in Figure (ref). As expected given Figure (ref), all three factors are highly persistent, although with clearly distinct trends. We take this as evidence against cointegration between $p_{i,t}$ and $w_{i,t}$, since under cointegration the factors should be stationary.

Because we estimate three time-varying factors, fixed effects OLS is invalid, as it only allows for a common time effect. Moreover, since the factors come from three distinct groups, and are not all stationary, PC is invalid, too. This leaves us with the IPC estimator, which again provides strong evidence against the theoretically predicted one-to-one cointegrated relationship between $p_{i,t}$ and $w_{i,t}$. Hollyetal2010 conclude that “[o]ur results support the hypothesis that real house prices have been rising in line with fundamentals (real incomes), and there seems little evidence of house price bubbles at the national level.” The results reported here reveal a completely different picture with housing prices being long run disconnected with real income.

Conclusion

The PC approach of Bai has attracted considerable interest in recent years, so much so that it has given rise to a separate PC literature. A key assumption in this literature is that both the unknown factors and regressors are stationary, which is rarely the case in practice. In the present paper, we relax this assumption by considering a very general specification in which the factors and regressors are essentially unrestricted. In spite of this generality, the proposed IPC estimator can be applied without any input from the practitioner, except for the maximum number of factors to be considered. The fact that in IPC there is no need to distinguish between deterministic and stochastic factors means that the usual problem in applied work of deciding on which deterministic terms to include in the model does not arise, as these are estimated along with the other factors of the model. There is also no need to pre-test the regressors for unit roots, which is otherwise standard practice when using procedures that do not require data to be stationary. In other words, the proposed IPC is not only very general but also extremely user-friendly. It should therefore be a valuable addition to the already existing menu of techniques for panel regression models with interactive effects.

table[table omitted — 1,849 chars of source]
table[table omitted — 2,523 chars of source]
table[table omitted — 669 chars of source]
center[center omitted — 376 chars of source]

\setcounter{page}{1}

\setcounter{equation}{0} \setcounter{section}{0} \setcounter{table}{0} \setcounter{figure}{0} \setcounter{assumption}{0}

We begin this appendix by laying out the notation that will be used throughout. This is done in Section (ref), which in Section (ref) is followed by a discussion of some conditions that ensure that the asymptotic distribution of $\widehat{\+\beta}$ is correctly centered at zero. Section (ref) provides an extra empirical study using the US bank data. Section (ref) describes the outline of the proofs. Sections (ref) and (ref) provide some auxiliary lemmas and their proofs, respectively. Proofs of the main results are provided in Section (ref).

Notation

The matrices $\+\Sigma_{F^0}$ and $\+\Sigma_{\Gamma^0}$ have been defined in Assumption (ref). In this appendix, we use $\+\Sigma_{F_g^0}$ and $\+\Sigma_{\Gamma_g^0}$ to denote the sub-matrices of $\+\Sigma_{F^0}$ and $\+\Sigma_{\Gamma^0}$ corresponding to $T^{-\nu_g}\*F_{g}^{0\prime}\*F_{g}^0$ and $N^{-1}\+\Gamma_g^{0\prime}\+\Gamma_g^0$, respectively, for $g = 1,\ldots,G$. We also define $\*V_{g} = \mathrm{diag}(\widehat{\lambda}_{g,1},\ldots, \widehat{\lambda}_{g,d_{max}})$, where $\widehat{\lambda}_{g,d}$ has been defined in Step (ref) of the IPC estimation procedure. We partition $\*F^0 = (\*F_{1}^0,\ldots,\*F_g^0,\*F_{+g}^0)$ and $\*C_T = \mathrm{diag} (T^{-\nu_1/2}\*I_{d_1},\ldots, T^{-\nu_g/2}\*I_{d_g}, \*C_{+g,T} )$, where $\*F_{+g}^0 = (\*F_{g+1}^0,\ldots,\*F_G^0)$ is $T\times (d_{g+1} +\cdots+d_G)$, and $\*C_{+g,T} = \mathrm{diag}( T^{ -\nu_{g+1}/2} \*I_{d_{g+1}} ,\ldots, T^{-\nu_G/2 } \*I_{d_G})$ is $(d_{g+1} +\cdots+d_G)\times (d_{g+1} +\cdots+d_G)$. We partition $\widehat{\*F}$, $\+\gamma_{i}^0$ and $\+\Gamma^0$ conformably as $\widehat{\*F} =(\widehat{\*F}_1,\ldots,\widehat{\*F}_g,\widehat{\*F}_{+g})$, $\+\gamma_{i}^0 = (\+\gamma_{1,i}^{0\prime},\ldots,\+\gamma_{g,i}^{0\prime}, \+\gamma_{+g,i}^{0\prime})'$ and $\+\Gamma^0 = (\+\Gamma_{1}^0,\ldots,\+\Gamma_{g}^0, \+\Gamma_{+g}^0)$, respectively.

We introduce $\lambda_{g,d} = T^{-\nu_g} \*h_{g,d}^{0\prime}\*F_{g}^{0\prime}\+\Sigma_{g}^0\*F_{g}^{0}\*h_{g,d}^0$, where $\+\Sigma_{g}^0= N^{-1} \*F_{g}^{0} \+\Gamma_{g}^{0\prime} \+\Gamma_{g}^0\*F_{g}^{0\prime}$ and $\*h_{g,d}^0$ is the $d$-th column of $\*H_{g}^0 = N^{-1}T^{(\nu_g-\delta)/2}\+\Gamma_{g}^{0\prime}\+\Gamma_{g}^0 \*F_{g}^{0\prime}\widehat{\*F}_{g}^0 (\*V_{g}^0)^{-1}$ with $\widehat{\*F}_{g}^0$ being the $T\times d_g$ matrix consisting of the first $d_g$ columns of $\widehat{\*F}_{g}$ and $\*V_{g}^0$ being the leading $d_g \times d_g$ principal submatrix of $\*V_{g}$. In other words, $\widehat{\*F}_{g}^0$ and $\*V_{g}^0$ are $\widehat{\*F}_{g}$ and $\*V_{g}$ based on treating the number of factors for each group $g$, $d_g$, as known. We also define $\*H_{g} = T^{-(\nu_g-\delta)/2}\*H_{g}^0$. In order to appreciate the implication of the difference in normalization with respect to $T$, let us consider $\*H_{g}^0$. By Assumption (ref), $N^{-1} \+\Gamma_{g}^{0\prime}\+\Gamma_{g}^0$ is asymptotically of full rank, and hence $\|N^{-1} \+\Gamma_{g}^{0\prime}\+\Gamma_{g}^0\| = O_P(1)$. Hence, since

align[align omitted — 213 chars of source]

and $\|(T^{-\nu_1}\*V_{1}^0)^{-1}\| = O_P(1)$ as explained under (ref), we can show that

align[align omitted — 182 chars of source]

which in turn implies

align[align omitted — 92 chars of source]

We further use $\widehat{\*F}_{g,d}^0$ to refer to the $d$-th column of $\widehat{\*F}_g^0$. In this notation, $\widehat{\lambda}_{g,d} = T^{-\delta}\widehat{\*F}_{g,d}^{0\prime} \widehat{\+\Sigma}_g \widehat{\*F}_{g,d}^0$ for $d = 1,\ldots,d_g$.

We also partition $\*X_i$ as $\*X_i= (\*X_{1,i},\ldots, \*X_{d_x,i})$ with $\*X_{j,i}$ being the $j$-th column of $\*X_i$. The $j$-th column of $\*X_i\*D_T$ is therefore given by $T^{-\kappa_j/2}\*X_{j,i}$. Moreover, $\mathrm{vec}\,\*A$, $\mathrm{rank}\,\*A$, $\mathrm{span}\,\*A$ and $\lambda(\*A)$ denote the vectorized version, rank, span and eigenvalues of $\*A$, respectively, $a\wedge b = \min\{a,b\}$ and $a\vee b = \max\{a,b\}$.

Conditions that ensure asymptotic unbiasedness

In this section, we provide a set of assumptions that ensure that the asymptotic distribution of $\sqrt{NT}\*D_T(\widehat{\+\beta} - \+\beta^0)$ given in Theorem (ref) is free of bias without for that matter requiring that $\varepsilon_{i,t}$ is serially and cross-sectionally independent. One way to accomplish this is to assume that $\rho_1 = \rho_2 = 0$, as in Corollary (ref). The assumptions considered here, which are stated in Assumption (ref), can be seen as alternatives to this last condition. In terms of the notation of Theorem (ref), they ensure that $\*A_{1}=\*A_{2} = \*0_{d_x\times 1}$.

assumption[No asymptotic bias]\ One of the following set of conditions is met: \begin{itemize} • $T^{1-\nu_g}E(\*f_{g,t}^{0\prime}\*f_{g,s}^0|\*x_t,\*x_s) = \phi_{ts}$ w.p.a.1 and $\sum_{t=1}^T\sum_{s=1}^T |\phi_{ts}| = O(T)$, where $\*x_{t} = (\*x_{1,t},\ldots, \*x_{N,t})'$. If $G\ge 2$, then $q<(\nu_G +\nu_{G-1})/2-1/4$ and $\nu_{g-1} - \nu_{g}>1/2$ for $g=2,\ldots,G$. • $T/N \to c_4 \in (0,\infty)$ and $\nu_G > 1$. If $G\ge 2$, then $q<(\nu_G +\nu_{G-1}-1)/2$ and $\nu_{g-1} - \nu_{g}>1/2$. • $T^{1-\nu_g-\kappa_j}E(\sum_{t=1}^T\sum_{s=1}^T x_{j,i,t}x_{j,k,s}\*f_{g,t}^{0\prime}\*f_{g,s}^0) = O(T^{2-r})$, where $r < 2$, $r+\nu_G -1 > 0$ and $x_{j,i,t}$ is the $j$-th row of $\*x_{i,t}$. If $G\ge 2$, then $\nu_{g-1} - \nu_{g}>1/2$. \end{itemize}

Assumption (ref) ensures that the asymptotic distribution of $\sqrt{NT}\*D_T(\widehat{\+\beta} - \+\beta^0)$ is bias-free. It is, however, not necessary, and (a)--(b) should therefore be viewed as examples of conditions under which there is no asymptotic distribution bias. These conditions all have their strengths and weaknesses, and so their suitability will in general depend on the context. Take as an example condition (b), which has the advantage of not requiring any more moment conditions than those that are already in Assumption (ref). It does, however, require that $\nu_G > 1$, which rules out both signal-weak and stationary factors ($\nu_G \leq 1$). Conditions (a) and (c) are more general in this regard, but then at the expense of requiring additional moment conditions. Condition (c) is a functional central limit theorem style moment condition.

The next corollary to Theorem (ref) verifies that the asymptotic distribution of $\sqrt{NT}\*D_T^{-1}(\widehat{\+\beta}-\+\beta^0)$ is indeed bias-free under Assumption (ref).

corollary[Unbiased asymptotic distribution]\ Suppose that Assumptions (ref)--(ref), (ref), and (ref) are met and that $N/T^{\nu_G}\to 0$. Then, as $N,\,T\to \infty$, \begin{eqnarray} \sqrt{NT}\*D_T^{-1}(\widehat{\+\beta}-\+\beta^0) \to_D MN(\*0_{d_x\times 1} ,\*B_0^{-1}\+\Omega\*B_0^{-1}). \end{eqnarray}

The asymptotic distribution in Corollary (ref) is the same as the one given in Corollary (ref).

If $\rho_1$, $\rho_2$, $\*A_{1}$ and $\*A_{2}$ are all different from zero, one possibility is to use bias correction. DJ2015 were first to bring attention to the relevance of the half-panel Jackknife approach for bias correction in panel data. These authors focus on the fixed effects case, but Chen2020 have shown that the Jackknife can be used to correct for bias also in models with interactive effects. In our setting, the value of $\nu_G$ is unknown, and therefore the standard Jackknife is not directly applicable. We thus turn to the hybrid half-panel Jackknife correction proposed by FW2018. The bias-corrected version of the IPC estimator is given by

eqnarray[eqnarray omitted — 182 chars of source]

where $\widehat{\+\beta}$ is the standard IPC estimator defined in (ref), $\widehat{\+\beta}_{1,N}$ and $\widehat{\+\beta}_{2,N}$ are defined in the same way but applied to cross-sectional units $\{1,\ldots, \lfloor N/2\rfloor \}$ and $\{\lfloor N/2\rfloor +1,\ldots, N\}$, respectively, and $\widehat{\+\beta}_{1,T}$ and $\widehat{\+\beta}_{2,T}$ are the IPC estimators applied to odd and even numbered time periods, respectively. The splitting in even and odd numbered time periods is needed in order to account for the behaviour of the factors. See FW2018 for a more detailed discussion of the hybrid Jackknife approach.

Extra empirical study --- Returns to scale

In this section, we estimate the returns to scale (RTS) of US banks, an area that has attracted considerable attention in the empirical production literature (see, for example, FengSerletis2008). Let us therefore denote by $CT_{i,t}$ the observed total cost of bank $i$ during period $t$. The total cost is a function $C(\cdot)$ of a $J\times 1$ vector of input prices $\*p_{i,t} = (p_{1,i,t},\ldots,p_{J,i,t})'$ and a $L\times 1$ vector of output quantities $\*q_{i,t} = (q_{1,i,t},\ldots,q_{L,i,t})'$. Hence, $CT_{i,t} = C(\*p_{i,t},\*q_{i,t})$, or when expressed in logs, $\ln CT_{i,t} = \ln C(\*p_{i,t},\*q_{i,t})$. The RTS is simply the reciprocal of the sum of the output elasticities, which can be estimated for a given $C(\cdot)$. In order to ensure that $C(\cdot)$ is linearly homogenous in prices, however, it is very common to first normalize costs and prices by one of the prices, $p_{J,i,t}$ say. Hence, $\ln(CT_{i,t}/p_{J,i,t}) = \ln C(\*p_{i,t}/p_{J,i,t},\*q_{i,t})$, which in turn implies that the RTS is given by

eqnarray[eqnarray omitted — 139 chars of source]

In order to estimate this quantity, estimates of the partial derivatives are needed. The standard way to obtain such estimates in the literature is to assume that $\ln C(\*p_{i,t}/p_{J,i,t},\*q_{i,t})$ is linear, to estimate the resulting linear relationship between $\ln (CT_{i,t}/p_{J,i,t})$, $\ln (\*p_{i,t}/p_{J,i,t})$ and $\ln \*q_{i,t}$ by OLS, and to estimate the partial derivatives by simply replacing the parameters by their OLS estimates. A number of modelling issues then arise. First, as pointed out by FengZhang, “[t]here are many ... types of ... unobserved heterogeneity in the US banking industry, which include, but are not limited to, location-related corporate income tax rate, property tax rate, and personal income tax rate.” Hence, there is a need to control for unobserved heterogeneity. Second, the literature typically assumes that the regressors in $\ln(\*p_{i,t}/p_{J,i,t})$ and $\ln\*q_{i,t}$ are stationary, which, as pointed out by DGP2020, is highly unlikely to be the case in practice. There is therefore a need to consider approaches that do not require the regressors to be stationary. Third, to account for the fact that costs are typically trending, many researchers fit their models with linear and quadratic trend terms. Of course, deterministic trends can account for some trending behaviour, but not all. Moreover, the results tend to be highly sensitive to how the trending is modelled, suggesting that this is an important issue (see FengSerletis2008, for a discussion).

The discussion of the last paragraph suggests that there is a need for an approach that is general enough to accommodate not only unobserved heterogeneity but also trending behaviour. The proposed IPC approach fits this bill and we will therefore use it in this empirical illustration to the RTS of US commercial banks. The data that we will use for this purpose, which are the same as in FGPZ2017, are quarterly and cover the period 1986--2005, which means that $T=80$. To avoid the impact of entry and exit, we only include continuously operating large banks with assets of at least one billion USD. This gives us a total of $N=466$ banks. The data set contains observations on three input prices and output quantities. They are the wage rate of labour ($p_{1,i,t}$), the price of deposits and purchased funds ($p_{2,i,t}$), the price of physical capital ($p_{3,i,t}$), consumer loans ($q_{1,i,t}$), non-consumer loans, which is composed of industrial, commercial and real estate loans ($q_{2,i,t}$), and securities, including non-loan financial assets ($q_{3,i,t}$). All outputs are deflated by the GDP deflator. Following the previous literature (see, for example, Stiroh), total cost ($CT_{i,t}$) is computed as the sum of total salaries and benefits divided by the number of full-time employees, the price of deposits and purchased funds equals total interest expense divided by total deposits and purchased funds, and the price of capital equals expenses on premises and equipment divided by premises and fixed assets. Hence, after normalization by $p_{3,i,t}$, there are six variables included in the model; $\ln (CT_{i,t}/p_{3,i,t})$, $\ln (p_{1,i,t}/p_{3,i,t})$, $\ln (p_{2,i,t}/p_{3,i,t})$, $\ln q_{1,i,t}$, $\ln q_{2,i,t}$ and $\ln q_{3,i,t}$.

figure[figure omitted — 179 chars of source]

In order to get a feeling for the trending behaviour of the data, in Figure (ref) we plot all $N$ series for each variable. A few observations are noteworthy. First, there is strong co-movement among the series. This is true for all six variables, including the dependent variable, $\ln (CT_{i,t}/p_{3,i,t})$, which we take as evidence in support of our interactive effects specification. Second, the variables are highly persistent and most are trending over time, suggesting that they are non-stationary. This last observation is supported by a formal augmented Dickey--Fuller (ADF) test (with intercept and linear trend) applied to each series. The highest rejection frequency across all $N$ series is obtained for $\ln q_{2,i,t}$ and is given by 6.7%, which in turn suggests that the evidence against unit root null hypothesis is weak. In fact, since the ADF test does not account for the multiplicity of the testing problem, the true proportion of (trend-)stationary series is likely to be even lower. Third, it is unclear if the trend in $\ln (CT_{i,t}/p_{3,i,t})$ is the same as those in the regressors, which means that it is important to allow the factors to be trending, as otherwise any unaccounted for trend in $\ln (CT_{i,t}/p_{3,i,t})$ will be pushed into the regression errors.

figure[figure omitted — 184 chars of source]

We now go on to discuss the actual estimation results, which are reported in Table (ref). The estimated model is given by (ref) with $y_{i,t}=\ln (CT_{i,t}/p_{3,i,t})$, $\+\beta^0 = (\beta_1^0,\ldots,\beta_5^0)'$ and $\*x_{i,t} = (\ln(p_{1,i,t}/p_{3,i,t}), \ln( p_{2,i,t}/p_{3,i,t}), \ln q_{1,i,t } , \ln q_{2,i,t}, \ln q_{3,i,t})'$. Hence, in this notation,

eqnarray[eqnarray omitted — 74 chars of source]

is the object of interest, which we estimate using

eqnarray[eqnarray omitted — 103 chars of source]

where $\widehat\beta_3$, $\widehat\beta_4$ and $\widehat\beta_5$ are from $\widehat{\+\beta} = (\widehat\beta_1,\ldots, \widehat\beta_5)'$. The values of $\delta$ and $d_{max}$ are set as in Section (ref). The value of $\delta$ is again arbitrary and does not require justification. As for $d_{max}$, 10 factors should be large enough to capture the true number, yet not “excessively” large (see AH13, for a discussion). In order to enable inference not only for $\+\beta^0$ but also for the RTS, we follow FGPZ2017 and report 95% bootstrap confidence intervals. The estimated factor groups are given by $\widehat{d}_1=\widehat{d}_2=1$, implying that the estimated number of factors equals $\widehat d_f = \widehat{d}_1+\widehat{d}_2=2$. In Figure (ref), we plot the estimated factors, denoted $\widehat f_{1,t}$ and $\widehat f_{2,t}$. We see that while both factor estimates are trending, $\widehat f_{2,t}$ is growing much faster. As a measure of this difference in the trends, we look at $\sum_{i=1}^N \widehat{\gamma}_{1,i}^2 /\sum_{i=1}^N \widehat{\gamma}_{2,i}^2$. By using the results provided in the appendix, we can show that this ratio should be $O_P(T^{\nu_1-\nu_2})$, implying $\ln(\sum_{i=1}^N \widehat{\gamma}_{1,i}^2 /\sum_{i=1}^N \widehat{\gamma}_{2,i}^2) /\ln T = \nu_1 - \nu_2 + o_P(1)$. By plugging in the known value of $T$ and $\sum_{i=1}^N \widehat{\gamma}_{1,i}^2 /\sum_{i=1}^N \widehat{\gamma}_{2,i}^2 = 123.4$, we get $\ln(\sum_{i=1}^N \widehat{\gamma}_{1,i}^2 /\sum_{i=1}^N \widehat{\gamma}_{2,i}^2) /\ln T = 1.1$, which is thus an estimate of $\nu_1 - \nu_2$. Hence, there is indeed a substantial difference in the degree of trending of the two factors. But it is not as large as it would have been if $\widehat f_{1,t}$ and $\widehat f_{2,t}$ were made up of a quadratic and a linear trend, say, in which case $\nu_1 = 5$ and $\nu_2 = 3$, such that $\nu_1 - \nu_2 = 2$. Hence, just as expected given Figure (ref), while trending, the behaviour of the factors is not just a simple deterministic function of time, which of course casts doubt on much of the existing research.

While our main interest is in the RTS, for completeness in Table (ref), we also report the estimated elasticities. The first thing to note is that the RTS is estimated to be significantly larger than one, suggesting that the large US banks considered here have had increasing returns to scale during the period of investigation. This is consistent with the results reported by, for example, FGPZ2017 and WheelockWilson2012. One explanation for this finding is that the banks have been adopting productivity-enhancing internet technologies, leading to increased RTS. The confidence intervals are very narrow not only for the RTS, but also for the elasticities, which are all significant.

table[table omitted — 1,036 chars of source]

Outline of the proofs

In this section, we describe the outline of the proofs. We assume throughout that $G\ge 2$. The proofs for the cases when $G\in \{0,1\}$ are much simpler, and can be obtained by manipulating the proofs for $G\ge 2$.

Lemma (ref) is a widely used result for studying the eigenvalues of large dimensional matrices (LYB2011), and is presented here for convenience.

Our own theoretical development is carried out stepwise, starting with Step (ref) of the IPC estimation procedure. Lemma (ref) is very important in this regard. It presents two order results that are used repeatedly throughout the proofs. Given Lemma (ref), we are able to establish Lemma (ref) of the main text, which provides a lower bound on the rate of convergence of the initial Step (ref) estimator $\widehat{\+\beta}_0$.

As a first step towards establishing the consistency of the Step (ref) estimators of $(d_2, \ldots, d_G)$, in Lemma (ref) we study limiting behaviour of the eigenvalues associated with $\widehat{\+\Sigma}_1$ in (ref). The consistency of $\widehat d_1$ reported in Lemma (ref) of the main text is a direct consequence of Lemma (ref). Lemmas (ref) and (ref) enable consistent estimation of $(d_2, \ldots, d_G)$, and are thus key in proving Lemma (ref) of the main text. For simplicity, in the proofs of Lemmas (ref) and (ref) we treat $d_1$ as known, which according to Lemma (ref) is justified in large samples.

The final step in our theoretical analysis is to prove the asymptotic mixed normality of the Step (ref) IPC estimator $\widehat{\+\beta}$, as stated in Theorem (ref). In so doing, we treat $(d_1,\ldots, d_G)$ as known, which is again justified in large samples. We begin by proving Lemma (ref), which is needed in the proof of Theorem (ref). The proofs of Corollaries (ref) and (ref) are immediate consequences of that of Theorem (ref). We also derive the rate of convergence of $\widehat{\+\beta}_0$ under the conditions of Theorem (ref). This rate is provided as a part of Lemma (ref).

Auxiliary lemmas

lemmaSuppose that $\*A$ and $\*A+\*E$ are $n\times n$ symmetric matrices and that $\*Q = (\*Q_1, \*Q_2)$, where $\*Q_1$ is $n\times r$ and $\*Q_2$ is $n\times (n-r)$, is an orthogonal matrix such that $\text{span}\,\*Q_1$ is an invariant subspace for $\*A$. Decompose $\*Q'\*A\*Q$ and $\*Q'\*E\*Q$ as $\*Q'\*A\*Q = \mathrm{diag}(\*D_1,\*D_2)$ and \begin{align} \*Q'\*E\*Q = \left( \begin{array}{cc} \*E_{11} & \*E_{12} \\ \*E_{21} & \*E_{22} \end{array} \right). \end{align} Let $\mathrm{sep}(\*D_1,\*D_2) = \min_{\lambda_1\in \lambda(\*D_1),\ \lambda_2\in \lambda(\*D_2)} |\lambda_1 -\lambda_2|$. If $\mathrm{sep}(\*D_1,\*D_2) > 0$ and $\|\*E\|_{2} \leq \mathrm{sep}(\*D_1,\*D_2)/5$, then there exists a $(n-r)\times r$ matrix $\*P$ with $\|\*P \|_{2} \leq 4 \|\*E_{21}\|_2/\mathrm{sep}(\*D_1,\*D_2)$, such that the columns of $\*Q_1^0 = (\*Q_1 + \*Q_2\*P)(\*I_r+\*P'\*P)^{-1/2}$ define an orthonormal basis for a subspace that is invariant for $\*A+\*E$.
lemmaUnder Assumption (ref), as $N,\,T\to\infty$, \begin{itemize} • $\sup_{\*F\in \mathbb{D}_F}(NT)^{-1}\sum_{i=1}^{N} \+\varepsilon_i^{\prime }\*P_{F} \+\varepsilon_i =O_P(N^{-1}\vee T^{-1} )$; • $\sup_{\*F\in \mathbb{D}_F} \|(NT)^{-1}\sum_{i=1}^{N}\*D_T \*X_{i}'\*P_F\+\varepsilon_i\| =O_P(N^{-1/2}\vee T^{-1/2})$. \end{itemize} If Assumption (ref) also holds, then \begin{itemize} • $(NT)^{-1} \|\+\varepsilon\+\varepsilon'\| =O_P(N^{-1/2}\vee T^{-1/2})$ and $(NT)^{-1} \|\+\varepsilon'\+\varepsilon\| = O_P(N^{-1/2}\vee T^{-1/2})$; • $\|\+\Gamma^{0\prime}\+\varepsilon\| =O_P(\sqrt{NT})$. \end{itemize}
lemmaLet Assumptions (ref)--(ref) hold. Then, as $N,\,T\to\infty$, \begin{itemize} • $T^{-\nu_1}|\widehat{\lambda}_{1,d} - \lambda_{1,d} | = O_P( T^{-(\nu_1-\nu_2)/2})$ for $d=1,\ldots, d_1$; • $T^{-\nu_1}|\widehat{\lambda}_{1,d} | =O_P( T^{-(\nu_1-\nu_2)})$ for $d=d_1+1,\ldots, d_{max}$. \end{itemize}
lemmaAs $N,\,T\to\infty$, the following hold under Assumptions (ref)--(ref): \begin{itemize} • $T^{-\delta} \|\*F_{2}^{0\prime} \widehat{\*F}_1\| = O_P( T^{-(\delta+\nu_1-\nu_2)/2}+N^{-1/2}T^{-(\delta+\nu_1 -\nu_2 -1)/2} + N^{-(1-p)}T^{-(\delta +\nu_1-2\nu_2)/2}+T^{-(\nu_1+\delta)/2-q})$; • $\sum_{i=1}^N \| \*F_{1}^{0\prime} \+\gamma_{1,i}^0 - \widehat{\*F}_1 \widehat{\+\gamma}_{1,i} \|^2 = O_P( N\vee T+NT^{2q -\nu_1}+N^{-(1-2p)}T^{\nu_2})$. \end{itemize}
lemmaLet $\tau_{NT} =N^{-1/2}T^{-(\nu_2-1)/2} +T^{-\nu_2/2} + T^{-(\nu_2-\nu_3)} + T^{q-(\nu_1+\nu_2)/2} + N^{-(1-p)}$. As $N,\,T\to\infty$, the following hold under the conditions of Lemma (ref): \begin{itemize} • $T^{-\nu_2}|\widehat{\lambda}_{2,d} - \lambda_{2,d} | = O_P( \tau_{NT} )$ for $d=1,\ldots, d_2$; • $T^{-\nu_2}|\widehat{\lambda}_{2,d} | =O_P( \tau_{NT}^2 )$ for $d=d_2+1,\ldots, d_{max}$. \end{itemize}
lemmaLet Assumptions (ref)--(ref) hold. In addition, let $ NT^{-\nu_G}<\infty $, as $N,\,T\to\infty$, \begin{itemize} • $\|\*D_T^{-1}( \widehat{\+\beta}_1-\widehat{\+\beta}_0 )\| = O_P((NT)^{-1/2}\vee \|\*D_T^{-1} ( \widehat{\+\beta}_0-\+\beta^0)\| )$; • $\|\*D_T^{-1} ( \widehat{\+\beta}_0-\+\beta^0)\| = O_P((NT)^{-1/2})$. \end{itemize}

Proofs of auxiliary lemmas

Proof of Lemma (ref).

This is Lemma 3 of LYB2011. The proof is therefore omitted.{$\blacksquare$}

Proof of Lemma (ref).

Consider (a). We have

align[align omitted — 404 chars of source]

where the first inequality follows from the fact that $|\text{tr}\,\*A| \leq \text{rank}\,\*A \, \| \*A\|_{2}$, the second inequality follows from the fact that $\|\*{P}_{{F}}\| _{2}=1$, and the second equality hold by Assumption (ref) (b).

The result in (b) is due to

align[align omitted — 498 chars of source]

where, with a slight abuse of notation and in this proof only, $\*X_j= (\*X_{j,1},\ldots, \*X_{j,N})'$, the second inequality follows from $|\text{tr}\,\*A| \leq \text{rank}\,\*A \,\| \*A\|_{2}$, while the first equality is due to Assumption (ref).

For (c), we use

align[align omitted — 828 chars of source]

where the third equality follows from using the mixing condition on $\varepsilon_{i,t}\varepsilon_{j,t}$ across $t$. The above result implies that $(NT)^{-1}\| \+\varepsilon'\+\varepsilon\| =O_P(N^{-1/2})+O_P(T^{-1/2})$ as required for the first result in (c). The second follows from

align[align omitted — 462 chars of source]

where the last step follows by the same arguments used to establish the first result of (c).

It remains to prove (d), which is a direct consequence of Assumptions (ref) and (ref), as seen from

eqnarray[eqnarray omitted — 237 chars of source]

This establishes (d) and hence the proof of the lemma is complete.{$\blacksquare$}

Proof of Lemma (ref).

Consider (a). As in Appendix (ref), decompose $\*F^0 = (\*F_{1}^0,\*F_{+1}^0)$ and $\*C_T = \mathrm{diag} (T^{-\nu_1/2}\*I_{d_1}, \*C_{+1,T} )$, where $\*F_{+1}^0 = (\*F_2^0,\ldots,\*F_G^0)$ is $T\times (d_f -d_1)$ and $\*C_{+1,T} = \mathrm{diag}( T^{ -\nu_2/2} \*I_{d_2} ,\ldots, T^{-\nu_G/2 } \*I_{d_G})$ is $(d_f -d_1)\times (d_f -d_1)$. We partition $\widehat{\*F}$, $\+\gamma_{i}^0$ and $\+\Gamma^0$ conformably as $\widehat{\*F} =(\widehat{\*F}_1,\widehat{\*F}_{+1})$, $\+\gamma_{i}^0 = (\+\gamma_{1,i}^{0\prime}, \+\gamma_{+1,i}^{0\prime})'$ and $\+\Gamma^0 = (\+\Gamma_{1}^0, \+\Gamma_{+1}^0)$, respectively.

By the definition of the eigenvectors and eigenvalues, $\widehat{\+\Sigma}_1 \widehat{\*F}_1 = \widehat{\*F}_1\*V_{1}$. By using this and the definition of $\widehat{\+\Sigma}_1$,

align[align omitted — 1,520 chars of source]

with implicit definitions of $\*J_{1},\ldots,\*J_{9}$. Note that

align[align omitted — 135 chars of source]

Hence, moving this term over to the left-hand side, the above expression for $T^{-(\nu_1+\delta)/2}\widehat{\*F}_1 \*V_{1}$ becomes

align[align omitted — 106 chars of source]

We now evaluate each of the terms on the right-hand side.

Because $T^{-\delta}\|\widehat{\*F}_1\|^2 = d_{max}$ and $(NT)^{-1}\sum_{i=1}^N \| \*X_i \*D_T\|^2 = O_p(1)$ by Assumption (ref), the order of $\*J_{1}$ is given by

align[align omitted — 442 chars of source]

Moreover, since

eqnarray[eqnarray omitted — 172 chars of source]

by Assumption (ref), we can show that

align[align omitted — 504 chars of source]

and by exactly the same arguments,

eqnarray[eqnarray omitted — 117 chars of source]

For $\*J_{4}$, we use

align[align omitted — 626 chars of source]

where the last equality makes use of the fact that $\|\*C_{+1,T}^{-1}\|= O(T^{\nu_2/2})$, as $\nu_2 > \cdots > \nu_G$ by Assumption (ref). We can further show that

align[align omitted — 602 chars of source]

where the last equality holds, because by Assumption (ref) and $(\mathrm{tr}\,\*A'\*B)^2 \le (\mathrm{tr}\,\*A'\*A )( \mathrm{tr}\,\*B'\*B)$, we have

align[align omitted — 736 chars of source]

Hence, by adding the results,

align[align omitted — 476 chars of source]

where we have used Assumptions (ref) and (ref) ($T/N^2=O(1)$ under $\nu_G < 1$) to show that $T^{(1-\delta - \nu_1 +\nu_2)/2}$ and $N^{-1/2}T^{(2-\nu_1-\delta)/2}$ are $o(1)$. The same steps can be used to show that

eqnarray[eqnarray omitted — 122 chars of source]

For $\*J_{6}$, we use

align[align omitted — 506 chars of source]

By Assumption (ref),

align[align omitted — 155 chars of source]

Another application of Assumption (ref) and Lemma (ref) gives

align[align omitted — 400 chars of source]

These results can be inserted into the expression for $T^{-\delta/2}\| \*J_{6}\|$, giving

eqnarray[eqnarray omitted — 114 chars of source]

Next up is $\*J_{7}$. By using Assumptions (ref) and (ref), and Lemma (ref), and the arguments use in evaluating $\*J_6$,

align[align omitted — 490 chars of source]

and we can similarly show that

eqnarray[eqnarray omitted — 73 chars of source]

By putting everything together, (ref) becomes

align[align omitted — 243 chars of source]

We now left multiply (ref) by $T^{-(\nu_1+\delta)/2}\widehat{\*F}_{1}^{\prime}$ to obtain that

align[align omitted — 510 chars of source]

where the third equality follows from Assumption (ref) and Lemma (ref). This implies that $\*V_1$ is at most of rank $d_1$. Similarly, we can left-multiply (ref) by $T^{-\nu_1}\*F_{1}^{0\prime}$ to obtain

align[align omitted — 159 chars of source]

which in turn implies that

align[align omitted — 205 chars of source]

Note that $T^{-(\nu_1+\delta)/2}\*F_{1}^{0\prime}\widehat{\*F}$ is of rank $d_1$, which then further indicates that $\*V_1$ has at least $d_1$ non-zero elements on the main diagonal which converge to the eigenvalues of $\+\Sigma_{F_{1}^0}\+\Sigma_{\Gamma_{1}^0}$. We now can conclude that $\*V_1$ is of rank $d_1$ in limit.

We are now ready to investigate $\widehat{\lambda}_{1,d}$ for $d\le d_1$. Because here $d\le d_1$, we then focus on $\widehat{\*F}_1^0$ and $\*V_{1}^0$, which are defined in Appendix (ref). Let us write (ref) as follows:

align[align omitted — 157 chars of source]

where $\*H_1^0$ is defined in Appendix (ref). The above expression implies that (ref) can be written as

align[align omitted — 135 chars of source]

where $\*J_{1}^0,\ldots,\*J_{9}^0$ are $\*J_{1},\ldots,\*J_{9}$ as defined in (ref), except that now $d_1$ is taken as known. Note that by the above development

align[align omitted — 143 chars of source]

Moreover, since $T^{-\nu_1}\*V_{1}^0$ converges to a full rank matrix by the argument under (ref), we have $\|(T^{-\nu_1}\*V_{1}^0)^{-1}\| = O_P(1)$, which in turn implies

align[align omitted — 206 chars of source]

This is an important result and in what follows we will use it frequently.

Let us now consider $T^{-\nu_1}(\widehat{\lambda}_{1,d} - \lambda_{1,d})$. By the definitions of $\lambda_{1,d}$ and $\widehat{\lambda}_{1,d}$ given in Section (ref),

align[align omitted — 1,322 chars of source]

with obvious definitions of $J_1,\ldots,J_5$. From (ref),

align[align omitted — 202 chars of source]

which is $o_P(1)$ under Assumption (ref). This implies $|J_1| = o_P(|J_5|)$, $|J_2| = o_P(|J_5|)$ and $|J_3| = o_P(|J_4|)$. It remains to consider $J_4$ and $J_5$. The order of the first of these terms is given by

align[align omitted — 417 chars of source]

where we have made use of the fact that $\|\*H_{g}^0\| = O_P(1)$ (see Appendix (ref)), which implies that $\|\*h_{1,d}^0\|$ is of the same order.

The order of $J_5$ is the same as that of $J_4$. In order to appreciate this, we begin by noting

align[align omitted — 1,096 chars of source]

By the proof for each term of (ref), it is easy to know that

eqnarray[eqnarray omitted — 99 chars of source]

and so

eqnarray[eqnarray omitted — 155 chars of source]

Hence, by putting everything together,

align[align omitted — 134 chars of source]

which establishes (a).

Consider (b). This proof is based on Lemma (ref). We therefore start by introducing some notation in order to make the problem here fit the one in Lemma (ref). Let us therefore denote by $\*F_1^{\perp}$ a $T\times (d_{max}- d_1)$ matrix such that $T^{-\nu_1} (\*F_1^{\perp}, \*F_1^0 \*R)'(\*F_1^{\perp}, \*F_1^0\*R) = \mathrm{diag}( \*I_{d_{max}-d_1} , \*I_{d_1} )$, where $\*R$ is a $d_1\times d_1$ rotation matrix. The matrices $T^{-\nu_1/2}\*F_1^{\perp}$, $T^{-\nu_1/2}\*F_1^0 \*R$, $\+\Sigma_{1}^0$ and $\widehat{\+\Sigma}_1 - \+\Sigma_{1}^0$ correspond to $\*Q_1$, $\*Q_2$, $\*A$ and $\*E$ of Lemma (ref). Our counterpart of the matrix $\*Q^0_1$ appearing in this other lemma is thus given by

eqnarray[eqnarray omitted — 117 chars of source]

where

align[align omitted — 265 chars of source]

Since $\widehat{\*F}^\perp$ is an orthonormal basis for a subspace that is invariant for $\widehat{\+\Sigma}_1$, we have $\widehat{\lambda}_{1,d_1+d} = \widehat{\*F}_d^{\perp\prime} \widehat{\+\Sigma}_1 \widehat{\*F}_{d}^\perp$, where $d = 1,\ldots,d_{max}-d_1$ and $\widehat{\*F}_{d}^\perp$ is the $d$-th column of $\widehat{\*F}^\perp$. Consider $\|\widehat{\*F}^\perp - T^{-\nu_1/2}\*F_1^{\perp} \|_{2}$. By the definition of $\widehat{\*F}^\perp$,

align[align omitted — 813 chars of source]

where the second and third inequalities follow from Magnus. This last result can be used to show that

align[align omitted — 847 chars of source]

where $\*F_{1,d}^{\perp}$ is the $d$-th column of $\*F_1^{\perp}$, and the last equality follows from (ref) and the proof of part (a). This completes the proof of the lemma.{$\blacksquare$}

Proof of Lemma (ref).

For (a), we take the same starting point as in the proof of part (a) in Lemma (ref), which is (ref) with $\widehat{\*F}_1$ and $\*V_{1}$ based on treating $d_1$ as known. The rationale for doing so is, as already explained in Appendix (ref), that $\widehat{d}_1$ is consistent. Pre-multiplying this equation through by $T^{-(\nu_1+\delta)/2}\*F_2^{0\prime}$ gives

align[align omitted — 417 chars of source]

or

align[align omitted — 478 chars of source]

Under Assumption (ref), the orders of $\*J_{1},\ldots,\*J_{6}$ are the same as in those in the proof of the first result of Lemma (ref). The stated orders of $\*J_{7}$ and $\*J_{8}$ are, however, not sharp and can be improved upon. The order of $T^{-(\delta+\nu_2)/2}\|\*F_2^{0\prime}\*J_{7}\|$ is given by

align[align omitted — 607 chars of source]

where the last equality follows from Assumption (ref) and Lemma (ref). We can similarly show that

align[align omitted — 761 chars of source]

where $\*F_{+2}^{0}$ and $\+\Gamma_{+2}^{0}$ are defined analogously to $\*F_{+1}^{0}$ and $\+\Gamma_{+1}^{0}$ in the proof of Lemma (ref). By using these last two results together with the orders of $\*J_{1},\ldots,\*J_{6}$ given in the proof of the first result of Lemma (ref),

align[align omitted — 439 chars of source]

where we have used $\|\*D_T^{-1}(\+\beta^0 - \widehat{\+\beta}_0 )\| =O_P(N^{-1/2} \vee T^{-1/2})$ of Lemma (ref), and Assumptions (ref) and (ref). Hence,

align[align omitted — 521 chars of source]

which together with Assumption (ref) yields

align[align omitted — 438 chars of source]

as was to be shown for (a).

Let us now consider (b). Analogously to the proof of (a), by invoking Assumption (ref) we can improve the orders of $\*J_{7}$ and $\*J_{8}$. For $\*J_7$,

align[align omitted — 317 chars of source]

where the equality follows from Assumption (ref) and Lemma (ref). For $\*J_8$,

eqnarray[eqnarray omitted — 118 chars of source]

This implies that the result in (ref) changes to (after replacing $T^{-\nu_1/2}$ by $T^{-\delta/2}$)

align[align omitted — 255 chars of source]

Hence, since $\*H_{1}^0$ is invertible with $\|\*H_{1}^0\| = O_P(1)$ and $\*H_1 = T^{-(\nu_1-\delta)/2}\*H_{1}^0$,

align[align omitted — 173 chars of source]

We are now ready to consider $\sum_{i=1}^N \| \*F_{1}^{0\prime} \+\gamma_{1,i}^0 - \widehat{\*F}_1 \widehat{\+\gamma}_{1,i}\|^2$.

align[align omitted — 490 chars of source]

We now evaluate each of the terms on the right-hand side one by one. Making use of Assumption (ref) and Lemma (ref), we get

align[align omitted — 233 chars of source]

and by another application of Lemma (ref),

align[align omitted — 119 chars of source]

For $\sum_{i=1}^N\| \*M_{\widehat{F}_1} \*F_1^0\+\gamma_{1,i}^0 \|^2$, we use (ref) from which it follows that

align[align omitted — 272 chars of source]

which in turn implies

eqnarray[eqnarray omitted — 126 chars of source]

For $\sum_{i=1}^N\|\*P_{\widehat{F}_1}\*F_{+1}^{0} \+\gamma_{+1,i}^0\|^2$, we use the result given in part (a), giving

align[align omitted — 584 chars of source]

The order of $\sum_{i=1}^N\|\*P_{\widehat{F}_1}\*F_{+2}^{0} \+\gamma_{+2,i}^0\|^2$ is the same. Hence, by adding the above results, (b) follows after simple algebra. The proof is now complete.{$\blacksquare$}

Proof of Lemma (ref).

Let $\*U_i = \*F_{+2}^0 \+\gamma_{+2,i}^0 +\*F_{1}^{0\prime} \+\gamma_{1,i}^0 - \widehat{\*F}_1 \widehat{\+\gamma}_{1,i}$. In this notation,

align[align omitted — 2,263 chars of source]

where $\*K_{1},\ldots,\*K_{16}$ are implicitly defined. Analogously to the proof of the first result of Lemma (ref) we move $\*K_{16} = \*F_2^0 (N^{-1}\+\Gamma_{2}^{0\prime} \+\Gamma_{2}^0)(T^{-(\nu_2+\delta)/2} \*F_2^{0\prime} \widehat{\*F}_2)$ over to the left, giving

eqnarray[eqnarray omitted — 213 chars of source]

By using the same steps employed in the proof of the first result of Lemma (ref), we can show that

eqnarray[eqnarray omitted — 141 chars of source]

For $\*K_{4}$,

align[align omitted — 891 chars of source]

where the second equality follows Lemma (ref), and the third follows from Assumptions (ref), (ref), and (ref). The same arguments can be used to show that

eqnarray[eqnarray omitted — 117 chars of source]

For $\*K_{6}$,

align[align omitted — 387 chars of source]

where the development is similar to (ref). The order of $T^{-\delta/2}\|\*K_{7}\|$ is the same.

For $\*K_{8}$,

align[align omitted — 337 chars of source]

where the last equality holds by Lemma (ref).

Further use of Lemma (ref) gives

align[align omitted — 334 chars of source]

and we can show that $T^{-\delta/2}\| \*K_{10} \|$ is of the same order.

$\*K_{11}$ requires more work. We begin by expanding it in the following way:

align[align omitted — 1,304 chars of source]

where, similarly to the analysis of $\*K_{6}$ and using $\|\*P_{\widehat F_1}\|_2 = 1$,

align[align omitted — 321 chars of source]

Also, making use of (ref), we can show that

align[align omitted — 251 chars of source]

Similarly, for $K_{113}$, we can show that

align[align omitted — 82 chars of source]

where we have used Lemmas (ref) and (ref), and Assumption (ref). Further use of Lemma (ref) shows that $K_{114}$ is of the following order:

eqnarray[eqnarray omitted — 152 chars of source]

while the order of $K_{115}$ is

eqnarray[eqnarray omitted — 156 chars of source]

By inserting the above results into (ref), we obtain

align[align omitted — 333 chars of source]

The order of $T^{-\delta/2}\| \*K_{12} \|$ is the same, which can be shown using the above steps.

We move on to $\*K_{13}$, whose order is given by

align[align omitted — 1,010 chars of source]

where the second equality follows from Lemma (ref) and Assumption (ref). The order of $T^{-\delta/2}\| \*K_{14} \|$ is the same.

For $\*K_{15}$,

align[align omitted — 582 chars of source]

where the second equality follows from Lemma (ref).

We now insert the above results for $\*K_1,\ldots,\*K_{15}$ into (ref). But first we left multiply by $T^{-\nu_2}\*F_2^{0\prime}$ and $T^{-(\nu_2+\delta)/2}\widehat{\*F}_2$ respectively. It then gives

align[align omitted — 382 chars of source]

and

align[align omitted — 347 chars of source]

The rest of the proofs of (a) and (b) follows from the same arguments used in Lemma (ref). It is therefore omitted.{$\blacksquare$}

Before continuing onto the proof of Lemma (ref), we note that the rates given in Lemma (ref) can actually be improved upon. Suppose that $d_2$ in known. Consider the next term which is a part of $\*K_{15}$:

align[align omitted — 883 chars of source]

where $\*V_2^0$ and $\*H_2$ are defined in Appendix (ref). Simple algebra shows that the term $T^{-(\nu_2-\nu_3)}$ in $\tau_{NT}$ of Lemma (ref) can be dropped.

Proof of Lemma (ref).

Consider (a). Let us assume without loss of generality that $\+\beta^0 = \*0_{d_x\times 1}$, as in Bai. It then follows that

align[align omitted — 1,744 chars of source]

Note that by expanding the term $\*M_{\widehat{F}}\*F^0$ as in the proof for the second result of this lemma, we can obtain that

align[align omitted — 467 chars of source]

It follows that

align[align omitted — 760 chars of source]

where the first term on the right is quadratic in $\|\*D_T^{-1}( \widehat{\+\beta}_1-\widehat{\+\beta}_0 )\|$. Hence, to ensure the right-hand side is non-positive, $\|\*D_T^{-1}( \widehat{\+\beta}_1-\widehat{\+\beta}_0 )\|$ cannot converge to zero at a rate faster than $(NT)^{-1/2} \vee \|\*D_T^{-1} ( \widehat{\+\beta}_0-\+\beta^0)\|$. It follows that

align[align omitted — 143 chars of source]

as was to be shown.

We now turn to (b). Note that

align[align omitted — 440 chars of source]

with obvious definitions of $\*L$ and $\*B$.

Consider $\*L$. Let $\*Q_{g} = (N^{-1}\+\Gamma_g^{0\prime}\+\Gamma_g^0)(T^{-(\nu_g+\delta)/2} \*F_{g}^{0\prime}\widehat{\*F}_g )$, such that $\*H_{g}^{-1} = T^{-(\nu_g+\delta)/2} \*V_{g}^0\*Q_{g}^{-1}$. We also introduce $\*e_{g,i}$, which is defined to be zero for $g=1$ and $\*e_{g,i}= \sum_{j=1}^{g-1} (\*F_{j}^0\gamma_{j,i}^0 - \widehat{\*F}_j \widehat{\+\gamma}_{j,i})$ for $g=2,\ldots,G$. In this notation,

align[align omitted — 262 chars of source]

where

align*[align* omitted — 3,335 chars of source]

We now evaluate each of these terms. From the analysis of $\*L_{2}$ below, it is easy to show that $\|\*L_{1} \| = o_P(\|\*D_T^{-1}( \widehat{\+\beta}_0 - \+\beta^0)\|)$. We therefore start from $\*L_{2}$, which we write as

align[align omitted — 1,013 chars of source]

We will use this expression for $\*L_{2}$ later.

Let us now move on to $\*L_{3}$.

align[align omitted — 537 chars of source]

where, by using arguments that are similar to those used in the proofs of Lemmas (ref) and (ref),

align[align omitted — 760 chars of source]

implying

align[align omitted — 459 chars of source]

The same steps can be used to show that $\*L_{4}$ and $\*L_{5}$ are of the same order.

For $\*L_{6}$, write

align[align omitted — 691 chars of source]

where the first term on the right is bounded by

align[align omitted — 897 chars of source]

The second term is of the same order. Thus, $\|\*L_{6} \| =o_P ( \| \*D_T^{-1}(\+\beta^0-\widehat{\+\beta}_0) \| )$. The same is true for $\|\*L_{7}\|$.

We now examine $\*L_{8}$. Let us define $\+\Sigma_{\varepsilon} = N^{-1}\sum_{i=1}^N \+\Sigma_{\varepsilon,i}$, in which $\+\Sigma_{\varepsilon,i} $ has been defined in Assumption (ref). $\*L_{8}$ can then be written as

align[align omitted — 632 chars of source]

For the first term on the right,

align[align omitted — 597 chars of source]

where the last equality holds by $NT^{-\nu_G}<\infty $. Let

align[align omitted — 246 chars of source]

Note how $\*A_{1} = \operatorname*{p\!\lim}_{N,T\to\infty}E(\overline{\*A}_1 | \mathcal{C})$. In this notation,

align[align omitted — 1,087 chars of source]

where the second equality follows from arguments similar to those used in (ref). Note also that $\lim \sqrt{N}T^{-\nu_G/2} <\infty$. Further use of the same steps used by JYGH establishes that the second and third terms of $\*L_{8}$ are $o_P(\|\*D_T^{-1}(\widehat{\+\beta}_0 -\+\beta^0) \|)+ o_P((NT)^{-1/2})$. Hence,

align[align omitted — 128 chars of source]

$\*L_{9}$ can be written as

align[align omitted — 512 chars of source]

where the first term on the right-hand side is bounded by

align[align omitted — 629 chars of source]

where the first inequality follows from $\|T^{-(\nu_g-1)/2} \+\Gamma_g^{0\prime}\+\varepsilon \*F_{g}^0 \|=O_P(\sqrt{NT})$. The second term on the right-hand side of $\*L_{9}$ is also $o_P( (NT)^{-1/2} )$. We therefore conclude that $\|\*L_{9}\| = o_P( (NT)^{-1/2} )$.

$\*L_{10}$ can be written more compactly as

align[align omitted — 342 chars of source]

which we will again make use of later.

Let us move on to $\*L_{11}$. We begin by rewriting $\*e_{g,j}$ as

align[align omitted — 465 chars of source]

where $\*F_{+d}^0 = (\*F_{d+1}^0,\ldots, \*F_{G}^0)$ and $\+\gamma_{+d,j}^0 = (\+\gamma_{d+1,j}^{0\prime},\ldots, \+\gamma_{G,j}^{0\prime})'$, as in Appendix (ref). By inserting this into $\*L_{11}$, we get

align[align omitted — 464 chars of source]

Hence, $\*L_{11}$ can be written as a sum of five terms. There is no need to study the forth term, the one due to $\*P_{\widehat{F}_d } \*e_{d,j}$, as we can keep expanding $\*e_{d,j}$ until we cannot. The first and fifth terms are $o_P(\|\*D_T^{-1}(\widehat{\+\beta}_0 -\+\beta^0) \|)$ and $o_P((NT)^{-1/2})$, respectively, by the same arguments used for evaluating $\*L_{7}$ and $\*L_{8}$. Moreover, the steps used for evaluating $\*L_{9}$ can be used to show that

align[align omitted — 1,128 chars of source]

The second term in $\*L_{11}$ is therefore negligible. It remains to consider the third term, which is

align[align omitted — 620 chars of source]

where we have used Assumption (ref). Therefore, $\|\*L_{11}\| = o_P( (NT)^{-1/2} )$.

For $\*L_{12}$,

align[align omitted — 433 chars of source]

where the second equality follows from the construction of $\*e_{g,j}$. Hence, since $\*L_{9}$ is negligible, $\*L_{12}$ is also negligible.

Next up is $\*L_{13}$, whose order can be worked out in the following way:

align[align omitted — 1,026 chars of source]

where the first inequality follows from the construction of $\*e_{g,j}$, and the last equality follows by going through a development similar to the first result of Lemma (ref) and further using a development similar to (ref). The same arguments show that $\|\*L_{14}\|$ and $\|\*L_{15}\|$ are negligible, too.

We now put everything together. This yields

align[align omitted — 491 chars of source]

which in turn implies

align[align omitted — 384 chars of source]

where $\*Z_i(\*F) $ is defined in Assumption (ref), and

align*[align* omitted — 107 chars of source]

This expression for $\*D_T^{-1}(\widehat{\+\beta}_1-\+\beta^0)$ can be inserted into $\*D_T^{-1}(\widehat{\+\beta}_0 -\+\beta^0)$, giving

align[align omitted — 440 chars of source]

which can be solved for $\*D_T^{-1}(\widehat{\+\beta}_0 -\+\beta^0)$

align[align omitted — 356 chars of source]

Also, making use of Assumption (ref), it is not difficult to show that $(NT)^{-1}\sum_{i=1}^N\*D_T \*Z_i(\widehat{\*F})'\+\varepsilon_i=O_P((NT)^{-1/2})$. Hence,

align[align omitted — 429 chars of source]

By using the fact that

align[align omitted — 184 chars of source]

the left-hand side of this last equation can be written as

align[align omitted — 387 chars of source]

It follows that

align[align omitted — 550 chars of source]

Consider the $o_P( \sqrt{NT}\|\*D_T^{-1} ( \widehat{\+\beta}_1-\widehat{\+\beta}_0)\| )$ reminder term. By the first result of this lemma, $\|\*D_T^{-1}( \widehat{\+\beta}_1-\widehat{\+\beta}_0 )\| = O_P((NT)^{-1/2}\vee \|\*D_T^{-1} ( \widehat{\+\beta}_0-\+\beta^0)\| )$. This can be inserted into (ref), giving

align[align omitted — 311 chars of source]

which in turn implies

align[align omitted — 80 chars of source]

The second result then follows.{$\blacksquare$}

Proofs of main results

Proof of Lemma (ref).

Without loss of generality, we assume that $\+\beta^0 = \*0_{d_x\times 1}$. This implies

align[align omitted — 1,149 chars of source]

where the second equality follows from Lemma (ref), and $\+\theta = \+\eta + \*B^{-1}\*C\+\beta$ and $\*B = N^{-1}\+\Gamma^{0\prime}\+\Gamma^0 \otimes \*I_T$ are defined as on page 1265 of Bai. Note that for $\+\beta\in \mathbb{R}^{d_x}$, we may have

eqnarray[eqnarray omitted — 176 chars of source]

In our proof of consistency, we consider two cases; (i) $\| \*D_T^{-1}\+\beta\| \le C$ and (ii) $\| \*D_T^{-1}\+\beta\| > C$, where $C$ is a large positive constant. Under (i),

eqnarray[eqnarray omitted — 196 chars of source]

by Lemma (ref). The expression given in (ref) for $(NT)^{-1}[\mathrm{SSR}( \+\beta, F) - \mathrm{SSR}(\+\beta^0, \*F^0)]$ therefore reduces to

align[align omitted — 306 chars of source]

where $\+\beta'\*D_T^{-1}\*B(\*F)\*D_T^{-1}\+\beta$ does not involve $\*F^0$ and $\sum_{i=1}^N\+\gamma_{i}^{0\prime} \*F^{0\prime} \*M_{F} \+\varepsilon_i$ is independent of $\+\beta$. Hence, provided $d_{max} \ge d_f$, the consistency of $\*D_T^{-1}\widehat{\+\beta}_0$ in case (i) follows from the same arguments as in Bai.

Under (ii), (ref) can be written as

align[align omitted — 866 chars of source]

where $c_0$ is defined in Assumption (ref), and the second inequality follows from the fact that the quadratic term dominates the linear one for large values of $C$. Hence, $(NT)^{-1}[\mathrm{SSR}( \+\beta, \*F) - \mathrm{SSR}(\+\beta^0, \*F^0)] > 0$, but from the definition of $\widehat{\+\beta}_0$ we also know that $\mathrm{SSR}( \widehat{\+\beta}_0, \widehat{\*F}) - \mathrm{SSR}(\+\beta^0, \*F^0) \le 0$, which means that $\*D_T^{-1}\widehat{\+\beta}_0$ cannot belong to (ii).

Note that since $\|(NT)^{-1}\sum_{i=1}^N \*D_T\*X_i'\*M_{F} \+\varepsilon_i \|=O_P( N^{-1/2}\vee T^{-1/2} )$ by Lemma (ref), all we need is $C = C^0 (N^{-1/2}\vee T^{-1/2})$ for some large constant $C^0$ in order to ensure that the last inequality of (ref) holds. This implies $\*D_T^{-1}(\widehat{\+\beta}_0-\+\beta^0) = O_P( N^{-1/2}\vee T^{-1/2} )$, so the proof is complete. {$\blacksquare$}

Proof of Lemma (ref).

Suppose first that $d_1=0$, such that $d_f=0$. In this case,

align[align omitted — 259 chars of source]

This implies $\tau_N \asymp 1/\ln (T\vee N)$, which is much larger than $\widehat{\lambda}_{1,1}/\widehat{\lambda}_{1,0}$. The result then follows immediately.

Suppose now instead that $d_1>0$. Straightforward algebra reveals that

align[align omitted — 423 chars of source]

which together with Assumption (ref) implies $\tau_N \asymp 1/\ln (T\vee N)$. Note that for $d=1,\ldots, d_1$,

align[align omitted — 707 chars of source]

where the fourth equality follows from (ref) and the last step is dues to (ref) and Assumption (ref). Using Lemma (ref), (ref) and (ref), we obtain that

align[align omitted — 361 chars of source]

for $d=1,\ldots, d_1-1$, and $\frac{\widehat{\lambda}_{1,d_1+1} }{\widehat{\lambda}_{1,d_1} } = O_P( T^{-(\nu_1-\nu_2)})$. Moreover, for $d=d_1+1,\ldots, d_{max}$, $\frac{T^{-\nu_1}\widehat{\lambda}_{1,d}}{T^{-\nu_1}\widehat{\lambda}_{1,0}} = O_P(T^{-(\nu_1-\nu_2)})$, which is less than $\tau_N $ by (ref) and Lemma (ref). The required result follows from this and the definition of $\widehat{d}_1$. {$\blacksquare$}

Proof of Lemma (ref).

Each step in the sequential procedure of Step (ref) introduces additional remainder terms that all converge to zero under Assumptions (ref) and (ref) by Lemmas (ref) and (ref). This proves (a). Part (b) follows from the (rotational) consistency of $\widehat{\*F}_g$ established as a part of the proofs of Lemmas (ref) and (ref).{$\blacksquare$}

Proof of Theorem (ref).

By Lemma (ref), the rate of convergence given in Lemma (ref) is therefore not the best one possible. The improved rate implies that

align[align omitted — 175 chars of source]

which can be inserted back into (ref), leading to

align[align omitted — 460 chars of source]

Note that $\*Z_i(\widehat{\*F})$ is $\widehat{\*Z}_i$ with $\widehat a_{ij}$ replaced by $a_{ij}$. We now show that the effect of the estimation of $a_{ij}$ is negligible. We begin by noting how

align[align omitted — 345 chars of source]

Here,

align[align omitted — 1,385 chars of source]

where the second equality here follows from the definition of $\*V_{g}^0$, which is again based on taking $d_g$ as known, while the last equality follows from direct calculation. The effect of the estimation of $a_{ij}$ in $\widehat{\*Z}_i$ is therefore negligible, which in turn implies that the right-hand side of (ref) becomes

align[align omitted — 573 chars of source]

It follows that

align[align omitted — 270 chars of source]

Consider $(NT)^{-1/2}\sum_{i=1}^N\*D_T \*Z_i(\widehat{\*F})'\+\varepsilon_i$. From the definition of $ \*Z_i(\*F)$,

align[align omitted — 373 chars of source]

where

align[align omitted — 222 chars of source]

Here,

align[align omitted — 538 chars of source]

with obvious implicit definitions of $\*R_{11},\ldots,\*R_{14}$ and $\*H = \mathrm{diag}(\*H_{1}, \ldots, \*H_{G} )$. Let $\*R_{1m,j}$ be the $j$-th row of $\*R_{1m}$ for $m\in\{1,\ldots,4\}$. In this notation,

align[align omitted — 306 chars of source]

where the equality follows from Lemma (ref) and the fact that $T^{(\nu_g-\delta)/2}\*H_g =O_P(1)$. Similarly, for $\*R_{12}$,

align[align omitted — 298 chars of source]

where the second equality is due to (ref).

Consider $\*R_{14}$. From $T^{-\delta/2}\|\widehat{\*F}-\*F^0\*H \|=o_P(1)$ and $\|T^{(\nu_g-\delta)/2}\*H_g\| =O_P(1)$, we obtain

align[align omitted — 191 chars of source]

Together with Lemma (ref) this implies

align[align omitted — 703 chars of source]

Finally, let us consider $\*R_{13}$.

align[align omitted — 409 chars of source]

Expanding $\widehat{\*F}_g \*H_g^{-1}-\*F_g^0 $ as we did for $\*L$ above, we then just need to focus on the leading terms equivalent to $\*L_2$, $\*L_{8}$, and $\*L_{10}$.

align[align omitted — 1,019 chars of source]

It is easy to see that $\| \*R_{131}\|=o_P(\|\+\beta^0 -\widehat{\+\beta}_0 \|)$, so negligible. We now examine $\*R_{133}$. Then we can write

align[align omitted — 847 chars of source]

where the first equality follows from (ref), and the last step follows from some routine analysis using Assumption (ref). Below, we investigate $\*R_{132}$, which is one source of the bias term.

align[align omitted — 1,197 chars of source]

where the second equality follows from the definition of $\*Q_g$, and the third equality follows from (ref). Then we are able to further write

align[align omitted — 709 chars of source]

Hence, by adding the results,

align[align omitted — 374 chars of source]

$\sqrt{NT} \*R_2$ can be evaluated in exactly the same way, and the limiting representation is analogous to the one given above for $\sqrt{NT} \*R_1$ with $\*X_i$ replaced by $-\sum_{j=1}^{N}\*X_ja_{ij}$. Moreover, $\|\*B(\widehat{\*F}) - \*B(\*F^0)\| = o_P(1).$ It follows that if we define

align[align omitted — 263 chars of source]

where $\*Z_i(0) = \*X_i - \sum_{j=1}^{N}\*X_ja_{ij}$, then (ref) reduces to

align[align omitted — 210 chars of source]

which in turn implies that (ref) becomes

align[align omitted — 309 chars of source]

The required result is now implied by Assumption (ref). {$\blacksquare$}

Proof of Corollary (ref).

Under $NT^{-\nu_G}\to 0$, (ref) in the proof of Theorem (ref) may be written as

align*[align* omitted — 169 chars of source]

Consider $(NT)^{-1} \sum_{i=1}^N \*D_T\*X_i'\*M_{\widehat{F}} \+\varepsilon_i$, which we can write as

align[align omitted — 313 chars of source]

As in the proof of Theorem (ref),

align[align omitted — 648 chars of source]

where $\|\*R_{11}\|$, $\|\*R_{12}\|$ and $\|\*R_{14}\|$ are all $o_P((NT)^{-1/2})$, just as before. Let us therefore consider $\*R_{13}$, which is the source of the bias in the asymptotic distribution of $\sqrt{NT}\*D_T^{-1} (\widehat{\+\beta} - \+\beta^0 )$. The purpose of Assumption (ref) is to control this term. Suppose first that condition (b) holds. Let $x_{j,k,t}$ denote the $j$-th row of $\*x_{k,t}$. In this notation,

align[align omitted — 666 chars of source]

Moreover, by applying Assumption (ref) (b) to the expression given for $\|T^{-\delta/2}\widehat{\*F}_{2,d} - T^{-\nu_2/2}\*F_2^0\*h_{2,d}^{0}\|$ in the proof of Lemma (ref), we can show that

align[align omitted — 87 chars of source]

Making use of these results, we obtain

align[align omitted — 237 chars of source]

Alternatively, we may invoke Assumption (ref) (a) to arrive at the same result. In this case, $\|T^{-\delta/2}\mathrm{vec}(\widehat{\*F}-\*F^0\*H)' \| = o_P( 1 )$, but we also have

align[align omitted — 965 chars of source]

and so $\|\*R_{3,j} \|$ is of the same order as before. The proof under Assumption (ref) (c) is simpler and is therefore omitted.

Hence, by adding the results,

align[align omitted — 114 chars of source]

We have therefore shown that

align[align omitted — 131 chars of source]

and we can similarly show that

align[align omitted — 146 chars of source]

These results can be inserted into (ref), giving

align[align omitted — 160 chars of source]

The sought result now follows from Assumption (ref).{$\blacksquare$}