Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
206,230 characters · 13 sections · 104 citation commands
\newtheorem{theorem}{Theorem}[section] \newtheorem{corollary}{Corollary}[section] \newtheorem{lemma}{Lemma}[section] \theoremstyle{definition} \newtheorem{assumption}{Assumption} \newtheorem{remark}{Remark} \newtheorem{step}{Step} \newtheorem{dgp}{DGP}
\numberwithin{corollary}{section} \numberwithin{equation}{section} \numberwithin{lemma}{section} \numberwithin{remark}{section} \numberwithin{theorem}{section}
\allowdisplaybreaks[4]
The use of panel data in regression analysis has attracted considerable attention in the empirical literature in economics and elsewhere. A major reason for this is the ability to deal with the presence of unobserved heterogeneity, and the problem that this causes when said heterogeneity is correlated with the regressors. The simplest approach is to assume that the unobserved heterogeneity is made up of additive individual- and time-specific constants, or “fixed effects”. Such effects can add hugely to the fit of the model, and are typically found to be much more important than the regressors predicted by economic theory. However, in many applications fixed effects are unlikely to be enough to capture the unobserved heterogeneity (see, for example, Lemmonetal2008, and DeAngeloRoll2015), and this has in turn spurred much interest in so-called “interactive effects” models, in which the individual and time effects enter in an multiplicative way. One of the most common approaches to such interactive effects models by far is the principal components (PC) approach of Bai. In fact, it is so common that it has given rise to a separate strand of literature (see Ando, bai2014, LiQianSu, MoonWeidner2015, to mention a few).\footnote{While popular, PC is not the only approach that can be used to estimate interactive effects models. One alternative is the so-called “common correlated effects” (CCE) estimator of Pesaran2006. The problem with this approach is that it requires that the number of time effects, or factors, is bounded by the number of observables, and that the regressors load on the same factors as the dependent variable, which is not necessary in PC. There is also the generalized method of moments (GMM) approach of ALS13. However, this approach supposes that the number of time periods is “small”, and is not suitable for the type of large panel data sets that we have in mind.} The present paper aims to contribute to this strand, and it does so in at least three ways.
The first contribution of the paper is to consider a very general data generating process (DGP) that includes most of the specifications considered previously in the literature as special cases. The only requirement is that suitably normalized sample second moment matrices of the factors and regressors have positive definite limits. This is noteworthy because the existing literature is almost exclusively based on the assumption that both the factors and regressors are stationary. The only exceptions known to us are BaiKaoNg, DGP2020 and HJPS, but they assume instead that either the regressors or the factors and the regressors are pure unit root processes, which is also not realistic. Indeed, regressors and factors of different order of magnitude are likely to be the rule rather than the exception, especially in economic and financial data, due to differences in persistence over time.
The unrestricted DGP is important in itself but also because it can be accommodated without requiring any knowledge thereof. Hence, not only do we treat the factors and their number as unknown, but we also do not require any knowledge of the order of magnitude of both factors and regressors. An important implication of this is that there is no need to distinguish between deterministic and stochastic factors, or stationary and non-stationary factors. In the existing literature, deterministic factors are often treated as known, and are projected out prior to the application of PC (see, for example, MoonWeidner2015). The problem here is that there is typically great uncertainty over which deterministic terms to include, which raises the issue of model misspecification. The fact that in the present paper deterministic terms are treated as additional factors means that the problem of deciding on which terms to include does not arise. Similarly, while the regressors can be tested for unit roots, and the estimation can be made conditional on the test outcome, this raises the issue of pre-testing bias. In the present paper we do not require any knowledge about the order of integration of the regressors, which means that there is no need for any pre-testing.
Equally as important as the general model formulation and its empirical appeal is the extension of the existing econometric theory, which has not yet ventured much outside the stationary or pure unit root environments. This is our second major contribution. The main difficulty here is not the unrestricted specification of the factors and regressors per se, but rather that the order of magnitude of the factors may differ. In particular, the problem is that the nonlinearity of the PC estimator distorts the signal coming from the factors, just as it does in estimation of nonlinear regression models with mixtures of integrated regressors (see, for example, ParkPhillips2000). This is true if both the number and order of the factors are known, and the problem does not become any simpler when these quantities are treated as unknown, as they are here. An additional problem that then arises is that existing studies on the selection of the number of factors all require that the data are stationary (see, for example, BaiNg2002, and AH13), and it is not obvious how one should go about this when the order of magnitude of the factor is unknown.
Intuitively, the factors whose order is largest should dominate the PC estimator. This motivates the use of an iterative estimation procedure in which the factors and their number are estimated in order according to their magnitude with relatively larger factors being estimated first. We begin by prescribing a large number of factors, and estimate the resulting model by PC. The estimated factors only capture the most dominating factors whose order of magnitude is largest. In spite of this, we can show that the estimator is consistent, albeit at a relatively low rate of convergence. The rate is, however, high enough to ensure that the number of dominating factors can be consistently estimated using a version of the eigenvalue ratio approach of LY12, and AH13. We then apply PC conditional on the first-step factor estimates, and estimate the second most dominating factors. This procedure continues until we cannot identify any more factors, and we can show that both the estimated factors and their number are consistent. Because of the iterative fashion in which the factors are estimated, we refer to the new estimation procedure as “iterative PC” (IPC), which is shown to be asymptotically normal and “oracle efficient”.
Our third contribution is to point to a “blessing” of trending factors. The blessing occurs if the magnitude of the factors is sufficiently large, in which case the otherwise so common asymptotic bias of the PC approach can be completely eliminated without imposing any additional restrictions on the cross-sectional and time series dependencies of the regression errors. This is noteworthy, because the conclusion made in the previous literature suggests that in order to eliminate the asymptotic bias, the errors have to be independent.
The reminder of the paper is organized as follows. We begin by describing the model that we will be considering and the proposed IPC approach that we will use to estimate it. This is done in Section (ref). Section (ref) presents the formal assumptions and our main asymptotic results, whose small-sample accuracy is evaluated using Monte Carlo simulations in Section (ref). Section (ref) reports the results obtained from an empirical application to the long-run relationship between US house prices and income. For the sake of space, we present an extra empirical study on the returns to scale in the US banking industry in the supplementary appendix. Section (ref) concludes. All proofs are relegated to the online appendix.
Consider the panel data variable $y_{i,t}$, observable for $i=1,\ldots,N$ cross-sectional units and $t=1,\ldots,T$ time periods. The model of this variable that we will be considering is given by
where $\*x_{i,t} = (x_{1,i,t},\ldots,x_{d_x,i,t})^{\prime }$ is a $d_x\times 1$ vector of regressors, $\*f_{t}^0=(f_{1,t}^{0},\ldots,f_{d_f,t}^{0})^{\prime }$ is a $d_f\times 1$ vector of unobservable common factors with $\+\gamma_{i}^0 = (\gamma_{1,i}^0,\ldots,\gamma_{d_f,i}^0)^{\prime }$ being a conformable vector of factor loadings, and $\varepsilon_{i,t}$ is an idiosyncratic error term. Moreover, the factors are divided into groups according to their order of magnitude. There are $G$ groups of size $d_1,\ldots,d_G$, which means that $d_1+\cdots+d_G = d_f$. Because the grouping is unknown, we may without loss of generality assume that the factors are ordered, such that the first $d_1$ factors have the highest order of magnitude, the next $d_2$ factors have the second highest order, and so on. Hence, if we denote by $\*f_{g,t}^0$ and $\+\gamma_{g,i}^0$ the $d_g\times 1$ vectors of factors and loadings associated with group $g$, respectively, then $\+\gamma_{i}^{0\prime} \*f_{t}^0 = \sum_{g=1}^G \+\gamma_{g,i}^{0\prime}\*f_{g,t}^0$, where $\*f_{t}^0 =(\* f_{1,t}^{0\prime},\ldots,\*f_{G,t}^{0\prime})^{\prime}=(f_{1,t}^{0},\ldots,f_{d_f,t}^{0})^{\prime }$ and $\+\gamma_{i}^0 = (\+\gamma_{1,i}^{0\prime},\ldots,\+\gamma_{G,i}^{0\prime})^{\prime }= (\gamma_{1,i}^0,\ldots,\gamma_{d_f,i}^0)^{\prime }$. If $d_f = 0$, then $G=0$ and $\+\gamma_{i}^{0\prime} \*f_{t}^0 = 0$.
It is useful to write (ref) on stacked vector form. Let us therefore introduce the $T\times 1$ vectors $\*y_i =(y_{i,1},\ldots, y_{i,T})^{\prime }$ and $\+\varepsilon_i = (\varepsilon_{i,1},\ldots, \varepsilon_{i,T})^{\prime }$, the $T\times d_f$ matrix $\*F^0 = (\*f_{1}^0,\ldots,\*f_{T}^0)^{\prime }$, and the $T\times d_x$ matrix $\*X_i =(\*x_{i,1},\ldots, \*x_{i,T})^{\prime }$. Analogous to $\*f_{t}^0$, $\*F^0$ is further partitioned as $\*F^0 =(\*F_{1}^{0},\ldots,\*F_{G}^{0})^{\prime }$, where the $T\times d_g$ matrix $\*F^0_g = (\*f_{g,1}^0,\ldots,\*f_{g,T}^0)^{\prime }$ contains the stacked factor observations for group $g$. In this notation, (ref) can be written as
The goal of this paper is to infer $\+\beta^0$. As mentioned in Section (ref), however, because of the generality of the model being considered, the main difficulty in the estimation process is how to control for $\*F^0$. Our proposed estimation procedure consists of three steps. We first initialize the estimation procedure by applying the PC estimator of Bai. However, because the first group of factors dominates all the other groups in terms of order of magnitude, the first-step PC factor estimator will only be estimating (the space spanned by) $\*F_{1}^0$. The second step of the procedure therefore involves iteratively applying PC conditional on previous factor estimates to estimate all subsequent groups of factors; hence, the “I” in IPC. In the third and final step, we estimate $\+\beta^0$ conditional on the second-step IPC estimator of $\*F^0$ and the first-step PC estimator of $\+\beta^0$.
Before we state our assumptions and asymptotic results, we introduce some notation. Specifically, if $\*A$ is a matrix, $\lambda_{min}(\*A)$ and $\lambda_{max}(\*A)$ signify its smallest and largest eigenvalues, respectively, $\text{tr}\,\*A$ signifies its trace, and $\|\*A\| = \sqrt{\text{tr}\,\*A'\*A}$ and $\|\*A\|_2 = \sqrt{\lambda_{max}(\*A)}$ signify its Frobenius and spectral norms, respectively. We write $\*A > 0$ to signify that $\*A$ is positive definite. If $\*B$ is also a matrix, then $\mathrm{diag}(\*A, \*B)$ denotes the block-diagonal matrix that takes $\*A$ ($\*B$) as the upper left (lower right) block. The symbols $\to_D$, $\to_P$ and $MN(\cdot, \cdot)$ signify convergence in distribution, convergence in probability and a mixed normal distribution, respectively. We use $N,\,T\to\infty$ to indicate that the limit has been taken while passing both $N$ and $T$ to infinity. We use w.p.a.1 to denote with probability approaching one.
Assumption (ref) is concerned with the order of magnitude of $\*f_t^0$ and $\*x_{i,t}$. It is therefore key. The assumption is stated in terms of the required moment conditions rather than primitive assumptions on $\*f_t^0$ and $\*x_{i,t}$. The reason is that we would like to be agnostic about these variables.
Assumption (ref) (a) is very general in that it imposes almost no restrictions on the type of trending behaviour that $\*x_{i,t}$ may have. The trending can be deterministic but it can also be stochastic, as in the presence of unit roots. Either way, the degree of the trending is not restricted in any way, provided that it is finite. The only requirement is that $0<\lambda _{min}(\+\Sigma _{X})\leq \lambda _{max}(\+\Sigma _{X})<\infty$ w.p.a.1, which means that the elements of $\*x_{i,t}$ cannot be asymptotically collinear. Note that $\+\Sigma _{X}$ is not required to be a constant matrix, as this would rule out regressors that are stochastically integrated. This is very different when compared to the bulk of the previous PC-based interactive effects literature in which $\*x_{i,t}$ is assumed to be either stationary (see, for example, Bai, and MoonWeidner2015), such that $\kappa _{1}=\cdots =\kappa _{d_{x}}=0$, or unit-root non-stationary (see BaiKaoNg, and DGP2020), such that $\kappa _{1}=\cdots =\kappa _{d_{x}}=1$. Assumption (ref) (a) is analogous to the work of, for example, DongLinton, wherein time series of different order of magnitude are considered.
The first requirement of Assumption (ref) (b) is quite mild and holds if a central limit theorem in only one of the two panel dimensions applies to the normalized sum of $\mathbf{D}_{T}\mathbf{X}_{i}^{\prime }\+\varepsilon_{i}$. The second requirement is quite common in the literature, and is expected to hold as long as $\varepsilon_{i,t}$ has zero mean, and weak serial and cross-sectional correlation (see MoonWeidner2015, for a discussion).
Assumption (ref) (c) is similar to (a) in that it leaves the trending behavior of the factors essentially unrestricted. A majority of previous PC-based works assume that $T^{-1}\*F^{0\prime }\*F^{0}$ converges to positive definite matrix (see, for example, Bai, and MoonWeidner2015). Notable exceptions include Bai2004, BaiKaoNg, and Choi2017, in which $\*f_{t}^{0}$ is assumed to follow pure unit root process, and BaiNg2004, who allow for a mix of stationary and unit root factors. The only study that comes close to ours in terms of the generality of the factors is that of Westerlund2018. However, he assumes that $\*x_{i,t}$ has a factor structure that loads on the same set of factors as $y_{i,t}$, which is not required here. Also, unlike Westerlund2018, we do not require $\nu _{G}\geq 1$, but also allow $1/2<\nu _{G}<1$, such that $T^{-1}\*F_{G}^{0\prime }\*F_{G}^{0}$ is asymptotically singular. The signal coming from $\*F_{G}^{0}$ is therefore even weaker than under stationarity, and we will therefore refer to these factors as “signal-weak", as opposed to “weak" in usual sense (see, for example, Chudiketal2011 and UY2020, for discussions). One example of such signal-weak factors is when $\*f_{G,t}^{0}$ is stationary and sparse.
Assumption (ref) (d) is standard and ensures that each factor has a nontrivial contribution to the variance of $y_{i,t}$.
The requirement that the smallest eigenvalue of $\*B(\*F)$ should be positive is equivalent to requiring that $\*B(\*F)$ be positive definite uniformly in $\*F$, which is a non-collinearity condition that rules out low-rank regressors (see MoonWeidner2015, for a discussion). It demands that the regressors in $\*X_{i}$ have enough variation after projecting out all variation that can be explained by $\*F^{0}$ and $\+\gamma _{i}^{0}$. Time-invariant regressors are therefore not allowed, which is similar to the conventional cross-section fixed effects-only OLS condition. Assumption (ref) is a high-level condition, just like Assumption (ref) (see, for example, Bai, and MoonWeidner2015, for similar high-level conditions). The reason for formulating it in this way is again that we want to avoid making highly specific and potentially invalid assumptions about the DGP of $\*x_{i,t}$, $\*f_{t}^{0}$ and $\+\gamma _{i}^{0}$. If all the regressors are stationary, Assumption (ref) reduces to Assumption A in Bai.
Assumptions (ref) and (ref) are enough to ensure that the initial estimator $\widehat{\+\beta}_0$ of $\+\beta^0$ is consistent.
Lemma (ref) establishes that $\widehat{\+\beta}_0$ is consistent for $\+\beta^0$ and that the rate of convergence is $\|\*D_T\|/\min\{\sqrt{N}, \sqrt{T}\} = \max\{T^{-\kappa_1/2}, \ldots, T^{-\kappa_{d_x}/2}\}/\min\{\sqrt{N}, \sqrt{T}\}$. To put this into perspective, suppose that $\*x_{i,t}$ is stationary, such that $\kappa_1 = \cdots = \kappa_{d_x} = 0$. In this case, $\*D_T = \*I_{d_x}$ and the rate of convergence is given by $1/\min\{\sqrt{N}, \sqrt{T}\}$, which is the slowest of the regular rates in pure time series and cross-section regressions. Still, the rate is fast enough for the estimation of the number of factors. This brings us to Step (ref) of the estimation procedure. In order to be able to show that $\widehat d_1,\ldots,\widehat d_{G+1}$ and $\widehat{\*F}$ are consistent, however, we need more structure.
Assumption (ref) is similar to Assumptions C and D in Bai. Assumption (ref) (b) and (c) ensure that the serial and cross-sectional dependencies of $\varepsilon_{i,t}$ are at most weak. Assumption (ref) (d) requires that the regressors are exogenous, which rules out the presence of lagged dependent variables in $\*x_{i,t}$. Assumption (ref) (d) can be relaxed by instead requiring that $\varepsilon_{i,t}$ satisfies a martingale type condition similar to Assumption 3 in DGP2020. However, since this will add to the already lengthy derivations and heavy notation, in the present paper we maintain Assumption (ref).
If $\nu_G \geq 1$, Assumptions (ref)--(ref) are enough to ensure that $\widehat d_1$ is consistent. If, however, $\nu_G<1$, so that some of the factors are signal-weak, then we also need Assumption (ref).
As already pointed out, $\*F_{1}^0$ dominates all other factors and is thus easiest to estimate. It is therefore not surprising to find that $\widehat{d}_1$ is consistent for $d_1$. Lemma (ref) states this result formally.
Hence, $\widehat{d}_1$ is consistent. In order to ensure that also $\widehat d_2,\ldots,\widehat d_{G+1}$ are consistent, the groups have to be “distinct”. The next assumption formalizes this requirement. The assumption is stated in terms of $\*F_{g}^0$ and $\+\Gamma_{g}^0$, where $\*F_{g}^0$ is as before and $\+\Gamma_{g}^0 = (\+\gamma_{g,1}^{0},\ldots,\+\gamma_{g,N}^{0})^{\prime }$ is the $N\times d_g$ matrix of stacked factor loadings for group $g$.
The condition that $\nu _{G-1}\geq 1$ means that we only allow for one group of weak factors. This can be seen as a form of normalization and is not particularly restrictive. Let us therefore instead consider $\max_{g\neq h}\Vert \*F_{g}^{0\prime }\*F_{h}^{0}\Vert =O_{P}(T^{q})$, which is less restrictive that the exact orthogonality condition typically required in papers on grouped factor structures (see, for example, Ando). As is well known, $\+\gamma _{i}^{0\prime }\*f_{t}^{0}=\+\gamma _{i}^{0\prime }\*W^{-1}\*W\*f_{t}^{0}$ for any positive definite matrix $\*W$. Now set $\*W=(\*F^{0\prime }\*F^{0}\*C_{T}^{2})^{-1/2}$. This implies
which means that the renormalized factors are exactly orthogonal, and hence that Assumption (ref) is satisfied for $\mathbf{F}^{0 }\*W'$ with $q=-\infty $. This is one way in which $\max_{g\neq h}\Vert \*F_{g}^{0\prime }\*F_{h}^{0}\Vert =O_{P}(T^{q})$ can be justified. Another way is through orthogonal basis functions on $[0,1]$. For example, if $\*f_{t}^{0}=(1,t)^{\prime }$ and $\*C_{T}=\mathrm{diag}(T^{-1/2},T^{-3/2})$, then $\sqrt{T}\*C_{T}\*f_{t}^{0}\rightarrow \*f^{0}(w)=(1,w)^{\prime }$ and $\*C_{T}\*F^{0\prime }\*F^{0}\*C_{T}\rightarrow \int_{0}^{1}\*f^{0}(w)\*f^{0}(w)^{\prime }dw$ as $T\rightarrow \infty $ with $w\in \lbrack 0,1]$. Because they are square integrable, the elements of $\*f^{0}(w)$ can be represented in terms of orthogonal basis functions (see, for example, DongLinton). This is true not only for the simple example given here but for most commonly used deterministic trend terms. The loading condition, $\max_{g\neq h}\Vert \+\Gamma _{g}^{0\prime }\+\Gamma _{h}^{0}\Vert =O_{P}(N^{p})$, can be rationalized in the same way as the factor condition. Another possibility is that $\+\Gamma _{1}^{0},\ldots ,\+\Gamma _{G}^{0}$ are independent and at most one of them has non-zero mean. Independence is often assumed and we therefore do not justify it here (see, for example, Chudiketal2011, and Pesaran2006). In order to justify the zero mean assumption, suppose for simplicity that $G=2$, that $d_{1}=\+\gamma _{1,i}^{0}=1$ and that $\+\gamma_{2,i}^{0}=\+\gamma _{2}^{0}+\+\eta _{i}$ with $E(\+\eta _{i})=\mathbf{0}_{d_{2}\times 1}$. Hence, $\+\gamma _{i}^{0}=(\+\gamma _{1,i}^{0}, \+\gamma_{2,i}^{0\prime })^{\prime }=(1,\+\gamma _{2}^{0\prime }+\+\eta _{i}^{\prime})^{\prime }$ and $\*f_{t}^{0}=(\*f_{1,t}^{0},\*f_{2,t}^{0\prime })^{\prime }$, which in turn implies that
The common component with $\+\gamma _{i}^{0}$ as loading and $\*f_{t}^{0}$ as factor can therefore be written equivalently as one in which $(\*f_{1,t}^{0}+\+\gamma _{2}^{0\prime }\*f_{2,t}^{0},\*f_{2,t}^{0\prime})^{\prime }$ is the factor and $(1,\+\eta _{i}^{\prime })^{\prime }$ is the loading.
Together with Lemma (ref), Lemma (ref) (a) implies that $\widehat d_1,\ldots,\widehat d_{G+1}$ are all consistent. The consistency of $\widehat d_1,\ldots,\widehat d_G$ is important for obvious reasons. The consistency of $\widehat d_{G+1}$ ensures that the stopping rule in Step (ref) of the estimation procedure is asymptotically valid, which in turn implies that $\widehat G$ is consistent. Hence, under the conditions of Lemma (ref), we have
as $N,\,T\to \infty$. The consistency of $\widehat G$ further implies that $d_f$ can be consistently estimated using $\widehat d_f = \widehat d_1 +\cdots+ \widehat d_{\widehat G}$.
As we alluded to in the discussion of Assumption (ref) above, $\*F^{0}$ and $\+\gamma _{i}^{0}$ are only identified up to a rotation matrix (see, for example, BaiNg2002). However, we cannot claim that $\widehat{\*F}$ is rotationally consistent for $\*F^{0}$, as the number of rows of both objects is growing with $T$. We therefore have to resort to alternative consistency concepts. This is where Lemma (ref) (b) comes in. It shows that the spaces spanned by $\widehat{\*F}$ and $\*F^{0}$ are asymptotically the same. This establishes that all the Step (ref) estimates are consistent. We therefore move on to investigate the asymptotic properties of the IPC estimator $\widehat{\+\beta }$ of $\+\beta ^{0}$ obtained in Step (ref).
It is important to note that if all the factors in $\*f_t^0$ are stationary such that $G=1$ and $\nu_G = \nu_1 = 1$, then Assumption (ref) (a) and (b) require that $N/T \to \rho_1 = 1/\rho_2 \in (0,\infty)$, which is the condition considered in Bai, and MoonWeidner2015. Assumption (ref) (a) and (b) are therefore more general than the condition considered in these other papers.
Assumption (ref) is a central limit theorem that is analogous to Assumption E of Bai. The reason for requiring that the asymptotic distribution is mixed normal as opposed to normal is that by doing so we can accommodate stochastically integrated factors (see, for example, BaiKaoNg, and HJPS). In the absence of such integrated factors, the mixed normal becomes normal. Either way, Assumption (ref) ensures that standard normal and chi-squared inference based on $\widehat{\+\beta }$ is possible.
According to Theorem (ref), the asymptotic bias is driven by the factors and loadings of group $G$, which is intuitive as the factors of this group are smallest in order of magnitude. They therefore dominate the asymptotic bias. By bounding $\nu_G$ from below Assumption (ref) ensures that $\*B_0^{-1}(\sqrt{\rho_1}\*A_{1} + \sqrt{\rho_2}\*A_{2})$ is not diverging. It is important to note that $\rho_1$ and $\rho_2$ may both be zero, which will be the case if $\nu_G>1$ and $T/N$ converges to a constant. In general, the larger $\nu_{G}$ is, the less restrictive the condition on $T/N$ has to be for $\rho_1$ and $\rho_2$ to be zero. For example, if $\nu _{G}=2$, then we only require $N/T^{2}\to 0$. In this sense, trending factors are a blessing. In order to put this into perspective, suppose again that all the factors are stationary, such that $G=1$, $\nu_G = \nu_1 = 1$ and $\rho_1 = 1/\rho_2 \in (0,\infty)$, and that the regressors are stationary, too, such that $\*D_T = \*I_{d_x}$. In this case, the bias in Theorem (ref) reduces to $\*B_0^{-1}(\sqrt{\rho_1}\*A_{1} + \rho_1^{-1/2}\*A_{2})$, which is identical to the bias reported in Theorem 3 of Bai. In fact, it is not difficult to show that under these conditions, $\sqrt{NT} (\widehat{\+\beta}-\widehat{\+\beta}_0) = o_P(1)$, so that the IPC estimator is asymptotically equivalent to the first step PC estimator. The point about the bias is that while $\sqrt{\rho_1}\*A_{1}$ can be made arbitrarily small (large) by just taking $\rho_1$ to zero (infinity), this will make $\rho_1^{-1/2}\*A_{2}$ divergent (negligible). It follows that unless $\rho_1 \in (0,\infty)$ the bias will diverge, which in turn means that there is no way to make the bias disappear by just manipulating $\rho_1$, which in practical terms means restricting $T/N$. By allowing $\nu_G > 1$, we break the inverse relationship between $\rho_1$ and $\rho_2$, which means that one can be zero without for that matter forcing the other to infinity.
In the appendix, we provide some alternative conditions that ensure $\*A_{1} = \*A_{2} = \*0_{d_x\times 1}$. If $\rho_1$, $\rho_2$, $\*A_{1}$ and $\*A_{2}$ are all different from zero, one possibility is to use bias correction. In the appendix, we explain how the Jackknife approach can be used for this purpose.
Theorem (ref) imposes only minimal conditions on the correlation and heteroskedasticity of $\varepsilon_{i,t}$, and is in this sense very general. Such generality is, however, not possible if we also want to ensure consistent estimation of $\+\Omega$. Let us therefore assume for a moment that $E(\varepsilon_{i,t}\varepsilon_{j,s}) = 0$ for all $(i,t)\ne (j,s)$, so that $\varepsilon_{i,t}$ is serially and cross-sectionally uncorrelated. In this case,
A natural estimator of this matrix is given by
where $\widehat{\sigma}_{\varepsilon,i}^2 = T^{-1}\sum_{t=1}^T (\*y_i -\*X_i\widehat{\+\beta})^{\prime }\*M_{\widehat F}(\*y_i -\*X_i\widehat{\+\beta})$ and $\widehat{\*Z}_i$ is as in the definition of $\widehat{\+\beta}$. It is not difficult to show that under the conditions of Theorem (ref),
Of course, in this paper we do not assume knowledge of the order of the regressors, which in practice means that the appropriate normalization matrix $\*D_T$ to use is unknown. This is not a problem, however, as the usual Wald and $t$-test statistics are self-normalizing. As an illustration, consider testing the null hypothesis of $H_0: \*R \+\beta^0 = \*r$, where $\*R$ is a $r_0 \times d_x$ matrix of rank $r_0 \leq d_x$ and $\*r$ is a $r_0 \times 1$ vector. The Wald test statistic for testing this hypothesis is given by
which has a limiting chi-squared distribution with $r_0$ degrees of freedom under $H_0$, as is clear from
In the next section, we use Monte Carlo simulations as a means to evaluate the accuracy of this last results in small samples.
The above results are for the case when $\varepsilon _{i,t}$ is serially and cross-sectionally uncorrelated. If $\varepsilon _{i,t}$ is serially and/or cross-sectionally correlated, we recommend following Bai, who discusses the issue of consistent covariance matrix estimation at length. The same arguments can be applied without change in current context.
This section reports the results obtained from a small-scale Monte Carlo simulation exercise. One important goal of the simulations is to examine the finite sample performance of $\widehat{\+\beta}_0$, $\widehat{\+\beta}_1$, and $\widehat{\+\beta}$, and justify the fact that $\widehat{\+\beta}$ is superior to $\widehat{\+\beta}_0$ and $\widehat{\+\beta}_1$.
We consider two DGPs, henceforth denoted “DGP (ref)” and “DGP (ref)”, which we now describe.
For each combination of $N$ and $T$, we report the correct selection frequency for $(\widehat d_1,\ldots,\widehat d_{\widehat G})$ when seen as an estimator of $(d_1,\ldots,d_{G})$ and for $\widehat d_g$ individually for each group $g = 1,\ldots,G$. Hence, while the former frequency captures the accuracy of the estimation of both $(d_1,\ldots,d_{G})$ and $G$, the latter frequency only captures the accuracy of the estimation of $d_g$ for each $g = 1,\ldots,G$. We also report the root mean squared error (RMSE) of $\widehat{\+\beta}$ and $\*P_{\widehat{F}}$, as measured by the square root of the average of $\|\widehat{\+\beta}-\+\beta^0\|^2$ and $\|\*P_{\widehat{F}} -\*P_{F^0}\|^2$, respectively, over the replications. In interest of comparison, RMSEs of $\widehat{\+\beta}$ is compared to that of $\widehat{\+\beta}_0$, $\widehat{\+\beta}_1$ and the infeasible OLS estimator of $\+\beta^0$ based on taking $\*F^0$ as known, $\widehat{\+\beta}(\*F^0)$. Some results on the 5% size of the Wald test based on $\*R =\*I_{d_x}$ and $\*r =\+\beta^0$ are also reported. In particular, we report the size of $W_{\widehat{\+\beta}}$, $W_{\widehat{\+\beta}_0}$ and $W_{\widehat{\+\beta}_1}$, which are computed in an obvious fashion by replacing $(\widehat{\+\beta},\widehat{\*F})$ with $(\widehat{\+\beta}_0,\widehat{\*F}_0)$ and $(\widehat{\+\beta}_1,\widehat{\*F})$, respectively, and $W_{\widehat{\+\beta} (\+F^0)}$, which is calculated in the same way as $W_{\widehat{\+\beta}}$ but with $\*M_F\*X_i$ in place of $\widehat{\*Z}_i$. The critical values are taken from $\chi^2(d_x)$. As pointed out in Section (ref), the IPC estimator is invariant with respect to $\delta$. The results reported here are based on $\delta=1$. In time series analysis, rules such as (the integer part of) $8 (T/100)^{1/4}$ are sometimes used for the maximum model dimension to be considered. In our theory, however, $d_{max}$ is fixed and some of the arguments used in the proofs fail if $d_{max}$ is allowed to increase with $N$ and $T$. In this section, we therefore follow the bulk of the previous literature and set $d_{max}$ to a fixed large number (see, for example, AH13, BaiNg2002, and MoonWeidner2015). We choose $d_{max}=10$, which led to the same results as some of the other values we tried. The number of replications is set to 1,000.
We begin by considering the results for DGP (ref), which are reported in Tables (ref) and (ref). In particular, while Table (ref) contains the results for the estimated common component, Table (ref) contains the results for the estimated slopes. We first look at the correct selection frequencies for all and each individual group of factors reported in Table (ref). The results for the individual groups suggest that the accuracy is generally degreasing in $g$, which is partly expected because the signal strength of the factors, as measured by $\nu_g$, decreases with increasing values of $g$. The accuracy of $(\widehat d_1,\ldots,\widehat d_{\widehat G})$ is almost identical to that of $\widehat d_{G}$, suggesting that the accuracy of $\widehat G$ is driven by the accuracy of the group whose factors has the weakest signal. Looking next at the RMSE results reported in Table (ref) for estimating $\+\beta^0$, we see that there is a clear improvement as the sample size increases. The best overall performance is generally obtained when taking $\*F^0$ as known, which is in accordance with our priori expectations. However, the improvement is not very large and it decreases with increases in $N$ and $T$. The reason for this is the accuracy of the estimated factors, which according to the RMSE of $\*P_{\widehat{F}}$ reported Table (ref) is high and increasing in $N$ and $T$. The second best performance is obtained by using $\widehat{\+\beta}$, followed by $\widehat{\+\beta}_1$, and then $\widehat{\+\beta}_0$. If our asymptotic theory is correct, while the size of $W_{\widehat{\+\beta}}$ and $W_{\widehat{\+\beta} (\+F^0)}$ should converge to 5% as the sample size increases, that of $W_{\widehat{\+\beta}_0}$ and $W_{\widehat{\+\beta}_1}$ should not, and this is exactly what we see in Table (ref). Note in particular how the size of $W_{\widehat{\+\beta}_0}$ and $W_{\widehat{\+\beta}_1}$ is not only nonconvergent but that it is in fact increasing in $N$ and $T$.
The RMSEs of $\widehat{\+\beta}_0$, $\widehat{\+\beta}_1$, $\widehat{\+\beta}$ and $\widehat{\+\beta}(\*F^0)$ in DGP (ref) are given by 0.0444, 0.0405, 0.0405 and 0.0391, respectively, and the sizes of the Wald tests associated with these estimators are given by 0.285, 0.103, 0.078 and 0.061, respectively. Hence, just as in DGP (ref), the best performance is obtained by using $\widehat{\+\beta}(\*F^0)$ with $\widehat{\+\beta}$, $\widehat{\+\beta}_1$ and $\widehat{\+\beta}_0$ ending up in second, third and fourth place, respectively.
Economists have become concerned that recently house prices have grown too quickly, and that prices are now too high relative to per capita incomes. If this is correct and there is any truth to the theory on the matter, prices should stagnate or fall until they are better aligned with income, which in statistical terms mean that house prices should be cointegrated with income. The validity of this assumption has important implications for policy, because a failure could be due to a housing bubble.
In this section, we revisit the real house price data set of Hollyetal2010, which comprises data on log real house prices ($p_{i,t}$) and log real per capita income ($w_{i,t}$) for 49 US states across the 1975--2003 period. According to theory, $p_{i,t}$ and $w_{i,t}$ should be cointegrated with cointegrating vector $(1,\,-1)'$. The previous empirical evidence of this prediction has, however, been mixed and far from convincing (see, for example, Gallin2006). Hollyetal2010 argue that this lack of empirical support can be attributed in part to a failure to account for cross-sectional dependence, leading to deceptive conclusions. The authors therefore apply the “CIPS” panel unit root test of Pesaran2007, which allow for cross-section dependence in the form of a common factor. The test is applied both to $p_{i,t}$ and $w_{i,t}$ separately, and to $p_{i,t}-w_{i,t}$. According to the results, while the variables are unit root non-stationary, their difference is not. Hollyetal2010 also report CCE results suggesting that the estimated income elasticity is indeed close to one. They therefore conclude that $p_{i,t}$ and $w_{i,t}$ are cointegrated with cointegrating vector $(1,\,-1)'$, just as predicted by theory.
Our interest in the work of Hollyetal2010 stems from their preference to apply the CIPS test, which tests for a unit root in the defactored data. This means that if $p_{i,t}$ and $w_{i,t}$ are not cointegrated by themselves, but only when conditioning on unit root common factors, because of the way that the data are defactored prior to the testing, the unit root null hypothesis is likely to be rejected by the CIPS test. That is, the test is likely to lead to the conclusion of cointegration when in fact there is none. In this section, we use IPC as a means to investigate this possibility.
As in the RTS illustration, we begin by plotting the variables. This is done in Figure (ref). As expected, both variables are highly persistent and the ADF test provides no evidence against the unit root null. This corroborates the unit root test results reported by Hollyetal2010. However, we also see that the trending behaviours of $p_{i,t}$ and $w_{i,t}$ are very different, suggesting that their stochastic trends are not the same, which they should be under cointegration. We also see that the trending behaviour is very similar across states, which is suggestive of non-stationary common factors. Of course, the IPC procedure does not require cointegration and it does allow for very general types of factors. We therefore proceed with the estimation of the model. Hence, in this illustration, $y_{i,t}=p_{i,t}$ and $\*x_{i,t} = w_{i,t}$. For comparison purposes, the IPC results are presented together with the results obtained by applying the PC estimator of Bai, as well as the usual OLS estimator with time and state fixed effects. The results reported in Table (ref). The first thing to note is that the estimated slopes vary a lot depending on the estimator used. Interestingly, the point estimates are increasing in the generality of the estimator with fixed effects OLS (IPC) leading to the lowest (highest) estimate. We therefore begin by considering the IPC results. The point estimate of 2.1024 is far from the theoretically predicted value of one, which is also not included in the reported 95% confidence interval. As for the factors, we estimate $\widehat d_1 = \widehat d_2 = \widehat d_3 = 1$, implying that $\widehat d_f =3$. The estimated factors, denoted $\widehat f_{1,t}$, $\widehat f_{2,t}$ and $\widehat f_{3,t}$, are plotted in Figure (ref). As expected given Figure (ref), all three factors are highly persistent, although with clearly distinct trends. We take this as evidence against cointegration between $p_{i,t}$ and $w_{i,t}$, since under cointegration the factors should be stationary.
Because we estimate three time-varying factors, fixed effects OLS is invalid, as it only allows for a common time effect. Moreover, since the factors come from three distinct groups, and are not all stationary, PC is invalid, too. This leaves us with the IPC estimator, which again provides strong evidence against the theoretically predicted one-to-one cointegrated relationship between $p_{i,t}$ and $w_{i,t}$. Hollyetal2010 conclude that “[o]ur results support the hypothesis that real house prices have been rising in line with fundamentals (real incomes), and there seems little evidence of house price bubbles at the national level.” The results reported here reveal a completely different picture with housing prices being long run disconnected with real income.
The PC approach of Bai has attracted considerable interest in recent years, so much so that it has given rise to a separate PC literature. A key assumption in this literature is that both the unknown factors and regressors are stationary, which is rarely the case in practice. In the present paper, we relax this assumption by considering a very general specification in which the factors and regressors are essentially unrestricted. In spite of this generality, the proposed IPC estimator can be applied without any input from the practitioner, except for the maximum number of factors to be considered. The fact that in IPC there is no need to distinguish between deterministic and stochastic factors means that the usual problem in applied work of deciding on which deterministic terms to include in the model does not arise, as these are estimated along with the other factors of the model. There is also no need to pre-test the regressors for unit roots, which is otherwise standard practice when using procedures that do not require data to be stationary. In other words, the proposed IPC is not only very general but also extremely user-friendly. It should therefore be a valuable addition to the already existing menu of techniques for panel regression models with interactive effects.
\setcounter{page}{1}
\setcounter{equation}{0} \setcounter{section}{0} \setcounter{table}{0} \setcounter{figure}{0} \setcounter{assumption}{0}
We begin this appendix by laying out the notation that will be used throughout. This is done in Section (ref), which in Section (ref) is followed by a discussion of some conditions that ensure that the asymptotic distribution of $\widehat{\+\beta}$ is correctly centered at zero. Section (ref) provides an extra empirical study using the US bank data. Section (ref) describes the outline of the proofs. Sections (ref) and (ref) provide some auxiliary lemmas and their proofs, respectively. Proofs of the main results are provided in Section (ref).
The matrices $\+\Sigma_{F^0}$ and $\+\Sigma_{\Gamma^0}$ have been defined in Assumption (ref). In this appendix, we use $\+\Sigma_{F_g^0}$ and $\+\Sigma_{\Gamma_g^0}$ to denote the sub-matrices of $\+\Sigma_{F^0}$ and $\+\Sigma_{\Gamma^0}$ corresponding to $T^{-\nu_g}\*F_{g}^{0\prime}\*F_{g}^0$ and $N^{-1}\+\Gamma_g^{0\prime}\+\Gamma_g^0$, respectively, for $g = 1,\ldots,G$. We also define $\*V_{g} = \mathrm{diag}(\widehat{\lambda}_{g,1},\ldots, \widehat{\lambda}_{g,d_{max}})$, where $\widehat{\lambda}_{g,d}$ has been defined in Step (ref) of the IPC estimation procedure. We partition $\*F^0 = (\*F_{1}^0,\ldots,\*F_g^0,\*F_{+g}^0)$ and $\*C_T = \mathrm{diag} (T^{-\nu_1/2}\*I_{d_1},\ldots, T^{-\nu_g/2}\*I_{d_g}, \*C_{+g,T} )$, where $\*F_{+g}^0 = (\*F_{g+1}^0,\ldots,\*F_G^0)$ is $T\times (d_{g+1} +\cdots+d_G)$, and $\*C_{+g,T} = \mathrm{diag}( T^{ -\nu_{g+1}/2} \*I_{d_{g+1}} ,\ldots, T^{-\nu_G/2 } \*I_{d_G})$ is $(d_{g+1} +\cdots+d_G)\times (d_{g+1} +\cdots+d_G)$. We partition $\widehat{\*F}$, $\+\gamma_{i}^0$ and $\+\Gamma^0$ conformably as $\widehat{\*F} =(\widehat{\*F}_1,\ldots,\widehat{\*F}_g,\widehat{\*F}_{+g})$, $\+\gamma_{i}^0 = (\+\gamma_{1,i}^{0\prime},\ldots,\+\gamma_{g,i}^{0\prime}, \+\gamma_{+g,i}^{0\prime})'$ and $\+\Gamma^0 = (\+\Gamma_{1}^0,\ldots,\+\Gamma_{g}^0, \+\Gamma_{+g}^0)$, respectively.
We introduce $\lambda_{g,d} = T^{-\nu_g} \*h_{g,d}^{0\prime}\*F_{g}^{0\prime}\+\Sigma_{g}^0\*F_{g}^{0}\*h_{g,d}^0$, where $\+\Sigma_{g}^0= N^{-1} \*F_{g}^{0} \+\Gamma_{g}^{0\prime} \+\Gamma_{g}^0\*F_{g}^{0\prime}$ and $\*h_{g,d}^0$ is the $d$-th column of $\*H_{g}^0 = N^{-1}T^{(\nu_g-\delta)/2}\+\Gamma_{g}^{0\prime}\+\Gamma_{g}^0 \*F_{g}^{0\prime}\widehat{\*F}_{g}^0 (\*V_{g}^0)^{-1}$ with $\widehat{\*F}_{g}^0$ being the $T\times d_g$ matrix consisting of the first $d_g$ columns of $\widehat{\*F}_{g}$ and $\*V_{g}^0$ being the leading $d_g \times d_g$ principal submatrix of $\*V_{g}$. In other words, $\widehat{\*F}_{g}^0$ and $\*V_{g}^0$ are $\widehat{\*F}_{g}$ and $\*V_{g}$ based on treating the number of factors for each group $g$, $d_g$, as known. We also define $\*H_{g} = T^{-(\nu_g-\delta)/2}\*H_{g}^0$. In order to appreciate the implication of the difference in normalization with respect to $T$, let us consider $\*H_{g}^0$. By Assumption (ref), $N^{-1} \+\Gamma_{g}^{0\prime}\+\Gamma_{g}^0$ is asymptotically of full rank, and hence $\|N^{-1} \+\Gamma_{g}^{0\prime}\+\Gamma_{g}^0\| = O_P(1)$. Hence, since
and $\|(T^{-\nu_1}\*V_{1}^0)^{-1}\| = O_P(1)$ as explained under (ref), we can show that
which in turn implies
We further use $\widehat{\*F}_{g,d}^0$ to refer to the $d$-th column of $\widehat{\*F}_g^0$. In this notation, $\widehat{\lambda}_{g,d} = T^{-\delta}\widehat{\*F}_{g,d}^{0\prime} \widehat{\+\Sigma}_g \widehat{\*F}_{g,d}^0$ for $d = 1,\ldots,d_g$.
We also partition $\*X_i$ as $\*X_i= (\*X_{1,i},\ldots, \*X_{d_x,i})$ with $\*X_{j,i}$ being the $j$-th column of $\*X_i$. The $j$-th column of $\*X_i\*D_T$ is therefore given by $T^{-\kappa_j/2}\*X_{j,i}$. Moreover, $\mathrm{vec}\,\*A$, $\mathrm{rank}\,\*A$, $\mathrm{span}\,\*A$ and $\lambda(\*A)$ denote the vectorized version, rank, span and eigenvalues of $\*A$, respectively, $a\wedge b = \min\{a,b\}$ and $a\vee b = \max\{a,b\}$.
In this section, we provide a set of assumptions that ensure that the asymptotic distribution of $\sqrt{NT}\*D_T(\widehat{\+\beta} - \+\beta^0)$ given in Theorem (ref) is free of bias without for that matter requiring that $\varepsilon_{i,t}$ is serially and cross-sectionally independent. One way to accomplish this is to assume that $\rho_1 = \rho_2 = 0$, as in Corollary (ref). The assumptions considered here, which are stated in Assumption (ref), can be seen as alternatives to this last condition. In terms of the notation of Theorem (ref), they ensure that $\*A_{1}=\*A_{2} = \*0_{d_x\times 1}$.
Assumption (ref) ensures that the asymptotic distribution of $\sqrt{NT}\*D_T(\widehat{\+\beta} - \+\beta^0)$ is bias-free. It is, however, not necessary, and (a)--(b) should therefore be viewed as examples of conditions under which there is no asymptotic distribution bias. These conditions all have their strengths and weaknesses, and so their suitability will in general depend on the context. Take as an example condition (b), which has the advantage of not requiring any more moment conditions than those that are already in Assumption (ref). It does, however, require that $\nu_G > 1$, which rules out both signal-weak and stationary factors ($\nu_G \leq 1$). Conditions (a) and (c) are more general in this regard, but then at the expense of requiring additional moment conditions. Condition (c) is a functional central limit theorem style moment condition.
The next corollary to Theorem (ref) verifies that the asymptotic distribution of $\sqrt{NT}\*D_T^{-1}(\widehat{\+\beta}-\+\beta^0)$ is indeed bias-free under Assumption (ref).
The asymptotic distribution in Corollary (ref) is the same as the one given in Corollary (ref).
If $\rho_1$, $\rho_2$, $\*A_{1}$ and $\*A_{2}$ are all different from zero, one possibility is to use bias correction. DJ2015 were first to bring attention to the relevance of the half-panel Jackknife approach for bias correction in panel data. These authors focus on the fixed effects case, but Chen2020 have shown that the Jackknife can be used to correct for bias also in models with interactive effects. In our setting, the value of $\nu_G$ is unknown, and therefore the standard Jackknife is not directly applicable. We thus turn to the hybrid half-panel Jackknife correction proposed by FW2018. The bias-corrected version of the IPC estimator is given by
where $\widehat{\+\beta}$ is the standard IPC estimator defined in (ref), $\widehat{\+\beta}_{1,N}$ and $\widehat{\+\beta}_{2,N}$ are defined in the same way but applied to cross-sectional units $\{1,\ldots, \lfloor N/2\rfloor \}$ and $\{\lfloor N/2\rfloor +1,\ldots, N\}$, respectively, and $\widehat{\+\beta}_{1,T}$ and $\widehat{\+\beta}_{2,T}$ are the IPC estimators applied to odd and even numbered time periods, respectively. The splitting in even and odd numbered time periods is needed in order to account for the behaviour of the factors. See FW2018 for a more detailed discussion of the hybrid Jackknife approach.
In this section, we estimate the returns to scale (RTS) of US banks, an area that has attracted considerable attention in the empirical production literature (see, for example, FengSerletis2008). Let us therefore denote by $CT_{i,t}$ the observed total cost of bank $i$ during period $t$. The total cost is a function $C(\cdot)$ of a $J\times 1$ vector of input prices $\*p_{i,t} = (p_{1,i,t},\ldots,p_{J,i,t})'$ and a $L\times 1$ vector of output quantities $\*q_{i,t} = (q_{1,i,t},\ldots,q_{L,i,t})'$. Hence, $CT_{i,t} = C(\*p_{i,t},\*q_{i,t})$, or when expressed in logs, $\ln CT_{i,t} = \ln C(\*p_{i,t},\*q_{i,t})$. The RTS is simply the reciprocal of the sum of the output elasticities, which can be estimated for a given $C(\cdot)$. In order to ensure that $C(\cdot)$ is linearly homogenous in prices, however, it is very common to first normalize costs and prices by one of the prices, $p_{J,i,t}$ say. Hence, $\ln(CT_{i,t}/p_{J,i,t}) = \ln C(\*p_{i,t}/p_{J,i,t},\*q_{i,t})$, which in turn implies that the RTS is given by
In order to estimate this quantity, estimates of the partial derivatives are needed. The standard way to obtain such estimates in the literature is to assume that $\ln C(\*p_{i,t}/p_{J,i,t},\*q_{i,t})$ is linear, to estimate the resulting linear relationship between $\ln (CT_{i,t}/p_{J,i,t})$, $\ln (\*p_{i,t}/p_{J,i,t})$ and $\ln \*q_{i,t}$ by OLS, and to estimate the partial derivatives by simply replacing the parameters by their OLS estimates. A number of modelling issues then arise. First, as pointed out by FengZhang, “[t]here are many ... types of ... unobserved heterogeneity in the US banking industry, which include, but are not limited to, location-related corporate income tax rate, property tax rate, and personal income tax rate.” Hence, there is a need to control for unobserved heterogeneity. Second, the literature typically assumes that the regressors in $\ln(\*p_{i,t}/p_{J,i,t})$ and $\ln\*q_{i,t}$ are stationary, which, as pointed out by DGP2020, is highly unlikely to be the case in practice. There is therefore a need to consider approaches that do not require the regressors to be stationary. Third, to account for the fact that costs are typically trending, many researchers fit their models with linear and quadratic trend terms. Of course, deterministic trends can account for some trending behaviour, but not all. Moreover, the results tend to be highly sensitive to how the trending is modelled, suggesting that this is an important issue (see FengSerletis2008, for a discussion).
The discussion of the last paragraph suggests that there is a need for an approach that is general enough to accommodate not only unobserved heterogeneity but also trending behaviour. The proposed IPC approach fits this bill and we will therefore use it in this empirical illustration to the RTS of US commercial banks. The data that we will use for this purpose, which are the same as in FGPZ2017, are quarterly and cover the period 1986--2005, which means that $T=80$. To avoid the impact of entry and exit, we only include continuously operating large banks with assets of at least one billion USD. This gives us a total of $N=466$ banks. The data set contains observations on three input prices and output quantities. They are the wage rate of labour ($p_{1,i,t}$), the price of deposits and purchased funds ($p_{2,i,t}$), the price of physical capital ($p_{3,i,t}$), consumer loans ($q_{1,i,t}$), non-consumer loans, which is composed of industrial, commercial and real estate loans ($q_{2,i,t}$), and securities, including non-loan financial assets ($q_{3,i,t}$). All outputs are deflated by the GDP deflator. Following the previous literature (see, for example, Stiroh), total cost ($CT_{i,t}$) is computed as the sum of total salaries and benefits divided by the number of full-time employees, the price of deposits and purchased funds equals total interest expense divided by total deposits and purchased funds, and the price of capital equals expenses on premises and equipment divided by premises and fixed assets. Hence, after normalization by $p_{3,i,t}$, there are six variables included in the model; $\ln (CT_{i,t}/p_{3,i,t})$, $\ln (p_{1,i,t}/p_{3,i,t})$, $\ln (p_{2,i,t}/p_{3,i,t})$, $\ln q_{1,i,t}$, $\ln q_{2,i,t}$ and $\ln q_{3,i,t}$.
In order to get a feeling for the trending behaviour of the data, in Figure (ref) we plot all $N$ series for each variable. A few observations are noteworthy. First, there is strong co-movement among the series. This is true for all six variables, including the dependent variable, $\ln (CT_{i,t}/p_{3,i,t})$, which we take as evidence in support of our interactive effects specification. Second, the variables are highly persistent and most are trending over time, suggesting that they are non-stationary. This last observation is supported by a formal augmented Dickey--Fuller (ADF) test (with intercept and linear trend) applied to each series. The highest rejection frequency across all $N$ series is obtained for $\ln q_{2,i,t}$ and is given by 6.7%, which in turn suggests that the evidence against unit root null hypothesis is weak. In fact, since the ADF test does not account for the multiplicity of the testing problem, the true proportion of (trend-)stationary series is likely to be even lower. Third, it is unclear if the trend in $\ln (CT_{i,t}/p_{3,i,t})$ is the same as those in the regressors, which means that it is important to allow the factors to be trending, as otherwise any unaccounted for trend in $\ln (CT_{i,t}/p_{3,i,t})$ will be pushed into the regression errors.
We now go on to discuss the actual estimation results, which are reported in Table (ref). The estimated model is given by (ref) with $y_{i,t}=\ln (CT_{i,t}/p_{3,i,t})$, $\+\beta^0 = (\beta_1^0,\ldots,\beta_5^0)'$ and $\*x_{i,t} = (\ln(p_{1,i,t}/p_{3,i,t}), \ln( p_{2,i,t}/p_{3,i,t}), \ln q_{1,i,t } , \ln q_{2,i,t}, \ln q_{3,i,t})'$. Hence, in this notation,
is the object of interest, which we estimate using
where $\widehat\beta_3$, $\widehat\beta_4$ and $\widehat\beta_5$ are from $\widehat{\+\beta} = (\widehat\beta_1,\ldots, \widehat\beta_5)'$. The values of $\delta$ and $d_{max}$ are set as in Section (ref). The value of $\delta$ is again arbitrary and does not require justification. As for $d_{max}$, 10 factors should be large enough to capture the true number, yet not “excessively” large (see AH13, for a discussion). In order to enable inference not only for $\+\beta^0$ but also for the RTS, we follow FGPZ2017 and report 95% bootstrap confidence intervals. The estimated factor groups are given by $\widehat{d}_1=\widehat{d}_2=1$, implying that the estimated number of factors equals $\widehat d_f = \widehat{d}_1+\widehat{d}_2=2$. In Figure (ref), we plot the estimated factors, denoted $\widehat f_{1,t}$ and $\widehat f_{2,t}$. We see that while both factor estimates are trending, $\widehat f_{2,t}$ is growing much faster. As a measure of this difference in the trends, we look at $\sum_{i=1}^N \widehat{\gamma}_{1,i}^2 /\sum_{i=1}^N \widehat{\gamma}_{2,i}^2$. By using the results provided in the appendix, we can show that this ratio should be $O_P(T^{\nu_1-\nu_2})$, implying $\ln(\sum_{i=1}^N \widehat{\gamma}_{1,i}^2 /\sum_{i=1}^N \widehat{\gamma}_{2,i}^2) /\ln T = \nu_1 - \nu_2 + o_P(1)$. By plugging in the known value of $T$ and $\sum_{i=1}^N \widehat{\gamma}_{1,i}^2 /\sum_{i=1}^N \widehat{\gamma}_{2,i}^2 = 123.4$, we get $\ln(\sum_{i=1}^N \widehat{\gamma}_{1,i}^2 /\sum_{i=1}^N \widehat{\gamma}_{2,i}^2) /\ln T = 1.1$, which is thus an estimate of $\nu_1 - \nu_2$. Hence, there is indeed a substantial difference in the degree of trending of the two factors. But it is not as large as it would have been if $\widehat f_{1,t}$ and $\widehat f_{2,t}$ were made up of a quadratic and a linear trend, say, in which case $\nu_1 = 5$ and $\nu_2 = 3$, such that $\nu_1 - \nu_2 = 2$. Hence, just as expected given Figure (ref), while trending, the behaviour of the factors is not just a simple deterministic function of time, which of course casts doubt on much of the existing research.
While our main interest is in the RTS, for completeness in Table (ref), we also report the estimated elasticities. The first thing to note is that the RTS is estimated to be significantly larger than one, suggesting that the large US banks considered here have had increasing returns to scale during the period of investigation. This is consistent with the results reported by, for example, FGPZ2017 and WheelockWilson2012. One explanation for this finding is that the banks have been adopting productivity-enhancing internet technologies, leading to increased RTS. The confidence intervals are very narrow not only for the RTS, but also for the elasticities, which are all significant.
In this section, we describe the outline of the proofs. We assume throughout that $G\ge 2$. The proofs for the cases when $G\in \{0,1\}$ are much simpler, and can be obtained by manipulating the proofs for $G\ge 2$.
Lemma (ref) is a widely used result for studying the eigenvalues of large dimensional matrices (LYB2011), and is presented here for convenience.
Our own theoretical development is carried out stepwise, starting with Step (ref) of the IPC estimation procedure. Lemma (ref) is very important in this regard. It presents two order results that are used repeatedly throughout the proofs. Given Lemma (ref), we are able to establish Lemma (ref) of the main text, which provides a lower bound on the rate of convergence of the initial Step (ref) estimator $\widehat{\+\beta}_0$.
As a first step towards establishing the consistency of the Step (ref) estimators of $(d_2, \ldots, d_G)$, in Lemma (ref) we study limiting behaviour of the eigenvalues associated with $\widehat{\+\Sigma}_1$ in (ref). The consistency of $\widehat d_1$ reported in Lemma (ref) of the main text is a direct consequence of Lemma (ref). Lemmas (ref) and (ref) enable consistent estimation of $(d_2, \ldots, d_G)$, and are thus key in proving Lemma (ref) of the main text. For simplicity, in the proofs of Lemmas (ref) and (ref) we treat $d_1$ as known, which according to Lemma (ref) is justified in large samples.
The final step in our theoretical analysis is to prove the asymptotic mixed normality of the Step (ref) IPC estimator $\widehat{\+\beta}$, as stated in Theorem (ref). In so doing, we treat $(d_1,\ldots, d_G)$ as known, which is again justified in large samples. We begin by proving Lemma (ref), which is needed in the proof of Theorem (ref). The proofs of Corollaries (ref) and (ref) are immediate consequences of that of Theorem (ref). We also derive the rate of convergence of $\widehat{\+\beta}_0$ under the conditions of Theorem (ref). This rate is provided as a part of Lemma (ref).
Proof of Lemma (ref).
This is Lemma 3 of LYB2011. The proof is therefore omitted.{$\blacksquare$}
Proof of Lemma (ref).
Consider (a). We have
where the first inequality follows from the fact that $|\text{tr}\,\*A| \leq \text{rank}\,\*A \, \| \*A\|_{2}$, the second inequality follows from the fact that $\|\*{P}_{{F}}\| _{2}=1$, and the second equality hold by Assumption (ref) (b).
The result in (b) is due to
where, with a slight abuse of notation and in this proof only, $\*X_j= (\*X_{j,1},\ldots, \*X_{j,N})'$, the second inequality follows from $|\text{tr}\,\*A| \leq \text{rank}\,\*A \,\| \*A\|_{2}$, while the first equality is due to Assumption (ref).
For (c), we use
where the third equality follows from using the mixing condition on $\varepsilon_{i,t}\varepsilon_{j,t}$ across $t$. The above result implies that $(NT)^{-1}\| \+\varepsilon'\+\varepsilon\| =O_P(N^{-1/2})+O_P(T^{-1/2})$ as required for the first result in (c). The second follows from
where the last step follows by the same arguments used to establish the first result of (c).
It remains to prove (d), which is a direct consequence of Assumptions (ref) and (ref), as seen from
This establishes (d) and hence the proof of the lemma is complete.{$\blacksquare$}
Proof of Lemma (ref).
Consider (a). As in Appendix (ref), decompose $\*F^0 = (\*F_{1}^0,\*F_{+1}^0)$ and $\*C_T = \mathrm{diag} (T^{-\nu_1/2}\*I_{d_1}, \*C_{+1,T} )$, where $\*F_{+1}^0 = (\*F_2^0,\ldots,\*F_G^0)$ is $T\times (d_f -d_1)$ and $\*C_{+1,T} = \mathrm{diag}( T^{ -\nu_2/2} \*I_{d_2} ,\ldots, T^{-\nu_G/2 } \*I_{d_G})$ is $(d_f -d_1)\times (d_f -d_1)$. We partition $\widehat{\*F}$, $\+\gamma_{i}^0$ and $\+\Gamma^0$ conformably as $\widehat{\*F} =(\widehat{\*F}_1,\widehat{\*F}_{+1})$, $\+\gamma_{i}^0 = (\+\gamma_{1,i}^{0\prime}, \+\gamma_{+1,i}^{0\prime})'$ and $\+\Gamma^0 = (\+\Gamma_{1}^0, \+\Gamma_{+1}^0)$, respectively.
By the definition of the eigenvectors and eigenvalues, $\widehat{\+\Sigma}_1 \widehat{\*F}_1 = \widehat{\*F}_1\*V_{1}$. By using this and the definition of $\widehat{\+\Sigma}_1$,
with implicit definitions of $\*J_{1},\ldots,\*J_{9}$. Note that
Hence, moving this term over to the left-hand side, the above expression for $T^{-(\nu_1+\delta)/2}\widehat{\*F}_1 \*V_{1}$ becomes
We now evaluate each of the terms on the right-hand side.
Because $T^{-\delta}\|\widehat{\*F}_1\|^2 = d_{max}$ and $(NT)^{-1}\sum_{i=1}^N \| \*X_i \*D_T\|^2 = O_p(1)$ by Assumption (ref), the order of $\*J_{1}$ is given by
Moreover, since
by Assumption (ref), we can show that
and by exactly the same arguments,
For $\*J_{4}$, we use
where the last equality makes use of the fact that $\|\*C_{+1,T}^{-1}\|= O(T^{\nu_2/2})$, as $\nu_2 > \cdots > \nu_G$ by Assumption (ref). We can further show that
where the last equality holds, because by Assumption (ref) and $(\mathrm{tr}\,\*A'\*B)^2 \le (\mathrm{tr}\,\*A'\*A )( \mathrm{tr}\,\*B'\*B)$, we have
Hence, by adding the results,
where we have used Assumptions (ref) and (ref) ($T/N^2=O(1)$ under $\nu_G < 1$) to show that $T^{(1-\delta - \nu_1 +\nu_2)/2}$ and $N^{-1/2}T^{(2-\nu_1-\delta)/2}$ are $o(1)$. The same steps can be used to show that
For $\*J_{6}$, we use
By Assumption (ref),
Another application of Assumption (ref) and Lemma (ref) gives
These results can be inserted into the expression for $T^{-\delta/2}\| \*J_{6}\|$, giving
Next up is $\*J_{7}$. By using Assumptions (ref) and (ref), and Lemma (ref), and the arguments use in evaluating $\*J_6$,
and we can similarly show that
By putting everything together, (ref) becomes
We now left multiply (ref) by $T^{-(\nu_1+\delta)/2}\widehat{\*F}_{1}^{\prime}$ to obtain that
where the third equality follows from Assumption (ref) and Lemma (ref). This implies that $\*V_1$ is at most of rank $d_1$. Similarly, we can left-multiply (ref) by $T^{-\nu_1}\*F_{1}^{0\prime}$ to obtain
which in turn implies that
Note that $T^{-(\nu_1+\delta)/2}\*F_{1}^{0\prime}\widehat{\*F}$ is of rank $d_1$, which then further indicates that $\*V_1$ has at least $d_1$ non-zero elements on the main diagonal which converge to the eigenvalues of $\+\Sigma_{F_{1}^0}\+\Sigma_{\Gamma_{1}^0}$. We now can conclude that $\*V_1$ is of rank $d_1$ in limit.
We are now ready to investigate $\widehat{\lambda}_{1,d}$ for $d\le d_1$. Because here $d\le d_1$, we then focus on $\widehat{\*F}_1^0$ and $\*V_{1}^0$, which are defined in Appendix (ref). Let us write (ref) as follows:
where $\*H_1^0$ is defined in Appendix (ref). The above expression implies that (ref) can be written as
where $\*J_{1}^0,\ldots,\*J_{9}^0$ are $\*J_{1},\ldots,\*J_{9}$ as defined in (ref), except that now $d_1$ is taken as known. Note that by the above development
Moreover, since $T^{-\nu_1}\*V_{1}^0$ converges to a full rank matrix by the argument under (ref), we have $\|(T^{-\nu_1}\*V_{1}^0)^{-1}\| = O_P(1)$, which in turn implies
This is an important result and in what follows we will use it frequently.
Let us now consider $T^{-\nu_1}(\widehat{\lambda}_{1,d} - \lambda_{1,d})$. By the definitions of $\lambda_{1,d}$ and $\widehat{\lambda}_{1,d}$ given in Section (ref),
with obvious definitions of $J_1,\ldots,J_5$. From (ref),
which is $o_P(1)$ under Assumption (ref). This implies $|J_1| = o_P(|J_5|)$, $|J_2| = o_P(|J_5|)$ and $|J_3| = o_P(|J_4|)$. It remains to consider $J_4$ and $J_5$. The order of the first of these terms is given by
where we have made use of the fact that $\|\*H_{g}^0\| = O_P(1)$ (see Appendix (ref)), which implies that $\|\*h_{1,d}^0\|$ is of the same order.
The order of $J_5$ is the same as that of $J_4$. In order to appreciate this, we begin by noting
By the proof for each term of (ref), it is easy to know that
and so
Hence, by putting everything together,
which establishes (a).
Consider (b). This proof is based on Lemma (ref). We therefore start by introducing some notation in order to make the problem here fit the one in Lemma (ref). Let us therefore denote by $\*F_1^{\perp}$ a $T\times (d_{max}- d_1)$ matrix such that $T^{-\nu_1} (\*F_1^{\perp}, \*F_1^0 \*R)'(\*F_1^{\perp}, \*F_1^0\*R) = \mathrm{diag}( \*I_{d_{max}-d_1} , \*I_{d_1} )$, where $\*R$ is a $d_1\times d_1$ rotation matrix. The matrices $T^{-\nu_1/2}\*F_1^{\perp}$, $T^{-\nu_1/2}\*F_1^0 \*R$, $\+\Sigma_{1}^0$ and $\widehat{\+\Sigma}_1 - \+\Sigma_{1}^0$ correspond to $\*Q_1$, $\*Q_2$, $\*A$ and $\*E$ of Lemma (ref). Our counterpart of the matrix $\*Q^0_1$ appearing in this other lemma is thus given by
where
Since $\widehat{\*F}^\perp$ is an orthonormal basis for a subspace that is invariant for $\widehat{\+\Sigma}_1$, we have $\widehat{\lambda}_{1,d_1+d} = \widehat{\*F}_d^{\perp\prime} \widehat{\+\Sigma}_1 \widehat{\*F}_{d}^\perp$, where $d = 1,\ldots,d_{max}-d_1$ and $\widehat{\*F}_{d}^\perp$ is the $d$-th column of $\widehat{\*F}^\perp$. Consider $\|\widehat{\*F}^\perp - T^{-\nu_1/2}\*F_1^{\perp} \|_{2}$. By the definition of $\widehat{\*F}^\perp$,
where the second and third inequalities follow from Magnus. This last result can be used to show that
where $\*F_{1,d}^{\perp}$ is the $d$-th column of $\*F_1^{\perp}$, and the last equality follows from (ref) and the proof of part (a). This completes the proof of the lemma.{$\blacksquare$}
Proof of Lemma (ref).
For (a), we take the same starting point as in the proof of part (a) in Lemma (ref), which is (ref) with $\widehat{\*F}_1$ and $\*V_{1}$ based on treating $d_1$ as known. The rationale for doing so is, as already explained in Appendix (ref), that $\widehat{d}_1$ is consistent. Pre-multiplying this equation through by $T^{-(\nu_1+\delta)/2}\*F_2^{0\prime}$ gives
or
Under Assumption (ref), the orders of $\*J_{1},\ldots,\*J_{6}$ are the same as in those in the proof of the first result of Lemma (ref). The stated orders of $\*J_{7}$ and $\*J_{8}$ are, however, not sharp and can be improved upon. The order of $T^{-(\delta+\nu_2)/2}\|\*F_2^{0\prime}\*J_{7}\|$ is given by
where the last equality follows from Assumption (ref) and Lemma (ref). We can similarly show that
where $\*F_{+2}^{0}$ and $\+\Gamma_{+2}^{0}$ are defined analogously to $\*F_{+1}^{0}$ and $\+\Gamma_{+1}^{0}$ in the proof of Lemma (ref). By using these last two results together with the orders of $\*J_{1},\ldots,\*J_{6}$ given in the proof of the first result of Lemma (ref),
where we have used $\|\*D_T^{-1}(\+\beta^0 - \widehat{\+\beta}_0 )\| =O_P(N^{-1/2} \vee T^{-1/2})$ of Lemma (ref), and Assumptions (ref) and (ref). Hence,
which together with Assumption (ref) yields
as was to be shown for (a).
Let us now consider (b). Analogously to the proof of (a), by invoking Assumption (ref) we can improve the orders of $\*J_{7}$ and $\*J_{8}$. For $\*J_7$,
where the equality follows from Assumption (ref) and Lemma (ref). For $\*J_8$,
This implies that the result in (ref) changes to (after replacing $T^{-\nu_1/2}$ by $T^{-\delta/2}$)
Hence, since $\*H_{1}^0$ is invertible with $\|\*H_{1}^0\| = O_P(1)$ and $\*H_1 = T^{-(\nu_1-\delta)/2}\*H_{1}^0$,
We are now ready to consider $\sum_{i=1}^N \| \*F_{1}^{0\prime} \+\gamma_{1,i}^0 - \widehat{\*F}_1 \widehat{\+\gamma}_{1,i}\|^2$.
We now evaluate each of the terms on the right-hand side one by one. Making use of Assumption (ref) and Lemma (ref), we get
and by another application of Lemma (ref),
For $\sum_{i=1}^N\| \*M_{\widehat{F}_1} \*F_1^0\+\gamma_{1,i}^0 \|^2$, we use (ref) from which it follows that
which in turn implies
For $\sum_{i=1}^N\|\*P_{\widehat{F}_1}\*F_{+1}^{0} \+\gamma_{+1,i}^0\|^2$, we use the result given in part (a), giving
The order of $\sum_{i=1}^N\|\*P_{\widehat{F}_1}\*F_{+2}^{0} \+\gamma_{+2,i}^0\|^2$ is the same. Hence, by adding the above results, (b) follows after simple algebra. The proof is now complete.{$\blacksquare$}
Proof of Lemma (ref).
Let $\*U_i = \*F_{+2}^0 \+\gamma_{+2,i}^0 +\*F_{1}^{0\prime} \+\gamma_{1,i}^0 - \widehat{\*F}_1 \widehat{\+\gamma}_{1,i}$. In this notation,
where $\*K_{1},\ldots,\*K_{16}$ are implicitly defined. Analogously to the proof of the first result of Lemma (ref) we move $\*K_{16} = \*F_2^0 (N^{-1}\+\Gamma_{2}^{0\prime} \+\Gamma_{2}^0)(T^{-(\nu_2+\delta)/2} \*F_2^{0\prime} \widehat{\*F}_2)$ over to the left, giving
By using the same steps employed in the proof of the first result of Lemma (ref), we can show that
For $\*K_{4}$,
where the second equality follows Lemma (ref), and the third follows from Assumptions (ref), (ref), and (ref). The same arguments can be used to show that
For $\*K_{6}$,
where the development is similar to (ref). The order of $T^{-\delta/2}\|\*K_{7}\|$ is the same.
For $\*K_{8}$,
where the last equality holds by Lemma (ref).
Further use of Lemma (ref) gives
and we can show that $T^{-\delta/2}\| \*K_{10} \|$ is of the same order.
$\*K_{11}$ requires more work. We begin by expanding it in the following way:
where, similarly to the analysis of $\*K_{6}$ and using $\|\*P_{\widehat F_1}\|_2 = 1$,
Also, making use of (ref), we can show that
Similarly, for $K_{113}$, we can show that
where we have used Lemmas (ref) and (ref), and Assumption (ref). Further use of Lemma (ref) shows that $K_{114}$ is of the following order:
while the order of $K_{115}$ is
By inserting the above results into (ref), we obtain
The order of $T^{-\delta/2}\| \*K_{12} \|$ is the same, which can be shown using the above steps.
We move on to $\*K_{13}$, whose order is given by
where the second equality follows from Lemma (ref) and Assumption (ref). The order of $T^{-\delta/2}\| \*K_{14} \|$ is the same.
For $\*K_{15}$,
where the second equality follows from Lemma (ref).
We now insert the above results for $\*K_1,\ldots,\*K_{15}$ into (ref). But first we left multiply by $T^{-\nu_2}\*F_2^{0\prime}$ and $T^{-(\nu_2+\delta)/2}\widehat{\*F}_2$ respectively. It then gives
and
The rest of the proofs of (a) and (b) follows from the same arguments used in Lemma (ref). It is therefore omitted.{$\blacksquare$}
Before continuing onto the proof of Lemma (ref), we note that the rates given in Lemma (ref) can actually be improved upon. Suppose that $d_2$ in known. Consider the next term which is a part of $\*K_{15}$:
where $\*V_2^0$ and $\*H_2$ are defined in Appendix (ref). Simple algebra shows that the term $T^{-(\nu_2-\nu_3)}$ in $\tau_{NT}$ of Lemma (ref) can be dropped.
Proof of Lemma (ref).
Consider (a). Let us assume without loss of generality that $\+\beta^0 = \*0_{d_x\times 1}$, as in Bai. It then follows that
Note that by expanding the term $\*M_{\widehat{F}}\*F^0$ as in the proof for the second result of this lemma, we can obtain that
It follows that
where the first term on the right is quadratic in $\|\*D_T^{-1}( \widehat{\+\beta}_1-\widehat{\+\beta}_0 )\|$. Hence, to ensure the right-hand side is non-positive, $\|\*D_T^{-1}( \widehat{\+\beta}_1-\widehat{\+\beta}_0 )\|$ cannot converge to zero at a rate faster than $(NT)^{-1/2} \vee \|\*D_T^{-1} ( \widehat{\+\beta}_0-\+\beta^0)\|$. It follows that
as was to be shown.
We now turn to (b). Note that
with obvious definitions of $\*L$ and $\*B$.
Consider $\*L$. Let $\*Q_{g} = (N^{-1}\+\Gamma_g^{0\prime}\+\Gamma_g^0)(T^{-(\nu_g+\delta)/2} \*F_{g}^{0\prime}\widehat{\*F}_g )$, such that $\*H_{g}^{-1} = T^{-(\nu_g+\delta)/2} \*V_{g}^0\*Q_{g}^{-1}$. We also introduce $\*e_{g,i}$, which is defined to be zero for $g=1$ and $\*e_{g,i}= \sum_{j=1}^{g-1} (\*F_{j}^0\gamma_{j,i}^0 - \widehat{\*F}_j \widehat{\+\gamma}_{j,i})$ for $g=2,\ldots,G$. In this notation,
where
We now evaluate each of these terms. From the analysis of $\*L_{2}$ below, it is easy to show that $\|\*L_{1} \| = o_P(\|\*D_T^{-1}( \widehat{\+\beta}_0 - \+\beta^0)\|)$. We therefore start from $\*L_{2}$, which we write as
We will use this expression for $\*L_{2}$ later.
Let us now move on to $\*L_{3}$.
where, by using arguments that are similar to those used in the proofs of Lemmas (ref) and (ref),
implying
The same steps can be used to show that $\*L_{4}$ and $\*L_{5}$ are of the same order.
For $\*L_{6}$, write
where the first term on the right is bounded by
The second term is of the same order. Thus, $\|\*L_{6} \| =o_P ( \| \*D_T^{-1}(\+\beta^0-\widehat{\+\beta}_0) \| )$. The same is true for $\|\*L_{7}\|$.
We now examine $\*L_{8}$. Let us define $\+\Sigma_{\varepsilon} = N^{-1}\sum_{i=1}^N \+\Sigma_{\varepsilon,i}$, in which $\+\Sigma_{\varepsilon,i} $ has been defined in Assumption (ref). $\*L_{8}$ can then be written as
For the first term on the right,
where the last equality holds by $NT^{-\nu_G}<\infty $. Let
Note how $\*A_{1} = \operatorname*{p\!\lim}_{N,T\to\infty}E(\overline{\*A}_1 | \mathcal{C})$. In this notation,
where the second equality follows from arguments similar to those used in (ref). Note also that $\lim \sqrt{N}T^{-\nu_G/2} <\infty$. Further use of the same steps used by JYGH establishes that the second and third terms of $\*L_{8}$ are $o_P(\|\*D_T^{-1}(\widehat{\+\beta}_0 -\+\beta^0) \|)+ o_P((NT)^{-1/2})$. Hence,
$\*L_{9}$ can be written as
where the first term on the right-hand side is bounded by
where the first inequality follows from $\|T^{-(\nu_g-1)/2} \+\Gamma_g^{0\prime}\+\varepsilon \*F_{g}^0 \|=O_P(\sqrt{NT})$. The second term on the right-hand side of $\*L_{9}$ is also $o_P( (NT)^{-1/2} )$. We therefore conclude that $\|\*L_{9}\| = o_P( (NT)^{-1/2} )$.
$\*L_{10}$ can be written more compactly as
which we will again make use of later.
Let us move on to $\*L_{11}$. We begin by rewriting $\*e_{g,j}$ as
where $\*F_{+d}^0 = (\*F_{d+1}^0,\ldots, \*F_{G}^0)$ and $\+\gamma_{+d,j}^0 = (\+\gamma_{d+1,j}^{0\prime},\ldots, \+\gamma_{G,j}^{0\prime})'$, as in Appendix (ref). By inserting this into $\*L_{11}$, we get
Hence, $\*L_{11}$ can be written as a sum of five terms. There is no need to study the forth term, the one due to $\*P_{\widehat{F}_d } \*e_{d,j}$, as we can keep expanding $\*e_{d,j}$ until we cannot. The first and fifth terms are $o_P(\|\*D_T^{-1}(\widehat{\+\beta}_0 -\+\beta^0) \|)$ and $o_P((NT)^{-1/2})$, respectively, by the same arguments used for evaluating $\*L_{7}$ and $\*L_{8}$. Moreover, the steps used for evaluating $\*L_{9}$ can be used to show that
The second term in $\*L_{11}$ is therefore negligible. It remains to consider the third term, which is
where we have used Assumption (ref). Therefore, $\|\*L_{11}\| = o_P( (NT)^{-1/2} )$.
For $\*L_{12}$,
where the second equality follows from the construction of $\*e_{g,j}$. Hence, since $\*L_{9}$ is negligible, $\*L_{12}$ is also negligible.
Next up is $\*L_{13}$, whose order can be worked out in the following way:
where the first inequality follows from the construction of $\*e_{g,j}$, and the last equality follows by going through a development similar to the first result of Lemma (ref) and further using a development similar to (ref). The same arguments show that $\|\*L_{14}\|$ and $\|\*L_{15}\|$ are negligible, too.
We now put everything together. This yields
which in turn implies
where $\*Z_i(\*F) $ is defined in Assumption (ref), and
This expression for $\*D_T^{-1}(\widehat{\+\beta}_1-\+\beta^0)$ can be inserted into $\*D_T^{-1}(\widehat{\+\beta}_0 -\+\beta^0)$, giving
which can be solved for $\*D_T^{-1}(\widehat{\+\beta}_0 -\+\beta^0)$
Also, making use of Assumption (ref), it is not difficult to show that $(NT)^{-1}\sum_{i=1}^N\*D_T \*Z_i(\widehat{\*F})'\+\varepsilon_i=O_P((NT)^{-1/2})$. Hence,
By using the fact that
the left-hand side of this last equation can be written as
It follows that
Consider the $o_P( \sqrt{NT}\|\*D_T^{-1} ( \widehat{\+\beta}_1-\widehat{\+\beta}_0)\| )$ reminder term. By the first result of this lemma, $\|\*D_T^{-1}( \widehat{\+\beta}_1-\widehat{\+\beta}_0 )\| = O_P((NT)^{-1/2}\vee \|\*D_T^{-1} ( \widehat{\+\beta}_0-\+\beta^0)\| )$. This can be inserted into (ref), giving
which in turn implies
The second result then follows.{$\blacksquare$}
Proof of Lemma (ref).
Without loss of generality, we assume that $\+\beta^0 = \*0_{d_x\times 1}$. This implies
where the second equality follows from Lemma (ref), and $\+\theta = \+\eta + \*B^{-1}\*C\+\beta$ and $\*B = N^{-1}\+\Gamma^{0\prime}\+\Gamma^0 \otimes \*I_T$ are defined as on page 1265 of Bai. Note that for $\+\beta\in \mathbb{R}^{d_x}$, we may have
In our proof of consistency, we consider two cases; (i) $\| \*D_T^{-1}\+\beta\| \le C$ and (ii) $\| \*D_T^{-1}\+\beta\| > C$, where $C$ is a large positive constant. Under (i),
by Lemma (ref). The expression given in (ref) for $(NT)^{-1}[\mathrm{SSR}( \+\beta, F) - \mathrm{SSR}(\+\beta^0, \*F^0)]$ therefore reduces to
where $\+\beta'\*D_T^{-1}\*B(\*F)\*D_T^{-1}\+\beta$ does not involve $\*F^0$ and $\sum_{i=1}^N\+\gamma_{i}^{0\prime} \*F^{0\prime} \*M_{F} \+\varepsilon_i$ is independent of $\+\beta$. Hence, provided $d_{max} \ge d_f$, the consistency of $\*D_T^{-1}\widehat{\+\beta}_0$ in case (i) follows from the same arguments as in Bai.
Under (ii), (ref) can be written as
where $c_0$ is defined in Assumption (ref), and the second inequality follows from the fact that the quadratic term dominates the linear one for large values of $C$. Hence, $(NT)^{-1}[\mathrm{SSR}( \+\beta, \*F) - \mathrm{SSR}(\+\beta^0, \*F^0)] > 0$, but from the definition of $\widehat{\+\beta}_0$ we also know that $\mathrm{SSR}( \widehat{\+\beta}_0, \widehat{\*F}) - \mathrm{SSR}(\+\beta^0, \*F^0) \le 0$, which means that $\*D_T^{-1}\widehat{\+\beta}_0$ cannot belong to (ii).
Note that since $\|(NT)^{-1}\sum_{i=1}^N \*D_T\*X_i'\*M_{F} \+\varepsilon_i \|=O_P( N^{-1/2}\vee T^{-1/2} )$ by Lemma (ref), all we need is $C = C^0 (N^{-1/2}\vee T^{-1/2})$ for some large constant $C^0$ in order to ensure that the last inequality of (ref) holds. This implies $\*D_T^{-1}(\widehat{\+\beta}_0-\+\beta^0) = O_P( N^{-1/2}\vee T^{-1/2} )$, so the proof is complete. {$\blacksquare$}
Proof of Lemma (ref).
Suppose first that $d_1=0$, such that $d_f=0$. In this case,
This implies $\tau_N \asymp 1/\ln (T\vee N)$, which is much larger than $\widehat{\lambda}_{1,1}/\widehat{\lambda}_{1,0}$. The result then follows immediately.
Suppose now instead that $d_1>0$. Straightforward algebra reveals that
which together with Assumption (ref) implies $\tau_N \asymp 1/\ln (T\vee N)$. Note that for $d=1,\ldots, d_1$,
where the fourth equality follows from (ref) and the last step is dues to (ref) and Assumption (ref). Using Lemma (ref), (ref) and (ref), we obtain that
for $d=1,\ldots, d_1-1$, and $\frac{\widehat{\lambda}_{1,d_1+1} }{\widehat{\lambda}_{1,d_1} } = O_P( T^{-(\nu_1-\nu_2)})$. Moreover, for $d=d_1+1,\ldots, d_{max}$, $\frac{T^{-\nu_1}\widehat{\lambda}_{1,d}}{T^{-\nu_1}\widehat{\lambda}_{1,0}} = O_P(T^{-(\nu_1-\nu_2)})$, which is less than $\tau_N $ by (ref) and Lemma (ref). The required result follows from this and the definition of $\widehat{d}_1$. {$\blacksquare$}
Proof of Lemma (ref).
Each step in the sequential procedure of Step (ref) introduces additional remainder terms that all converge to zero under Assumptions (ref) and (ref) by Lemmas (ref) and (ref). This proves (a). Part (b) follows from the (rotational) consistency of $\widehat{\*F}_g$ established as a part of the proofs of Lemmas (ref) and (ref).{$\blacksquare$}
Proof of Theorem (ref).
By Lemma (ref), the rate of convergence given in Lemma (ref) is therefore not the best one possible. The improved rate implies that
which can be inserted back into (ref), leading to
Note that $\*Z_i(\widehat{\*F})$ is $\widehat{\*Z}_i$ with $\widehat a_{ij}$ replaced by $a_{ij}$. We now show that the effect of the estimation of $a_{ij}$ is negligible. We begin by noting how
Here,
where the second equality here follows from the definition of $\*V_{g}^0$, which is again based on taking $d_g$ as known, while the last equality follows from direct calculation. The effect of the estimation of $a_{ij}$ in $\widehat{\*Z}_i$ is therefore negligible, which in turn implies that the right-hand side of (ref) becomes
It follows that
Consider $(NT)^{-1/2}\sum_{i=1}^N\*D_T \*Z_i(\widehat{\*F})'\+\varepsilon_i$. From the definition of $ \*Z_i(\*F)$,
where
Here,
with obvious implicit definitions of $\*R_{11},\ldots,\*R_{14}$ and $\*H = \mathrm{diag}(\*H_{1}, \ldots, \*H_{G} )$. Let $\*R_{1m,j}$ be the $j$-th row of $\*R_{1m}$ for $m\in\{1,\ldots,4\}$. In this notation,
where the equality follows from Lemma (ref) and the fact that $T^{(\nu_g-\delta)/2}\*H_g =O_P(1)$. Similarly, for $\*R_{12}$,
where the second equality is due to (ref).
Consider $\*R_{14}$. From $T^{-\delta/2}\|\widehat{\*F}-\*F^0\*H \|=o_P(1)$ and $\|T^{(\nu_g-\delta)/2}\*H_g\| =O_P(1)$, we obtain
Together with Lemma (ref) this implies
Finally, let us consider $\*R_{13}$.
Expanding $\widehat{\*F}_g \*H_g^{-1}-\*F_g^0 $ as we did for $\*L$ above, we then just need to focus on the leading terms equivalent to $\*L_2$, $\*L_{8}$, and $\*L_{10}$.
It is easy to see that $\| \*R_{131}\|=o_P(\|\+\beta^0 -\widehat{\+\beta}_0 \|)$, so negligible. We now examine $\*R_{133}$. Then we can write
where the first equality follows from (ref), and the last step follows from some routine analysis using Assumption (ref). Below, we investigate $\*R_{132}$, which is one source of the bias term.
where the second equality follows from the definition of $\*Q_g$, and the third equality follows from (ref). Then we are able to further write
Hence, by adding the results,
$\sqrt{NT} \*R_2$ can be evaluated in exactly the same way, and the limiting representation is analogous to the one given above for $\sqrt{NT} \*R_1$ with $\*X_i$ replaced by $-\sum_{j=1}^{N}\*X_ja_{ij}$. Moreover, $\|\*B(\widehat{\*F}) - \*B(\*F^0)\| = o_P(1).$ It follows that if we define
where $\*Z_i(0) = \*X_i - \sum_{j=1}^{N}\*X_ja_{ij}$, then (ref) reduces to
which in turn implies that (ref) becomes
The required result is now implied by Assumption (ref). {$\blacksquare$}
Proof of Corollary (ref).
Under $NT^{-\nu_G}\to 0$, (ref) in the proof of Theorem (ref) may be written as
Consider $(NT)^{-1} \sum_{i=1}^N \*D_T\*X_i'\*M_{\widehat{F}} \+\varepsilon_i$, which we can write as
As in the proof of Theorem (ref),
where $\|\*R_{11}\|$, $\|\*R_{12}\|$ and $\|\*R_{14}\|$ are all $o_P((NT)^{-1/2})$, just as before. Let us therefore consider $\*R_{13}$, which is the source of the bias in the asymptotic distribution of $\sqrt{NT}\*D_T^{-1} (\widehat{\+\beta} - \+\beta^0 )$. The purpose of Assumption (ref) is to control this term. Suppose first that condition (b) holds. Let $x_{j,k,t}$ denote the $j$-th row of $\*x_{k,t}$. In this notation,
Moreover, by applying Assumption (ref) (b) to the expression given for $\|T^{-\delta/2}\widehat{\*F}_{2,d} - T^{-\nu_2/2}\*F_2^0\*h_{2,d}^{0}\|$ in the proof of Lemma (ref), we can show that
Making use of these results, we obtain
Alternatively, we may invoke Assumption (ref) (a) to arrive at the same result. In this case, $\|T^{-\delta/2}\mathrm{vec}(\widehat{\*F}-\*F^0\*H)' \| = o_P( 1 )$, but we also have
and so $\|\*R_{3,j} \|$ is of the same order as before. The proof under Assumption (ref) (c) is simpler and is therefore omitted.
Hence, by adding the results,
We have therefore shown that
and we can similarly show that
These results can be inserted into (ref), giving
The sought result now follows from Assumption (ref).{$\blacksquare$}