EconBase
← Back to paper

Dynamic Linear Panel Regression Models with Interactive Fixed Effects

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

108,937 characters · 16 sections · 69 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Dynamic Linear Panel Regression Models with Interactive Fixed Effects

abstractWe analyze linear panel regression models with interactive fixed effects and predetermined regressors, for example lagged-dependent variables. The first-order asymptotic theory of the least squares (LS) estimator of the regression coefficients is worked out in the limit where both the cross-sectional dimension and the number of time periods become large. We find two sources of asymptotic bias of the LS estimator: bias due to correlation or heteroscedasticity of the idiosyncratic error term, and bias due to predetermined (as opposed to strictly exogenous) regressors. We provide a bias-corrected LS estimator. We also present bias-corrected versions of the three classical test statistics (Wald, LR, and LM test) and show their asymptotic distribution is a ${\Greekmath 011F}^2$-distribution. Monte Carlo simulations show the bias correction of the LS estimator and of the test statistics also work well for finite sample sizes.

Introduction

In this paper, we study a linear panel regression model in which the individual fixed effects ${\Greekmath 0115}_{i}$, called factor loadings, interact with common time-specific effects $f_{t}$, called factors. This interactive fixed effect specification contains the conventional individual specific effects and time-specific effects as special cases but is significantly more flexible because it allows the factors $f_t$ to affect each individual with a different loading ${\Greekmath 0115}_i$.

Factor models have been widely studied in various economics disciplines, for example, in asset pricing, forecasting, empirical macro, and empirical labor economics.\footnote{See, e.g., Chamberlain and Rothschild CR1983, Ross Ross1976, and Fama and French FamaFrench1993 for asset pricing; Stock and Watson StockWatson2002 and Bai and Ng BaiNg2006 for forecasting; Bernanke, Boivin, and Eliasz BernankeBoivinEliasz2005 for empirical macro; and Holtz-Eakin, Newey, and Rosen HoltzEakin-Newey-Rosen1988 for empirical labor economics.} The panel literature often uses factor models to represent time-varying individual effects (or heterogenous time effects), so-called interactive fixed effects. For panels with a large cross-sectional dimension ($N$) but a short time dimension ($T$), Holtz-Eakin, Newey, and Rosen HoltzEakin-Newey-Rosen1988 (hereafter HNR) study a linear panel regression model with interactive fixed effects and lagged dependent variables. To solve the incidental parameter problem caused by the ${\Greekmath 0115}_i$'s, they estimate a quasi-differenced version of the model using appropriate lagged variables as instruments, and treating $f_t$'s as a fixed number of parameters to estimate. Ahn, Lee, and Schmidt AhnLeeSchmidt2001 also consider large $N$ but short $T$ panels. Instead of eliminating the individual effects ${\Greekmath 0115}_i$ by transforming the panel data, they impose various second-moment restrictions including the correlated random effects ${\Greekmath 0115}_i$, and derive moment conditions to estimate the regression coefficients. The more recent literature considers panels with comparable size of $N$ and $T$. The interactive fixed effect panel regression model of Pesaran Pesaran2006 allows heterogenous regression coefficients. Pesaran's estimator is the common correlated effect (CCE) estimator that uses the cross-sectional averages of the dependent variable and the independent variables as control functions for the interactive fixed effects.\footnote{ The theory of the CCE estimator was further developed in, e.g., Harding and Lamarche HardingLamarche2009,HardingLamarche2011, Kapetanios, Pesaran, and Yamagata KapetaniosPesaranYamagata2011, Pesaran and Tosetti PesaranTosetti2011, Chudik, Pesaran, and Tosetti ChudikPesaranTosetti2011, and Chudik and Pesaran ChudikPesaran2015. }

Among the interactive fixed effect panel literature, most closely related to our paper is Bai Bai2009. Bai assumes the regressors are strictly exogenous and the number of factors is known. The estimator he investigates is the least squares (LS) estimator, which minimizes the sum of squared residuals of the model jointly over the regression coefficients and the fixed effect parameters ${\Greekmath 0115}_i$ and $f_t$.\footnote{ The LS estimator is sometimes called “concentrated” least squares estimator in the literature, and in an earlier version of the paper, we referred to it as the “Gaussian Quasi Maximum Likelihood Estimator”, because LS estimation is equivalent to maximizing a conditional Gaussian likelihood function.} Using alternative asymptotics where $N,T\rightarrow \infty$ at the same rate,\footnote{ Hahn and Kuersteiner HahnKuersteiner2002 introduced the alternative asymptotics to characterize the asymptotic bias due to incidental parameter problems in fixed effect dynamic panel data models. See also Arellano and Hahn ArellanoHahn2007 and Moon, Perron, and Phillips MoonPerronPhillips2014 and references therein. } Bai shows the LS estimator is $\sqrt{NT}$-consistent and asymptotically normal, but may have an asymptotic bias. The bias in the normal limiting distribution occurs when the regression errors are correlated or heteroscedastic. Bai also shows how to estimate the bias, and proposes a bias-corrected estimator.

Following the methodology in Bai Bai2009, we investigate the LS estimator for a linear panel regression with a known number of interactive fixed effects. The main difference from Bai is that we consider predetermined regressors, thus allowing feedback of past outcomes to future regressors. One of the main findings of the present paper is that the limit distribution of the LS estimator has two types of biases: one type of bias due to correlated or heteroscedastic errors (the same bias as in Bai) and the other type of bias due to the predetermined regressors. This additional bias term is analogous to the incidental parameter bias of Nickell Nickell1981 in finite $T$ and the bias in Hahn and Kuersteiner HahnKuersteiner2002 in large $T$.

In addition to allowing for predetermined regressors, we also extend Bai's results to models in which both “low-rank regressors” (e.g., time-invariant and common regressors, or interactions of those two) and “high-rank-regressors” (almost all other regressors that vary across individuals and over time) are present simultaneously, wheras Bai (2009) only considers the low-rank regressors separately and in a restrictive setting (in particular, not allowing for regressors that are obtained by interacting time-invariant and common variables). A general treatment of low-rank regressors is desirable because they often occur in applied work, for example, Gobillon and Magnac GobillonMagnac2013. The analysis of those regressors is challenging, however, because the unobserved interactive fixed effects also represent a low-rank $N \times T$ matrix, thus posing a non-trivial identification problem for low-rank regressors, which needs to be addressed. We provide conditions under which the different types of regressors are identified jointly, and under which they can be estimated consistently as $N$ and $T$ grow large.

Another contribution of this paper is to establish the asymptotic theory of the three classical test statistics (Wald test, LR test, and LM (or score) test) for testing restrictions on the regression coefficients in a large $N$, $T$ panel framework.\footnote{ The “likelihood ratio” and the score used in the tests are based on the LS objective function, which can be interpreted as the (misspecified) conditional Gaussian likelihood function. } Regarding testing for coefficient restrictions, Bai Bai2009 investigates the Wald test based on the bias-corrected LS estimator, and HNR consider the LR test in their 2SLS estimation framework with fixed $T$.\footnote{ Another type of widely studied tests in the interactive fixed effect panel literature are panel unit root tests, e.g., Bai and Ng BaiNg2004, Moon and Perron MoonPerron2004, and Phillips and Sul PhillipsSul2003. } What we show is that the conventional LR and LM test statistics based on the LS profile objective function have non-central chi-square limits due to incidental parameters in the interactive fixed effects. We therefore propose modified LR and LM tests whose asymptotic distributions are conventional chi-square distributions.

To establish the asymptotic theories of the LS estimator and the three classical tests, we use the quadratic approximation of the profile LS objective function derived in Moon and Weidner MoonWeidner2015. This method is different from Bai Bai2009, who uses the first-order condition of the LS optimization problem as the starting point of his analysis. One advantage of our methodology is that it can also directly be applied to derive the asymptotic properties of the LR and LM test statistics.

In this paper, we assume the regressors are not endogenous and the number of factors is known, which might be restrictive in some applications. In other papers, we study how to relax these restrictions. Moon and Weidner MoonWeidner2015 investigates the asymptotic properties of the LS estimator of the linear panel regression model with factors when the number of factors is unknown and extra factors are included unnecessarily in the estimation. We find that under suitable conditions,\footnote{In Moon and Weidner MoonWeidner2015 we do not consider low-rank regressors or testing problems, and we impose more restrictive assumptions on the error term of the model implying that some leading bias terms of the LS estimator are not present.} the limit distribution of the LS estimator is unchanged when the number of factors is overestimated. The extension to allow for endogenous regressors is very briefly discussed in section 6 of the current paper, and is closely related to the results in Moon, Shum, and Weidner MoonShumWeidner2012 (hereafter MSW). MSW's main purpose is to extend the random coefficient multinomial logit demand model (known as the BLP demand model from Berry, Levinsohn, and Pakes BerryLevinsohnPakes1995) by allowing for interactive product and market specific fixed effects. Although the main model of interest is quite different from the linear panel regression model of the current paper, MSW's econometrics framework is directly applicable to the model of the current paper with endogenous regressors.\footnote{Lee, Moon, and Weidner LeeMoonWeidner2012 also apply the MSW estimation method to estimate a simple dynamic panel regression with interactive fixed effect and classical measurement errors.}

Comparing the different estimation approaches for interactive fixed effect panel regressions proposed in the literature, it seems fair to say that the LS estimator in Bai Bai2009 and our paper, the CCE estimator of Pesaran Pesaran2006, and the IV estimator based on quasi-differencing in HNR all have their own relative advantages and disadvantages. These three estimation methods handle the interactive fixed effects quite differently. The LS method concentrates out the interactive fixed effects by taking out the principal components. The CCE method controls the factor (or time effects) using the cross-sectional averages of the dependent and independent variables. The HNR's approach quasi-differences out the individual effects, treating the remaining time effects as parameters to estimate. The IV estimator of HNR should work well when $T$ is short, but is expected to also suffer from an incidental parameter problem when $T$ becomes large, because then many factors need to be estimated as parameters that enter the model non-linearly. Pesaran's CCE estimation method does not require the number of factors to be known and does not require the strong factor assumption that we will impose below, but for the CCE estimator to work, not only the DGPs of the dependent variable (e.g., the regression model) but also the DGPs of the explanatory variables need to be restricted such that their cross-sectional average can control for unobserved factors. The LS estimator and its bias-corrected version perform well under relatively weak restrictions on the regressors, but requires that $T$ should not be too small and that the factors should be sufficiently strong to be correctly picked up as the leading principal components.

The paper is organized as follows. In section (ref), we introduce the interactive fixed effect model and provide conditions for identifying the regression coefficients in the presence of the interactive fixed effects. In section (ref), we define the LS estimator of the regression parameters and provide a set of assumptions that are sufficient to show consistency of the LS estimator. In section (ref), we work out the asymptotic distribution of the LS estimator under alternative asymptotics. We also provide a consistent estimator for the asymptotic bias and a bias-corrected LS estimator. In section (ref), we consider the Wald, LR, and LM tests for testing restrictions on the regression coefficients of the model. We present bias-corrected versions of these tests and show that they have chi-square limiting distributions. In section (ref), we briefly discuss how to estimate the interactive fixed effect linear panel regression when the regressors are endogenous. In section (ref), we present Monte Carlo simulation results for an ${\rm AR}(1)$ model with interactive fixed effects. The simulations show the LS estimator for the ${\rm AR}(1)$ coefficient is biased, and the tests based on it can have severe size distortions and power asymmetries, wheras the bias-corrected LS estimator and test statistics have better properties. We conclude in section (ref). We present all proofs of theorems and some technical details in the appendix or supplementary material.

A few words on notation are due. For a column vector $v$, the Euclidean norm is defined by $\| v \| = \sqrt{v^{\prime}v}$. For the $n$-th largest eigenvalues (counting multiple eigenvalues multiple times) of a symmetric matrix $B$, we write ${\Greekmath 0116}_n(B)$. For an $m\times n$ matrix $A$, the Frobenius norm is $\| A \|_{F} = \sqrt{{\rm Tr} (AA^{\prime})}$, and the spectral norm is $\| A \| = \max_{0 \neq v \in \mathbb{R}^n} \, \frac{ \| A v \|} {\| v\|}$, or equivalently $\| A \| = \sqrt{ {\Greekmath 0116}_1(A^{\prime}A) }$. Furthermore, we define $P_A = A (A^{\prime}A)^{\dagger} A'$ and $M_A = \mathbb{I} - A (A^{\prime}A)^{\dagger} A'$, where $\mathbb{I}$ is the $m\times m$ identity matrix, and $(A^{\prime}A)^{\dagger}$ is the Moore-Penrose pseudoinverse, to allow for the case that $A$ is not of full column rank. For square matrices $B$, $C$, we write $B>C$ (or $B\geq C$) to indicate $B-C$ is positive (semi) definite. For a positive definite symmetric matrix $A$, we write $A^{1/2}$ and $A^{-1/2}$ for the unique symmetric matrices that satisfy $A^{1/2} A^{1/2}=A$ and $A^{-1/2} A^{-1/2} =A^{-1}$. We use ${\Greekmath 0272}$ for the gradient of a function; that is, ${\Greekmath 0272} f(x)$ is the column vector of partial derivatives of $f$ with respect to each component of $x$. We use “wpa1” for “with probability approaching one”.

Model and Identification

We study the following panel regression model with cross-sectional size $N$, and $T$ time periods:

align[align omitted — 155 chars of source]

where $X_{it}$ is a $K\times 1$ vector of observable regressors, ${\Greekmath 010C}^{0}$ is a $K\times 1$ vector of regression coefficients, ${\Greekmath 0115}_{i}^{0}$ is an $R\times 1$ vector of unobserved factor loadings, $f_{t}^{0}$ is an $R\times1$ vector of unobserved common factors, and $e_{it}$ are unobserved errors. The superscript zero indicates the true parameters. We write $f_{tr}^{0}$ and ${\Greekmath 0115}^0_{ir}$, where $r=1,\ldots,R$, for the components of ${\Greekmath 0115}_{i}^{0}$ and $f_{t}^{0}$, respectively. $R$ is the number of factors. Note that we can have $f_{tr}^{0}=1$ for all $t$ and a particular $r$, in which case the corresponding ${\Greekmath 0115}^0_{ir}$ become standard individual-specific effects. Analogously, we can have ${\Greekmath 0115}^0_{ir}=1$ for all $i$ and a particular $r$, so that the corresponding $f_{tr}^{0}$ become standard time-specific effects.

Throughout this paper, we assume the true number of factors $R$ is known.\footnote{ To remove this restriction, one could estimate $R$ consistently in the presence of the regressors. In the literature so far, however, consistent estimation procedures for $R$ are established mostly in pure factor models (e.g., Bai and Ng BaiNg2002, Onatski Onatski2010 and Harding Harding2007). Alternatively, one could rely on Moon and Weidner MoonWeidner2015 who consider a regression model with interactive fixed effects when only an upper bound on the number of factors is known --- but extending those results to the more general setup considered here is mathematically challenging.} We introduce the notation ${\Greekmath 010C}^0 \cdot X = \sum_{k=1}^{K}\,{\Greekmath 010C}_{k}^{0}\,X_{k}$. In matrix notation, the model can then be written as

align*[align* omitted — 87 chars of source]

where $Y$, $X_{k}$, and $e$ are $N\times T$ matrices, ${\Greekmath 0115}^0$ is an $N\times R$ matrix, and $f^0$ is a $T\times R$ matrix. The elements of $X_k$ are denoted by $X_{k,it}$.

We separate the $K$ regressors into $K_1$ “low-rank regressors” $X_{l}$, $l=1,\ldots,K_1$, and $K_2=K-K_1$ “high-rank regressors” $X_{m}$, $m=K_1+1,\ldots,K$. Each low-rank regressor $l=1,\ldots,L$ is assumed to satisfy ${\rm rank}(X_l) = 1$. Therefore, we can write $X_l=w_l v_l'$, where $w_l$ is an $N$-vector and $v_l$ is a $T$-vector, and we also define the $N \times K_1$ matrix $w=(w_1,\ldots,w_{K_1})$ and the $T \times K_1$ matrix $v=(v_1,\ldots,v_{K_1})$.

Let $l=1,\ldots,K_1$. The two most prominent types of low-rank regressors are time-invariant regressors, which satisfy $X_{l,it}=Z_{i}$ for all $i,t$, and common (or cross-sectionally invariant) regressors, in which case $X_{l,it}=W_{t}$ for all $i,t$. Here, $Z_{i}$ and $W_{t}$ are some observed variables, which only vary over $i$ or $t$, respectively. A more general low-rank regressor can be obtained by interacting $Z_i$ and $W_t$ multiplicatively, namely, $X_{l,it}=Z_{i} W_t$, an empirical example of which is given in Gobillon and Magnac GobillonMagnac2013. In these examples, and probably for the vast majority of applications, the low-rank regressors all satisfy ${\rm rank}(X_{l})=1$, but our results can easily be extended to more general low-rank regressors.\footnote{ If we have low-rank regressors with rank larger than one, then we write $X_l=w_l v_l'$, where $w_l$ is an $N\times {\rm rank}(X_l)$ matrix and $v_l$ is a $T \times {\rm rank}(X_l)$ matrix, and we define $w=(w_1,\ldots,w_{K_1})$ as a $N \times \sum_{l=1}^L {\rm rank}(X_l)$ matrix, and $v=(v_1,\ldots,v_{K_1})$ ae a $T \times \sum_{l=1}^L {\rm rank}(X_l)$ matrix. All our results are then unchanged, as long as ${\rm rank}(X_l)$ is a finite constant for all $l=1,\ldots,K_1$, and we replace $2 R + K_1$ by $ 2 R + {\rm rank}\left( w \right)$ in Assumption (ref)$(v)$ and Assumption (ref)$(ii)(a)$. }

High-rank regressors are those whose distribution guarantees they have high rank (usually full rank) when considered as an $N \times T$ matrix. For example, a regressor whose entries satisfy $X_{m,it} \sim iid \, {\cal N}({\Greekmath 0116},{\Greekmath 011B})$, with ${\Greekmath 0116} \in \mathbb{R}$ and ${\Greekmath 011B}>0$, satisfies ${\rm rank}(X_m)=\min(N,T)$ with probability one.

This separation of the regressors into low- and high-rank regressors is important to formulate our assumptions for identification and consistency, but actually plays no role in the estimation and inference procedures for $\widehat {\Greekmath 010C}$ discussed below.

samepage\newtheorem{IDassumption}{Assumption} \begin{IDassumption}[\bf Assumptions for Identification] \begin{itemize} • {\bf Existence of Second Moments:} \\ The second moments of $X_{k,it}$ and $e_{it}$ conditional on ${\Greekmath 0115}^0$, $f^0$, $w$ exist for all $i$, $t$, $k$. • {\bf Mean Zero Errors and Exogeneity:} \\ $\mathbb{E}\left( e_{it} | {\Greekmath 0115}^0, f^0,w \right)=0$, \, and \, $\mathbb{E}(X_{k, it} e_{it} | {\Greekmath 0115}^0, f^0,w)=0$, \, a.s., \, for all $i$, $t$, $k$. \end{itemize} The following two assumptions only need to be imposed if $K_1>0$, that is, if low-rank regressors are present: \begin{itemize} • {\bf Non-collinearity of Low-Rank Regressors:} \\ Consider linear combinations ${\Greekmath 010B} \cdot X_{{\rm low}} = \sum_{l=1}^{K_1} {\Greekmath 010B}_l X_l$ of the low-rank regressors $X_l$ with ${\Greekmath 010B} \in \mathbb{R}^{K_1}$. For all ${\Greekmath 010B} \neq 0$, we assume \begin{align*} \mathbb{E}\left[ ({\Greekmath 010B} \cdot X_{{\rm low}}) M_{f^0} ({\Greekmath 010B} \cdot X_{{\rm low}})' \big| {\Greekmath 0115}^0, f^0,w \right] \neq 0 \, , \qquad \relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{a.s.} \end{align*} • {\bf No Collinearity between Factor Loadings and Low-Rank Regressors:} \\ ${\rm rank}( M_{w} {\Greekmath 0115}^0 ) = {\rm rank}({\Greekmath 0115}^0)$.\footnote{ Note that ${\rm rank}({\Greekmath 0115}^0) = R$ if $R$ factors are present. Our identification results are consistent with the possibility that ${\rm rank}({\Greekmath 0115}^0) < R$, i.e., that $R$ only represents an upper bound on the number of factors, but later we assume ${\rm rank}({\Greekmath 0115}^0) = R$ to show consistency. } \end{itemize} The following assumption only needs to be imposed if $K_2>0$, that is, if high-rank regressors are present: \begin{itemize} • {\bf Non-collinearity of High-Rank Regressors:} \\ Consider linear combinations ${\Greekmath 010B} \cdot X_{{\rm high}} = \sum_{m=K_1+1}^K {\Greekmath 010B}_m X_m$ of the high-rank regressors $X_m$ for ${\Greekmath 010B} \in \mathbb{R}^{K_2}$, where the components of the $K_2$-vector ${\Greekmath 010B}$ are denoted by ${\Greekmath 010B}_{K_1+1}$ to ${\Greekmath 010B}_{K}$. For all ${\Greekmath 010B} \neq 0$, we assume \begin{align*} {\rm rank}\left\{ \mathbb{E}\left[ ({\Greekmath 010B} \cdot X_{{\rm high}}) ({\Greekmath 010B} \cdot X_{{\rm high}})' \big| {\Greekmath 0115}^0, f^0,w \right] \right\} > 2 R + K_1 \, , \qquad \relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{a.s.} \end{align*} \end{itemize} \end{IDassumption}

All expectations in the assumptions are conditional on ${\Greekmath 0115}^0$, $f^0$, and $w$; in particular, $e_{it}$ is not allowed to be correlated with ${\Greekmath 0115}^0$, $f^0$, and $w$. However, $e_{it}$ is allowed to be correlated with $v$ (i.e., predetermined low-rank regressors are allowed). If desired, one can interchange the role of $N$ and $T$ in the assumptions, by using the formal symmetry of the model under exchange of the panel dimensions ($N \leftrightarrow T$, ${\Greekmath 0115}^0 \leftrightarrow f^0$, $Y \leftrightarrow Y'$, $X_k \leftrightarrow X_k'$, $w \leftrightarrow v$).

Assumptions (ref)$(i)$ and $(ii)$ have standard interpretations, but the other assumptions require some further discussion.

Assumption (ref)$(iii)$ states the low-rank regressors are non-collinear even after projecting out all variation that is explained by the true factors $f^0$. This assumption would, for example, be violated if $v_l = f^0_r$ for some $l=1,\ldots,K_1$ and $r=1,\ldots,R$, because then $X_l M_{f^0} = 0$ and we can choose ${\Greekmath 010B}$ such that $X_{\rm low} = X_l$. Similarly, Assumption (ref)$(iv)$ rules out, for example, that $w_l = {\Greekmath 0115}^0_r$ for some $l=1,\ldots,K_1$ and $r=1,\ldots,R$, because then ${\rm rank}( M_{w} {\Greekmath 0115}^0 ) < {\rm rank}({\Greekmath 0115}^0)$, in general. It ought to be expected that ${\Greekmath 0115}^0$ and $f^0$ have to feature in the identification conditions for the low-rank regressors, because the interactive fixed effects structure and the low-rank regressors represent similar types of low-rank $N \times T$ structures.

Assumption (ref)$(v)$ is a generalized non-collinearity assumption for the high-rank regressors, which guarantees any linear combination ${\Greekmath 010B} \cdot X_{{\rm high}}$ of the high-rank regressors is sufficiently different from the low-rank regressors and from the interactive fixed effects. A standard non-collinearity assumption can be formulated by demanding the $N \times N$ matrix $\mathbb{E}\big[ \allowbreak ({\Greekmath 010B} \cdot X_{{\rm high}}) ({\Greekmath 010B} \cdot X_{{\rm high}})' \big| \allowbreak {\Greekmath 0115}^0, f^0,w \big]$ is non-zero for all non-zero ${\Greekmath 010B} \in \mathbb{R}^{K_2}$, which can be equivalently expressed as ${\rm rank}\big\{ \allowbreak \mathbb{E}\left[ ({\Greekmath 010B} \cdot X_{{\rm high}}) ({\Greekmath 010B} \cdot X_{{\rm high}})' \big| {\Greekmath 0115}^0, f^0,w \right] \big\} \allowbreak > 0$ for all non-zero ${\Greekmath 010B} \in \mathbb{R}^{K_2}$. Assumption (ref)$(v)$ strengthens this standard non-collinearity assumption by demanding the rank not only to be positive, but larger than $2R+K_1$. This also explains the name “high-rank regressors,” because their rank has to be sufficiently large to satisfy this assumption. Note also that only the number of factors $R$, but not ${\Greekmath 0115}^0$ and $f^0$, features in Assumption (ref)$(v)$. The sample version of this assumption is given by Assumption (ref)$(ii)(a)$ below, which is also very closely related to Assumption A in Bai Bai2009.

theorem[\bf Identification] Suppose the Assumptions (ref) are satisfied. Then, the minima of the expected objective function $\mathbb{E}\left(\left\| Y \, - \, {\Greekmath 010C} \cdot X \, - \, {\Greekmath 0115} \, f' \right\|^2_F \Big| {\Greekmath 0115}^0, f^0,w \right)$ over $({\Greekmath 010C},{\Greekmath 0115},f) \in \mathbb{R}^{K + N \times R + T \times R}$ satisfy ${\Greekmath 010C}={\Greekmath 010C}^0$ and ${\Greekmath 0115} f' = {\Greekmath 0115}^0 f^{0 \prime}$. This shows that ${\Greekmath 010C}^0$ and ${\Greekmath 0115}^0 f^{0 \prime}$ are identified.

The theorem shows the true parameters are identified as minima of the expected value of $\left\| Y \, - \, {\Greekmath 010C} \cdot X \, - \, {\Greekmath 0115} \, f' \right\|^2_F = \sum_{i,t} ( Y_{it} {\Greekmath 010C}^{\prime } - X_{it}- {\Greekmath 0115}_{i}^{\prime }f_{t} )^2$, which is the sum of squared residuals. We use the same objective function, to define the estimators $\widehat {\Greekmath 010C}$, $\widehat {\Greekmath 0115}$ and $\widehat f$ below. Without further normalization conditions, the parameters ${\Greekmath 0115}^0$ and $f^0$ are not separately identified, because the outcome variable $Y$ is invariant under transformations ${\Greekmath 0115}^0 \rightarrow {\Greekmath 0115}^0 A'$ and $f^0 \rightarrow f^0 A^{-1}$, where $A$ is a non-singular $R\times R$ matrix. However, the product ${\Greekmath 0115}^0 f^{0 \prime}$ is uniquely identified according to the theorem. Because our focus is on identification and estimation of ${\Greekmath 010C}^0$, we do not need to discuss those additional normalization conditions for ${\Greekmath 0115}^0$ and $f^0$ in this paper.

Estimator and Consistency

The objective function of the model is simply the sum of squared residuals, which in matrix notation can be expressed as

align[align omitted — 437 chars of source]

The estimator we consider is the LS estimator that jointly minimizes ${\cal L}_{NT}({\Greekmath 010C},{\Greekmath 0115},f)$ over ${\Greekmath 010C}$, ${\Greekmath 0115}$ and $f$. Our main objects of interest are the regression parameters ${\Greekmath 010C}=\left( {\Greekmath 010C}_{1},...,{\Greekmath 010C}_{K}\right)^{\prime}$, whose estimator is given by

align[align omitted — 160 chars of source]

where $\mathbb{B} \subset \mathbb{R}^{K}$ is a compact parameter set that contains the true parameter, namely, ${\Greekmath 010C}^0 \in \mathbb{B}$, and the objective function is the profile objective function

align[align omitted — 564 chars of source]

Here, the first expression for $L_{NT}({\Greekmath 010C})$ is its definition as the minimum value of ${\cal L}_{NT}({\Greekmath 010C},{\Greekmath 0115},f)$ over ${\Greekmath 0115}$ and $f$. We denote the minimizing incidental parameters by $\widehat {\Greekmath 0115}({\Greekmath 010C})$ and $\widehat f({\Greekmath 010C})$, and we define the estimators $\widehat {\Greekmath 0115} = \widehat {\Greekmath 0115}(\widehat {\Greekmath 010C})$ and $\widehat f = \widehat f(\widehat {\Greekmath 010C})$. Those minimizing incidental parameters are not uniquely determined -- for the same reason that ${\Greekmath 0115}^0$ and $f^0$ are non uniquely identified -- but the product $\widehat {\Greekmath 0115}({\Greekmath 010C}) \widehat f'({\Greekmath 010C})$ is unique.

The second expression for $L_{NT}\left( {\Greekmath 010C} \right)$ in equation (ref) is obtained by concentrating out ${\Greekmath 0115}$ (analogously, one can concentrate out $f$ to obtain a formulation whereby only the parameter ${\Greekmath 0115}$ remains). The optimal $f$ in the second expression is given by the $R$ eigenvectors that correspond to the $R$ largest eigenvalues of the $T\times T$ matrix $\left( Y-{\Greekmath 010C} \cdot X \right)^{\prime}\left( Y- {\Greekmath 010C} \cdot X \right)$. This insight leads to the third line that presents the profile objective function as the sum over the $T-R$ smallest eigenvalues of this $T\times T$ matrix. Lemma (ref) in the appendix shows equivalence of the three expressions for $L_{NT}({\Greekmath 010C})$ given above.

Multiple local minima of $L_{NT}({\Greekmath 010C})$ may exist, and one should use multiple starting values for the numerical optimization of ${\Greekmath 010C}$ to guarantee the true global minimum $\widehat {\Greekmath 010C}$ is found.

To show consistency of the LS estimator $\widehat {\Greekmath 010C}$ of the interactive fixed effect model, and also later for our first-order asymptotic theory, we consider the limit $N,T\rightarrow \infty$. In the following we present assumptions on $X_k$, $e$, ${\Greekmath 0115}$, and $f$ that guarantee consistency.\footnote{ We could write $X_k^{(N,T)}$, $e^{(N,T)}$, ${\Greekmath 0115}^{(N,T)}$, and $f^{(N,T)}$, because all these matrices, and even their dimensions, are functions on $N$ and $T$, but we suppress this dependence throughout the paper.}

assumption(i) $\limfunc{plim}_{N,T \rightarrow \infty}\left({\Greekmath 0115}^{0\prime} {\Greekmath 0115}^0/N\right) > 0$, (ii) $\limfunc{plim}_{N,T \rightarrow \infty} \left( f^{0\prime} f^0 / T \right) > 0$.
assumption$\limfunc{plim}_{N,T \rightarrow \infty} \left[ (NT)^{-1} {\rm Tr}(X_k \, e^{\prime}) \right] = 0$, for all $k=1,\ldots,K$.
assumption$\limfunc{plim}_{N,T \rightarrow \infty} \left( \| e \| / \sqrt{NT} \right) = 0$.

Assumption (ref) guarantees the matrices $f^0$ and ${\Greekmath 0115}^0$ have full rank, that is, that $R$ distinct factors and factor loadings exist asymptotically, and that the norm of each factor and factor loading grows at a rate of $\sqrt{T}$ and $\sqrt{N}$, respectively. Assumption (ref) demands the regressors are weakly exogenous. Assumption (ref) restricts the spectral norm of the $N \times T$ error matrix $e$. We discuss this assumption in more detail in the next section, and we give examples of error distributions that satisfy this condition in section (ref) of the supplementary material. The final assumption needed for consistency is an assumption on the regressors $X_k$. We already introduced the distinction between the $K_1$ “low-rank regressors” $X_{l}$, $l=1,\ldots,K_1$, and the $K_2=K-K_1$ “high-rank regressors” $X_{m}$, $m=K_1+1,\ldots,K$ above.

assumption$\phantom{a}$ \begin{itemize} • $\limfunc{plim}_{N,T \rightarrow \infty} \left[ (NT)^{-1} \, \sum_{i=1}^N \, \sum_{t=1}^T \, X_{it} X_{it}' \right] > 0$. • The two types of regressors satisfy: \begin{itemize} • Consider linear combinations ${\Greekmath 010B} \cdot X_{{\rm high}}=\sum_{m=K_1+1}^K {\Greekmath 010B}_m X_m$ of the high-rank regressors $X_m$ for $K_2$-vectors ${\Greekmath 010B}$ with $\|{\Greekmath 010B}\|=1$, where the components of the $K_2$-vector ${\Greekmath 010B}$ are denoted by ${\Greekmath 010B}_{K_1+1}$ to ${\Greekmath 010B}_{K}$. We assume a constant $b>0$ exists such that \begin{align*} \min_{\{{\Greekmath 010B} \in \mathbb{R}^{K_2}, \|{\Greekmath 010B}\|=1\}} \, \sum_{r=2R+K_1+1}^N \, {\Greekmath 0116}_{r} \left[ \frac{ ({\Greekmath 010B} \cdot X_{{\rm high}}) ({\Greekmath 010B} \cdot X_{{\rm high}})' } {NT} \right] \; &\geq b \; \qquad \relax\ifmmode\expandafter\text@\else\expandafter\mbox\fi{wpa1.} \end{align*} • For the low-rank regressors, we assume ${\rm rank}(X_l)=1$, $l=1,\ldots,K_1$; that is, they can be written as $X_l=w_l v_l'$ for $N$-vectors $w_l$ and $T$-vectors $v_l$, and we define the $N\times K_1$ matrix $w=(w_1,\ldots,w_{K_1})$ and the $T\times K_1$ matrix $v=(v_1,\ldots,v_{K_1})$. We assume a constant $B>0$ exists such that $N^{-1} \, {\Greekmath 0115}^{0\prime} \, M_{w} \, {\Greekmath 0115}^0 \, > \, B \, \mathbb{I}_R$ and $T^{-1} \, f^{0\prime} \, M_{v} \, f^0 \, > \, B \, \mathbb{I}_R$, wpa1. \end{itemize} \end{itemize}

Assumption (ref)$(i)$ is a standard non-collinearity condition for all the regressors. Assumption (ref)$(ii)(a)$ is an appropriate sample analog of the identification Assumption (ref)$(v)$. If the sum in Assumption (ref)$(ii)(a)$ were to start from $r=1$, we would have $\sum_{r=1}^N {\Greekmath 0116}_{r} \left[ \frac{ ({\Greekmath 010B} \cdot X_{{\rm high}}) ({\Greekmath 010B} \cdot X_{{\rm high}})' } {NT} \right] = \frac 1 {NT} {\rm Tr}[ ({\Greekmath 010B} \cdot X_{{\rm high}}) ({\Greekmath 010B} \cdot X_{{\rm high}})' ]$, so that the assumption would become a standard non-collinearity condition. Not including the first $2R+K_1$ eigenvalues in the sum implies the $N \times N$ matrix $({\Greekmath 010B} \cdot X_{{\rm high}}) ({\Greekmath 010B} \cdot X_{{\rm high}})'$ needs to have rank larger than $2R + K_1$.

Assumption (ref)$(ii)(b)$ is closely related to the identification Assumptions (ref)$(iii)$ and $(iv)$. The appearance of the factors and factor loadings in this assumption on the low-rank regressors is inevitable to guarantee consistency. For example, consider a low-rank regressor that is cross-sectionally independent and proportional to the $r$'th unobserved factor, for example, $X_{l,it}=f_{tr}$. The corresponding regression coefficient ${\Greekmath 010C}_l$ is then not identified, because the model is invariant under a shift ${\Greekmath 010C}_l \mapsto {\Greekmath 010C}_l + a$, ${\Greekmath 0115}_{ir} \mapsto {\Greekmath 0115}_{ir} - a$, for an arbitrary $a\in \mathbb{R}$. This phenomenon is well known from ordinary fixed effect models, where the coefficients of time-invariant regressors are not identified. Assumption (ref)(ii)(b) therefore guarantees for $X_l=w_l v_l'$ that $w_l$ is sufficiently different from ${\Greekmath 0115}^0$, and $v_l$ is sufficiently different from $f^0$.

theorem[\bf Consistency] Let Assumptions (ref), (ref), (ref), and (ref) be satisfied; let the parameter set $\mathbb{B}$ be compact; and let ${\Greekmath 010C}^0 \in \mathbb{B}$. In the limit $N,T \rightarrow \infty$, we then have \begin{align*} \widehat {\Greekmath 010C} \; \limfunc{\longrightarrow}_p \; {\Greekmath 010C}^0 \; . \end{align*}

We assume compactness of $\mathbb{B}$ to guarantee existence of the minimizing $\widehat {\Greekmath 010C}$. We also use boundedness of $\mathbb{B}$ in the consistency proof, but only for those parameters ${\Greekmath 010C}_l$, $l=1\ldots K_1$, that correspond to low-rank regressors, that is, if only high-rank regressors ($K_1=0$) are present, the compactness assumption can be omitted, as long as existence of $\widehat {\Greekmath 010C}$ is guaranteed (e.g., for $\mathbb{B}=\mathbb{R}^K$).

Bai Bai2009 also proves consistency of the LS estimator of the interactive fixed effect model, but under somewhat different assumptions. He also employs what we call Assumptions (ref) and (ref), and he uses a low-level version of Assumption (ref). He demands the regressors to be strictly exogenous. Regarding consistency, the main difference between our assumptions and his is the treatment of high- and low-rank regressors. He first gives a condition on the regressors (his Assumption A) that rules out low-rank regressors, and later discusses the case in which all regressors are either time-invariant or common regressors (i.e., are all low rank). By contrast, our Assumption (ref) allows for a combination of high- and low-rank regressors, and for low-rank regressors that are more general than time-invariant and common regressors.

Asymptotic Distribution and Bias Correction

Because we have already shown consistency of the LS estimator $\widehat {\Greekmath 010C}$, it is sufficient to study the local properties of the objective function $L_{NT}({\Greekmath 010C})$ around ${\Greekmath 010C}^0$ to derive the first-order asymptotic theory of $\widehat {\Greekmath 010C}$. Moon and Weidner MoonWeidner2015 derived a useful approximation of $L_{NT}({\Greekmath 010C})$ around ${\Greekmath 010C}^0$, and we briefly summarize the ideas and results of this approximation in the following subsection. We then apply those results to derive the asymptotic distribution of the LS estimator, including working out the asymptotic bias, which was not done previously. Afterward, we discuss bias correction and inference.

Expansion of the Profile Objective Function

The last expression in equation (ref) for the profile objective function is convenient because it does not involve any minimization over the parameters ${\Greekmath 0115}$ or $f$. On the other hand, this expression cannot be easily discussed by analytic means, because in general, no explicit formula exists for the eigenvalues of a matrix. The conventional method that involves a Taylor series expansion in the regression parameters ${\Greekmath 010C}$ {\it alone} seems infeasible here. In Moon and Weidner MoonWeidner2015, we showed how to overcome this problem by expanding the profile objective function {\it jointly} in ${\Greekmath 010C}$ and $\|e\|$. The key idea is the following decomposition:

align*[align* omitted — 368 chars of source]

If the perturbation term is zero, the profile objective $L_{NT}({\Greekmath 010C})$ is also zero, because the leading term ${\Greekmath 0115}^0 f^{0\prime}$ has rank $R$, so that the $T-R$ smallest eigenvalues of $f^0 {\Greekmath 0115}^{0\prime} {\Greekmath 0115}^0 f^{0\prime}$ all vanish. One may thus expect that small values of the perturbation term should correspond to small values of $L_{NT}({\Greekmath 010C})$. This idea can indeed be made mathematically precise. By using the perturbation theory of linear operators (see, e.g., Kato Kato), one can work out an expansion of $L_{NT}({\Greekmath 010C})$ in the perturbation term, and one can show this expansion is convergent as long as the spectral norm of the perturbation term is sufficiently small.

The assumptions on the model made so far are in principle already sufficient to apply this expansion of the profile objective function, but to truncate the expansion at an appropriate order and to provide a bound on the remainder term that is sufficient to derive the first-order asymptotic theory of the LS estimator, we need to strengthen Assumption (ref) as follows.

\newtheorem*{assumption3p}{Assumption (ref)$^*$} \begin{assumption3p} $\|e\|=o_p(N^{2/3})$. \end{assumption3p}

In the rest of the paper, we only consider asymptotics in which $N$ and $T$ grow at the same rate; that is, we could equivalently write $o_p(T^{2/3})$ instead of $o_p(N^{2/3})$ in Assumption (ref)$^*$. In section (ref) of the supplementary material, we provide examples of error distributions that satisfy Assumption (ref)$^*$. In fact, for these examples, we have $\|e\|={\cal O}_p(\sqrt{\max(N,T)})$. A large literature studies the asymptotic behavior of the spectral norm of random matrices; see, for example, Geman Geman1980, Silverstein Silverstein1989, Bai, Silverstein, and Yin BaiSilvYin1988, Yin, Bai, and Krishnaiah BaiKrishYin1988, and Latala Latala2006. Loosely speaking, we expect the result $\|e\|={\cal O}_p(\sqrt{\max(N,T)})$ to hold as long as the errors $e_{it}$ have mean zero, uniformly bounded fourth moment, and weak time-serial and cross-sectional correlation (in some well-defined sense, see the examples).

We can now present the quadratic approximation of the profile objective function $L_{NT}({\Greekmath 010C})$ that we derived in Moon and Weidner MoonWeidner2015.

theorem[\bf Expansion of Profile Objective Function] Let Assumption (ref), (ref)$^*$, and (ref)(i) be satisfied, and consider the limit $N,T \rightarrow \infty$ with $N/T \rightarrow {\Greekmath 0114}^2$, $0<{\Greekmath 0114}<\infty$. Then, the profile objective function satisfies $L_{NT}({\Greekmath 010C}) = L_{q,NT}({\Greekmath 010C}) + (NT)^{-1} \, R_{NT}({\Greekmath 010C})$, where the remainder $R_{NT}({\Greekmath 010C})$ is such that for any sequence ${\Greekmath 0111}_{NT}\rightarrow 0$, we have \begin{align*} \sup_{\{{\Greekmath 010C} :\left\| {\Greekmath 010C} -{\Greekmath 010C}^{0} \right\| \leq {\Greekmath 0111}_{NT}\}} \frac{ \left| R_{NT}({\Greekmath 010C}) \right| } { \left( 1 + \sqrt{NT} \, \left\| {\Greekmath 010C} -{\Greekmath 010C}^{0} \right\| \right)^2 } = o_{p}\left( 1 \right) , \end{align*} and $L_{q,NT}({\Greekmath 010C})$ is a second-order polynomial in ${\Greekmath 010C}$; namely, \begin{align*} L_{q,NT}({\Greekmath 010C}) \, &= L_{NT}({\Greekmath 010C}^0) \, - \, \frac 2 {\sqrt{NT}} \, ({\Greekmath 010C}-{\Greekmath 010C}^0)' \, C_{NT} + \, ({\Greekmath 010C}-{\Greekmath 010C}^0)' \, W_{NT} \, ({\Greekmath 010C}-{\Greekmath 010C}^0) \; , \end{align*} with $K\times K$ matrix $W_{NT}$ defined by $W_{NT,k_1 k_2} = (NT)^{-1} \, {\rm Tr}(M_{f^0} \, X^{\prime}_{k_1} \, M_{{\Greekmath 0115}^0} \, X_{k_2})$, and $K$-vector $C_{NT}$ with entries $C_{NT,k} = C^{(1)}\left({\Greekmath 0115}^0 \, ,f^0 \, ,X_k \, e \right) +C^{(2)}\left({\Greekmath 0115}^0 \, ,f^0 \, ,X_k \, e \right)$, where \begin{align*} C^{(1)}\left({\Greekmath 0115}^0,\, f^0,\, X_{k},\, e \right) &= \frac 1 {\sqrt{NT}} \, {\rm Tr}(M_{f^0} \, e^{\prime}\, M_{{\Greekmath 0115}^0} \, X_k) \; , \notag \\ C^{(2)}\left({\Greekmath 0115}^0,\, f^0,\, X_{k},\, e \right) &= - \, \frac 1 {\sqrt{NT}} \, \bigg[ {\rm Tr}\left(e M_{f^0} \, e' \, M_{{\Greekmath 0115}^0} \, X_k \, f^0 \, (f^{0\prime}f^0)^{-1} \, ({\Greekmath 0115}^{0\prime}{\Greekmath 0115}^0)^{-1} \, {\Greekmath 0115}^{0\prime} \right) \nonumber \\ & \qquad \qquad \quad +{\rm Tr}\left(e^{\prime}M_{{\Greekmath 0115}^0} \, e \, M_{f^0} \, X^{\prime}_k \, {\Greekmath 0115}^0 \, ({\Greekmath 0115}^{0\prime}{\Greekmath 0115}^0)^{-1} \, (f^{0\prime}f^0)^{-1} \, f^{0\prime} \right) \nonumber \\ & \qquad \qquad \quad +{\rm Tr}\left(e^{\prime}M_{{\Greekmath 0115}^0} \, X_k \, M_{f^0} \, e^{\prime} \, {\Greekmath 0115}^0 \, ({\Greekmath 0115}^{0\prime}{\Greekmath 0115}^0)^{-1} \, (f^{0\prime}f^0)^{-1} \, f^{0\prime} \right) \bigg] \; . \end{align*}

We refer to $W_{NT}$ and $C_{NT}$ as the approximated Hessian and the approximated score (at the true parameter ${\Greekmath 010C}^0$). The exact Hessian and the exact score (at the true parameter ${\Greekmath 010C}^0$) contain higher-order expansion terms in $e$, but the expansion up to the particular order above is sufficient to work out the first-order asymptotic theory of the LS estimator, as the following corollary shows.

corollaryLet the assumptions of Theorem (ref) and (ref) hold; let ${\Greekmath 010C}^0$ be an interior point of the parameter set $\mathbb{B}$; and assume $C_{NT}={\cal O}_p(1)$. We then have $\sqrt{NT} \big( \widehat {\Greekmath 010C} - {\Greekmath 010C}^0 \big) = W^{-1}_{NT} C_{NT} + o_p(1) = {\cal O}_p(1)$.

Combining consistency of the LS estimator and the expansion of the profile objective function in Theorem (ref), one obtains $\sqrt{NT} \, W_{NT} \big( \widehat {\Greekmath 010C} - {\Greekmath 010C}^0 \big) = C_{NT} + o_p(1)$; see, for example, Andrews Andrews1999. To obtain the corollary, one needs in addition that $W_{NT}$ does not become degenerate as $N,T \rightarrow \infty$; that is, the smallest eigenvalue of $W_{NT}$ should be bounded from below by a positive constant. Our assumptions already guarantee existence of such a lower bound, as is shown in the supplementary material.

Analogous to the expansions of the profile objective function $L_{NT}({\Greekmath 010C})$, one can also derive expansions of the projectors $M_{\widehat {\Greekmath 0115}}$ and $M_{\widehat f}$, and those can be used to show consistency of $\widehat {\Greekmath 0115}$ and $\widehat f$, up to normalization; see Lemma (ref) in the supplementary material.

Asymptotic Distribution

We now apply Corollary (ref) to work out the asymptotic distribution of the LS estimator $\widehat {\Greekmath 010C}$. For this purpose, we need more specific assumptions on ${\Greekmath 0115}^0$, $f^0$, $X_k$, and $e$.

assumptionA sigma algebra ${\cal C} = {\cal C}_{NT}$ (which in the following we will refer to as the conditioning set) exists that contains the sigma algebra generated by ${\Greekmath 0115}^0$ and $f^0$, such that \begin{itemize} • $\mathbb{E}\left[ e_{it} \, \big| \, {\cal C} \vee {\Greekmath 011B}( \{ (X_{is}, e_{i,s-1}), s \leq t \}) \right] = 0$, for all $i,t$.\footnote{ Here and in the following, we write ${\Greekmath 011B}(A)$ for the sigma algebra generated by the (collection of) random variable(s) $A$, and we write ${\cal A} \vee {\cal B}$ for the sigma algebra generated by the unions of all elements in the sigma algebra ${\cal A}$ and ${\cal B}$, so that in the conditional expectation in Assumption (ref)(ii), we condition jointly on ${\cal C}$ and $\{ (X_{is}, e_{i,s-1}), s \leq t \}$. } • $e_{it}$ is independent over $t$, conditional on ${\cal C}$, for all $i$. • $\{ (X_{it}, e_{it}) , t=1,\ldots,T \}$ is independent across $i$, conditional on ${\cal C}$. • $\frac 1 {NT} \sum_{i=1}^N \sum_{t,s=1}^T \left| {\rm Cov}\left( X_{k,it}, X_{\ell,is} \Big| \, {\cal C} \right) \right| = {\cal O}_p(1)$, for all $k,\ell=1,\ldots,K$. • $\frac 1 {N T^2} \sum_{i=1}^N \sum_{t,s,u,v=1}^T \left| {\rm Cov}\left( e_{it} \widetilde X_{k,is}, \, e_{iu} \widetilde X_{\ell,iv} \Big| \, {\cal C} \right) \right| = {\cal O}_p(1)$, where $\widetilde X_{k,it} = X_{k,it} - \mathbb{E}\left[ X_{k,it} \big| {\cal C} \right]$, \\ for all $k,\ell=1,\ldots,K$. • An ${\Greekmath 010F}>0$ exists such that $\mathbb{E}\left( e_{it}^8 \big| \, {\cal C} \right)$ and $\mathbb{E}\left( \| X_{it} \|^{8+{\Greekmath 010F}} \big| \, {\cal C} \right)$ and $\mathbb{E} \| {\Greekmath 0115}^0_i \|^4$ and $\mathbb{E} \| f^0_t \|^{4+{\Greekmath 010F}}$ are bounded by a non-random constant, uniformly over $i,t$ and $N,T$. • ${\Greekmath 010C}^0$ is an interior point of the compact parameter set $\mathbb{B}$. \end{itemize}

Remarks on Assumption (ref)

itemize• Part $(i)$ of Assumption (ref) imposes that $e_{it}$ is a martingale difference sequence over time for a particular filtration. Conditioning on ${\cal C}$, the time series of $e_{it}$ is independent over time (part $(ii)$ of the assumption) and the error term $e_{it}$ and regressors $X_{it}$ are cross-sectionally independent (part $(iii)$ of the assumption), but unconditional correlation is allowed. Part $(iv)$ imposes weak time-serial correlation of $X_{it}$. Part $(v)$ demands weak time-serial correlation of $\widetilde X_{k,it} = X_{k,it} - \mathbb{E}\left[ X_{k,it} \big| {\cal C} \right]$ and $e_{it}$. Finally, parts $(vi)$ and $(vii)$ require bounded higher moments of the error term, regressors, factors and factor loadings, and a compact parameter set with an interior true parameter. • Assumption (ref)$(i)$ implies $\mathbb{E} \left(X_{k,it}e_{it} | \mathcal{C} \right) = 0$ and $\mathbb{E} \left(X_{k,it}e_{it} X_{\ell,is}e_{is} | \mathcal{C} \right) = 0$ for $t \neq s$. Thus, the assumption guarantees $X_{it}e_{it}$ is mean zero and uncorrelated over $t$, and independent across $i$, conditional on ${\cal C}$. Notice the conditional mean independence restriction in Assumption (ref)$(i)$ is weaker than Assumption D of Bai (2009), besides sequential exogeneity. Bai imposes independence between $e_{it}$ and $(\{ X_{js},{\Greekmath 0115}_j,f_s \}_{j,s})$. • Assumption (ref) is sufficient for Assumption (ref). To see this, notice $ {\rm Tr}(X_k \, e^{\prime}) = \sum_{i,t} X_{k,it}e_{it}$, and also that the sequential exogeneity and the cross-sectional independence assumption imply $\mathbb{E}\left[ \left( (NT)^{-1} \sum_{i,t} X_{k,it}e_{it} \right)^2 \Big| {\cal C} \right] =(NT)^{-2} \sum_{i,t} \mathbb{E}\left[ \left( X_{k,it}e_{it} \right)^2\Big| {\cal C} \right]$. Then, together with the assumption of bounded moments, we have $(NT)^{-1} \sum_{i,t} X_{k,it}e_{it} = o_p(1)$. • Assumption (ref) is also sufficient for Assumption (ref)$^*$ (and thus for Assumption (ref)), because $e_{it}$ is assumed independent over $t$ and across $i$ and has a bounded fourth moment, conditional on ${\cal C}$, which by using results in Latala Latala2006, implies the spectral norm satisfies $\| e \| = \sqrt{\max(N,T)}$ as $N$ and $T$ become large; see the supplementary material. • Examples of regressor processes, which satisfy Assumptions (ref)$(iv)$ and $(v)$, are discussed in the following. These examples also illuminate the role of the conditioning sigma field ${\cal C}$.

Examples of DGPs for $X_{it}$

Here we provide examples of the DGPs of the regressors $X_{it}$ that satisfy the conditions in Assumption (ref). Proofs for these examples are provided in the supplementary material.

exampleThe first example is a simple AR(1) interactive fixed effect regression: \[ Y_{it} = {\Greekmath 010C}^0 Y_{i,t-1} + {\Greekmath 0115}_{i}^{0\prime} f_{t}^0 + e_{it}, \] where $e_{it}$ is mean zero, independent across $i$ and $t$, and independent of ${\Greekmath 0115}^0$ and $f^0$. Assume $|{\Greekmath 010C}^0| < 1$ and that $e_{it}$, ${\Greekmath 0115}_{i}^{0}$, and $f_t^0$ all possess uniformly bounded moments of order $8+{\Greekmath 010F}$. In this case, the regressor is $X_{it} = Y_{it-1} = {\Greekmath 0115}_{i}^{0\prime} F_{t}^{0} + U_{it}$, where $F_{t}^{0} = \sum_{s=0}^{\infty} ({\Greekmath 010C}^{0})^{s} f_{t-1-s}^{0}$ and $U_{it} = \sum_{s=0}^{\infty} ({\Greekmath 010C}^{0})^{s} e_{i,t-1-s}$. For the conditioning sigma field $\mathcal{C}$ in Assumption (ref), we choose $\mathcal{C} = {\Greekmath 011B} \left( \{ {\Greekmath 0115}^0_{i} : 1 \leq i \leq N \}, \{ f^0_t : 1 \leq t \leq T \} \right) $. Conditional on ${\cal C}$, the only variation in $X_{it}$ stems from $U_{it}$, which is independent across $i$ and weakly correlated over $t$, so that Assumption (ref)$(iv)$ holds. Furthermore, we have $\mathbb{E} \left( X_{it} | \mathcal{C} \right) = {\Greekmath 0115}_{i}^{0\prime} F_{t}^{0}$ and $\widetilde X_{it} = U_{it}$, which allows us to verify Assumption (ref)$(v)$. This example can be generalized to a ${\rm VAR}(1)$ model as follows: \begin{align} {Y_{it} \choose Z_{it}} &= {\cal B} \, \underbrace{ {Y_{i,t-1} \choose Z_{i,t-1}} }_{=X_{it}} + {{\Greekmath 0115}^{0 \prime}_i f^{0}_t \choose d_{it}} + \underbrace{ {e_{it} \choose u_{it}} }_{=E_{it}} \; , \end{align} where $Z_{it}$ is an $m \times 1$ vector of additional variables and ${\cal B}$ is an $(m+1)\times(m+1)$ matrix of VAR parameters whose eigenvalues lie within the unit circle. The $m\times 1$ vector $d_{it}$ and the factors $f^0_t$ and factor loadings ${\Greekmath 0115}^0_i$ are assumed to be independent of the $(m+1) \times 1$ vector of innovations $E_{it}$. Suppose our interest is to estimate the first row in equation (ref), which corresponds exactly to our interactive fixed effects model with regressors $Y_{i,t-1}$ and $Z_{i,t-1}$. Choosing $\mathcal{C}$ to be the sigma field generated by all $f^0_t$, ${\Greekmath 0115}^0_i$, $d_{it}$, we obtain $\widetilde X_{it} = \sum_{s=0}^{\infty} {\cal B}^{s} E_{i,t-1-s}$. Analogous to the AR(1) case, we then find Assumption (ref)$(iv)$ and $(v)$ are satisfied in this example if the innovations $E_{it}$ are independent across $i$ and over $t$, and have appropriate bounded moments.
exampleConsider a scalar $X_{it}$ for simplicity, and let $X_{it}=g\left( v_{it}, {\Greekmath 010E}_{i},h_{t}\right)$. We assume (i) $\left\{ \left( e_{it},v_{it}\right) _{i=1,...,N;t=1,...,T}\right\} \bot \left\{ \left( {\Greekmath 0115}_{i}^{0},{\Greekmath 010E} _{i}\right)_{i=1,...,N},\left( f_{t}^{0},h_{t}\right)_{t=1,...,T}\right\} , $ (ii) $\left( e_{it},v_{it},{\Greekmath 010E}_{i}\right) $ are independent across $i$ for all $t,$ and (iii) $v_{is}\perp e_{it}$ for $s\leq t$ and all $i$. Furthermore, assume $\sup_{it}\mathbb{E} |X_{it}|^{8+{\Greekmath 010F}} < \infty$ for some positive ${\Greekmath 010F}$. For the conditioning sigma field $\mathcal{C}$ in Assumption (ref), we choose $\mathcal{C}={\Greekmath 011B} \left( \left\{ {\Greekmath 0115}_{i}^{0}:1\leq i\leq N\right\} , \relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ }\left\{ {\Greekmath 010E}_{i}:1\leq i\leq N\right\} ,\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ }\left\{ f_{t}^{0}: -\infty \leq t\leq \infty \right\} ,\relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ }\left\{ h_{t}: -\infty \leq t\leq \infty\right\} \right) .$ Furthermore, as in Hahn and Kuersteiner HahnKuersteiner2011, let $\mathcal{F}_{{\Greekmath 011C}}^{t}(i) = {\cal C} \vee {\Greekmath 011B} \left( \{ (e_{is},v_{is}) : {\Greekmath 011C} \leq s \leq t \} \right)$, and define the conditional ${\Greekmath 010B}$-mixing coefficient on $\mathcal{C}$: \[ {\Greekmath 010B}_{m}(i) = \sup_{A \in \mathcal{F}_{-\infty}^{t}(i), B \in \mathcal{F}_{t+m}^{\infty}(i)} \left[ \mathbb{P} \left(A \cap B \right) - \mathbb{P} \left(A \right) \mathbb{P} \left(B \right) | \mathcal{C} \right]. \] Let ${\Greekmath 010B}_m = \sup_{i} {\Greekmath 010B}_{m}(i)$, and assume $ {\Greekmath 010B}_m = O \left( m^{-{\Greekmath 0110}} \right), \relax\protect\ifmmode\expandafter\text@\else\expandafter\mbox\fi{ where } {\Greekmath 0110} > \frac{12 \, p}{4p-1}$ for $p>4$. Then, Assumptions (ref)$(iv)$ and $(v)$ are satisfied. In this example, the shocks $h_t$ (which may contain the factors $f^0_t$), ${\Greekmath 010E}_i$ (which may contain the factor loadings ${\Greekmath 0115}^0_i$), and $v_{it}$ (which may contain past values of $e_{it}$) can enter in a general non-linear way into the regressor $X_{it}$.

The following assumption guarantees the limiting variance and the asymptotic bias converge to constant values.

assumptionLet ${\cal X}_k=M_{{\Greekmath 0115}^0} \, X_{k} \, M_{f^0}$, which is an $N \times T$ matrix with entries ${\cal X}_{k,it}$. For each $i$ and $t$, define the $K$-vector ${\cal X}_{it}=({\cal X}_{1,it},\ldots,{\cal X}_{K,it})'$. We assume existence of the following probability limits for all $k=1,\ldots,K$: \begin{align*} W &= \limfunc{plim}_{N,T \rightarrow \infty} \frac 1 {NT} \, \sum_{i=1}^N \, \sum_{t=1}^T \, {\cal X}_{it} \, {\cal X}_{it}' \, , \nonumber \\ \Omega &= \limfunc{plim}_{N,T \rightarrow \infty} \frac 1 {NT} \, \sum_{i=1}^N \, \sum_{t=1}^T \, e_{it}^2 {\cal X}_{it} \, {\cal X}_{it}' \; , \nonumber \\ B_{1,k} &= \limfunc{plim}_{N,T \rightarrow \infty} \frac 1 {N} \, {\rm Tr} \left[ P_{f^0} \mathbb{E}\left( e' X_k \, \big| \, {\cal C} \right) \right] \; , \nonumber \\ B_{2,k} &= \limfunc{plim}_{N,T \rightarrow \infty} \frac 1 T {\rm Tr}\left[ \mathbb{E}\left( e e' \, \big| \, {\cal C} \right) \, M_{{\Greekmath 0115}^0} \, X_k \, f^0 \, (f^{0\prime}f^0)^{-1} \, ({\Greekmath 0115}^{0\prime}{\Greekmath 0115}^0)^{-1} \, {\Greekmath 0115}^{0\prime} \right] \; , \nonumber \\ B_{3,k} &= \limfunc{plim}_{N,T \rightarrow \infty} \frac 1 N {\rm Tr}\left[ \mathbb{E}\left( e' e \, \big| \, {\cal C} \right) \, M_{f^0} \, X^{\prime}_k \, {\Greekmath 0115}^0 \, ({\Greekmath 0115}^{0\prime}{\Greekmath 0115}^0)^{-1} \, (f^{0\prime}f^0)^{-1} \, f^{0\prime} \right] \; , \end{align*} where ${\cal C}$ is the same conditioning set that appears in Assumption (ref).

Here, $W$ and $\Omega$ are $K\times K$ matrices, and we define the $K$-vectors $B_1$, $B_2$, and $B_3$ with components $B_{1,k}$, $B_{2,k}$ and $B_{3,k}$, $k=1,\ldots,K$.

theorem[\bf Asymptotic Distribution] Let Assumptions (ref), (ref), (ref), and (ref) be satisfied,\footnote{ Assumption (ref) and (ref)$^*$ are implied by Assumption (ref) and therefore need not be explicitly assumed here.} and consider the limit $N,T \rightarrow \infty$ with $N/T \rightarrow {\Greekmath 0114}^2$, where $0<{\Greekmath 0114}<\infty$. Then we have \begin{align*} \sqrt{NT} \left( \widehat {\Greekmath 010C} - {\Greekmath 010C}^0 \right) \, \limfunc{\rightarrow}_d \, {\cal N} \left( W^{-1} B , \; W^{-1} \, \Omega \, W^{-1} \right) \; , \end{align*} where $B=-{\Greekmath 0114} B_1-{\Greekmath 0114}^{-1} B_2 - {\Greekmath 0114} B_3$.

From Corollary (ref), we already know the limiting distribution of $\widehat {\Greekmath 010C}$ is given by the limiting distribution of $W_{NT}^{-1} C_{NT}$. Note $W_{NT} = \frac 1 {NT} \, \sum_{i=1}^N \, \sum_{t=1}^T \, {\cal X}_{it} \, {\cal X}_{it}'$; that is, $W$ is simply defined as the probability limit of $W_{NT}$. Assumption (ref) guarantees $W$ is positive definite, as shown in the supplementary material.

Thus, the main task in proving Theorem (ref) is to show the approximated score at the true parameter satisfies $C_{NT} \limfunc{\rightarrow}_d \, {\cal N} \left( B , \Omega \right)$. We find the asymptotic variance $\Omega$ and the asymptotic bias $B_1$ originate from the $C^{(1)}$ term, whereas the two further bias terms $B_2$ and $B_3$ originate from the $C^{(2)}$ term of $C_{NT}$.

The bias $B_1$ is due to correlation of the errors $e_{it}$ and the regressors $X_{k,i{\Greekmath 011C}}$ in the time direction (for ${\Greekmath 011C}>t$). This bias term generalizes the Nickell Nickell1981 bias that occurs in dynamic models with standard fixed effects, and it is not present in Bai Bai2009, where only strictly exogenous regressors are considered.

The other two bias terms $B_2$ and $B_3$ are already described in Bai Bai2009. If $e_{it}$ is homoscedastic, that is, if $\mathbb{E}(e_{it} | {\cal C}) = {\Greekmath 011B}^2$, then $\mathbb{E}\left( e e' | {\cal C} \right) = {\Greekmath 011B}^2 \mathbb{I}_N$ and $\mathbb{E}\left( e' e | {\cal C} \right) = {\Greekmath 011B}^2 \mathbb{I}_T$, so that $B_2 = 0$ and $B_3=0$ (because the trace is cyclical and $ f^{0\prime} M_{f^0} =0$ and ${\Greekmath 0115}^{0\prime} M_{{\Greekmath 0115}^0} = 0$). Thus, $B_2$ is only non-zero if $e_{it}$ is heteroscedastic across $i$, and $B_3$ is only non-zero if $e_{it}$ is heteroscedastic over $t$. Correlation in $e_{it}$ across $i$ or over $t$ would also generate non-zero bias terms of exactly the form $B_2$ and $B_3$, but is ruled out by our assumptions.

Bias Correction

Estimators for $W$, $\Omega$, $B_1$, $B_2$, and $B_3$ are obtained by forming suitable sample analogs and replacing the unobserved ${\Greekmath 0115}^0$, $f^0$, and $e$ by the estimates $\widehat {\Greekmath 0115}$, $\widehat f$, and the residuals $\widehat e$.

definitionLet $\widehat{{\cal X}}_k=M_{\widehat {\Greekmath 0115}} \, X_{k} \, M_{\widehat f}$. For each $i$ and $t$, define the $K$-vector $\widehat{{\cal X}}_{it}=(\widehat{{\cal X}}_{1,it},\ldots,\widehat{{\cal X}}_{K,it})'$. Let $\Gamma: \mathbb{R} \rightarrow \mathbb{R}$ be the truncation kernel defined by $\Gamma(x)=1$ for $|x|\leq 1$, and $\Gamma(x)=0$ otherwise. Let $M$ be a bandwidth parameter that depends on $N$ and $T$. We define the $K\times K$ matrices $\widehat W$ and $\widehat \Omega$, and the $K$-vectors $\widehat B_{1}$, $\widehat B_{2}$, and $\widehat B_{3}$ as follows: \begin{align*} \widehat W &= \frac 1 {NT} \, \sum_{i=1}^N \, \sum_{t=1}^T \, \widehat{{\cal X}}_{it} \, \widehat{{\cal X}}_{it}' \; , \nonumber \\ \widehat \Omega &= \frac 1 {NT} \, \sum_{i=1}^N \, \sum_{t=1}^T \, (\widehat e_{it})^2 \, \widehat{{\cal X}}_{it} \, \widehat{{\cal X}}_{it}' \; , \nonumber \\ \widehat B_{1,k} &= \frac 1 N \, \sum_{i=1}^N \sum_{t=1}^{T-1} \sum_{s=t+1}^T \Gamma\left( \frac{s-t} M \right) \, \big[ P_{\widehat f} \big]_{t s} \; \widehat e_{it} \; X_{k,is} \; , \nonumber \\ \widehat B_{2,k} &= \frac 1 T \sum_{i=1}^N \sum_{t=1}^T (\widehat e_{it})^2 \left[ M_{\widehat {\Greekmath 0115}} \, X_k \, \widehat f \, (\widehat f^{\prime} \widehat f)^{-1} \, (\widehat {\Greekmath 0115}^{\prime}\widehat {\Greekmath 0115})^{-1} \, \widehat {\Greekmath 0115}^{\prime} \right]_{ii} \; , \nonumber \\ \widehat B_{3,k} &= \frac 1 N \sum_{i=1}^N \sum_{t=1}^T (\widehat e_{it})^2 \left[ M_{\widehat f} \, X^{\prime}_k \, \widehat {\Greekmath 0115} \, (\widehat {\Greekmath 0115}^{\prime} \widehat {\Greekmath 0115})^{-1} \, (\widehat f^{\prime}\widehat f)^{-1} \, \widehat f^{\prime} \right]_{tt} \; , \end{align*} where $\widehat e = Y \, - \, \widehat {\Greekmath 010C} \cdot X \, - \, \widehat {\Greekmath 0115} \, \widehat f'$, and $\widehat e_{it}$ denotes the elements of $\widehat e$, $\left[ A \right]_{ts}$ denotes the (t,s)th element of the matrix $A$.

Notice the estimators $\widehat\Omega$, $\widehat B_2$, and $\widehat B_3$ are similar to White's standard error estimator under heteroskedasticity, and the estimator $\widehat B_1$ is similar to the HAC estimator with a kernel. To show consistency of these estimators, we impose some additional assumptions.

samepage\begin{assumption} $\phantom{a}$ \begin{itemize} • $\| {\Greekmath 0115}^0_i \|$ and $\|f^0_t\|$ are uniformly bounded over $i$, $t$, and $N$, $T$. • There exist $c>0$ and ${\Greekmath 010F} >0$ such that for all $i,t,m,N$, and $T$, we have \\ $ \left| \frac 1 N \sum_{i=1}^N \mathbb{E} (e_{it} X_{k,it+m} \, \big| \, {\cal C}) \right| \leq c \, m^{- (1+ {\Greekmath 010F})}$. \end{itemize} \end{assumption}

Assumption (ref)$(i)$ is made for convenience to simplify the consistency proof for the estimators in Definition (ref). Weakening this assumption is possible by only assuming suitable bounded moments of $\| {\Greekmath 0115}^0_i \|$ and $\|f^0_t\|$. To show consistency of $\widehat B_1$, we need to control how strongly $e_{it}$ and $X_{k,i {\Greekmath 011C}}$, $t<{\Greekmath 011C}$, are allowed to be correlated, which is done by Assumption (ref)$(ii)$. It is straightforward to verify Assumption (ref)$(ii)$ is satisfied in the two examples of regressor processes presented after Assumption (ref).

theorem[\bf Consistency of Bias and Variance Estimators] Let Assumptions (ref), (ref), (ref), (ref), and (ref) hold, and consider a limit $N,T \rightarrow \infty$ with $N/T \rightarrow {\Greekmath 0114}^2$, $0<{\Greekmath 0114}<\infty$, such that the bandwidth $M=M_{NT}$ satisifies $M \rightarrow \infty$ and $M^5 / T \rightarrow 0$. We then have $\widehat W=W + o_p(1)$, $\widehat \Omega=\Omega + o_p(1)$, $\widehat B_1=B_1 + o_p(1)$, $\widehat B_2=B_{2} + o_p(1)$, and $\widehat B_3=B_{3} + o_p(1)$.

The assumption $M^5/T \rightarrow 0$ can be relaxed if additional higher- moment restrictions on $e_{it}$ and $X_{k,it}$ are imposed. Note also that for the construction of the estimators $\widehat W$, $\widehat \Omega$, and $\widehat B_i$, $i=1,2,3$, knowing whether the regressors are strictly exogenous or predetermined is unnecessary; in both cases, the estimators for $W$, $\Omega$, and $B_i$, $i=1,2,3$, are consistent. We can now present our bias-corrected estimator and its limiting distribution.

corollaryUnder the assumptions of Theorem (ref), the bias-corrected estimator \begin{align*} \widehat {\Greekmath 010C}^* &= \widehat {\Greekmath 010C} + \widehat W^{-1} \left( T^{-1} \widehat B_1 + N^{-1} \widehat B_2 + T^{-1} \widehat B_3 \right) \end{align*} satisfies $\sqrt{NT} \left( \widehat {\Greekmath 010C}^* - {\Greekmath 010C}^0 \right) \, \limfunc{\rightarrow}_d \, {\cal N} \left(0 , \; W^{-1} \, \Omega \, W^{-1} \right)$.

According to Theorem (ref), a consistent estimator of the asymptotic variance of $\widehat {\Greekmath 010C}^*$ is given by $\widehat W^{-1} \, \widehat \Omega \, \widehat W^{-1}$.

An alternative to the analytical bias-correction result given by Corollary (ref) is to use Jackknife bias correction to eliminate the asymptotic bias. For panel models with incidental parameters only in the cross-sectional dimensions, one typical finds a large $N,T$ leading incidental parameter bias of order $1/T$ for the parameters of interest. To correct for this $1/T$ bias, one can use the delete-one Jackknife bias correction if observations are iid over $t$ HahnNewey2004 and the split-panel Jackknife bias-correction if observations are correlated over $t$ DhaeneJochmans2015. In our current model, we have incidental parameters in both panel dimensions (${\Greekmath 0115}^0_i$ and $f^0_t$), resulting in leading bias terms of order $1/T$ (bias term $B_1$ and $B_3$) and of order $1/N$ (bias term $B_2$). Fern{\'a}ndez-Val and Weidner FernandezValWeidner2013 discuss the generalizations of the split-panel Jackknife bias-correction to that case.

The corresponding bias-corrected split-panel Jackknife estimator reads $\widehat {\Greekmath 010C}^J = 3 \widehat {\Greekmath 010C}_{NT} - \overline {\Greekmath 010C}_{N,T/2} - \overline {\Greekmath 010C}_{N/2,T}$, where $\widehat {\Greekmath 010C}_{NT} = \widehat {\Greekmath 010C}$ is the LS estimator obtained from the full sample, $\overline {\Greekmath 010C}_{N,T/2}$ is the average of the two LS estimators that leave out the first and second halves of the time periods, and $\overline {\Greekmath 010C}_{N/2,T}$ is the average of the two LS estimators that leave out half of the individuals. Jackknife bias correction is convenient because only the order of the bias, and not the structure of the terms $B_1$, $B_2$, and $B_3$, needs not be known in detail. However, one requires additional stationarity assumptions over $t$ and homogeneity assumptions across $i$ to justify the Jackknife correction and to show that $\widehat {\Greekmath 010C}^J$ has the same limiting distribution as $\widehat {\Greekmath 010C}^*$ in Corollary (ref); see Fern{\'a}ndez-Val and Weidner FernandezValWeidner2013 for more details. They also observe through Monte Carlo simulations that the finite sample variance of the Jackknife-corrected estimator is often larger than of the analytically corrected estimator. We do not explore Jackknife bias-correction further in this paper.

Testing Restrictions on ${\Greekmath 010C}^0$

In this section, we discuss the three classical test statistics for testing linear restrictions on ${\Greekmath 010C}^0$. The null hypothesis is $H_0: \; H {\Greekmath 010C}^0 = h$, and the alternative is $H_a: \; H {\Greekmath 010C}^0 \neq h$, where $H$ is an $r \times K$ matrix of rank $r \leq K$, and $h$ is an $r \times 1$ vector. We restrict the presentation to testing a linear hypothesis for ease of exposition. One can generalize the discussion to the testing of non-linear hypotheses, under conventional regularity conditions. Throughout this subsection, we assume ${\Greekmath 010C}^0$ is an interior point of $\mathbb{B}$; that is, no local restrictions are on ${\Greekmath 010C}$ as long as the null hypothesis is not imposed. Using the expansion of $L_{NT}({\Greekmath 010C})$, one could also discuss testing when the true parameter is on the boundary, as shown in Andrews Andrews2001.

The restricted estimator is defined by

align[align omitted — 177 chars of source]

where $\widetilde{\mathbb{B}} = \{ {\Greekmath 010C} \in \mathbb{B}| \, H{\Greekmath 010C} = h \}$ is the restricted parameter set. Analogous to Theorem (ref) for the unrestricted estimator $\widehat {\Greekmath 010C}$, we can use the expansion of the profile objective function to derive the limiting distribution of the restricted estimator. Under the assumptions of Theorem (ref), we have

align*[align* omitted — 233 chars of source]

where $\mathfrak{W}^{-1} = W^{-1} - W^{-1} H' (H W^{-1} H')^{-1} H W^{-1}$. The $K\times K$ covariance matrix in the limiting distribution of $\widetilde {\Greekmath 010C}$ is not full rank, but satisfies ${\rm rank}(\mathfrak{W}^{-1} \, \Omega \, \mathfrak{W}^{-1})=K-r$, because $H \mathfrak{W}^{-1}=0$ and thus ${\rm rank}(\mathfrak{W}^{-1})=K-r$. The asymptotic distribution of $\sqrt{NT}(\widetilde {\Greekmath 010C} - {\Greekmath 010C}^0)$ is therefore $K-r$ dimensional, as it should be for the restricted estimator.

Wald Test

Using the result of Theorem (ref), we find that under the null hypothesis, $\sqrt{NT} \left( H \widehat {\Greekmath 010C} - h \right)$ is asymptotically distributed as ${\cal N} \left( H W^{-1} B , \; H W^{-1} \, \Omega \, W^{-1} H' \right)$. Thus, due to the presence of the bias $B$, the standard Wald test statistic $WD_{NT} = NT \left( H \widehat {\Greekmath 010C} - h \right)' \left( H \widehat W^{-1} \, \widehat \Omega \, \widehat W^{-1} H' \right)^{-1} \left( H \widehat {\Greekmath 010C} - h \right)$ is not asymptotically ${\Greekmath 011F}^2_r$ distributed. Using the estimator $\widehat B = - \sqrt{\frac N T} \, \widehat B_{1} - \sqrt{\frac T N} \, \widehat B_{2} - \sqrt{\frac N T} \, \widehat B_{3}$ for the bias, we can define the bias-corrected Wald test statistic as

align[align omitted — 319 chars of source]

where $\widehat {\Greekmath 010C}^* = \widehat {\Greekmath 010C} - \widehat W^{-1} \widehat B$ is the bias-corrected estimator. $WD^*_{NT}$ is just the standard Wald test statistics applied to $\widehat {\Greekmath 010C}^*$. Under the null hypothesis and the Assumptions of Theorem (ref), we find $WD^*_{NT} \, \limfunc{\rightarrow}_d \, {\Greekmath 011F}^2_r$.

Likelihood Ratio Test

To implement the LR test, we need the relationship between the asymptotic Hessian $W$ and the asymptotic score variance $\Omega$ of the profile objective function to be of the form $\Omega=c W$, where $c>0$ is a scalar constant. This condition is satisfied in our interactive fixed effect model if $\mathbb{E}(e_{it}^2 | {\cal C}) = c$, that is, if the error is homoskedastic. A consistent estimator for $c$ is then given by $\widehat c = (NT)^{-1} \sum_{i=1}^N \sum_{t=1}^T \widehat e_{it}^2$, where $\widehat e = Y \, - \, \widehat {\Greekmath 010C} \cdot X \, - \, \widehat {\Greekmath 0115} \, \widehat f'$. Because the likelihood function for the interactive fixed effect model is just the sum of squared residuals, we have $\widehat c = L_{NT}( \widehat {\Greekmath 010C} )$. The likelihood ratio test statistic is defined by

align*[align* omitted — 166 chars of source]

Under the assumption of Theorem (ref), we then have

align*[align* omitted — 130 chars of source]

where $C \sim {\cal N}(B,\Omega)$, i.e. $C_{NT} \limfunc{\rightarrow}_d C$. It is the same limiting distribution that one finds for the Wald test if $\Omega=c W$ (in fact, one can show $WD_{NT}=LR_{NT}+o_p(1)$). Therefore, we need to do a bias-correction for the LR test to achieve a ${\Greekmath 011F}^2$ limiting distribution. We define

align[align omitted — 376 chars of source]

where $\widehat B$ and $\widehat W$ do not depend on the parameter ${\Greekmath 010C}$ in the minimization problem.\footnote{ Alternatively, one could use $\widehat B(\widetilde {\Greekmath 010C})$ and $\widehat W(\widetilde {\Greekmath 010C})$ as estimates for $B$ and $W$, and would obtain the same limiting distribution of $LR_{NT}^*$ under the null hypothesis $H_0$. These alternative estimators are not consistent if $H_0$ is false, i.e. the power-properties of the test would be different. The question of which specification should be preferred is left for future research. } Asymptotically, we have $\min_{{\Greekmath 010C} \in \mathbb{B}} L_{NT}\left( {\Greekmath 010C} + (NT)^{-1/2} \widehat W^{-1} \widehat B \right) = L_{NT}(\widehat {\Greekmath 010C})$, because ${\Greekmath 010C} \in \mathbb{B}$ does not impose local constraints; in other words, close to ${\Greekmath 010C}^0$, whether one minimizes over ${\Greekmath 010C}$ or over ${\Greekmath 010C} + (NT)^{-1/2} \widehat W^{-1} \widehat B$ does not matter for the value of the minimum. The correction to the LR test therefore originates from the first term in $LR_{NT}^{*}$. For the minimization over the restricted parameter set, whether the argument of $L_{NT}$ is ${\Greekmath 010C}$ or ${\Greekmath 010C} + (NT)^{-1/2} \widehat W^{-1} \widehat B$ matters, because generically, we have $H W^{-1} B \neq 0$ (otherwise, no correction would be necessary for the LR statistics). One can show that

align*[align* omitted — 135 chars of source]

that is, we obtain the same formula as for $LR_{NT}$, but the bias-corrected term $C-B$ replaces the limit of the score $C$. Under the Assumptions of Theorem (ref), if $H_0$ is satisfied, and for homoscedastic errors $e_{it}$, we have $LR^*_{NT} \, \limfunc{\rightarrow}_d \, {\Greekmath 011F}^2_r$. In fact, one can show $LR^*_{NT} = WD^*_{NT}+o_p(1)$.

Lagrange Multiplier Test

Let $\widetilde {\Greekmath 0272} {\cal L}_{NT}$ be the gradient of the LS objective function (ref) with respect to ${\Greekmath 010C}$, evaluated at the restricted parameter estimates; that is,

align*[align* omitted — 824 chars of source]

where $\widetilde {\Greekmath 0115} = \widehat {\Greekmath 0115}(\widetilde {\Greekmath 010C})$, $\widetilde f = \widehat f(\widetilde {\Greekmath 010C})$, and $\widetilde e = Y \, - \, \widetilde {\Greekmath 010C} \cdot X \, - \, \widetilde {\Greekmath 0115} \, \widetilde f'$. Under the assumptions of Theorem (ref), and if the null hypothesis $H_0: \; H {\Greekmath 010C}^0 = h$ is satisfied, one finds that\footnote{The proof of the statement is given in the supplementary material as part of the proof of Theorem (ref).}

align[align omitted — 194 chars of source]

Due to this equation, one can base the Lagrange multiplier test on the gradient of ${\cal L}_{NT}(\widetilde {\Greekmath 010C},\widetilde {\Greekmath 0115},\widetilde f)$, or on the gradient of the profile quasi-likelihood function $L_{NT}(\widetilde {\Greekmath 010C})$, and obtain the same limiting distribution.

Using the bound on the remainder $R_{NT}({\Greekmath 010C})$ given in Theorem (ref), one cannot infer any properties of the score function, that is, of the gradient ${\Greekmath 0272} L_{NT}({\Greekmath 010C})$, because nothing is said about ${\Greekmath 0272} R_{NT}({\Greekmath 010C})$. The following theorem gives a bound on ${\Greekmath 0272} R_{NT}({\Greekmath 010C})$ that is sufficient to derive the limiting distribution of the Lagrange multiplier.

theoremUnder the assumptions of Theorem (ref), and with $W_{NT}$ and $C_{NT}$ as defined there, the score function satisfies \begin{align*} {\Greekmath 0272} L_{NT}({\Greekmath 010C}) \, &= \, 2 \, W_{NT} \, ({\Greekmath 010C}-{\Greekmath 010C}^0) \, - \, \frac 2 {\sqrt{NT}} \, C_{NT} \, + \frac 1 {NT} \, {\Greekmath 0272} R_{NT}({\Greekmath 010C}) \; , \end{align*} where the remainder ${\Greekmath 0272} R_{NT}({\Greekmath 010C})$ satisfies for any sequence ${\Greekmath 0111}_{NT}\rightarrow 0$: \begin{align*} \sup_{\{{\Greekmath 010C} :\left\| {\Greekmath 010C} -{\Greekmath 010C}^{0} \right\| \leq {\Greekmath 0111}_{NT}\}} \frac{ \left\| {\Greekmath 0272} R_{NT}({\Greekmath 010C}) \right\| } { \sqrt{NT} \left(1 + \sqrt{NT} \, \left\| {\Greekmath 010C} -{\Greekmath 010C}^{0} \right\| \right) } = o_{p}\left( 1 \right) . \end{align*}

From this theorem, and the fact that $\widetilde {\Greekmath 010C}$ is $\sqrt{NT}$-consistent under $H_0$, we obtain

align*[align* omitted — 321 chars of source]

Remember $\widetilde {\Greekmath 010C}$ is the restricted estimator defined in equation (ref). Using this result and the known limiting distribution of $\widetilde {\Greekmath 010C}$, we now find

align[align omitted — 187 chars of source]

Note also that $\sqrt{NT} H W^{-1} {\Greekmath 0272} L_{NT}(\widetilde {\Greekmath 010C}) \, \limfunc{\rightarrow}_d \, - 2 H W^{-1} C$. We define $\widetilde B$, $\widetilde W$, and $\widetilde \Omega$, analogous to $\widehat B$, $\widehat W$, and $\widehat \Omega$, but with unrestricted parameter estimates replaced by restricted parameter estimates. The LM test statistic is then given by

align*[align* omitted — 257 chars of source]

One can show the LM test is asymptotically equivalent to the Wald test: $LM_{NT}=WD_{NT}+o_p(1)$; that is, again, bias-correction is necessary. We define the bias-corrected LM test statistic as

align[align omitted — 388 chars of source]

The following theorem summarizes the main results of the present subsection.

theorem[\bf Chi-Square Limit of Bias-Corrected Test Statistics] Let the assumptions of Theorem (ref) and the null hypothesis $H_0: \; H {\Greekmath 010C}^0 = h$ be satisfied. For the bias-corrected Wald and LM test statistics introduced in equation (ref) and (ref), we then have \begin{align*} WD^*_{NT} \; \; &\limfunc{\longrightarrow}_d \; \; {\Greekmath 011F}^2_r \; , & LM^*_{NT} \; \; &\limfunc{\longrightarrow}_d \; \; {\Greekmath 011F}^2_r \; . \end{align*} If, in addition, we assume $\mathbb{E}(e_{it}^2 | {\cal C}) = c$, that is, the idiosyncratic errors are homoscedastic, and we use $\widehat c = L_{NT}(\widehat {\Greekmath 010C})$ as an estimator for $c$, the LR test statistic defined in equation (ref) satisfies \begin{align*} LR^{*}_{NT} \; \; &\limfunc{\longrightarrow}_d \; \; {\Greekmath 011F}^2_r \; . \end{align*}

Extension to Endogenous Regressors

In this section, we briefly discuss how to estimate the regression coefficient ${\Greekmath 010C}^0$ of Model (ref) when some of the regressors in $X_{it}$ are endogenous with respect to the regression error $e_{it}$. The question is how instrumental variables can be used to estimate the regression coefficients of the endogenous regressor in the presence of the interactive fixed effects ${\Greekmath 0115}_{i}^{0\prime }f_{t}^{0}$.

The existing literature has already investigated similar questions under various setups. Harding and Lamarche HardingLamarche2009,HardingLamarche2011 investigate the problem of estimating an endogenous panel (quantile) regression with interactive fixed effects, and show how to use IVs in the CCE estimation framework. Moon, Shum, and Weidner MoonShumWeidner2012 (hereafter MSW) estimate a random coefficient multinomial demand model (as in Berry, Levinsohn, and Pakes BerryLevinsohnPakes1995) when the unobserved product-market characteristics have interactive fixed effects. The IVs are required to identify the parameters of the random coefficient distribution and to control for price endogeneity. They suggested a multi-step “least squares-minimum distance” (LS-MD) estimator.\footnote{ Chernazhukov and Hansen ChernozhukovHansen2005 also used a similar method for estimating endogenous quantile regression models.} The LS-MD approach is also applicable to linear panel regression models with endogenous regressors and interactive fixed effects, as demonstrated in Lee, Moon, and Weidner LeeMoonWeidner2012 for the case of a dynamic linear panel regression model with interactive fixed effects and measurement error.

We now discuss how to implement the LS-MD estimation in our setup. Let $X_{it}^{\rm end}$ be the vectors of endogenous regressors, and let $X_{it}^{\rm exo}$ be the vector of exogenous regressors, with respect to $e_{it}$, such that $X_{it}= (X_{it}^{\rm end \prime},X_{it}^{\rm exo \prime})'$. The model then reads \[ Y_{it}={\Greekmath 010C}_{\rm end}^{0\prime }X_{it}^{\rm end}+{\Greekmath 010C}_{\rm exo}^{0\prime }X_{it}^{\rm exo}+{\Greekmath 0115}_{i}^{0\prime }f_{t}^{0}+e_{it}, \] where $X_{it}^{\rm exo}$ denotes the exogenous and $X_{it}^{\rm end}$ denotes the endogenous regressors (wrt to $e_{it}$). Suppose $Z_{it}$ is an additional $L$-vector of exogenous instrumental variables (IVs), but $Z_{it}$ may be correlated with ${\Greekmath 0115}^0_i$ and $f^0_t$. The LS-MD estimator of ${\Greekmath 010C} ^{0}=\left( {\Greekmath 010C}_{\rm end}^{0\prime },{\Greekmath 010C} _{\rm exo}^{0\prime }\right) ^{\prime }$ can then be calculated by the following three steps:

itemize• For given ${\Greekmath 010C}_{\rm end}$, we run the least squares regression of $ Y_{it}-{\Greekmath 010C}_{\rm end}^{\prime }X_{it}^{\rm end}$ on the included exogeneous regressors $X_{it}^{\rm exo}$, the interactive fixed effects ${\Greekmath 0115} _{i}^{\prime }f_{t}$, and the IVs $Z_{it}:$ \begin{align*} & \left( \widetilde{{\Greekmath 010C}}_{\rm exo}\left( {\Greekmath 010C}_{\rm end}\right) ,\widetilde{{\Greekmath 010D}}\left( {\Greekmath 010C}_{\rm end}\right) ,\widetilde{{\Greekmath 0115}}\left( {\Greekmath 010C}_{\rm end}\right) ,\widetilde{f} \left( {\Greekmath 010C}_{\rm end}\right) \right) \\ & \qquad = \; \limfunc{argmin}_{\{ {\Greekmath 010C}_{\rm exo},{\Greekmath 010D} ,{\Greekmath 0115} ,f \} } \; \sum_{i=1}^{N}\sum_{t=1}^{T}\left( Y_{it}-{\Greekmath 010C}_{\rm end}^{\prime }X_{it}^{\rm end}-{\Greekmath 010C}_{\rm exo}^{\prime }X_{it}^{\rm exo}-{\Greekmath 010D} ^{\prime }Z_{it}-{\Greekmath 0115}_{i}^{\prime }f_{t}\right) ^{2}. \end{align*} • We estimate ${\Greekmath 010C}_{\rm end}$ by finding $\widetilde{{\Greekmath 010D}}\left( {\Greekmath 010C} _{\rm end}\right) $, obtained by step (1), that is closest to zero. To do so, we choose a symmetric positive definite $L \times L$ weight matrix $W_{NT}^{{\Greekmath 010D} }$ and compute \[ \widehat{{\Greekmath 010C}}_{\rm end}=\limfunc{argmin}_{{\Greekmath 010C}_{\rm end}}\widetilde{{\Greekmath 010D}}\left( {\Greekmath 010C} _{\rm end}\right) ^{\prime } \, W_{NT}^{{\Greekmath 010D} } \, \widetilde{{\Greekmath 010D}}\left( {\Greekmath 010C} _{\rm end}\right) . \] • We estimate ${\Greekmath 010C}_{\rm exo}$ (and ${\Greekmath 0115} ,f)$ by running the least squares regression of $Y_{it}-\widehat{{\Greekmath 010C}}_{\rm end}^{\prime }X_{it}^{\rm end}$ on the included exogeneous regressors $X_{it}^{\rm exo}$ and the interactive fixed effects ${\Greekmath 0115}_{i}^{\prime }f_{t}$: \begin{align*} \left( \widehat{{\Greekmath 010C}}_{\rm exo} ,\widehat {\Greekmath 0115} ,\widehat{f} \right) &= \; \limfunc{argmin}_{\{ {\Greekmath 010C}_{\rm exo},{\Greekmath 010D} ,{\Greekmath 0115} ,f \} } \; \sum_{i=1}^{N}\sum_{t=1}^{T}\left( Y_{it}- \widehat {\Greekmath 010C}_{\rm end}^{\prime }X_{it}^{\rm end}-{\Greekmath 010C}_{\rm exo}^{\prime }X_{it}^{\rm exo} -{\Greekmath 0115}_{i}^{\prime }f_{t}\right) ^{2}. \end{align*}

The idea behind this estimation procedure is that valid instruments are excluded from the model for $Y_{it}$, so that their first-step regression coefficients $\widetilde{{\Greekmath 010D}}\left( {\Greekmath 010C}_{\rm end}\right)$ should be close to zero if ${\Greekmath 010C}_{\rm end}$ is close to its true value ${\Greekmath 010C}_{\rm end}^0$. Thus, as long as $X_{it}^{\rm exo}$ and $Z_{it}$ jointly satisfy the assumptions of the current paper, we obtain $\widetilde{{\Greekmath 010D}}\left( {\Greekmath 010C}_{\rm end}^0 \right) = o_p(1)$ for the first-step LS estimator, and we also obtain the asymptotic distribution of $\widetilde{{\Greekmath 010D}}\left( {\Greekmath 010C}_{\rm end}^0 \right)$ from the results derived in section (ref).

However, to justify the second-step minimization formally, one needs to study the properties of $\widetilde{{\Greekmath 010D}}\left( {\Greekmath 010C}_{\rm end} \right)$ also for $ {\Greekmath 010C}_{\rm end} \neq {\Greekmath 010C}_{\rm end}^0$. To do so, we refer to MSW. Our ${\Greekmath 010C}_{\rm end},{\Greekmath 010C}_{\rm exo},$ and $Y_{it}-{\Greekmath 010C}_{\rm end}^{\prime }X_{it}^{\rm end}$ correspond to their ${\Greekmath 010B}$, ${\Greekmath 010C}$, and ${\Greekmath 010E}_{jt}\left( {\Greekmath 010B} \right)$, respectively. Assumptions 1 - 5 in MSW can be translated accordingly, and the results in MSW show large $N,T$ consistency and asymptotic normality of the LS-MD estimator.

The final step of the LS-MD estimation procedure is essentially a repetition of the first step, but without including $Z_{it}$ in the set of regressors, which results in some efficiency gains for $\widehat{{\Greekmath 010C}}_{\rm exo} $ compared to the first step.

Monte Carlo Simulations

We consider an ${\rm AR}(1)$ model with $R=1$ factors:

align*[align* omitted — 128 chars of source]

We estimate the model as an interactive fixed effect model; that is, no distributional assumptions on ${\Greekmath 0115}^0_{i}$ and $f^0_{t}$ are made in estimation. The parameter of interest is ${\Greekmath 011A}^0$. The estimators we consider are the OLS estimator (which completely ignores the presence of the factors), the least squares estimator with interactive fixed effects (denoted FLS in this section to differentiate from OLS) defined in equation (ref),\footnote{ Here we can either use $\mathbb{B}=(-1,1)$, or $\mathbb{B}=\mathbb{R}$. In the present model, we only have high-rank regressors; i.e., the parameter space need not be bounded to show consistency.} and its bias-corrected version (denoted BC-FLS), defined in Theorem (ref).

For the simulation, we draw the $e_{it}$ independently and identically distributed from a t-distribution with five degrees of freedom, the ${\Greekmath 0115}^0_{i}$ independently distributed from ${\cal N}(1,1)$, and we generate the factors from an ${\rm AR}(1)$ specification, namely, $f^0_{t}={\Greekmath 011A}_f \, f^0_{t-1} + u_{t}$, where $u_{t} \sim {\rm iid} {\cal N}(0,(1-{\Greekmath 011A}_f^2){\Greekmath 011B}_f^2)$, and ${\Greekmath 011B}_f$ is the standard deviation of $f^0_t$. For all simulations, we generate 1,000 initial time periods for $f^0_t$ and $Y_{it}$ that are not used for estimation. This approach guarantees the simulated data used for estimation are distributed according to the stationary distribution of the model.

This setup contains no correlation and heteroscedasticity in $e_{it}$; that is, only the bias term $B_1$ of the FLS estimator is non-zero, but we ignore this information in the estimation; that is, we correct for all three bias terms ($B_1$, $B_2$, and $B_3$, as introduced in Assumption (ref)) in the bias-corrected FLS estimator.

Table (ref) shows the simulation results for the bias, standard error, and root mean square error of the three different estimators for the case $N=100$, ${\Greekmath 011A}_f=0.5$, and ${\Greekmath 011B}_f=0.5$, and different values of ${\Greekmath 011A}^0$ and $T$. The OLS estimator, the FLS estimator (computed with correct $R=1$), and the corresponding bias-corrected FLS estimator with factors (BC-FLS) were computed for 10,000 simulation runs. The table lists the mean bias, the standard deviation (std), and the square root of the mean square error (rmse) for the three estimators. As expected, the OLS estimator is biased because of the factor structure and its bias does not vanish (it actually increases) as $T$ increases. The FLS estimator is also biased, but as theory predicts its bias vanishes as $T$ increases. The bias-corrected FLS estimator performs better than the non-corrected FLS estimator, in particular, its bias vanishes faster. Because we only correct for the first-order bias of the FLS estimator, we could not expect the bias-corrected FLS estimator to be unbiased. However, as $T$ gets larger, more and more of the FLS estimator bias is corrected for; for example, for ${\Greekmath 011A}^0=0.3$, we find that at $T=5$, the bias correction only corrects for about half of the bias, whereas at $T=80$, it already corrects for about $90 \%$ of it.

Table (ref) is similar to Table (ref), with the only difference being that we allow for misspecification in the number of factors $R$, namely, the true number of factors is assumed to be $R=1$ (i.e., same DGP as for Table (ref)), but we incorrectly use $R=2$ factors when calculating the FLS and BC-FLS estimator. By comparing Table (ref) with Table (ref), we find this type of misspecification of the number of factors increases the bias and the standard deviation of both the FLS and the BC-FLS estimator in finite samples. That increase, however, is comparatively small once both $N$ and $T$ are large. According to the results in Moon and Weidner MoonWeidner2015, we expect the limiting distribution of the correctly specified ($R=1$) and incorrectly specified ($R=2$) FLS estimator to be identical when $N$ and $T$ grow at the same rate. Our simulations suggest the same is true for the BC-FLS estimator. The remaining simulation all assume correctly specified $R=1$.

An import issue is the choice of bandwidth $M$ for the bias correction. Table (ref) gives the fraction of the FLS estimator bias that is captured by the estimator for the bias in a model with $N=100$, $T=20$, ${\Greekmath 011A}_f=0.5$, ${\Greekmath 011B}_f=0.5$ and different values for ${\Greekmath 011A}$ and $M$. The table shows the optimal bandwidth (in the sense that most of the bias is corrected for) depends on ${\Greekmath 011A}^0$: it is $M=1$ for ${\Greekmath 011A}=0$, $M=2$ for ${\Greekmath 011A}=0.3$, $M=3$ and ${\Greekmath 011A}=0.6$, and $M=5$ for ${\Greekmath 011A}=0.9$. Choosing too large or too small a bandwidth results in a smaller fraction of the bias to be corrected. Table (ref) also reports the properties of the BC-FLS estimator for different values of ${\Greekmath 011A}^0$, $T$, and $M$. It shows the effect of the bandwidth choice on the standard deviation of the BC-FLS estimator is relatively small at $T=40$, but is more pronounced at $T=20$. The issue of optimal bandwidth choice is therefore an important topic for future research. In the simulation results presented here, we tried to choose reasonable values for $M$, but made no attempt to optimize the bandwidth.

In our setup, we have $\|{\Greekmath 0115}^0 f^{0 \prime}\| \approx \sqrt{2 N T} {\Greekmath 011B}_f$ and $\|e\| \approx \sqrt{N} + \sqrt{T}$.\footnote{ To be precise, we have $\|{\Greekmath 0115}^0 f^{0 \prime}\| / (\sqrt{2 N T} {\Greekmath 011B}_f) \, \rightarrow_p \, 1$, and $\|e\|/(\sqrt{N} + \sqrt{T}) \, \rightarrow_p \, 1$.} Assumptions (ref) and (ref) imply $\|{\Greekmath 0115}^0 f^{0 \prime}\| \gg \|e\|$ asymptotically. We can therefore only be sure our asymptotic results for the FLS estimator distribution are a good approximation of the finite sample properties if $\|{\Greekmath 0115}^0 f^{0 \prime}\| \gtrsim \|e\|$, that is, if $\sqrt{2 N T} {\Greekmath 011B}_f \, \gtrsim \, \sqrt{N} + \sqrt{T}$. To explore this further, we present in Table (ref) simulation results for $N=100$, $T=20$, ${\Greekmath 011A}^0=0.6$, and different values of ${\Greekmath 011A}_f$ and ${\Greekmath 011B}_f$. For ${\Greekmath 011B}_f=0$, we have $0=\|{\Greekmath 0115}^0 f^{0 \prime}\| \ll \|e\|$, and this case is equivalent to $R=0$ (no factor at all). In this case, the OLS estimator estimates the true model and is almost unbiased, and correspondingly, the FLS estimator and the bias-corrected FLS estimator perform worse than OLS in finite samples (though we expect all three estimators are asymptotically equivalent), but the bias-corrected FLS estimator has a lower bias and a lower variance than the non-corrected FLS estimator. The case ${\Greekmath 011B}_f = 0.2$ corresponds to $\|{\Greekmath 0115}^0 f^{0 \prime}\| \approx \|e\|$, and one finds the bias and the variance of the OLS estimator and of the FLS estimator are of comparable size. However, the bias-corrected FLS estimator already has much smaller bias and a bit smaller variance in this case. Finally, in the case ${\Greekmath 011B}_f = 0.5$, we have $\|{\Greekmath 0115}^0 f^{0 \prime}\| > \|e\|$, and we expect our asymptotic results to be a good approximation of this situation. Indeed, one finds that for ${\Greekmath 011B}_f=0.5$, the OLS estimator is heavily biased and very inefficient compared to the FLS estimator, whereas the bias-corrected FLS estimator performs even better in terms of bias and variance.

In Table (ref), we present simulation results for the size of the various tests discussed in the last section when testing the null hypothesis $H_0: {\Greekmath 011A}={\Greekmath 011A}^0$. We choose a nominal size of $5\%$, ${\Greekmath 011A}_f=0.5$, ${\Greekmath 011B}_f=0.5$, and different values for ${\Greekmath 011A}^0$, $N$, and $T$. In all cases, the size distortions of the uncorrected Wald, LR, and LM test are rather large, and the size distortion of these tests do not vanish as $N$ and $T$ increase: the size for $N=100$ and $T=20$ is about the same as for $N=400$ and $T=80$, and the size for $N=400$ and $T=20$ is about the same as for $N=1600$ and $T=80$. By contrast, the size distortions for the bias-corrected Wald, LR, and LM test are much smaller, and tend toward zero (i.e., the size becomes closer to $5\%$) as $N,T$ increase, holding the ratio $N/T$ constant. For fixed $T$, an increase in $N$ results in a larger size distortion, whereas for fixed $N$, an increase in $T$ results in a smaller size distortion (both for the non-corrected and for the bias-corrected tests).

In Table (ref) and (ref), we present the power and the size-corrected power when testing the left-sided alternative $H^{\rm left}_a: {\Greekmath 011A} = {\Greekmath 011A}^0-(NT)^{-1/2}$ and the right-sided alternative $H^{\rm right}_a: {\Greekmath 011A} = {\Greekmath 011A}^0+(NT)^{-1/2}$. The model specifications are the same as for the size results in Table (ref). Because both the FLS estimator and the bias-corrected FLS estimator for ${\Greekmath 011A}$ have a negative bias, one finds the power for the left-sided alternative to be much smaller than the power for the right-sided alternative. For the uncorrected tests, this effect can be extreme and the size-corrected power of these tests for the left-sided alternative is below $2 \%$ in all cases and does not improve as $N$ and $T$ become large, holding $N/T$ fixed. By contrast, the power for the bias-corrected tests becomes more symmetric as $N$ and $T$ become large, and the size-corrected power for the left-sided alternative is much larger than for the uncorrected tests, whereas the size-corrected power for the right-sided alternative is about the same.

Conclusions

This paper studies the least squares estimator for dynamic linear panel regression models with interactive fixed effects. We provide conditions under which the estimator is consistent, allowing for predetermined regressors and for a general combination of “low-rank” and “high-rank” regressors. We then show how a quadratic approximation of the profile objective function $L_{NT}({\Greekmath 010C})$ can be used to derive the first-order asymptotic theory of the LS estimator of ${\Greekmath 010C}$ under the alternative asymptotic $N,T \rightarrow \infty$. We find the asymptotic distribution of the LS estimator can be asymptotically biased (i) because of weak exogeneity of the regressors and (ii) because of heteroscedasticity (and correlation) of the idiosyncratic errors $e_{it}$. Consistent estimators for the asymptotic covariance matrix and for the asymptotic bias of the LS estimator are provided, and thus a bias-corrected LS estimator is given. We furthermore study the asymptotic distributions of the Wald, LR, and LM test statistics for testing a general linear hypothesis on ${\Greekmath 010C}$. The uncorrected test statistics are not asymptotically chi-square because of the asymptotic bias of the score and of the LS estimator, but bias-corrected test statistics that are asymptotically chi-square distributed can be constructed. We also discussed a possible extension of the estimation procedure to the case of endogeneous regressors. The findings of our Monte Carlo simulations show our asymptotic results on the distribution of the (bias-corrected) LS estimator and of the (bias-corrected) test statistics provide a good approximation of their finite sample properties. Although the bias-corrected LS estimator has a non-zero bias in finite samples, this bias is much smaller than that of the LS estimator. Analogously, the size distortions and power asymmetries of the bias-corrected Wald, LR, and LM test are much smaller than for the non-bias-corrected versions.