EconBase
← Back to paper

Efficiency of QMLE for dynamic panel data models with interactive effects

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

98,095 characters · 9 sections · 50 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Efficiency of QMLE for dynamic panel data models with interactive effects

frontmatter\runtitle{Dynamic panel data models} \begin{aug} , \address[A]{Department of Economics, Columbia University\printead[presep={,\ }]{e1}} \end{aug} \begin{abstract} This paper studies the problem of efficient estimation of panel data models in the presence of an increasing number of incidental parameters. We formulate the dynamic panel as a simultaneous equations system, and derive the efficiency bound under the normality assumption. We then show that the Gaussian quasi-maximum likelihood estimator (QMLE) applied to the system achieves the normality efficiency bound without the normality assumption. Comparison of QMLE with the fixed effects approach is made. \end{abstract} \begin{keyword}[class=MSC] \kwd[Primary ]{62H12} \kwd[; secondary ]{62F12} \end{keyword} \begin{keyword} \kwd{Fixed effects} \kwd{incidental parameters} \kwd{local likelihood ratios} \kwd{local parameter space} \kwd{regular estimators} \kwd{efficiency bound} \kwd{factor models} \end{keyword}

Introduction

Consider the dynamic panel data model with interactive effects

equation[equation omitted — 104 chars of source]

\[ i=1,2,...,N; t=1,...,T\] where $y_{it}$ is the outcome variable, $\lambda_i$ and $f_t$ are each $r\times 1$ and both are unobservable, $\delta_t$ is the time effect, and $\eps_{it}$ is the error term. Only $y_{it}$ are observable. The above model is increasingly used for empirical studies in social sciences. The purpose of this paper is about efficient estimation of the model by deriving the efficiency bound. We show that quasi-maximum likelihood estimation (QMLE) achieves the efficiency bound. But first, we explain the meaning of QMLE for this model and its motivations.

The index $i$ is referred to as individuals (e.g., households) and $t$ as time. If $f_t=1$ for all $t$, and $\lambda_i$ is scalar, then $\delta_t +\lambda_i'f_t =\delta_t+\lambda_i$, we obtain the usual additive individual and time fixed effects model. Dynamic panel models with additive fixed effects remain the workhorse for empirical research. The product $\lambda_i'f_t$ is known as the interactive effects Bai2009, and is more general than additive fixed effects models. The models allow the individual heterogeneities (such as unobserved innate ability, captured by $\lambda_i$) to have time varying impact (through $f_t$) on the outcome variable $y_{it}$. In a different perspective, the models allow common shocks (modeled by $f_t)$ to have heterogeneous impact (through $\lambda_i$) on the outcome. For many panel data sets, $N$ is usually much larger than $T$ because it is costly to keep track of the same individuals over time. Under fixed $T$, typical estimation methods such as the least squares do not yield consistent estimation of the model parameters. Consider the special case

equation[equation omitted — 84 chars of source]

where $c_i$ are fixed effects. This corresponds to $\delta_t=0$ and $f_t=1$ for all $t$. Even if $c_i$ are iid, zero mean and finite variance, and independent of $\eps_{it}$, the least squares estimator $\hat \alpha =(\sum_i\sum_t y_{it-1}^2)^{-1} \sum_i\sum_t y_{it-1} y_{it}$ is easily shown to be inconsistent. When $c_i$ are treated as parameters to be estimated along with $\alpha$, the least squares method is still biased, and the order of bias is $O(1/T)$, no matter how large is $N$, see Nickell1981. So unless $T$ goes to infinity, the least squares method remains inconsistent, an issue known as the incidental parameters problem, for example, Kiviet1995 and Lancaster2000.

However, provided $T\ge 3$, consistent estimation of $\alpha$ is possible with the instrumental variables (IV) method. Anderson and Hsiao AndersonHsiao1982 suggested the IV estimator by solving $ \sum_{i=1}^N y_{i1} (\Delta y_{i3}-\alpha \Delta y_{i2}) =0$, where $\Delta y_{it}=y_{it}-y_{it-1}$. Differencing the data purges $c_i$, but introduces correlation between the regressor and the resulting errors, which is why the IV method is used. With $T$ strictly greater than 3, more efficient IV estimator is suggested by Arellano and Bond ArellanoBond1991. Model ((ref)) can also be estimated by the Gaussian quasi-maximum likelihood method, for example, AlvarezArellano2022, Bai2013, and Moreira2009.

For model ((ref)), differencing cannot remove the interactive effects since $ \Delta y_{it}=\alpha \, \Delta y_{it-1} +\Delta\delta_t + \lambda_i'\Delta f_t + \Delta\eps_{it}$. The model can be estimated by the fixed effects approach, treating both $\lambda_i$ and $f_t$ as parameters. Just like the least squares for the earlier additive effects model, the fixed effects method will produce bias. Below we introduce the quasi-likelihood approach, similar to BaiLi2012, BaiLi2014 for non-dynamic models.

Project the first observation $y_{i1}$ on $[1,\lambda_i]$ and write $ y_{i1} =\delta_1^* + \lambda_i'f_1^* + \eps_{i1}^*$, where $(\delta_1^*, f_1^*)$ is the projection coefficients, and $\eps_{i1}^*$ is the projection error. The asterisk variables are different from the true $(\delta_1,f_1,\eps_{i1})$ that generates $y_{i1}$. This projection is called upon because $y_{i0}$ is not observable.\footnote{The first observation starts at $t=1$, $y_{i0}$ is not available. If $y_{i0}$ were observable we would have $y_{i1}=\alpha y_{i0} +\delta_1 +\lambda_i'f_1 + \eps_{i1}$. But then a projection of $y_{i0}$ on $[1,\lambda_i]$ would be required.} Note that we can drop the superscript $*$ to simplify the notation. This is because we will treat $\delta_t$ and $f_t$ as (nuisance and free) parameters, and we do not require $\eps_{it}$ to have the same distribution over time. This means we can rewrite $y_{i1}$ as $y_{i1}= \delta_1+ \lambda_i'f_1 + \eps_{i1}$.

The following notation will be used

equation[equation omitted — 576 chars of source]

together with the following $T\times T$ matrices,

equation[equation omitted — 609 chars of source]

Note that $L= J B^{-1}$. With these notations, we can write the model as \[ By_i = \delta +F \lambda_i + \eps_i \] This gives a simultaneous equations system with $T$ equations. We assume $\lambda_i$ are iid, independent of $\eps_i$. Without loss of generality, we assume $\mathbb E(\lambda_i)=0$, otherwise, absorb $F \lambda$ (where $\lambda=\mathbb E\lambda_i )$ into $\delta$. Further assume $ \mathbb E(\lambda_i\lambda_i')=I_r$ (a normalization restriction, where $I_r$ is an identity matrix). We also assume $\eps_i$ are iid with zero mean, and \[ D =\var(\eps_i) =\diag(\sigma_1^2, \sigma_2^2,...,\sigma_T^2) \] These assumptions imply that $By_i$ are iid with mean $\delta$ and covariance matrix $FF'+D$. Consider the Gaussian quasi likelihood function \[ \ell_{NT} (\theta) =-\frac N 2 \log|FF'+D|-\frac 1 2 \sumiN (By_i-\delta)'(FF'+D)^{-1} (By_i-\delta) \] where the Jacobian does not enter since the determinant of $B$ is 1, where $\theta=(\alpha, \delta, F, \sigma_1^2,...,\sigma_T^2)$. The quasi-maximum likelihood estimator (QMLE) is defined as \[ \hat \theta =\argmax_\theta \ell_{NT}(\theta) \] The asymptotic distribution of this estimator is studied by Bai2024.

An alternative estimator, the fixed effects estimator, treats both $\lambda_i$ and $f_t$ as parameters, in addition to $\alpha$ and $\delta_t$. The corresponding likelihood function under normality of $\eps_{it}$ is given in ((ref)) below. The fixed effects framework estimates more nuisance parameters (can be substantially more under large $N$), the source of incidental parameters problem. Our analysis focuses on QMLE. Comparison of the two approaches will be made.

The objectives of the present paper are threefold. First, what is the efficiency bound for the system maximum likelihood estimator under normality assumption? Second, does the QMLE attain the normality efficiency bound with or without the normality assumption? Third, how does QMLE fare in comparison to the fixed effects estimator?

We approach these questions with Le Cam's type of analysis. The difficulty lies in the increasing dimension of the parameter space as $T$ goes to infinity because the number of parameters is of order $T$. No sparsity in parameters is assumed. With sparsity, JankovaGeer2018 derived efficiency bounds and constructed efficient estimators via regularization for various models. The ability to deal with non-sparsity in the current model relies on panel data.

On notation: $\|A\|$ denotes the Frobenius norm for matrix (or vector) $A$, that is, $\|A\|= (\mathrm{tr}(A'A))^{1/2}$, and $\|A\|_2$ denotes the spectral norm of $A$, that is, the square root of the largest eigenvalue of $A'A$. Notice $\|AB\|\le \|A\|_2 \|B\|$. The transpose of $A$ is denoted by $A'$; $|A|$ and $\mathrm{tr}(A)$ denote, respectively, its determinant and trace for a square matrix $A$.

Assumptions for QMLE

We assume $|\alpha|<1$ for asymptotic analysis. The following assumptions are made for the model.

{\bf Assumption A}

(i) $\eps_i$ are iid over $i$; $\mathbb E(\eps_{it})=0$, $\var(\eps_{it})=\sigma_t^2>0$, and $\eps_{it}$ are independent over $t$; $\mathbb E\eps_{it}^4 \leq M<\infty$ for all $i$ and $t$.

(ii) The $\lambda_i$ are iid, independent of $\eps_i$, with $\mathbb E\lambda_i=0$, $\mathbb E(\lambda_i \lambda_i')=I_r$, and $\mathbb E\|\lambda_i\|^4 \le M$.

(iii) There exist constants $a$ and $b$ such that $0 < a < \sigma_t^2 < b< \infty$ for all $t$; $ \frac 1 T F' D^{-1} F =\frac 1 T \sumtT \sigma_t^{-2} f_t f_t' \rightarrow \Sigma_{ff}>0$.

Two comments are in order for this model. First, Assumption A(ii) assumes $\lambda_i$ are random variables, but they can be fixed bounded constant. All needed is that $\Psi_N:=\frac 1 N \sumiN (\lambda_i-\bar \lambda) (\lambda_i-\bar \lambda)' \rightarrow \Psi >0$ (an $r\times r$ positive definite matrix), where $\bar \lambda$ is the sample average of $\lambda_i$. One can normalize the matrix $\Psi$ to be an identity matrix. Second, $F$ is determined up to an orthogonal rotation since $FF'=F R (FR)'$ for $RR'=I_r$. the rotational indeterminacy can be removed by the normalization that $F'D^{-1}F$ is a diagonal matrix (with distinct elements), see Anderson2003 (p.573) and LawleyMaxwell1971 (p.8). Rotational indeterminacy does not affect the estimate for $\alpha$, $D$, and $\delta$.

Under Assumption A, Bai2024 showed that the QMLE for $\hat \alpha$ has the following asymptotic representation under $N,T\rightarrow \infty$ with $N/T^3\rightarrow 0$,

equation[equation omitted — 204 chars of source]

where $L$, $D$, and $\eps_i$ are defined earlier. Note that $\mathbb E[(L \eps_i)'D ^{-1} \eps_i] =\mathrm{tr}(L')=0$, and $\mathbb E[(L \eps_i)'D ^{-1} \eps_i]^2= \mathrm{tr}(LD L' D ^{-1})$ with the expression

equation*[equation* omitted — 188 chars of source]

We assume the above converges to $\gamma>0$, as $T\rightarrow \infty$, that is

equation[equation omitted — 95 chars of source]

Then we have, \[ \sqrt{NT}(\hat \alpha-\alpha) \stackrel{d}{\longrightarrow} \mathcal{N}(0, 1/\gamma). \] For the special case of homoskedasticity, ($\sigma_t^2=\sigma^2$ for all $t$), $\gamma=1/(1-\alpha^2)$, and hence $\sqrt{NT}(\hat \alpha-\alpha) \stackrel{d}{\longrightarrow} \mathcal{N}(0, 1-\alpha^2)$.

QMLE requires no bias correction, unlike the fixed effects regression. The latter is considered by Bai2009 and MoonWeidner2017. Our objective is to show that $1/\gamma$ is the efficiency bound under normality assumption, and QMLE attains the normality efficiency bound. This result is obtained in the presence of increasing number of incidental parameters. The estimator $\hat \alpha$ is also consistent under fixed $T$ in contrast to the fixed effects estimator. The estimated $ f_t, \sigma_t^2$ are all $\sqrt{N}$ consistent and asymptotically normal. In particular, the estimated factors $\hat f_t$ have the asymptotic representation, for each $t=1,2,...,T$

equation[equation omitted — 116 chars of source]

This implies $\sqrt{N} (\hat f_t -f_t)\stackrel{d}{\longrightarrow} \mathcal{N}(0,\sigma_t^2 I_r)$. Details are given in Bai Bai2024. Also see BaiLi2012 for non-dynamic factor models.

Local likelihood ratios and efficiency bound

Related literature

A closely related work is that of Iwakura and Okui IwakuraOkui2014. They consider the fixed effects framework instead of the QMLE. The fixed effects estimation procedure treats both $\lambda_i$ and $f_t$ as parameters $(i=1,2,...,N; t=1,2,...,T$), along with $\alpha$ and $\delta$. The corresponding likelihood function under normality of $\eps_{it}$ is\footnote{The fixed effects likelihood does not have a global maximum under heteroskedasticity, for example, Anderson2003 (p.587), but local maximization is still meaningful. Another solution is to impose homoskedasticity.}

equation[equation omitted — 214 chars of source]

The fixed effects estimator for $\alpha$ will generate bias, similar to the fixed effects estimator for dynamic panels with additive effects. The bias is studied by MoonWeidner2017.\footnote{ In contrast, the QMLE does not generate bias under fixed $T$.} Iwakura and Okui IwakuraOkui2014 derive the efficiency bound for the fixed effects estimators under homoskedasticity ($\sigma_t^2=\sigma^2$ for all $t$). Another closely related work is that of Hahn and Kuersteiner HahnKuersteiner2002. They consider the efficiency bound problem under the fixed effects framework for the additive effects model described by ((ref)) without a factor structure. Throughout this paper, the fixed effects framework refers to methods that also estimate the factor loadings $\lambda_i$ in addition to $f_t$.

In contrast, we consider the likelihood function for the system of equations

equation[equation omitted — 113 chars of source]

QMLE does not estimate $\lambda_i$ (even if they are fixed constants as explained earlier), thus eliminating the incidental parameters in the cross-section dimension. The incidental parameters are now $\delta$, $F$ and $D$, and the number of parameters increases with $T$. Despite fewer number of incidental parameters, the analysis of local likelihood is more demanding than that of the fixed effects likelihood ((ref)). Intuitively, the fixed effects likelihood ((ref)) is quadratic in $F$, but the QMLE likelihood $\ell(\theta)$ in ((ref)) depends on $F$ through the inverse matrix $(FF'+D)^{-1}$ and through the log-determinant of this matrix. The high degree of nonlinearity makes the perturbation analysis more challenging. As demonstrated later, the local analysis brings insights regarding the relative merits of the QMLE and the fixed effects estimators.

Notice $By_i-\delta = B(y_i-\delta^\dag)$, where $\delta^\dag=B^{-1}\delta$ is a vector of free parameters because $\delta$ is a vector of free parameters. The concentrated likelihood function by concentrating out $\delta^\dag$ (its maximum likelihood estimator is simply $\bar y =\frac 1 N \sumiN y_i$) is given by

equation[equation omitted — 134 chars of source]

This is the likelihood function studied by Bai2024 and BaiMones2025 for the QMLE, with the latter paper focusing on global identification. Our efficiency bound analysis is based on ((ref)).

There is a substantial body of research on dynamic models with interactive effects. For example, BaiLi2021 and shilee2017 study spatial models, miaophillipssu2023 examine high-dimensional vector autoregressions, and lamyao2012 focus on high-dimensional time series models. However, unlike the present work, these studies do not address efficiency bounds. Throughout this paper, we assume that the number of factors is known. In practice, the number of factors can be estimated using methods such as the one proposed by lamyao2012.

The \texorpdfstring{$\ell^\infty$}{Lg} local parameter space

Local likelihood ratio processes are indexed by local parameters. Since the convergence rate for the estimated parameter of $\alpha^0$ is $\sqrt{NT}$, that is, $\sqrt{NT} (\hat \alpha-\alpha^0) =O_p(1)$, it is natural to consider the local parameters of the form \[ \alpha^0+ \frac 1 {\sqrt{NT} } \tilde \alpha \] where $\tilde \alpha \in \mathbb R$. However, the consideration of local parameters for $f_t^0$ is non-trivial, as explained by Iwakura and Okui IwakuraOkui2014 for the fixed effects likelihood ratio. We consider the following local parameters

equation[equation omitted — 112 chars of source]

where $\|\tilde f_t\| \le M <\infty$ for all $t$ with $M$ arbitrarily given; $\|\cdot\|$ denotes the r-dimensional Euclidean norm. In view that the estimated factor $\hat f_t$ is $\sqrt{N}$ consistent, that is, $\sqrt{N}(\hat f_t-f_t^0)=O_p(1)$, one would expect local parameters in the form $f_t^0+ N^{-1/2} \tilde f_t$, the extra scale factor $T^{-1/2}$ in the above local rate looks rather unusual. However, ((ref)) is the suitable local rate for the local likelihood ratio to be $O_p(1)$, as is shown in both the statement and the proof of Theorem (ref) below. Without the scale factor $T^{-1/2}$, the local likelihood ratio will diverge to infinity (in absolute values) if no restrictions are imposed on $\tilde f_t$ other than its boundedness. This type of local parameters was used in earlier work by HahnKuersteiner2002 for additive fixed effects estimator. Later we shall consider a different type of local parameters without the extra scale factor $T^{-1/2}$, but other restrictions on $\tilde f_t$ will be needed.

Additionally, even if one regards $\frac 1 {\sqrt{NT} } \tilde f_t$ to be small (relative to $\frac 1 {\sqrt{N}} \tilde f_t$), it is for the better provided that the associated efficiency bound is achievable by an estimator. This is because the smaller the perturbation, the lower the efficiency bound, and hence harder to attain by any estimator.

Consider the space \[ \ell_r^\infty: =\{ (\tilde f_t)_{t=1}^\infty \, \Big| \tilde f_t \in \mathbb R^r, \, \sup_s \|\tilde f_s\| <\infty \}\] the space of bounded sequences, each coordinate is $\mathbb R^r$-valued. Let \[ \tilde f =(\tilde f_1,\tilde f_2, ...) \in \ell_r^\infty \] and define $\tilde F=(\tilde f_1,\tilde f_2,...,\tilde f_T)'$, the projection of $\tilde f$ onto the first $T$ coordinates. The matrix $\tilde F$ is $T\times r$, but we suppress its dependence on $T$ for notational simplicity. Since $\tilde f \in \ell_r^\infty$, it follows that

equation[equation omitted — 113 chars of source]

To simplify the analysis, we assume that $D$ is known. This assumption does not affect the efficiency bound for $\alpha$, but it does simplify the derivation considerably. Nonetheless, it may be worthwhile to rigorously analyze the case where $D$ is unknown. In general, treating some parameters as known can result in a lower efficiency bound, which may be difficult to attain. In our setting, however, we demonstrate that the efficiency bounds are achievable.

Let $\theta^0=(\alpha^0, F^0)$ and $\tilde \theta=(\tilde \alpha, \tilde F)$, we study the asymptotic behavior of \[ \ell(\theta^0+\frac 1 {\sqrt{NT} } \tilde \theta) -\ell(\theta^0) \] under the normality of $\eps_{it}$ and $\lambda_i$. The normality assumption allows us to derive the parametric efficiency bound in the presence of increasing number of nuisance parameters. We then show the QMLE without normality attains the efficiency bound. In the rest of the paper, we use $(\alpha^0, F^0)$ and $(\alpha, F)$ interchangeably (they represent the true parameters); $(\tilde \alpha, \tilde F)$ represent local parameters, and $\hat \alpha$ and $\hat F=(\hat f_1,...,\hat f_T)'$ are the QMLE estimated parameters.

{\bf Assumption B}

(i): $\eps_{it}$ are iid over $i$, and independent over $t$ such that $\eps_{it}\sim \mathcal{N}(0,\sigma_t^2)$.

(ii): $\lambda_i$ are iid $\mathcal{N}(0,I_r)$, independent of $\eps_{it}$ for all $i$ and $t$.

(iii): $\sigma_t^2 \in [a, b]$ with $ 0< a < b < \infty$ for all $t$.

(iv): $\|f_t\|\le M<\infty$ for all $t$, and $\frac 1 T F'D^{-1}F\rightarrow \Sigma_{ff}>0$, where $D=\diag(\sigma_1^2,\sigma_2^2,...,\sigma_T^2)$.

(v): As $T\rightarrow \infty$,\\ (a) $\frac 1 T \mathrm{tr}(L'D^{-1} L D) \rightarrow \gamma >0$,\\ (b) $\frac 1 T \mathrm{tr}[ (LF)' (D^{-1/2} M_{D^{-1/2}F} D^{-1/2} )(LF)] \rightarrow \nu \ge 0$\\ where $M_{D^{-1/2}F}$ denote the projection matrix orthogonal to $D^{-1/2} F$. Specifically, \[ D^{-1/2} M_{D^{-1/2}F} D^{-1/2} =D^{-1}-D^{-1}F(F'D^{-1}F)^{-1}F'D^{-1}. \]

Under assumptions B(i) and (ii), $F \lambda_i +\eps_i$ are iid $\mathcal{N}(0, FF'+D)$, implying a parametric model with an increasing dimension of incidental parameters. Normality for $\lambda_i$ and $\eps_{it}$ is a standard assumption in factor analysis, see, e.g., Anderson Anderson2003 (p.576). Here we switch the role of $\lambda_i$ and $f_t$. Note in classical factor analysis, the time dimension $T$ (in our notation) is fixed, there is no incidental parameters problem since the number of parameters is fixed. But we consider $T$ that goes to infinity. The following theorem gives the asymptotic representation for the local likelihood ratios.

theoremUnder Assumption B, for $\tilde \alpha\in \mathbb R$, $\tilde f\in \ell_r^\infty$, $\tilde F=(\tilde f_1,\tilde f_2,...,\tilde f_T)'$, we have as $N,T\rightarrow \infty$, with $N/T^3\rightarrow 0$, \[ \ell(\theta^0+\frac 1 {\sqrt{NT} } \tilde \theta) -\ell(\theta^0) =\Delta_{NT}(\tilde \theta)-\frac 1 2 \mathbb E[\Delta_{NT}(\tilde \theta)]^2 + o_p(1)\] where \begin{align} \Delta_{NT}(\tilde \theta) & = \frac 1 {\sqrt{NT} } \sumiN \lambda_i'\tilde F'[D^{-1/2} M_{D^{-1/2}F} D^{-1/2} ]\eps_i \notag\\ & + \tilde \alpha \, \frac 1 {\sqrt{NT} } \sumiN (L\eps_i)'D^{-1}\eps_i \\ & +\tilde \alpha \, \frac 1 {\sqrt{NT} } \sumiN\lambda_i' (LF)'[D^{-1/2} M_{D^{-1/2}F} D^{-1/2} ] \eps_i \notag \end{align} \begin{align} \mathbb E[\Delta_{NT}(\tilde \theta)]^2= & \frac 1 T \mathrm{tr} \Big[\tilde F'[D^{-1/2} M_{D^{-1/2}F} D^{-1/2}] \tilde F \Big] \notag \\ & + \tilde \alpha^2 \Big[ \frac 1 T \mathrm{tr}(L'D^{-1}LD ) \Big] \\ & + \tilde \alpha^2 \mathrm{tr} \Big[ \frac 1 T (LF)' [D^{-1/2} M_{D^{-1/2}F} D^{-1/2} ](LF)\Big] \notag\\ & + 2\tilde \alpha \, \mathrm{tr} \Big[ \frac 1 T (LF)'[D^{-1/2} M_{D^{-1/2}F} D^{-1/2} ] \tilde F \Big] \notag \end{align} where $o_p(1)$ is uniform over $\tilde \theta$ such that $|\tilde \alpha|\le M$, $ \frac 1 T \|\tilde F'\tilde F \|=\frac 1 T \sum_{t=1}^T \|\tilde f_t\|^2 \le M$, for any given $M<\infty$.

The proof of Theorem (ref) is given in the \hyperref[appn]{Appendix}.

Note that the expected value of $\Delta_{NT}(\tilde \theta)$ is zero, so $\mathbb E[\Delta_{NT}(\tilde \theta)]^2$ is the variance.

All terms in $\Delta_{NT}(\tilde \theta)$ are stochastically bounded, they have expressions of the form $\frac 1 {\sqrt{NT} } \sumiN\sumtT \xi_{it}$, where $\xi_{it}$ have zero mean, and finite variance (and in fact finite moments of any order under Assumption B). Assume $\tilde f$ satisfies

align*[align* omitted — 154 chars of source]

as well as existence of limits for $\frac 1 T \tilde F'D^{-1} F$, then the variance $\mathbb E [\Delta_{NT}(\tilde \theta)]^2$ has a limit. Let $\mathbb E [\Delta_{NT}(\tilde \theta)]^2 \rightarrow \tau^2$ for some $\tau^2$ depending on $(\tilde \alpha, \tilde f)$. We can further show \[ \Delta_{NT}(\tilde \theta) \stackrel{d}{\longrightarrow} \mathcal{N}(0, \tau^2). \] Thus the local likelihood ratio can be rewritten as \[ \ell(\theta^0+\frac 1 {\sqrt{NT} } \tilde \theta) -\ell(\theta^0) =\Delta_{NT}(\tilde \theta)-\frac 1 2 \tau^2 + o_p(1).\]

We next consider the asymptotic efficiency bound for regular estimators. Regular estimators rule out Hodges type “superefficient” and James-Stein type estimators. A regular estimator sequence converges locally uniformly (under the local laws) to a limiting distribution that is free from the local parameters (van der Vaart vanderVaart1998, p.115, p.365).

Efficient scores and efficiency bound

In the likelihood ratio expansion, the term $\Delta_{NT}(\tilde \theta)$ contains the scores of the likelihood function. The coefficient of $\tilde \alpha$ gives the score for $\alpha^0$, and the coefficient of $\tilde f_t$ gives the score of $f_t^0$. The efficient score for $\alpha^0$ is the projection residual of its own score onto the scores of $f_1^0,...,f_T^0$. Moreover, the inverse of the variance of the efficient score gives the efficiency bound (Bickel et al, bickel1993, p. 28).

To derive the efficient score for $\alpha^0$, rewrite $\Delta_{NT}(\tilde \theta)$ of Theorem 1 as \[ \Delta_{NT}(\tilde \theta)=\Delta_{NT1} + \tilde \alpha\, [ \Delta_{NT2}+\Delta_{NT3}] \] where $\Delta_{NT1}$ denotes the first term of $\Delta_{NT}(\tilde \theta)$, see ((ref)), and $\Delta_{NTj}$ $(j=2,3$) denote the last two terms of ((ref)), but taking out $\tilde \alpha$. So the score for $\alpha^0$ is the sum $\Delta_{NT2}+\Delta_{NT3}$. Next, rewrite

align*[align* omitted — 171 chars of source]

where $v_t= \frac 1 {\sqrt{N}} \sumiN \lambda_i v_{it} $ $(r\times 1)$ and $v_{it}$ is the $t$-th element of the vector $[D^{-1/2} M_{D^{-1/2}F} D^{-1/2} ]\eps_i$. Thus $v_t$ is the score of $f^0_t$ ($t=1,2,...,T$). To obtain the efficient score for $\alpha^0$, we project $\Delta_{NT2}+\Delta_{NT3}$ onto the scores of $f_t^0$, that is onto $[v_1,v_2....,v_T]$to get the projection residuals. Let $V_T=(v_1',v_2',...,v_T')'$. The projection residual is given by

equation[equation omitted — 153 chars of source]

Notice $\Delta_{NT2}$ is uncorrelated with the scores of $f_t^0$, i.e. $\mathbb E(V_T \Delta_{NT2})=0$. This follows because $L\eps_i$ contains the lags of $\eps_i$, so $\Delta_{NT2}$ is composed of terms $\eps_{it-s}\eps_{it}$ (with $s\ge 1$), and $\mathbb E (\eps_{it-s}\eps_{it} \eps_{ik})=0$ for any $k$. Next, $\Delta_{NT3}$ is simply a linear combination of $V_T= [v_1,v_2,...,v_T]$ since $\Delta_{NT3}$ can be written as $T^{-1/2} \sum_{t=1}^T p_t' v_t$, where $p_t'$ is the $t$-th row of the matrix $L F$. Thus $V_T'[ \mathbb E(V_T V_T')]^{-1} \mathbb E [V_T (\Delta_{NT3})]\equiv \Delta_{NT3}$. In summary, we have \[ V_T'[ \mathbb E(V_T V_T')]^{-1} \mathbb E \left[V_T (\Delta_{NT2} +\Delta_{NT3}\right] =\Delta_{NT3} \] It follows that the projection residual in ((ref)) is equal to $\Delta_{NT2}$. Hence the efficient score for $\alpha^0$ is $\Delta_{NT2}$. Notice,

equation[equation omitted — 223 chars of source]

its limit is $\gamma$ by Assumption B(v), so $1/\gamma$ gives the asymptotic efficiency bound.

We summarize the result in the following corollary.

corollaryUnder Assumption B, the asymptotic efficiency bound for regular estimators of $\alpha^0$, is $1/\gamma$, with $\gamma$ being the limit of ((ref)).

Since Assumption B is stronger than Assumption A, the asymptotic representation in ((ref)) holds under Assumption B. That is, under the normality assumption, the system maximum likelihood estimator satisfies \[ \sqrt{NT}(\hat \alpha-\alpha) = \Big(\frac 1 T \mathrm{tr}(LD L' D ^{-1}) \bigg)^{-1} \Big[ \frac 1 {\sqrt{NT}} \sumiN (L \eps_i)'D ^{-1} \eps_i \Big ]+o_p(1). \] We see that $\sqrt{NT}(\hat \alpha-\alpha^0)$ is expressed in terms of the efficient influence functions, thus $\hat \alpha$ is regular and asymptotically efficient (van der Vaart vanderVaart1998, p.121 and p.369). We state the result in the following corollary.

corollaryUnder Assumption B, the system maximum likelihood estimator $\hat \alpha$ is a regular estimator and achieves the asymptotic efficiency bound (in spite of an increasing number of incidental parameters).

The preceding corollaries imply that, under normality, we are able to establish the asymptotic efficiency bound in the presence of increasing number of nuisance parameters. Further, the system maximum likelihood estimator achieves the efficiency bound. These results are not obvious owing to the incidental parameters problem.

QMLE in Section 2 does not require normality. But it achieves the normality efficiency bound, see equation ((ref)). So QMLE is robust to the normality assumption. If $\lambda_i$ and $\eps_i$ are non-normal, and their distributions are known, one should be able to construct a more efficient estimator than the QMLE. But for panel data analysis in practice, researchers usually do not impose distributional assumptions other than some moment conditions such as those in Assumption A. Thus QMLE presents a viable estimation procedure, knowing that it achieves the normality efficiency bound. Furthermore, QMLE does not need bias correction, unlike the fixed effects estimator.

Constructing a Hilbert subspace

The result of Corollary 1 is not directly obtained via a limit experiment and the convolution theorem (e.g., van der Vaart and Wellner, vanderVaartWellner1996, chapter 3.11). Since $\ell_r^\infty$ is not a Hilbert space, the convolution theorem is not directly applicable. However, using the line of argument in IwakuraOkui2014 it is possible to construct a Hilbert subspace with an appropriate inner product in which the efficiency bound for the low dimensional parameter $\alpha^0$ can be shown to be $1/\gamma$. That is, Corollary 1 can be obtained via the convolution theorem.

We now construct a Hilbert subspace so that convolution theory can be directly applied. For this purpose, we adopt the method proposed by IwakuraOkui2014.

{\bf Assumption C.} Let $\psi$ be a vector of continuous function from [0,1] to $\mathbb R^r$, and $f_t =\psi(t/T)$. Let $\sigma$ be a continuous function from [0,1] to $\mathbb R$ and $b\ge\sigma(s)\ge a>0$ for all $s \in [0,1]$, and $\sigma_t^2 =\sigma^2(t/T)$. They satisfy $ \frac 1 T \sumtT \frac 1 {\sigma_t^2} f_t f_t'\rightarrow \int_0^1 \frac 1 {\sigma(s)^2} \psi(s) \psi(s)' ds >0$.

Let $C([0,1]; \mathbb{R}^r)$ be the space of continuous \(\mathbb{R}^r\)-valued functions on the interval \([0,1]\). That is, \[ C([0,1]; \mathbb{R}^r) = \left\{ g : [0,1] \to \mathbb{R}^r \mid g \text{ is continuous} \right\} \] Let $C_f$ be the subspace of $C([0,1]; \mathbb{R}^r)$ that is orthogonal to $\psi$ weighted by $\sigma \in C[0,1]$. That is, for each $\tilde \psi \in C_f$, we have $\int_0^1 \frac 1 {\sigma(s)^2} \tilde \psi(s) \psi'(s) ds =0$. Clearly, $C_f$ is a linear space. If we define $\langle g_1, g_2\rangle =\int_0^1 \frac 1 {\sigma(s)^2} g_1(s)' g_2(s) ds$ for any $g_1,g_2\in C_f$, then $\langle \cdot, \cdot\rangle$ forms an inner product. In particular, $\langle g, g\rangle =0$ if and only if $g=0$.

We consider the local parameter space \[ \mathbb H= \mathbb R \times C_f \] For any $h_1=(\tilde \alpha_1, \tilde \psi_1)\in \mathbb H$ and $h_2=(\tilde \alpha_2, \tilde \psi_2)\in \mathbb H$, define \[ \langle h_1, h_2\rangle_{\mathbb H}= \tilde \alpha_1 \tilde \alpha_2 \gamma + \int_0^1 \frac 1 {\sigma(s)^2} \tilde \psi_1(s)' \tilde \psi_2(s) ds \] where $\gamma>0$ is given in ((ref)), then $\langle \cdot,\cdot\rangle_{\mathbb H}$ is an inner product on $\mathbb H$, and equipped with it, $\mathbb H$ becomes a Hilbert space.

For each $\tilde \psi \in C_f$, we define $\tilde f_t =\tilde \psi(t/T)$, and the local parameter $f_t^0 + \tilde f_t/\sqrt{NT}$, $t=1,2,...,T$. Let $\tilde F =(\tilde f_1,\tilde f_2,...,\tilde f_T)'$ be a $T\times r$ matrix.

theoremUnder Assumptions Theorem 1 and Assumption C, we have \[ \ell(\theta^0+\frac 1 {\sqrt{NT} } \tilde \theta) -\ell(\theta^0) =\Delta_{NT}(\tilde \theta)-\frac 1 2 \mathbb E[\Delta_{NT}(\tilde \theta)]^2 + o_p(1)\] where \begin{align} \Delta_{NT}(\tilde \theta) & = \frac 1 {\sqrt{NT} } \sumiN \lambda_i'\tilde F' D^{-1}\eps_i + \tilde \alpha \, \frac 1 {\sqrt{NT} } \sumiN (L\eps_i)'D^{-1}\eps_i \end{align} \begin{align*} \mathbb E[\Delta_{NT}(\tilde \theta)]^2= & \frac 1 T \mathrm{tr} \Big[\tilde F'D^{-1} \tilde F \Big] + \tilde \alpha^2 \Big[ \frac 1 T \mathrm{tr}(L'D^{-1}LD ) \Big] \end{align*} where $o_p(1)$ is uniform over $\tilde \theta$ such that $|\tilde \alpha|\le M$, $ \frac 1 T \|\tilde F'\tilde F \|=\frac 1 T \sum_{t=1}^T \|\tilde f_t\|^2 \le M$, for any given $M<\infty$.

The proof of Theorem (ref) is in the Appendix. From \[ \frac 1 T \mathrm{tr} (\tilde F' D^{-1}\tilde F) = \frac 1 T \sum_{t=1}^T \frac 1 {\sigma(t/T)^2} \tilde \psi(t/T)'\tilde \psi(t/T) \rightarrow \int_0^1 \frac 1 {\sigma(s)^2} \tilde \psi(s)' \tilde \psi(s) ds \] and by ((ref)), $\frac 1 T \mathrm{tr}(L'D^{-1}LD )\rightarrow \gamma>0$, we have

\[ \mathbb E[\Delta_{NT}(\tilde \theta)]^2 \longrightarrow \tilde \alpha^2 \gamma + \int_0^1 \frac 1 {\sigma(s)^2} \tilde \psi(s)' \tilde \psi(s) ds \] The right hand side is simply $\|h\|_{\mathbb H}^2$ for $h=(\tilde \alpha, \tilde \psi) \in \mathbb H$. That is, \[ \|h\|^2_{\mathbb H} =\langle h, h\rangle= \tilde \alpha^2 \gamma + \int_0^1 \frac 1 {\sigma(s)^2} \tilde \psi(s)' \tilde \psi(s) ds \] We can write \[ \mathbb E[\Delta_{NT}(\tilde \theta)]^2 = \| h\|^2_{\mathbb H} +o(1). \]

Also note that $\Delta_{NT}(\tilde \theta)$ consists of two terms, and each of them is asymptotically normal, and the two terms are also asymptotically independent, thus \[ \Delta_{NT}(\tilde \theta) \stackrel{d}{\longrightarrow} N(0, \|h\|_{\mathbb H}^2). \] Summarizing results, we have

corollaryUnder Assumptions B and C, we have \[ \ell(\theta^0+\frac 1 {\sqrt{NT} } \tilde \theta) -\ell(\theta^0) = \Delta_{NT}(\tilde \theta) -\frac 1 2 \|h\|_{\mathbb H}^2 + o_p(1) \] and $\Delta_{NT}(\tilde \theta) \stackrel{d}{\longrightarrow} N(0, \|h\|_{\mathbb H}^2).$

It is also straightforward to show that the likelihood ratio process is LAN (locally asymptotically normal).

The convolution theorem (vanderVaartWellner1996, chapter 3.11) implies that the efficiency bound for any regular estimators of $\alpha^0$ is $1/\gamma$, which is the inverse of the coefficients of $\tilde \alpha^2$ in $\|h\|_{\mathbb H}^2$ (see the detailed argument for the application of the convolution theorem following Corollary (ref) below). The efficiency bound coincides with the result in Corollary (ref).

As noted by IwakuraOkui2014, however, the parameter sequence $f_t^0+\tilde f_t/\sqrt{N}$ is not regular, the convolution theory is not applicable for deriving the efficiency bound for regular estimators of $f_t^0$. The underlying reason for the non-regularity is that the linear functional, say, $\dot \kappa: C_f \rightarrow \mathbb R$ such that $\dot \kappa(\tilde \psi) =\tilde \psi(x_0)$, for an $x_0\in [0,1]$ is not a continuous functional. The non-continuity is because convergence in the $C_f$ space with the given norm does not imply pointwise convergence. This motivates us to consider the $\ell^2$ norm in the next section, in the hope that the convolution theorem will be applicable for all parameters, including the incidental ones.

The \texorpdfstring{$\ell^2$}{Lg} local parameter space

The previous result is only applicable for low dimensional parameters. To apply the convolution theorem (vanderVaartWellner1996, chapter 3.11) to all parameters including the incidental ones, we consider the second type of local parameter space, which is also used by IwakuraOkui2014 for the fixed effects estimators:

equation[equation omitted — 87 chars of source]

with $\tilde f=(\tilde f_1,\tilde f_2,...)$ being required to be in $\ell_r^2$: \[ \ \ell_r^2 := \Big\{ (\tilde f_t)_{t=1}^\infty \, \Big| \tilde f_t \in \mathbb R^r, \sum_{s=1}^\infty \|\tilde f_s\|^2 <\infty \Big\}. \] For this type of local parameters, we can remove the scale factor $T^{-1/2}$, (cf. ((ref))). Since $\tilde f \in \ell_r^2$, we have, for $\tilde F=(\tilde f_1,\tilde f_2,...,\tilde f_T)'$ (projection of $\tilde f$ on the first $T$ coordinates), \[ \tilde F'\tilde F = \sum_{t=1}^T \tilde f_t \tilde f_t' = O(1). \]

theoremUnder Assumption B, for $\tilde \alpha \in \mathbb R$, $\tilde f\in \ell_r^2$, $\tilde F=(\tilde f_1,\tilde f_2,...,\tilde f_T)'$, as $N,T\rightarrow \infty$, with $N/T^3\rightarrow 0$, we have \[ \ell(\alpha^0+\frac {\tilde \alpha} {\sqrt{NT}}, F^0+\frac 1 {\sqrt{N}} \tilde F) -\ell(\theta^0) =\Delta_{NT}^\dag(\tilde \theta)-\frac 1 2 \mathbb E[\Delta_{NT}^\dag(\tilde \theta)]^2 + o_p(1)\] where \begin{align} \Delta_{NT}^\dag(\tilde \theta) & = \frac 1 {\sqrt{N}} \sumiN \lambda_i'\tilde F'D^{-1}\eps_i \notag \\ & + \tilde \alpha \, \frac 1 {\sqrt{NT} } \sumiN (L\eps_i)'D^{-1}\eps_i \\ & +\tilde \alpha \, \frac 1 {\sqrt{NT} } \sumiN\lambda_i' (LF)'[D^{-1/2} M_{D^{-1/2}F} D^{-1/2}] \eps_i \notag \\ \mathbb E[\Delta_{NT}^\dag(\tilde \theta)]^2 & = \mathrm{tr} \Big[\tilde F'D^{-1} \tilde F \Big] \notag \\ & + \tilde \alpha^2 \Big[ \frac 1 T \mathrm{tr}(L'D^{-1}LD ) \Big] \\ & + \tilde \alpha^2 \mathrm{tr} \Big[ \frac 1 T (LF)' [D^{-1/2} M_{D^{-1/2}F} D^{-1/2} ](LF)\Big] \notag \end{align} where $o_p(1)$ is uniform over $|\tilde \alpha|\le M$, and $\|\tilde F\|\le M$ for any given $M<\infty$.

In comparison with Theorem (ref), Theorem (ref) has simpler expressions, due to the smaller local parameter space. The first two terms in $\Delta_{NT}(\tilde \theta)$ are simplified, with the corresponding simplification in the variance, and in addition, the covariance term in $\mathbb E[\Delta_{NT}(\tilde \theta)]^2$ is dropped. The proof of Theorem (ref) is given in the appendix.

We next establish the local asymptotic normality (LAN) property for the local likelihood ratios.

For each $(\tilde \alpha,\tilde f)=(\tilde \alpha, \tilde f_1,\tilde f_2,....) \in \mathbb R\times \ell_{r}^2$, we introduce a new sequence

equation[equation omitted — 185 chars of source]

so $h_0= \tilde \alpha (\gamma+\nu)^{1/2}$, and $h_s= \frac 1 {\sigma_s} \tilde f_s$ for $s\ge 1$, where $\gamma$, $\nu$, are defined in Assumption B(v), and $\sigma_s^2$ is the variance of $\eps_{is}$. Hence, $h(\tilde \alpha,\tilde f)$ is a scaled version of $(\tilde \alpha,\tilde f)$. By Assumption B(iii), $\min_s \sigma_s^2\ge a >0$, it follows that $h(\tilde \alpha,\tilde f) \in \mathbb H := \mathbb R\times \ell_{r}^2$. For any $h,g\in \mathbb H$, define the inner product, $\langle h,g\rangle = h_0 g_0 + \sum_{s=1}^\infty h_s'g_s$, then $\mathbb H$ is a Hilbert space. Let $\|h\|_{\mathbb H}^2=\langle h, h\rangle$. In particular, for $h=h(\tilde \alpha,\tilde f)$ in ((ref)), we have

equation[equation omitted — 150 chars of source]

Notice $ \mathrm{tr}(\tilde F' D^{-1} \tilde F)=\sum_{s=1}^T\frac 1 {\sigma_s^2} \tilde f_s'\tilde f_s = \sum_{s=1}^\infty \frac 1 {\sigma_s^2} \tilde f_s'\tilde f_s +o(1)$ because the series is convergent. By Assumption B(v), we can write ((ref)) as

equation[equation omitted — 207 chars of source]

where $h =h(\tilde \alpha,\tilde f)$ is given in ((ref)). Next, rewrite ((ref)) as \[ \Delta_{NT}^\dag(\tilde \theta) = \frac 1 {\sqrt{N}} \sumiN \lambda_i'\tilde F'D^{-1}\eps_i +\tilde \alpha ( \Delta_{NT2}+\Delta_{NT3}) \] where $\Delta_{NTj}$ ($j=2,3)$ are defined earlier. The first term \[ \frac 1 {\sqrt{N}} \sumiN \lambda_i'\tilde F'D^{-1}\eps_i =\sumtT \frac 1 {\sigma_t^2} \tilde f_t' \Big(\frac 1 {\sqrt{N}} \sumiN \lambda_i \eps_{it}\Big) \stackrel{d}{\longrightarrow} \mathcal{N}\Big(0, \sum_{t=1}^\infty \tilde f_t'\tilde f_t/\sigma_t^2 \Big) \] because $N^{-1/2}\sumiN \lambda_i \eps_{it} \stackrel{d}{\longrightarrow} \mathcal{N}(0,I_r)$. The LHS above is asymptotically independent of $\Delta_{NT2}+\Delta_{NT3}$ (their covariance being zero, as is shown in the appendix). From \[ \Delta_{NT2}+\Delta_{NT3}\stackrel{d}{\longrightarrow} \mathcal{N}(0, \gamma+\nu ) \] where $\gamma$ and $\nu$ are given in Assumption B, we have \[ \Delta_{NT}^\dag(\tilde \theta) \stackrel{d}{\longrightarrow} \mathcal{N}\left(0, \|h\|_{\mathbb H}^2 \, \right). \]

Moreover, it is not difficult to establish the finite dimensional weak convergence. Let $\tilde \alpha^{j}\in \mathbb R$, $\tilde f^{j}\in \ell_{r}^2$ for $j=1,2,..,q$. Let $h^j=h(\tilde \alpha^j,\tilde f^{j})$ and $\tilde\theta^j=(\tilde \alpha^j, \tilde F^{j})$, for any finite integer $q$, \[ (\Delta_{NT}^\dag(\tilde \theta^1),...,\Delta_{NT}^\dag(\tilde \theta^q))' \stackrel{d}{\longrightarrow} \mathcal{N}\Big(0, ( \langle h^j, h^k\rangle)_{j,k=1}^q \Big). \] Summarizing the above, we have

corollaryUnder the assumption of Theorem (ref), \[ \ell(\alpha^0+\frac {\tilde \alpha} {\sqrt{NT}}, F^0+\frac 1 {\sqrt{N}} \tilde F) -\ell(\theta^0) =\Delta_{NT}^\dag(\tilde \theta) -\frac 1 2 \|h\|_{\mathbb H}^2 +o_p(1) \] and \[ \Delta_{NT}^\dag(\tilde \theta)\stackrel{d}{\longrightarrow} \mathcal{N}\left(0, \|h\|_{\mathbb H}^2\right) \] where $h=h(\tilde \alpha,\tilde f)$ and $\|h\|_{\mathbb H}^2$ are defined in ((ref)) and ((ref)), respectively. Furthermore, the likelihood ratio is locally asymptotically normal (LAN).

Using the convolution theorem for locally asymptotically normal (LAN) experiments, the implied efficiency bound for regular estimators of $\alpha^0$ is $1/(\gamma+\nu)$. The implied efficiency bound for regular estimators of $f_t^{0}$ is $\sigma_t^2 I_{r}$ for each $t$. These bounds are, respectively, the inverse of the coefficient of $\tilde \alpha^2$, and the inverse of the matrix in the quadratic form $\tilde f_t'\tilde f_t/\sigma_t^2 = \tilde f_t'(I_{r}/\sigma_t^2) \tilde f_t$ in the expression for $\|h\|_{\mathbb H}^2$.

To see this, fix $s\in \mathbb N$, with $s\ge 1$. For $h=(h_0, h_1, ..., h_s,...)\in \mathbb H$, consider the parameter sequence, \[ \phi_{NT,s}(h) := f_s^{0} + N^{-1/2} \tilde f_s = f_s^{0}+ N^{-1/2} \sigma_s h_s, \quad \phi_{NT,s}(0) =f_s^{0} \] so $\sqrt{N} [\phi_{NT,s}(h)-\phi_{NT,s}(0)] = \sigma_s h_s$. If we define $\dot \phi_s (h)= \sigma_s h_s$, then \[ \sqrt{N} [\phi_{NT,s}(h)-\phi_{NT,s}(0)] = \dot \phi_s(h). \] Since $\dot \phi_s$ is a coordinate projection map (multiplied by a positive constant $\sigma_s$), it is a continuous linear map, $\dot \phi_s: \mathbb H \rightarrow \mathbb R^{r}$. Its adjoint map $\dot \phi_s^*: \mathbb R^{r} \rightarrow \mathbb H$ (both spaces are self-dual) is the inclusion map (i.e., embedding): $\dot \phi_s^* x= (0,...,0, \sigma_s x, 0,...)\in \mathbb H$, for all $x\in \mathbb R^{r}$. The adjoint map satisfies $\langle \dot \phi_s^* x, h\rangle =\sigma_s x' h_s = x' \dot \phi_s(h)=\langle x, \dot \phi_s(h)\rangle$. Let $Z$ denote the limiting distribution of efficient estimators of $f_s^{0}$. Theorem 3.11.2 in van der Vaart and Wellner (vanderVaartWellner1996, p.414) show that $x'Z\sim \mathcal{N}(0, \|\dot \phi_s^*x\|_{\mathbb H}^2)$ for all $x\in \mathbb R^{r}$. But $\|\dot \phi_s^*x\|_{\mathbb H}^2 = \sigma_s^2 x'x$. It follows that $Z \sim \mathcal{N}(0,\sigma_s^2 I_{r})$. Thus the efficiency bound for regular estimators of $f_s^{0}$ is $\sigma_s^2 I_{r}$.

For $s=0$, the same argument shows that the efficiency bound for regular estimators of $\alpha^0$ is $1/(\gamma+\nu)$. In summary, we have

corollaryUnder the assumptions of Theorem (ref), the asymptotic efficiency bound for regular estimators of $\alpha^0$ is $1/(\gamma+\nu)$, and the efficiency bound for regular estimators of $f_t^{0}$ is $\sigma_t^2 I_{r}$.

It can be shown that the efficiency bound $1/(\gamma+\nu)$ corresponds to the case in which the incidental parameters $F^0=(f_1^0,f_2^0,...,f_T^0)'$ are known, thus the implied efficiency bound is likely too low to be attainable. The implication is that the $\ell^2$ perturbation is “too small”. Intuitively, the smaller the local parameter space, the lower the efficiency bound, making it harder to achieve the implied bound. Interestingly, however, the $\ell^2$ local parameter space is appropriately sized for the fixed effects estimator, where both loadings and factors are estimated, as shown by IwakuraOkui2014.

However, if $ F^0$ satisfies Assumption C, as in Theorem (ref), then Lemma (ref) in the appendix shows that $ \nu = 0$. As a result, the efficiency bound simplifies to $1/\gamma$. This means that $\ell^2$ is a suitable local parameter space for the maximum likelihood estimator when combined with Assumption C. In this case, the efficiency bound is attainable by the QMLE.

commentWhen the same model is estimated by the fixed effects method (that is, the $\lambda_i$'s are also treated as parameters), Iwakura and Okui IwakuraOkui2014 show that the $\ell^2$ perturbation is a suitable choice, and that the corresponding efficiency bound for $\alpha^0$ is $1/\gamma$ (the authors confined their analysis to the homoskedastic case so $1/\gamma=1-\alpha^2$). To the method of QMLE, however, there is no sufficient variation in the $\ell^2$ perturbation, thus implying a smaller bound. Which is to say that QMLE is a more efficient estimation procedure than the fixed effects approach. This finding is consistent with the result that even under fixed $T$, QMLE provides a consistent estimator for $\alpha^0$, but fixed effects estimator is not consistent, see Moreira2009 and Bai2013. To recap, the $\ell^2$ local parameter space is “too small” for the QMLE, but is the suitable local parameter space for the fixed effects approach. This implies that, as explained earlier, QMLE is a better procedure than the fixed effects method. By analyzing the local likelihood ratios, we are able to inform the merits of different estimators that are otherwise hard to discern based on the usual asymptotics alone (e.g., limiting distributions).

Conclusion

We derive the efficiency bound under normality for estimating dynamic panel models with interactive effects by formulating the model as a simultaneous equations system. We show that quasi-maximum likelihood method applied to the system attains the efficiency bound. These results are obtained under an increasing number of incidental parameters.

appendix\section*{Proof of Results} We first consider the local parameter space for $\tilde f \in \ell_r^\infty$. Let $G=F+\frac 1 {\sqrt{NT} } \tilde F$. We drop the superscript 0 associated with true parameters to make the notation less cumbersome. When evaluated at the true parameters, the likelihood function is, see ((ref)) \begin{equation} \ell(\theta^0) =-\frac N 2 \log|FF'+D|-\frac 1 2 \sumiN (y_i-\bar y)'B'(FF'+D)^{-1} B(y_i-\bar y) \end{equation} When evaluated at the local parameters \begin{align*} \ell(\theta^0+\frac 1 {\sqrt{NT} } \tilde \theta) & = -\frac N 2 \log|GG'+D| \\ & -\frac 1 2 \sumiN (y_i-\bar y)'(B -\tilde \alpha \frac 1 {\sqrt{NT} } J)' (GG'+D)^{-1} (B -\tilde \alpha \frac 1 {\sqrt{NT} } J) (y_i-\bar y) \\ & = -\frac N 2 \log|GG'+D|-\frac 1 2 \sumiN (y_i-\bar y)'B'(GG'+D)^{-1} B (y_i-\bar y)\\ & + \tilde \alpha \frac 1 {\sqrt{NT} } \sumiN (y_i-\bar y)'J'(GG'+D)^{-1} B(y_i-\bar y)\\ & -\frac 1 2 \tilde \alpha^2 \frac 1 {NT} \sumiN (y_i-\bar y)'J'(GG'+D)^{-1} J (y_i-\bar y) \end{align*} Thus, the difference is \begin{align} \ell(\theta^0+\frac 1 {\sqrt{NT} } \tilde \theta)& -\ell(\theta^0) = \notag\\ - & \frac N 2 \Big[\log|GG'+D|-\log|FF'+D| \Big] \notag \\ - & \frac 1 2\sumiN (y_i-\bar y)'B' \Big[(GG'+D)^{-1}-(FF'+D)^{-1}\Big] B (y_i-\bar y) \\ + & \tilde \alpha \frac 1 {\sqrt{NT} } \sumiN (y_i-\bar y)'J'(GG'+D)^{-1} B(y_i-\bar y) \notag \\ - & \frac 1 2 \tilde \alpha^2 \frac 1 {NT} \sumiN (y_i-\bar y)'J'(GG'+D)^{-1} J (y_i-\bar y) \notag \end{align} Notice that $ B(y_i-\bar y)= F(\lambda_i-\bar \lambda) + \eps_i-\bar \eps$, where $\bar \lambda =\frac 1 N \sumiN \lambda_i$ and $\bar \eps=\frac 1 N \sumiN \eps_i$. Furthermore, from $y_i-\bar y= B^{-1} F(\lambda_i-\bar \lambda) + B^{-1}( \eps_i-\bar \eps)$, right multiply by $J$ and notice $J B^{-1}= L$ (see equation ((ref))), we have $J(y_i-\bar y) =L F(\lambda_i-\bar \lambda) +L( \eps_i-\bar \eps)$. We also write $J(y_i-\bar y) = y_{i,-1} -\bar y_{-1},$ where $y_{i,-1}=(0, y_{i1},...,y_{i,T-1})'$ and $ \bar y_{-1}=(0, \bar y_1,...,\bar y_{T-1})'$. Thus we can rewrite the log-likelihood ratio as \begin{align} \ell(\theta^0+\frac 1 {\sqrt{NT} } \tilde \theta)& -\ell(\theta^0) = \notag\\ - & \frac N 2 \Big[\log|GG'+D|-\log|FF'+D| \Big] \notag \\ - & \frac 1 2\sumiN \Big[ F(\lambda_i-\bar \lambda) + \eps_i-\bar \eps\Big]' \Big[(GG'+D)^{-1}-(FF'+D)^{-1}\Big] \Big[ F(\lambda_i-\bar \lambda) + \eps_i-\bar \eps\Big] \\ + & \tilde \alpha \frac 1 {\sqrt{NT} } \sumiN (y_{i,-1}-\bar y_{-1})'(GG'+D)^{-1} \Big[ F(\lambda_i-\bar \lambda) + \eps_i-\bar \eps\Big] \notag \\ - & \frac 1 2 \tilde \alpha^2 \frac 1 {NT} \sumiN (y_{i,-1}-\bar y_{-1})'(GG'+D)^{-1} (y_{i,-1}-\bar y_{-1}). \notag \end{align} We shall derive the limit of the log likelihood ratio. Throughout, we use the matrix inversion formula (Woodbury formula) \[ (FF'+D)^{-1}=D^{-1} -D^{-1}F(I_r+F'D^{-1} F)^{-1} F'D^{-1} \] \[ (FF'+D)^{-1}F = D^{-1}F(I_r + F'D^{-1}F)^{-1} \] and the matrix determinant result \[ |FF'+D|=|D||I_r +F'D^{-1}F| \] From now on, we assume $r=1$ to simplify the derivation. We define $\omega_F^2$ and $\eta_F^2$ as \[ \omega_F^2=\frac 1 T F'D^{-1} F, \quad \eta_F^2 =\frac 1 T (1+F'D^{-1} F) = \frac 1 T+ w_F^2 \] \begin{lemma} For $G= F +\frac 1 {\sqrt{NT} } \tilde F$, \begin{equation} -\frac N 2 \Big[\log|GG'+D|-\log|FF'+D| \Big] = -\sqrt{\frac N T} (F'D^{-1}\tilde F/T) \frac 1 {\eta_F^2} + O(\frac 1 T) \end{equation} \end{lemma} Proof: Notice \begin{equation*} \begin{split} |GG'+D| = & |(F+\frac 1 {\sqrt{NT} } \tilde F)(F+\frac 1 {\sqrt{NT} } \tilde F)'+D| \\ = & |D| [1+ (F+\frac 1 {\sqrt{NT} } \tilde F)'D^{-1}(F+\frac 1 {\sqrt{NT} } \tilde F)] \\ = & |D|\Big( 1+F'D^{-1}F + 2 \frac 1 {\sqrt{NT} } F'D^{-1}\tilde F + \frac 1 {NT} \tilde F'D^{-1}\tilde F \Big) \\ = & |D| (1+F'D^{-1}F) \Big[ 1+ 2 \frac 1 T \frac 1 {\sqrt{NT} } F'D^{-1}\tilde F \frac 1 {\eta_F^2} + \frac 1 {NT^2} \tilde F'D^{-1}\tilde F \frac 1 {\eta_F^2} \Big] \\ \end{split} \end{equation*} From $\log(1+x)= x +O(x^2)$ for small $x$ (big $O$) \[ \log|GG'+D|-\log|FF'+D| = \log\Big[ 1+ 2 \frac 1 T \frac 1 {\sqrt{NT} } F'D^{-1}\tilde F/\eta_F^2 + \frac 1 {NT^2} \tilde F'D^{-1}\tilde F/\eta_F^2 \Big] \] \[ =2 \frac 1 T \frac 1 {\sqrt{NT} } F'D^{-1}\tilde F/\eta_F^2 + \frac 1 {NT^2} \tilde F'D^{-1}\tilde F/\eta_F^2 +O(\frac 1 {NT})\] where $O(x^2)=O(1/(NT))$. Thus \begin{align*} & -\frac N 2 \Big[\log|GG'+D|-\log|FF'+D| \Big] \\ & = - \frac N T \frac 1 {\sqrt{NT} } F'D^{-1}\tilde F/\eta_F^2 - \frac 1 {2T^2} \tilde F'D^{-1}\tilde F/\eta_F^2 + O(\frac 1 T) \end{align*} The second term on the right is also $O(1/T)$. This proves Lemma (ref). $\Box$ \begin{lemma} Let $H:= (1+G'D^{-1}G)^{-1} -(1+F'D^{-1}F)^{-1}$. Then \begin{align} H = &- 2 \frac 1 {T^2} \frac 1 {\sqrt{NT} } F'D^{-1}\tilde F \frac {1}{\eta_F^4}-\frac 1 {NT^3} \tilde F'D^{-1}\tilde F \frac {1}{\eta_F^4} \\ & +4(\frac{F'D^{-1}\tilde F} T)^2 \frac {1} {NT^2\eta_F^6} + O(\frac 1 {N^{3/2}T^{5/2}}) \nonumber \end{align} \end{lemma} Proof: \begin{align*} (1+G'D^{-1}G)& = \Big(1+F'D^{-1}F + 2 \frac 1 {\sqrt{NT} } F'D^{-1}\tilde F + \frac 1 {NT} \tilde F'D^{-1}\tilde F \Big) \\ & =(1+F'D^{-1}F) \Big( 1+ \frac{ 2 \frac 1 {\sqrt{NT} } F'D^{-1}\tilde F + \frac 1 {NT} \tilde F'D^{-1}\tilde F }{1+F'D^{-1}F} \Big) \\ (1+G'D^{-1}G)^{-1} & =(1+F'D^{-1}F)^{-1} \Big( 1+ \frac{ 2 \frac 1 {\sqrt{NT} } F'D^{-1}\tilde F + \frac 1 {NT} \tilde F'D^{-1}\tilde F }{1+F'D^{-1}F} \Big)^{-1} \end{align*} Let $A= 2 \frac 1 {\sqrt{NT} } F'D^{-1}\tilde F + \frac 1 {NT} \tilde F'D^{-1}\tilde F$ and use the expansion $1/(1+x) =1-x +x^2 +O(x^3)$ for small $x$, we have \begin{align*} & (1+G'D^{-1}G)^{-1}\\ & = (1+F'D^{-1}F)^{-1} \Big( 1 -\frac A {1+F'D^{-1}F} +\frac {A^2}{(1+F'D^{-1}F)^2} + O(\frac { A^3} {(1+F'D^{-1}F)^3}) \Big) \\ & =(1+F'D^{-1}F)^{-1} -\frac A {(1+F'D^{-1}F)^2} +\frac {A^2}{(1+F'D^{-1}F)^3 } + O(\frac { A^3} {(1+F'D^{-1}F)^4}) \\ \end{align*} Now \begin{align*} -\frac A {(1+F'D^{-1}F)^2} & = - 2 \frac 1 {T^2} \frac 1 {\sqrt{NT} } F'D^{-1}\tilde F \frac {1}{\eta_F^4}-\frac 1 {NT^3} \tilde F'D^{-1}\tilde F \frac {1}{\eta_F^4} \\ \frac {A^2}{(1+F'D^{-1}F)^3 } & =\frac {A^2}{T^3 \eta_F^6} = 4(\frac{F'D^{-1}\tilde F} T)^2 \frac {1} {NT^2\eta_F^6} +O(\frac 1 {N^{3/2} T^{5/2}}) \\ \frac {A^3}{(1+F'D^{-1}F)^4 } & = O(\frac 1 {N^{3/2}T^{5/2}}) \end{align*} This proves the lemma. $\Box$. A remark is in order. In the above analysis, if we write $A=a+b$, then $A^2 =a^2 +2 ab +b^2$. Only $a^2/(T^3\eta_F^6)$ is non-negligible, $ab$ and $b^2$ are of smaller order of magnitude. Finally, $A^3/(T^4\eta_F^8)$ is also a smaller magnitude. \begin{lemma} The following $T\times T$ matrix $\Xi$ satisfies \[ \Xi:=(GG'+D)^{-1}- (FF'+D)^{-1} = -\Xi_a-\Xi_b-\Xi_c-\Xi_d +R\] where \begin{align} &\Xi_a= H D^{-1} FF'D^{-1} \\ & \Xi_b =\Big[ \frac 1 {\sqrt{NT} } \Big( \frac 1 {T \eta_F^2} \Big)-2 \frac 1 {NT^3}F'D^{-1} \tilde F \frac 1 {\eta_F^4} \Big] D^{-1} F \tilde F'D^{-1} \\ & \Xi_c =\Big[ \frac 1 {\sqrt{NT} } \Big( \frac 1 {T\eta_F^2} \Big)-2 \frac 1 {NT^3}F'D^{-1} \tilde F \frac 1{\eta_F^4} \Big] D^{-1} \tilde F F'D^{-1} \\ &\Xi_d= \frac 1 {NT^2} \frac 1{\eta_F^2}D^{-1} \tilde F\tilde F'D^{-1} \end{align} where $H$ is defined in Lemma (ref), and $R$ satisfies $\|R\|_2 =O(1/(NT)^{3/2})$. \end{lemma} Proof: From the Woodbury formula \[ -\Xi= D^{-1}G(1+ G'D^{-1}G)^{-1} G'D^{-1}-D^{-1}F(1+F'D^{-1}F)^{-1}F'D^{-1} \] By Lemma (ref), we can write \begin{equation} G(1+G'D^{-1}G)^{-1}G'=(1+F'D^{-1}F)^{-1} GG' + H GG' \end{equation} where $H$ is defined in Lemma (ref). The first term on the right hand side above is \begin{align*} (1+F'D^{-1}F)^{-1}GG' &=(1+F'D^{-1}F)^{-1} \Big(FF'+\frac 1 {\sqrt{NT} } F\tilde F' + \frac 1 {\sqrt{NT} } \tilde F F' +\frac 1 {NT} \tilde F\tilde F'\Big). \end{align*} Thus, pre- and post-multiplying by $D^{-1}$ \begin{align} (1+ & F'D^{-1}F)^{-1} D^{-1} GG D^{-1}- (1+F'D^{-1}F)^{-1} D^{-1} FF' D^{-1} \nonumber \\ &= (1+F'D^{-1}F)^{-1} \Big[\frac 1 {\sqrt{NT} } D^{-1} F\tilde F' D^{-1} + \frac 1 {\sqrt{NT} } D^{-1} \tilde F F' D^{-1} +\frac 1 {NT} D^{-1} \tilde F\tilde F' D^{-1} \Big] \nonumber\\ &=\frac 1 { T\eta_F^2} \frac 1 {\sqrt{NT} } D^{-1} F\tilde F' D^{-1} + \frac 1 { T\eta_F^2} \frac 1 {\sqrt{NT} } D^{-1} \tilde F F' D^{-1} +\frac 1 {NT^2\eta_F^2 } D^{-1} \tilde F\tilde F' D^{-1} \end{align} Next, consider $ H GG'$ in ((ref)), using $GG'=FF'+\frac 1 {\sqrt{NT} } F \tilde F'+\frac 1 {\sqrt{NT} } \tilde F F'+\frac 1 {NT} \tilde F\tilde F' $, \begin{equation} H GG' = H FF'+ H \frac 1 {\sqrt{NT} } F \tilde F' + H \frac 1 {\sqrt{NT} } \tilde F F' + H \frac 1 {NT} \tilde F \tilde F' \end{equation} where $ H $ is given in ((ref)). All of the four terms in $H$ are non-negligible for the matrix $HFF'$; only the first term in $H$ is non-negligible for $H \frac 1 {\sqrt{NT} } F\tilde F'$ and $H \frac 1 {\sqrt{NT} } \tilde F F'$. More specifically, notice \[ H = - 2 \frac 1 {T^2} \frac 1 {\sqrt{NT} } F'D^{-1}\tilde F \frac {1}{\eta_F^4} + O(\frac 1 {NT^2}) \] Left multiplying $H$ by $\frac 1 {\sqrt{NT} } F \tilde F'$, \begin{align*} H \frac 1 {\sqrt{NT} } F \tilde F' & = - 2 \frac 1 {NT^3} F'D^{-1}\tilde F \frac {1}{\eta_F^4} F \tilde F' + O(\frac 1 {NT^2}) \frac 1 {\sqrt{NT} } F \tilde F' \\ & = - 2 \frac 1 {NT^3} F'D^{-1}\tilde F \frac {1}{\eta_F^4} F \tilde F' + R_1 \end{align*} where $\|R_1||_2= O(1/( NT)^{3/2})$ because the spectrum norm of $F\tilde F'$ is $O(T)$. Similarly, \begin{align*} H \frac 1 {\sqrt{NT} } \tilde F F' & = - 2 \frac 1 {NT^3} F'D^{-1}\tilde F \frac {1}{\eta_F^4} \tilde F F' + R_2 \end{align*} where $\|R_2||_2= O(1/( NT)^{3/2})$. Finally write \[ R_3 := H \frac 1 {NT} \tilde F \tilde F' \] then $\|R_3 \|_2 = O(1/( NT)^{3/2})$ because $H=O(N^{-1/2} T^{-3/2})$, and $\|\tilde F\tilde F'\|_2=O(T)$. Pre- and post-multiply both sides of ((ref)) by $D^{-1}$ to obtain \begin{align} H D^{-1} GG'D^{-1} & =H D^{-1} F F' D^{-1} \notag \\ & - 2 \Big(\frac 1 {NT^3} F'D^{-1}\tilde F \frac {1}{\eta_F^4}\Big) D^{-1} F \tilde F' D^{-1} \\ & - 2 \Big(\frac 1 {NT^3} F'D^{-1}\tilde F \frac {1}{\eta_F^4}\Big) D^{-1} \tilde F F' D^{-1} \notag\\ & +D^{-1}(R_1+R_2 +R_3)D^{-1} \notag \end{align} Let $R=D^{-1}(R_1+R_2 +R_3)D^{-1} $, then $\|R\|_2=\|R_j\|_2= O(1/( NT)^{3/2})$ for $j=1,2,3$ since $\|D^{-1}\|_2= O_p(1)$. The sum of ((ref)) and ((ref)) is equal to $-\Xi$. This proves the lemma. $\Box$. \begin{lemma} \begin{align*} -\frac 1 2 \sumiN & [F(\lambda_i-\bar \lambda)]'\Big[(GG'+D)^{-1}-(FF'+D)^{-1}\Big] F(\lambda_i-\bar \lambda) \\ & =\sqrt{\frac N T} (F'D^{-1}\tilde F/T) \frac 1{\eta_F^2} - \frac 1 2 \frac 1 T \tilde F' (FF'+D)^{-1} \tilde F \\ & +O_p( N^{1/2} {T^{-3/2}}) + O_p(1/T) +O_p( N^{-1/2}). \end{align*} \end{lemma} Note the first term has an opposite sign with ((ref)). Proof: By Lemma (ref), \begin{equation}\begin{split} \sumiN [F(\lambda_i-\bar \lambda)]' & \Xi_a F(\lambda_i-\bar \lambda) = H \sumiN [F(\lambda_i-\bar \lambda)]' D^{-1}FF'D^{-1} F(\lambda_i-\bar \lambda)\\ = & H (T^2 \omega_F^4) \sumiN (\lambda_i-\bar \lambda)^2 \\ = & \bigg[ -2\sqrt{NT}(F'D^{-1}\tilde F/T) \frac {\omega_F^4}{\eta_F^4}\\ & -(\tilde F'D^{-1}\tilde F/T) \frac {\omega_F^4}{\eta_F^4} \\ & + 4(F'D^{-1}\tilde F/T)^2 \frac {\omega_F^4}{\eta_F^6}\bigg] \frac 1 N \sumiN (\lambda_i-\bar \lambda)^2 +O_p(\frac 1 {\sqrt{NT} } ), \end{split} \end{equation} and \begin{align*} \sumiN & [F(\lambda_i-\bar \lambda)]'(\Xi_b +\Xi_c)F(\lambda_i-\bar \lambda) \\ &= \left[ 2\sqrt{NT}(\tilde F'D^{-1}F/T) \frac {\omega_F^2}{\eta_F^2} - 4 (F'D^{-1}\tilde F/T)^2 \frac {\omega_F^2}{\eta_F^4} \right ] \frac 1 N \sumiN (\lambda_i-\bar \lambda)^2. \end{align*} Adding the two expressions \begin{align*} \sumiN & [F(\lambda_i-\bar \lambda)]' (\Xi_a+\Xi_b+\Xi_c) F(\lambda_i-\bar \lambda)\\ & = \left[ 2\sqrt{NT}(F'D^{-1}\tilde F/T) [ \frac {\omega_F^2}{\eta_F^2}- \frac {\omega_F^4}{\eta_F^4} ] -(\tilde F'D^{-1}\tilde F/T)\frac {\omega_F^4}{\eta_F^4}\right] \frac 1 N \sumiN (\lambda_i-\bar \lambda)^2\\ & + \left[4(F'D^{-1}\tilde F/T)^2 \Big(\frac {\omega_F^4}{\eta_F^6}- \frac {\omega_F^2}{\eta_F^4}\Big) \right] \frac 1 N \sumiN (\lambda_i-\bar \lambda)^2. \end{align*} Note \[ \frac {\omega_F^2}{\eta_F^2}- \frac {\omega_F^4}{\eta_F^4}\equiv \frac {\omega_F^2}{\eta_F^4}\frac 1 T = \frac 1 {T \eta_F^2} + O(\frac 1 {T^2}). \] Hence \[ 2\sqrt{NT}(F'D^{-1}\tilde F/T) [ \frac {\omega_F^2}{\eta_F^2}- \frac {\omega_F^4}{\eta_F^4} ] = 2\sqrt{N/T}(F'D^{-1}\tilde F/T)\frac 1 {\eta_F^2} + O_p(N^{1/2} T^{-3/2}). \] Using the fact that $\omega_F^j/\eta_F^k= 1+O(1/T)$ for $j,k=2,4,6$, \[ (\tilde F'D^{-1}\tilde F/T)\frac {\omega_F^4}{\eta_F^4}= (\tilde F'D^{-1}\tilde F/T) + O(1/T), \] \[ 4(F'D^{-1}\tilde F/T)^2 \Big(\frac {\omega_F^4}{\eta_F^6}- \frac {\omega_F^2}{\eta_F^4}\Big) = O(1/T). \] Combined with $\frac 1 N \sumiN (\lambda_i-\bar \lambda)^2= 1+O_p(N^{-1/2})$, we obtain \begin{align} \sumiN & [F(\lambda_i-\bar \lambda)]' (\Xi_a+\Xi_b+\Xi_c) F(\lambda_i-\bar \lambda) \nonumber \\ & =2\sqrt{N/T}(F'D^{-1}\tilde F/T)\frac 1 {\eta_F^2} -(\tilde F'D^{-1}\tilde F/T) \\ & + O_p(N^{1/2} T^{-3/2}) +O_p(\frac 1 {\sqrt{NT} } ) +O(\frac 1 T). \nonumber \end{align} Next, \begin{align} \sumiN & [F(\lambda_i-\bar \lambda)]' \Xi_d F(\lambda_i-\bar \lambda) \nonumber \\ & =(F'D^{-1}\tilde F/T)^2 \frac 1 {\eta_F^2} \frac 1 N \sumiN (\lambda_i-\bar \lambda)^2 = (F'D^{-1}\tilde F/T)^2 \frac 1 {\eta_F^2} + O_p(N^{-1/2}). \end{align} By summing ((ref)) and ((ref)) and multiplying by 1/2, we obtain \begin{align*} -\frac 1 2 \sumiN & [F(\lambda_i-\bar \lambda)]' \Big[(GG'+D)^{-1}-(FF'+D)^{-1}\Big] F(\lambda_i-\bar \lambda) \\ & =\sqrt{\frac N T} (F'D^{-1}\tilde F/T) \frac 1{\eta_F^2} -\frac 1 2 (\tilde F'D^{-1}\tilde F/T) + \frac 1 2 (F'D^{-1}\tilde F/T)^2 \frac 1 { \eta_F^2} \\ & + O_p(N^{1/2} T^{-3/2}) +O(1/T)+O_p(N^{-1/2}) \end{align*} The lemma is obtained by noting \begin{align*} (\tilde F'D^{-1}\tilde F/T) & -(F'D^{-1}\tilde F/T)^2 \frac 1 { \eta_F^2}=\frac 1 T \tilde F'(FF'+D)^{-1} \tilde F. \end{align*} $\Box$. \begin{lemma} \begin{equation} -\frac 1 2 \sumiN (\eps_i-\bar \eps)' \Big[(GG'+D)^{-1}-(FF'+D)^{-1}\Big] (\eps_i-\bar \eps) =o_p(1) \end{equation} \end{lemma} Proof: First note that $\bar \eps =\frac 1 N \sum_{k=1}^N \eps_k$. It can be shown that $\bar \eps$ is a dominated term and can be ignored. By the notation of Lemma (ref), we evaluate $\sumiN \eps_i'\Xi \eps_i$. Note that $\Xi$ consists of four parts. For this lemma, it is sufficient to approximate $\Xi$ by $\Xi_1+\Xi_2+\Xi_3$, where \begin{align} \Xi_1 &= \Big(2 \frac 1 {T^2} \frac 1 {\sqrt{NT} } F'D^{-1}\tilde F \frac 1 {\eta_F^4}\Big) D^{-1}FF'D^{-1} \nonumber \\ \Xi_2 & = - \Big( \frac 1 {\sqrt{NT} } \frac 1 T \frac 1 {\eta_F^2}\Big) D^{-1}F\tilde F'D^{-1}\\ \Xi_3 & = - \Big( \frac 1 {\sqrt{NT} } \frac 1 T \frac 1 {\eta_F^2} \Big)D^{-1} \tilde F F'D^{-1} \nonumber \end{align} In the above approximation, we kept the first term of $H$ in ((ref)) ($H$ has four terms), and kept the very first term inside the brackets in ((ref)) and ((ref)). All other terms are negligible in the evaluation of $\sumiN \eps_i'\Xi\eps_i$. Using trace, it is easy to obtain the expected values \[ \mathbb E\sumiN \eps_i'\Xi_1\eps_i= 2 \sqrt{\frac N T} (F'D^{-1}\tilde F/T) \frac 1{\eta_F^2} \] And \[ \mathbb E\sumiN \eps_i'(\Xi_2+\Xi_3)\eps_i= - 2 \sqrt{\frac N T} (F'D^{-1}\tilde F/T) \frac 1{\eta_F^2} \] Thus the sum of the expected values is zero. In addition, the deviation of each term from its expected value is $o_p(1)$. This is because the variance of each term is $O(1/T)=o(1)$. For example, $ \var(\sumiN \eps_i'\Xi_2\eps_i)=\mathrm{tr}(\Xi_2 D \Xi_2 D)+\mathrm{tr}(\Xi_2'D \Xi D)=\frac 1 T (\frac 1 T \tilde F'D^{-1}F)^2 \frac 1 {\eta_F^4} + \frac 1 T (\frac 1 T \tilde F'D^{-1}F)(\frac 1 T F'D^{-1}F) \frac 1 {\eta_F^4} =O(1/T) $. We have used the fact that for normal $\eps_i\sim N(0,D)$, $\var (\eps_i'A\eps_i)=\mathrm{tr} (A D A D)+\mathrm{tr}(A'D A D)$ for any $A$, This proves the lemma. $\Box$. \begin{lemma} \begin{align*} \sumiN [F(\lambda_i-\bar \lambda)]' & \Big[(GG'+D)^{-1}-(FF'+D)^{-1}\Big] (\eps_i-\bar \eps)\\ & = - \frac 1 {\sqrt{NT} } \sumiN \lambda_i'\tilde F'(FF'+D)^{-1} \eps_i +o_p(1) \end{align*} \end{lemma} {\bf Proof:} Recall $\Xi=(GG'+D)^{-1}-(FF'+D)^{-1}$. The preceding approximation of $\Xi$ by $\Xi_1+\Xi_2+\Xi_3$ in ((ref)) is sufficient (other terms are negligible). We evaluate $\sumiN (F\lambda_i)'\Xi_k \eps_i$ $(k=1,2,3)$. Here we ignore $\bar \lambda$ and $\bar \eps$, the associated terms are negligible. \begin{align*} \sumiN (F\lambda_i)'\Xi_1 \eps_i & = 2 (F'D^{-1}\tilde F/T)(F'D^{-1}F/T)\frac 1 {\sqrt{NT} } \sumiN F'D^{-1}\eps_i \lambda_i \frac 1 {\eta_F^4} \\ & =2 \frac {\omega_F^2} {\eta_F^4} (F'D^{-1}\tilde F/T) \frac 1 {\sqrt{NT} } \sumiN F'D^{-1}\eps_i \lambda_i\\ & =2 \frac 1 {\eta_F^2} (F'D^{-1}\tilde F/T) \frac 1 {\sqrt{NT} } \sumiN F'D^{-1}\eps_i \lambda_i +O_p(1/T) \end{align*} \begin{align*} \sumiN (F\lambda_i)'\Xi_2 \, \eps_i & = -\frac 1 {\eta_F^2} (F'D^{-1}F/T) \frac 1 {\sqrt{NT} } \sumiN \tilde F'D^{-1}\eps_i \lambda_i \\ & =- \frac 1 {\sqrt{NT} } \sumiN \tilde F'D^{-1}\eps_i \lambda_i+O_p(1/T) \end{align*} and \[ \sumiN (F\lambda_i)'\Xi_3 \, \eps_i= - \frac 1 {\eta_F^2} (F'D^{-1}\tilde F/T) \frac 1 {\sqrt{NT} } \sumiN F'D^{-1}\eps_i \lambda_i\] Combining the three expressions \begin{align*} \sumiN (F\lambda_i)'\Xi \eps_i & =- \frac 1 {\sqrt{NT} } \sumiN \lambda_i' \tilde F'D^{-1}\eps_i +\frac 1 {\eta_F^2} (F'D^{-1}\tilde F/T) \frac 1 {\sqrt{NT} } \sumiN F'D^{-1}\eps_i\lambda_i +O_p(1/T) \\ & = -\frac 1 {\sqrt{NT} } \sumiN \lambda_i'\tilde F'(FF'+D)^{-1} \eps_i +O_p(1/T). \end{align*} This proves the lemma. $\Box$ \begin{corollary} Under the Assumptions of Theorem 1, \begin{align*} -\frac 1 2 \sumiN [F(\lambda_i-\bar \lambda) +\eps_i-\bar \eps] & \Big[(GG'+D)^{-1}-(FF'+D)^{-1}\Big] [F(\lambda_i-\bar \lambda) +\eps_i-\bar \eps]\\ & = \sqrt{\frac N T} (F'D^{-1}\tilde F/T) \frac 1{\eta_F^2}\\ & + \frac 1 {\sqrt{NT} } \sumiN \lambda_i'\tilde F'(FF'+D)^{-1} \eps_i \\ & -\frac 1 2 \frac 1 T \tilde F' (FF'+D)^{-1} \tilde F +o_p(1) \end{align*} \end{corollary} Proof: This follows by combining the results of Lemmas (ref), (ref), and (ref). $\Box$ \begin{lemma} \begin{align} & \frac 1 {\sqrt{NT} } \sumiN (y_{i,-1}-\bar y_{-1})' (FF'+D)^{-1} [F(\lambda_i-\bar \lambda) +\eps_i-\bar \eps] \notag \\ & = \frac 1 {\sqrt{NT} } \sumiN (L\eps_i)'D^{-1}\eps_i + \frac 1 {\sqrt{NT} } \sumiN\lambda_i' (LF)'[(FF'+D)^{-1} ] \eps_i +o_p(1) \end{align} \end{lemma} Proof: Using $ y_{i,-1}-\bar y_{-1} =LF(\lambda_i-\bar \lambda) +L (\eps_i-\eps)$, we have \begin{align*} LHS &= \frac 1 {\sqrt{NT} } \sumiN[ LF(\lambda_i-\bar \lambda) +L (\eps_i-\eps)]'(FF'+ D)^{-1} F(\lambda_i-\bar \lambda) \\ & +\frac 1 {\sqrt{NT} } \sumiN[ LF(\lambda_i-\bar \lambda) +L (\eps_i-\eps)]' (FF'+D)^{-1}(\eps_i-\bar \eps) \\ & :=a+b \end{align*} where $a$ is defined as the frist term, $b$ as the second term. From the formula, \[ (FF'+D)^{-1} F= D^{-1} F ( {1+F'D^{-1}F})^{-1}=D^{-1}F (1+T\omega_F^2)^{-1} \] term $a$ equals \begin{align*} a &= \frac {\sqrt{N/T}} {1+T\omega_F^2} [\frac 1 N \sumiN (\lambda_i-\bar \lambda)^2] (F'L'D^{-1} F) + \frac 1 {1+T\omega_F^2} \frac 1 {\sqrt{NT} } \sumiN (\eps_i-\bar \eps)'L'D^{-1} F (\lambda_i-\bar \lambda) \\ & = \frac {\sqrt{N/T}} {1+T\omega_F^2} (F'L'D^{-1} F) +O_p(N^{-1/2})+O_p(1/T) \end{align*} where we used $\frac 1 N \sumiN (\lambda_i-\bar \lambda)^2=1+O_p(N^{-1/2})$, and the second term is $O_p(1/T)$. For term $b$, we shall ignore $\bar \lambda$ and $\bar \eps$ (the associated terms are negligible). By the Woodbury formula, \begin{align*} b & = \frac 1 {\sqrt{NT} } \sumiN (LF \lambda_i)' D^{-1} \eps_i + \frac 1 {\sqrt{NT} } \sumiN (L\eps_i)'D^{-1} \eps_i \\ & - \frac 1 {\sqrt{NT} } \sumiN \frac {\lambda_i' F'L'D^{-1}F}{(1+F'D^{-1}F)} F'D^{-1}\eps_i-\frac 1 {(1+F'D^{-1}F)} \frac 1 {\sqrt{NT} } \sumiN \eps_i' L' D^{-1}F F'D^{-1}\eps_i \end{align*} Note the expected value of the last term in the preceding equation is \[ -\frac {F'L'D^{-1} F}{1+F'D^{-1}F} \sqrt{\frac N T} \] and the deviation from its expected value is negligible (because its variance is $O(1/T)=o(1)$, following from the same argument in Lemma (ref)). The above expected value cancels out with term $a$. Thus, we can rewrite \begin{align*} a+ b & = \frac 1 {\sqrt{NT} } \sumiN (L\eps_i)'D^{-1} \eps_i +\frac 1 {\sqrt{NT} } \sumiN\lambda_i' (LF)' (FF'+D)^{-1} \eps_i + o_p(1) \end{align*} proving Lemma (ref). \begin{lemma} \begin{align} \frac 1 {\sqrt{NT} } \sumiN & (y_{i,-1}-\bar y_{-1})' \Big[(GG'+D)^{-1}-(FF'+D)^{-1}\Big] [F(\lambda_i-\bar \lambda)+ \eps_i-\bar\eps] \\ &=- \frac 1 T (LF)'[(FF'+D)^{-1} ] \tilde F +o_p(1) \notag \end{align} \end{lemma} Proof: It is not difficult to show \[ \frac 1 {\sqrt{NT} } \sumiN (y_{i,-1}-\bar y_{-1})' \Big[(GG'+D)^{-1}-(FF'+D)^{-1}\Big] ( \eps_i-\bar \eps) =o_p(1)\] we thus focus on $\frac 1 {\sqrt{NT} } \sumiN (y_{i,-1}-\bar y_{-1})' \Xi F(\lambda_i-\bar \lambda)$, where $\Xi=(GG'+D)^{-1}-(FF'+D)^{-1}$. Approximating $\Xi$ by ((ref)) is sufficient. Using $ y_{i,-1}-\bar y_{-1} =LF(\lambda_i-\bar \lambda) +L (\eps_i-\eps)$, \begin{align*} & \frac 1 {\sqrt{NT} } \sumiN (y_{i,-1}-\bar y_{-1})' \Xi_1 F(\lambda_i-\bar \lambda) \\ & = \frac 1 {\sqrt{NT} } \sumiN [LF(\lambda_i-\bar \lambda) +L (\eps_i-\eps)]' \Big(2\frac 1 {T^2}\frac 1 {\sqrt{NT} } (F'D^{-1}\tilde F) \frac 1 {\eta_F^4}\Big)D^{-1}(FF') D^{-1} F(\lambda_i-\bar \lambda)\\ & = 2 (F'L'D^{-1}F/T) (F'D^{-1}\tilde F/T) \frac 1 { \eta_F^2} +o_p(1) \end{align*} where the term involving $L(\eps_i-\bar \eps)$ is $O_p(N^{-1/2})$, thus negligible. Here we have used $F'D^{-1}F=T\omega_F^2$, $\omega_F^2/\eta_F^4=1/\eta_F^2 +O(T^{-1})$, and $\frac 1 N \sumiN (\lambda_i-\bar \lambda)^2=1+o_p(1)$. Next, \begin{align*} & \frac 1 {\sqrt{NT} } \sumiN (y_{i,-1}-\bar y_{-1})' \Xi_2 F(\lambda_i-\bar \lambda) \\ &=- \frac 1 {\sqrt{NT} } \sumiN [LF(\lambda_i-\bar \lambda) +L (\eps_i-\eps)]' \Big( \frac 1 {\sqrt{NT} } \frac 1 T \frac 1 {\eta_F^2}\Big) D^{-1}(F\tilde F') D^{-1}F(\lambda_i-\bar \lambda)\\ &=-\frac 1 {\eta_F^2} (F'L'D^{-1}F/T)(\tilde F'D^{-1}F/T) +o_p(1), \end{align*} where term involving $L(\eps_i-\bar \eps)$ is negligible. Next \begin{align*} & \frac 1 {\sqrt{NT} } \sumiN (y_{i,-1}-\bar y_{-1})' \Xi_3 F(\lambda_i-\bar \lambda)\\ & =-\frac 1 {\sqrt{NT} } \sumiN [LF(\lambda_i-\bar \lambda) +L (\eps_i-\eps)]' \Big( \frac 1 {\sqrt{NT} } \frac 1 T \frac 1 {\eta_F^2}\Big)D^{-1} (\tilde F F')D^{-1} F(\lambda_i-\bar\lambda) \\ & =- (F'L'D^{-1}\tilde F/T) +o_p(1) \end{align*} which follows from the same reasoning for the term involving $\Xi_1$. Summing up, \begin{align*} & \frac 1 {\sqrt{NT} } \sumiN (y_{i,-1}-\bar y_{-1})' \Xi [F(\lambda_i-\bar \lambda)+ \eps_i-\bar \eps] \\ &=\frac 1 {\eta_F^2} (F'L'D^{-1}F/T)(\tilde F'D^{-1}F/T) -(F'L'D^{-1}\tilde F/T) +o_p(1) \\ & =-\frac 1 T (LF)'(FF'+D)^{-1} \tilde F+o_p(1) \end{align*} proving the lemma. \begin{corollary} \begin{equation} \begin{split} \tilde \alpha \frac 1 {\sqrt{NT} } \sumiN & (y_{i,-1}-\bar y_{-1})' (GG'+D)^{-1}[F(\lambda_i-\bar \lambda)+ \eps_i-\bar \eps] \\ & = \tilde \alpha \, \frac 1 {\sqrt{NT} } \sumiN (L\eps_i)'D^{-1}\eps_i \\ & +\tilde \alpha \, \frac 1 {\sqrt{NT} } \sumiN\lambda_i' (LF)'[(FF'+D)^{-1} ] \eps_i \\ & -\tilde \alpha \, \frac 1 T (LF)'[(FF'+D)^{-1} ] \tilde F +o_p(1) \end{split} \end{equation} \end{corollary} Proof: adding and subtracting terms \begin{align*} \tilde \alpha \frac 1 {\sqrt{NT} } \sumiN & (y_{i,-1}-\bar y_{-1})' (GG'+D)^{-1}[F(\lambda_i-\bar \lambda)+ \eps_i-\bar \eps] \\ &=\tilde \alpha \frac 1 {\sqrt{NT} } \sumiN (y_{i,-1}-\bar y_{-1})' (FF'+D)^{-1}[F(\lambda_i-\bar \lambda)+ \eps_i-\bar \eps] \\ &+\tilde \alpha \frac 1 {\sqrt{NT} } \sumiN (y_{i,-1}-\bar y_{-1})' \Xi [F(\lambda_i-\bar \lambda)+ \eps_i-\bar \eps] \end{align*} where, by definition, $\Xi=(GG'+D)^{-1} - (FF'+D)^{-1}$. The corollary follows from Lemmas (ref) and (ref), that is, by summing ((ref)) and ((ref)), where every term is multiplied by $\tilde \alpha$. $\Box$ \begin{lemma} \begin{align} \frac 1 {NT} \sumiN & (y_{i,-1}-\bar y_{-1})' (GG'+D)^{-1} (y_{i,-1}-\bar y_{-1}) \\ & = \frac 1 T \mathrm{tr}(L'D^{-1}LD ) + \frac 1 T \mathrm{tr}[ (LF)' [(FF'+D)^{-1} ](LF) ] +o_p(1) \nonumber \end{align} \end{lemma} Proof: Here it is sufficient to approximate $(GG'+D)^{-1}$ by $ (FF'+D)^{-1}$ (recall $(GG'+D)^{-1}=(FF'+D)^{-1}+\Xi$, terms involving $\Xi$ are negligible because of the factor $(NT)^{-1}$). Using $ y_{i,-1}-\bar y_{-1} =LF(\lambda_i-\bar \lambda) +L (\eps_i-\bar \eps)$, we rewrite \begin{align*} \frac 1 {NT}& \sumiN (y_{i,-1}-\bar y_{-1})' (FF'+D)^{-1} (y_{i,-1}-\bar y_{-1})\\ & =\frac 1 {NT} \sumiN[LF(\lambda_i-\bar \lambda) +L (\eps_i-\bar \eps)]'(FF'+D)^{-1}[LF(\lambda_i-\bar \lambda) +L (\eps_i-\bar \eps)] \\ & = \frac 1 {NT} \sumiN[LF(\lambda_i-\bar \lambda)]'(F'F+D)^{-1} LF(\lambda_i-\bar \lambda)\\ & + 2 \frac 1 {NT} \sumiN [LF(\lambda_i-\bar \lambda)]'(F'F+D)^{-1} L (\eps_i-\bar \eps)\\ & + \frac 1 {NT} \sumiN [L (\eps_i-\bar \eps)]'(FF'+D)^{-1} L (\eps_i-\bar \eps) = a +b +c \end{align*} where the three terms are denoted by $a,b,c$, respectively. First, \begin{align*} a & = \frac 1 T \mathrm{tr}[ (LF)'(FF'+D)^{-1} LF \frac 1 N \sumiN(\lambda_i-\bar \lambda) (\lambda_i-\bar \lambda)' ]\\ & =\frac 1 T \mathrm{tr}[ (LF)'(FF'+D)^{-1} LF] +O_p(N^{-1/2}) \end{align*} where the second equality is due to $N^{-1}\sumiN(\lambda_i-\bar \lambda) (\lambda_i-\bar \lambda)'=I_r+O_p(N^{-1/2})$. Term $b$ can be shown to be $O_p(1/\sqrt{NT})$ and thus negligible. For $c$, we use the Woodbury formula, \begin{align*} c & = \frac 1 {NT} \sumiN (\eps_i-\bar \eps)'L' D^{-1}L (\eps_i-\bar \eps)\\ & -\frac 1 {(1+F'D^{-1}F)} \frac 1 {NT} \sumiN (\eps_i-\bar \eps)'L'D^{-1}F F'D^{-1}L (\eps_i-\bar \eps) =c1-c2. \end{align*} The mean of $c1$, $\mathbb E (c1)= [(N-1)/N]\frac 1 T\mathrm{tr}(L'D^{-1}LD)= \frac 1 T\mathrm{tr}(L'D^{-1}LD)+O(1/N)$. The deviation from the mean is negligible, because the variance of $c1$ is $O(1/N)$. Thus \[ c1= \frac 1 T\mathrm{tr}(L'D^{-1}LD) +o_p(1) \] Consider $c2$. Its expected value \[ \mathbb E (c2) = \frac 1 {(1+F'D^{-1}F)} (\frac {N-1}N) \frac 1 T \mathrm{tr} [ F'D^{-1} L DL' D^{-1} F] =O(1/T) \] Because $c2$ is nonnegative and its expected value converges to zero, we have $c2=o_p(1)$. Thus the LHS of ((ref)) is determined by $a$ and $c1$. This proves the lemma. $\Box$ {\bf Proof of Theorem (ref).} The local likelihood ratio is given by ((ref)). Using Lemma (ref), Corollary (ref), Corollary (ref), Lemma (ref) (multiplied by $-\frac 1 2 \tilde \alpha^2$), we obtain \begin{equation} \ell(\theta^0+\frac 1 {\sqrt{NT} } \tilde \theta) -\ell(\theta^0) = A(\tilde \theta)+ o_p(1) \end{equation} where \addtocounter{equation}{1} \begin{align*} A(\tilde \theta) & = \frac 1 {\sqrt{NT} } \sumiN \lambda_i'\tilde F'(FF'+D)^{-1} \eps_i\\ & + \tilde \alpha \, \frac 1 {\sqrt{NT} } \sumiN (L\eps_i)'D^{-1}\eps_i \\ & +\tilde \alpha \, \frac 1 {\sqrt{NT} } \sumiN\lambda_i' (LF)'[(FF'+D)^{-1} ] \eps_i \\ &-\frac 1 2 \frac 1 T \mathrm{tr}[\tilde F'(FF'+D)^{-1} \tilde F ] \tag{\theequation} \\ & -\frac 1 2 \tilde \alpha^2 \mathrm{tr} \Big[ \frac 1 T (L'D^{-1}LD ) \Big] \\ & -\frac 1 2 \tilde \alpha^2 \mathrm{tr} \Big[ \frac 1 T (LF)' (FF'+D)^{-1} (LF)\Big]\\ & -\tilde \alpha \, \frac 1 T \mathrm{tr}[ (LF)'(FF'+D)^{-1} \tilde F]. \end{align*} Next, we show that $(FF'+D)^{-1}$ can be replaced by $D^{-1/2} M_{D^{-1/2}F} D^{-1/2}$. In particular, \begin{equation} A(\tilde \theta) =\Delta_{NT}(\tilde \theta)-\frac 1 2 \mathbb E[\Delta_{NT}(\tilde \theta)]^2 + O_p(T^{-1}) \end{equation} where $\Delta_{NT}(\tilde \theta)$ is defined in Theorem 1. The replacement ensures that the last four terms of $A(\tilde \theta)$ are exactly the variance of the first three terms of $A(\tilde \theta)$ (multiplied by -1/2). Let $\Upsilon= (FF'+D)^{-1}-D^{-1/2} M_{D^{-1/2}F} D^{-1/2}$. We first show, concerning the first term of $A(\tilde \theta)$, \[ \frac 1 {\sqrt{NT} } \sumiN \lambda_i'\tilde F'\Upsilon \eps_i = O_p(1/T) \] For simplicity, assume $r=1$, we can write $\Upsilon = \frac 1 {T^2 \omega_F^2 \eta_F^2} D^{-1} F F D^{-1}$. Hence, \[ \frac 1 {\sqrt{NT} } \sumiN \lambda_i'\tilde F'\Upsilon \eps_i =\frac 1 {T \omega_F^2 \eta_F^2} (\frac {\tilde F'D^{-1}F} T) \frac 1 {\sqrt{NT} } \sumiN F' D^{-1} \eps_i\lambda_i =O_p(1/T). \] The same analysis is applicable to the third term of $A(\tilde \theta)$, that is, $\frac 1 {\sqrt{NT} } \sumiN \lambda_i' (LF)'\Upsilon \eps_i = O_p(1/T)$. Consider the fourth term, \[ \frac 1 T \tilde F'\Upsilon \tilde F = \frac 1 {T \omega_F^2 \eta_F^2} (\frac {\tilde F'D^{-1}F} T) ( \frac {F'D^{-1} \tilde F} T) =O(\frac 1 T). \] Each of the last two terms of $A(\tilde \theta)$ being $O(1/T)$ follows from the same analysis. This establishes ((ref)). The proof of Theorem 1 is complete. $\Box$ \begin{comment} Inspecting $A(\tilde \theta)$, the variances of the first three terms are given by the fourth to the sixth terms (times -1/2), respectively. The last term is twice of the covariance between the first and the third term (times -1/2), thus the last term is simply the negative covariance. The second term is uncorrelated with the first and third terms (since $L\eps_i$ depends on the lags of $\eps_{it}$, and $\eps_{it-1} \eps_{it}$ is uncorrelated with $\eps_{ik}$ for all $t$ and $k$). Finally, $\Delta_{NT}(\tilde \theta)$ of Theorem 1 is composed of the first three terms of $A(\tilde \theta)$ (these are the random terms). All the remaining terms constitute $-(1/2) \mathbb E \Delta_{NT}(\tilde \theta)^2$. This completes the proof of Theorem 1. $\Box$ \end{comment} To prove Theorem (ref), we need the following result. \begin{lemma} Under Assumption C, we have \[ \tag{a} \frac 1 T F'L'D^{-1} F \rightarrow \frac 1 {(1-\alpha)} \int_0^1 \psi(s)\psi(s)' /\sigma(s)^2 ds \] \[ \tag{b} \frac 1 T F'L'D^{-1} LF \rightarrow \frac 1 {(1-\alpha)^2} \int_0^1 \psi(s) \psi(s)'/\sigma(s)^2 ds \] \end{lemma} This lemma extends Lemma 1 of Bai (2013), where the proof was much simpler. Proof of (a). For simplicity, assume a single factor, and recall $F=(f_1,f_2,...,f_T)'$, and $D=\diag(\sigma_1^2,...,\sigma_T^2)$. \[ (LF)'D^{-1} F = \sum_{k=0}^{T-2} \alpha^k \sum_{t=k+2}^{T-k} f_{t-k-1} f_t/\sigma_t^2 \] For any $\epsilon>0$, choose a finite $K$ such that $|\sum_{k=0}^K \alpha^k -1/(1-\alpha)|<\epsilon$ and $|\alpha|^{K}\le \epsilon$. Under Assumption C, $f_t =\psi(t/T)$ and $\psi(s)$ is continuous on [0,1], thus uniformly continuous on [0,1]. For the previous $\epsilon>0$, there is an $\delta>0$ such that for all $|x-y|<\delta$, we have $|\psi(x)-\psi(y)|<\epsilon$. \[ \frac 1 T (LF)'D^{-1} F = \sum_{k=0}^{K} \alpha^k \frac 1 T \sum_{t=k+2}^{T-k} f_{t-k-1} f_t/\sigma_t^2 + \sum_{k=K+1}^{T-2} \alpha^k \frac 1 T \sum_{t=k+2}^{T-k} f_{t-k-1} f_t/\sigma_t^2 \] Note the second term is small. Since $f_t$ and $1/\sigma_t^2$ are bounded by assumption, the second term is bounded by $|\alpha|^K M /(1-|\alpha|) = \epsilon M /(1-|\alpha|) \le \epsilon M'$. Consider the first term. \[ \frac 1 T \sum_{t=k+2}^{T-k} f_{t-k-1} f_t/\sigma_t^2= \frac 1 T \sum_{t=k+2}^{T-k} (f_{t-k-1}-f_t) f_t/\sigma_t^2 + \frac 1 T \sum_{t=k+2}^{T-k} f_t f_t/\sigma_t^2 \] Again by the boundedness of $f_t$ and $1/\sigma_t^2$, we have \[ |\frac 1 T \sum_{t=k+2}^{T-k} (f_{t-k-1}-f_t) f_t/\sigma_t^2| \le M \frac 1 T \sum_{t=k+2}^{T-k} |f_{t-k-1}-f_t| = M \frac 1 T \sum_{t=k+2}^{T-k} |\psi(\frac{t-k-1} T)-\psi(\frac t T)|\] For large enough $T$, for all $k \le K$, $(k+1)/T \le (K+1)/T < \delta$, thus $|\psi(\frac{t-k-1} T)-\psi(\frac t T)|<\epsilon$, for all $t$ and all $k\le K$. Thus, for all large $T$, \[ \left| \sum_{k=0}^{K} \alpha^k \frac 1 T \sum_{t=k+2}^{T-k} f_{t-k-1} f_t/\sigma_t^2- \sum_{k=0}^{K} \alpha^k \frac 1 T \sum_{t=k+2}^{T-k} f_t f_t/\sigma_t^2\right| \le M \epsilon /(1-|\alpha|) \] Next, \begin{align*} \Big| \sum_{k=0}^{K} & \alpha^k \frac 1 T \sum_{t=k+2}^{T-k} f_t f_t/\sigma_t^2 - \frac 1 {1-\alpha} \int_0^1 \psi(s)^2/\sigma(s)^2 ds \Big|\\ & \le \left|\sum_{k=0}^{K} \alpha^k \Big(\frac 1 T \sum_{t=k+2}^{T-k} f_t f_t/\sigma_t^2 - \frac 1 T \sum_{t=1}^{T} f_t f_t/\sigma_t^2\Big) \right| \\ & + \left|\Big( \sum_{k=0}^{K} \alpha^k -\frac 1 {1-\alpha}\Big) \Big( \frac 1 T \sum_{t=1}^{T} f_t f_t/\sigma_t^2\Big) \right| \\ & + \left| \frac 1 {1-\alpha} \Big( \frac 1 T \sum_{t=1}^{T} f_t f_t/\sigma_t^2 -\int_0^1 \psi(s)^2/\sigma(s)^2 \Big) \right| \end{align*} The first term on the right hand side is bounded by $M/(1-|\alpha|) (2K+1)/T$, which is less than $\epsilon$ for large $T$. The second term is bounded by $\epsilon M$ by the choice of $K$, and third term is bounded by $\epsilon$ for all large $T$. In summary, there exists an $M<\infty$, independent of $T$ such that for every $\epsilon>0$, and for all sufficiently large $T$, we have \[ \left \| \frac 1 T F'L'D^{-1} F - \frac 1 {(1-\alpha)} \int_0^1 \psi(s)\psi(s)' /\sigma(s)^2 ds \right\|\le M \epsilon \] This proves part (a). Proof of (b). Notice \[ \frac 1 T F'L' D^{-1} LF = \frac 1 T \sum_{t=2}^{T} \frac 1 {\sigma_t^2} \left( \sum_{j=2}^{t} \alpha^{t - j} f_{j - 1} \right)^2 \] \begin{equation} =\sum_{j=0}^{T-2} \sum_{k=0}^{T-2} \alpha^{j+k} \left( \frac 1 T \sum_{t=\max(j,k)+2}^{T} \frac{f_{t - j - 1} f_{t - k - 1}}{\sigma_t^2}\right) \end{equation} Although we could make the proof rigorous by using the argument in part a), this would be repetitive. We therefore focus on the key insight. For any fix $j$ and $k$, \[ \frac 1 T \sum_{t=\max(j,k)+2}^{T} \frac{f_{t - j - 1} f_{t - k - 1}}{\sigma_t^2} \rightarrow \int_0^1 \psi(s)^2/\sigma(s)^2 ds \] which does not depend on $j$ and $k$. For large enough J and $K$, $ \sum_{j=0}^J \sum_{k=0}^K \alpha^{j+k} =(\sum_{j=0}^J \alpha^j) (\sum_{k=0}^K \alpha^k)$, which is close to $1/(1-\alpha)^2$. The tail part of ((ref)) is negligible. This gives (b). $\Box$ \begin{lemma} Under Assumption C, and for $\tilde f \in C_f$, and $\tilde F=(\tilde f_1,...,\tilde f_T)'$. we have \[ \tag{a} \frac 1 T (LF)' [D^{-1/2} M_{D^{-1/2}F} D^{-1/2}] LF \rightarrow 0 \] \[ \tag{b} \frac 1 T (LF)' [D^{-1/2} M_{D^{-1/2}F} D^{-1/2}] \tilde F \rightarrow 0 \] \end{lemma} Proof of (a). By definition, \begin{align*} \frac 1 T (LF)'[D^{-1/2} M_{D^{-1/2}F} D^{-1/2}] & LF = \frac 1 T F'L' D^{-1} LF \\ & - \frac 1 T (F'L'D^{-1}F) \Big(\frac{F'D^{-1}F} T\Big)^{-1} \frac 1 T (F'D^{-1} LF). \end{align*} Part (a) then follows from Lemma (ref) and $(F'D^{-1}F/T)^{-1} \rightarrow [\int_0^1 \psi(s)\psi(s)'/\sigma(s)^2 ds]^{-1}$. Proof of (b). The LHS of (b) is equal to \[ \frac { (LF)' D^{-1} \tilde F} T - \Big(\frac{(LF)'D^{-1}F} T \Big)\Big( \frac{F'D^{-1}F} T \Big)^{-1} \Big( \frac{ F'D^{-1}\tilde F} T\Big). \] For the first term, using the same argument for Lemma (ref) part (a), we can show \[ \frac 1 T (LF)'D^{-1}\tilde F \rightarrow \frac 1 {1-\alpha} \int_0^1 \frac 1 {\sigma(s)^2} \psi(s) \tilde \psi(s)' ds, \] which is equal to zero on $C_f$. The second term also converges to zero on $C_f$, \begin{equation} \frac{ F'D^{-1}\tilde F} T \rightarrow \int_0^1 \frac 1 {\sigma(s)^2} \psi(s) \tilde \psi(s)' ds =0. \end{equation} This proves the lemma. $\Box$ {\bf Proof of Theorem (ref)}. Since the local parameter space $\mathbb H=\mathbb R \times C_f$ is a subspace of $\ell^\infty$ for Theorem (ref), the proof of Theorem (ref) holds on $\mathbb H$. However, some of the expressions in Theorem (ref) are simplified under $\mathbb H$. We begin by showing the first term of $\Delta_{NT}(\tilde \theta)$ (see ((ref))) can be simplified \begin{align*} \frac 1 {\sqrt{NT} } \sumiN &\lambda_i'\tilde F'D^{-1/2} M_{D^{-1/2}F} D^{-1/2} \eps_i \\ & = \frac 1 {\sqrt{NT} } \sumiN \lambda_i'\tilde F'D^{-1}\eps_i -\frac 1 {\sqrt{NT} } \sumiN \lambda_i'\tilde F' D^{-1} F ( F'D^{-1}F)^{-1} F' D^{-1}\eps_i \end{align*} We show that the second term is $o_p(1)$. It suffices to show its variance converges to zero since the mean is zero. Its variance is \[ \mathrm{tr} [\frac 1 T (\tilde F'D^{-1} F) (F'D^{-1} F)^{-1} (F'D^{-1}\tilde F)] \] But on $C_f$, ((ref)) holds. Hence the said variance converges to zero on $C_f$. Next we show the third term of $\Delta_{NT}(\tilde \theta)$ is $o_p(1)$. That is, \[ \frac 1 {\sqrt{NT} } \sumiN\lambda_i' (LF)'[D^{-1/2} M_{D^{-1/2}F} D^{-1/2} ] \eps_i =o_p(1). \] We also show its variance converges to zero. Its variance is given by \[ \mathrm{tr} \Big[ \frac 1 T (LF)' D^{-1/2} M_{D^{-1/2}F} D^{-1/2} (LF)\Big] \] which converges to zero by Lemma (ref) part (a). Next, consider the first term of $\mathbb E [ \Delta_{NT}(\tilde \theta)]^2$ (see ((ref))) . We show \begin{equation} \frac 1 T \mathrm{tr}[\tilde F'D^{-1/2} M_{D^{-1/2}F} D^{-1/2} \tilde F ] = \frac 1 T \mathrm{tr}( \tilde F' D^{-1} \tilde F) + o(1). \end{equation} By definition, \[ \frac 1 T \tilde F'D^{-1/2} M_{D^{-1/2}F} D^{-1/2} \tilde F =\frac 1 T \tilde F'D^{-1} \tilde F - \frac 1 T \tilde F'D^{-1} F (F'D^{-1}F)^{-1} F'D^{-1} \tilde F \] The second term in the above equation converges to zero on $C_f$ is argued earlier. Hence ((ref)) holds. The third term of $\mathbb E [ \Delta_{NT}(\tilde \theta)]^2$ is $o(1)$ by Lemma (ref) part (a). The last term of $\mathbb E [ \Delta_{NT}(\tilde \theta)]^2$ converges to zero by Lemma (ref) part (b). In summary, Theorem (ref) is simplified to \begin{align*} \Delta_{NT}(\tilde \theta) -\frac 1 2 \mathbb E [ \Delta_{NT}(\tilde \theta)]^2 & = \frac 1 {\sqrt{NT} } \sumiN \lambda_i'\tilde F' D^{-1}\eps_i + \tilde \alpha \, \frac 1 {\sqrt{NT} } \sumiN (L\eps_i)'D^{-1}\eps_i \\ & -\frac 1 2 \frac 1 T \mathrm{tr} \Big[\tilde F'D^{-1} \tilde F \Big] -\frac 1 2 \tilde \alpha^2 \Big[ \frac 1 T \mathrm{tr}(L'D^{-1}LD ) \Big] +o_p(1). \end{align*} This proves Theorem (ref). $\Box$ {\bf Proof of Theorem (ref)}. With respect to the local parameters $\tilde F$, the proof of Theorem (ref) only uses $\|T^{-1/2}\tilde F\|=O(1)$. If $\tilde f \in \ell_r^2$, then $\|\tilde F\|=O(1)$. The entire proof of Theorem (ref) holds with $T^{-1/2}\tilde F$ replaced by $\tilde F$. In particular, equations ((ref)) and ((ref)) hold with $T^{-1/2}\tilde F$ replaced by $\tilde F$ (that is, omitting $T^{-1/2}$) due to $\|\tilde F\|=O(1)$. Notice $\tilde F$ appears in three places in ((ref)) and ((ref)). We analyze each of them. The first term of ((ref)) after replacing $T^{-1/2}\tilde F$ with $\tilde F$ is written as \begin{align*} \frac 1 {\sqrt{N}} \sumiN \lambda_i'\tilde F' & D^{-1/2} M_{D^{-1/2}F} D^{-1/2} \eps_i =\frac 1 {\sqrt{N}} \sumiN \lambda_i'\tilde F' D^{-1} \eps_i\\ & - \mathrm{tr}[\tilde F'D^{-1}F (F'D^{-1}F)^{-1} \frac 1 {\sqrt{N}} \sumiN F'D^{-1}\eps_i\lambda_i']\end{align*} But the second term on the righthand side is $o_p(1)$, because it can be written as (ignore the trace) \[ \frac 1 {\sqrt{T}} (\tilde F'D^{-1}F) (F'D^{-1}F/T)^{-1} \frac 1 {\sqrt{NT} } \sumiN F'D^{-1}\eps_i\lambda_i' \] Now $\|F'D^{-1}F/T]^{-1}\|=O(1)$, $\frac 1 {\sqrt{NT} } \sumiN F'D^{-1}\eps_i\lambda_i'=\frac 1 {\sqrt{NT} } \sumtT \sumiN \frac 1 {\sigma^2} f_t\lambda_i' \eps_{it} =O_p(1)$, but \begin{equation} \|\frac 1 {\sqrt{T}} (\tilde F'D^{-1}F)\|=\|\frac 1 {\sqrt{T}} \sumtT \frac 1 {\sigma_t^2} \tilde f_t f_t' \|\le \frac 1 a M \frac 1 {\sqrt{T}} \sum_{t=1}^T \|\tilde f_t\| =o(1) \end{equation} we have used $\sigma_t^2\ge a>0$, and $\|f_t\|\le M$. To see $\frac 1 {\sqrt{T}} \sum_{t=1}^T \|\tilde f_t\| =o(1)$ for $\tilde f\in \ell_r^2$, notice $\tilde f_s\rightarrow 0$ as $s\rightarrow \infty$, and by the Toeplitz lemma, $\frac 1 {\sqrt{T}} \sum_{t=1}^{\sqrt{T}}\|\tilde f_t\| \rightarrow 0$. By the Cauchy-Schwarz, $\frac 1 {\sqrt{T}} \sum_{t=\sqrt{T}+1}^T \|\tilde f_t\| \le (\sum_{t=\sqrt{T}+1}^T \|\tilde f_t\|^2)^{1/2} \le (\sum_{t=\sqrt{T}}^\infty \|\tilde f_t\|^2 )^{1/2}\rightarrow 0.$ (Also see IwakuraOkui2014 for a similar result.) Thus, \[ \frac 1 {\sqrt{N}} \sumiN \lambda_i'\tilde F'D^{-1/2} M_{D^{-1/2}F} D^{-1/2} \eps_i=\frac 1 {\sqrt{N}} \sumiN \lambda_i'\tilde F' D^{-1} \eps_i +o_p(1). \] The first term of ((ref)) after replacing $\frac 1 T \tilde F$ by $\tilde F$ becomes (ignore the -1/2 and the trace), \begin{align*} \tilde F'D^{-1/2} M_{D^{-1/2}F} D^{-1/2} \tilde F & =\tilde F' D^{-1} \tilde F -\tilde F' D^{-1}F ( F'D^{-1}F)^{-1} F'D^{-1}\tilde F \\ & = \tilde F' D^{-1} \tilde F +o(1) \end{align*} owing to $(F'D^{-1}F/T)^{-1}=O(1)$ and $\tilde F' D^{-1}F/T^{1/2}=o(1)$ due to ((ref)). The last term of ((ref)) being $o(1)$ with $\tilde F$ in place of $\frac 1 {\sqrt{T}} \tilde F$ follows from the same argument. Collecting the simplified and the non-negligible terms, we obtain the expressions in Theorem (ref). $\Box$
thebibliography{24} \bibitem {AlvarezArellano2022} Alvarez, J. and M. Arellano (2022). Robust likelihood estimation of dynamic panel data models. {\em Journal of Econometrics}, 226, 21-61. \bibitem {Anderson2003} Anderson, T. W. (2003). {\em An Introduction to Multivariate Statistical Analysis}, 3rd ed. Wiley, Hoboken, NJ. \bibitem {AndersonHsiao1982} Anderson, T.W., and C. Hsiao (1982). Formulation and estimation of dynamic Models with Error Components, {\em Journal of Econometrics}, 76, 598-606. \bibitem {ArellanoBond1991} Arellano, M. and S. Bond (1991). Some tests of specification for panel data: Monte Carlo evidence and an application to employment equations, {\em The Review of Economic Studies} 58 (2), 277-297. \bibitem {Bai2009} Bai, J. (2009). Panel data models with interactive fixed effects, {\em Econometrica}, 77 1229-1279. \bibitem {Bai2013} Bai, J. (2013). Fixed effects dynamic panel data models, a factor analytical method, {\em Econometrica,} 81, 285-314. \bibitem {Bai2024} Bai, J. (2024). Likelihood approach to dynamic panel models with interactive effects. {\em Journal of Econometrics.} Vol 240, Issue 1. https://doi.org/10.1016/j.jeconom.2023.105636 \bibitem {BaiLi2012} Bai, J. and K.P. Li (2012). Statistical analysis of factor models of high dimension. {\em Annals of Statistics,} 40, 436-465. \bibitem {BaiLi2014} Bai, J. and K.P. Li (2014). Theory and methods for panel data models with interactive effects. {\em The Annals of Statistics} Vol. 42, No. 1, 142-170. \bibitem {BaiLi2021} Bai, J. and K.P. Li (2021). Dynamic spatial panel data models with common shocks, {\em Journal of Econometrics,} 224(1) 134-160. \bibitem {BaiMones2025} Bai, J. and P. Mones (2025). Global identification of dynamic panel data models with interactive effects. arXiv:2504.14354 \bibitem {bickel1993} Bickel, P.J., C.A Klaassen, Y. Ritov, Jon A Wellner (1993). Efficient and adaptive estimation for semiparametric models. Springer. \bibitem {HahnKuersteiner2002} Hahn, J., and G. Kuersteiner (2002). “Asymptotically Unbiased Inference for a Dynamic Panel Model with Fixed Effects when Both n and T Are Large," {\em Econometrica}, 70, 1639-1657. \bibitem {IwakuraOkui2014} Iwakura, H. and R. Okui (2014). Asymptotic efficiency in factor models and dynamic panel data models, Institute of Economic Research, Kyoto University, Available at SSRN 2395722 \bibitem {JankovaGeer2018} Jankova J. and S. van de Geer (2018). Semiparametric efficiency bounds for high-dimensional models. {\em The Annals of Statistics} 46 (5), 2336-2359. \bibitem {Kiviet1995} Kiviet, J. (1995). “On Bias, Inconsistency, and Efficiency of Various Estimators in Dynamic Panel Data Models", {\em Journal of Econometrics,} 68, 53-78. \bibitem {Lancaster2000} Lancaster, T. (2000). “The incidental parameter problem since 1948," {\em Journal of Econometrics}, 95 391-413. \bibitem {LawleyMaxwell1971} Lawley, D.N. and A.E. Maxwell (1971). {\em Factor Analysis as a Statistical Method}, London: Butterworth. \bibitem{lamyao2012} Lam, C. and Q. Yao (2012). Factor modeling for high-dimensional time series: Inference for the number of factors, The Annals of Statistics 40(2), 694-726. \bibitem{miaophillipssu2023} Miao, K., PCB Phillips, and L. Su (2023), ‘High-dimensional vars with common factors’, Journal of Econometrics 233(1), 155-183. \bibitem {MoonWeidner2017} Moon H.R. and M. Weidner (2017). “Dynamic Linear Panel Regression Models with Interactive Fixed Effects", Econometric Theory 33, 158-195. \bibitem {Moreira2009} Moreira, M.J. (2009), A Maximum Likelihood Method for the Incidental Parameter Problem, {\em Annals of Statistics}, 37, 3660-3696. \bibitem {Nickell1981} Nickell, S. (1981). Biases in Dynamic Models with Fixed Effects, {\em Econometrica}, 49, 1417-1426. \bibitem{shilee2017} Shi, W. and Lee, L. F. (2017), Spatial dynamic panel data models with interactive fixed effects, Journal of Econometrics 197(2), 323–347. \bibitem {vanderVaart1998} van der Vaart, A.W. (1998). {\em Asymptotics Statistics}, Cambridge University Press. \bibitem {vanderVaartWellner1996} van der Vaart, A.W. and J.A. Wellner (1996). {\em Weak Convergence and Empirical Processes with Applications to Statistics}, Springer-Verlag.