Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
116,356 characters · 16 sections · 53 citation commands
Global identification of dynamic panel models with interactive effects
\def\spacingset#1{ {#1}} \spacingset{1}
\if11 \fi
\if01 {
} \fi
{\it Keywords: } Dynamic panel models, global identification, factor models, panel data, interactive effects
\spacingset{1.9}
Consider the dynamic panel data model with interactive effects
where $i=1,2,\dots,N$ and $t=1,2,\dots, T$. $y_{it}$ is the outcome variable, $\alpha$ is the autoregressive parameter, whereas $\delta_t$ captures time-specific fixed effects. $f_{t}$ is an $\bar{r}\times 1$ vector containing the value of the $\bar{r}$ latent factors at time $t$, and $\lambda_{i}$ is an $\bar{r}\times 1$ vector of factor loadings for individual $i$. $\varepsilon_{it}$ is the idiosyncratic error term, which we assume to be i.i.d. across $i$, heteroskedastic across $t$ but serially uncorrelated, i.e. $\mathbb E[\varepsilon_{it}^{2}]=\sigma_{t}^{2} \; \text{and} \; \mathbb E[\varepsilon_{it}\varepsilon_{is}]=0 \;$ for $t\neq s$, and independent from the interactive effects. In this model, $y_{it}$ is the only observable variable.
This framework provides a flexible way to model unobservable common trends and heterogeneous responses. The model allows the common shocks (modeled by $f_t$) to have a heterogeneous effect in individuals (modeled by $\lambda_i$) on the outcome variable $y_{it}$. A notable special case is the standard additive fixed-effects model, where $y_{it} = \delta_t + \gamma_i + \varepsilon_{it}$, which corresponds to $\alpha=0$ and a single constant factor $f_t=1$. Thus, this model generalizes the additive framework by allowing for richer dynamics and heterogeneity in the factor structure. Parameter estimation in this model can be conducted via quasi maximum likelihood based methods as in hayakawa2023short and bai2024likelihood.
This paper studies the problem of global identification of the model parameters when the number of cross-sectional units $(N)$ is large, but the time dimension $(T)$ is fixed. This case is particularly relevant in applied economics because many panel datasets, such as those involving countries, firms, or households, span a large number of cross-sectional units but only cover a limited time horizon. For instance, macroeconomic panels often include many countries but only a few decades of annual data, making the short-$T$, large-$N$ framework empirically realistic.
Following rothenberg1971identification, we say that a parameter point $\theta^0$ in the parametric space $\Theta$ is globally identifiable if there is no other $\tilde{\theta}\in\Theta$ which is observationally equivalent. Establishing the identifiability of a model’s parameters is crucial, as it ensures that the estimation procedure yields meaningful and interpretable results.
Despite the apparent simplicity of the model, demonstrating parameter identifiability is a nontrivial task. The seminal work of anderson1956statistical established that in the absence of the autoregressive term ($\alpha=0$), the parameters of the model are globally identified up to an indeterminacy given by an $\bar{r}\times\bar{r}$ rotation matrix, provided that $T \geq 2\bar{r}+1$.
To the best of our knowledge, no existing study has established the global identifiability of the parameters in the model described by (ref). bai2024likelihood demonstrates that the parameters of this model are locally identified when $\bar{r}$ does not exceed the modified Ledermann bound ledermann1937rank. In contrast, hayakawa2023short argues that global identification is unattainable in dynamic panel data models with interactive effects.
We demonstrate that under standard assumptions in the factor analysis literature, the parameters of the model are almost surely globally identifiable when $T\geq 2(\bar{r}+1)$. The global identification is achieved using the first two moments. Since the quasi-Gaussian likelihood depends only on the first two moments, this equivalence implies global identification for quasi-Gaussian maximum likelihood estimation. This result is of particular importance, as it establishes a one-to-one correspondence between the population moments and the model parameters, despite high degree of nonlinearity of the problem.
This paper contributes to the literature on global identification and panel data models. The question of global identification in parametric models has been the focus of extensive research (see, for example, shapiro1985identifiability, komunjer2011dynamic, komunjer2012global, and kociecki2023solution), while other authors have focused on trying to understand when and why identification fails forneron2024detecting. A comprehensive review of this literature is provided in lewbel2019identification.
Panel data models have been a core focus in econometrics for many years, with extensive study and application over the past few decades. Notable monographs and textbooks on panel data include works by arellano2003panel, baltagi2008econometric, hsiao2022analysis, and wooldridge2010econometric, among others. Moreover, the study of panel data models with interactive fixed effects\footnote{Sometimes referred to in the literature as multiple time-varying individual effects ahn2013panel or multifactor error structure pesaran2006estimation.} has gained notable attention during the recent years. While there is an extensive literature in this area, some of the most noteworthy papers include holtz1988estimating, ahn2001gmm, pesaran2006estimation, bai2009panel, and ahn2013panel. More recently, there has been significant research on dynamic panel models with interactive effects, including studies by moon2017dynamic, hayakawa2023short, and bai2024likelihood, to name a few. However, these studies have not addressed the global identifiability of the model parameters for the short-$T$ case, a gap this paper aims to fill.
We also argue that the techniques developed in this paper can be used to prove identification in a broad class of dynamic models featuring time, individual and interactive effects. For example, we show that in a dynamic panel with individual effects arellano1991some, the autoregressive coefficient can be globally identified even when $\alpha=1$, a case previously thought to be non-identifiable (see blundell1998initial; sentana2024finite).
The remainder of the paper is organized as follows: Section 2 formalizes the identification problem in our context. Section 3 outlines the assumptions required for our proof and presents some useful mathematical results that will be instrumental in our analysis. Section 4 derives intermediate results and provides the main proof of global identification under the stated assumptions. Section 5 discusses identification in the presence of individual fixed effects. Finally, Section 6 provides some concluding remarks with a discussion of potential extensions of the paper.
Before formally defining global identification, it is useful to express the model in matrix notation. Suppose we observe $y_i = (y_{i1},y_{i2},\dots, y_{iT})'$. We begin by projecting the initial observation $y_{i1}$ onto $[1,\lambda_{i}]$, which gives
$$y_{i1} = \delta_1^*+\lambda_{i}'f_1^*+\varepsilon_{i1}^*$$
where $\delta_1^*$ and $f_1^*$ are the projection coefficients, and $\varepsilon_{i1}^*$ is the projection residual. Since $\delta_1^*$ and $f_1^*$ are free parameters, and $\varepsilon_{it}$ in the model are not required to have identical distributions, we can simply rename $\delta_1^*, f_1^*$, and $\varepsilon_{i1}^*$ as $\delta_1$, $f_1$, and $\varepsilon_{i1}$, respectively, without any loss of generality. So we can write $y_{i1} = \delta_1+\lambda_{i}'f_1 +\varepsilon_{i1}$.
Next, stacking observations over $t$, we obtain
$$\mathbf{B} y_{i} = \delta + \mathbf{F} \lambda_i + \varepsilon_i$$
where $y_i = [y_{i1},y_{i2},\dots,y_{iT}]'$ is a $T\times1$ vector of individual observations for $t=1,\dots,T$, $\delta = [\delta_1 , \dots, \delta_T]'$ is a $T\times1$ vector, and $\mathbf{F}=[f_{1},f_{2}, \dots, f_{T}]'$ is a $T\times \bar{r}$ matrix containing the $\bar{r}$ factors, $\lambda_i = [\lambda_{i1},\lambda_{i2},\dots,\lambda_{i\bar{r}}]'$ is an $\bar{r}\times1$ vector of factor loadings for individual $i$, and $\varepsilon_i = [\varepsilon_{i1},\varepsilon_{i2},\dots,\varepsilon_{iT}]'$ is a $T\times 1$ vector of idiosyncratic errors. Furthermore,
is a nonsingular lower-triangular $T\times T$ matrix, hence invertible. Let $\mathbf{\Gamma}=\mathbf{B}^{-1}.$ Pre-multiplying by $\mathbf{B}^{-1}$ allows us to rewrite the model as
where $\mathbf{\Gamma}$ is $T\times T$ and is given by $$\mathbf{\Gamma} = \mathbf{B}^{-1} =
$$
In the original model, the errors had a diagonal covariance matrix given by $\mathbf{D} =\diag(d_1,d_2,\dots,d_T) := \diag(\sigma_{1}^{2}, \dots, \sigma_{T}^{2})$ (for notational simplicity, we write $\sigma_t^2$ as $d_t$). After transformation, the idiosyncratic error $\mathbf{\Gamma}\varepsilon_i$ has covariance matrix $\mathbf{\Gamma}\mathbf{D}\mathbf{\Gamma}'$, which is no longer diagonal. Let
$$\mathbf{\Psi}_{N} = \frac{1}{N-1}\sum_{i=1}^{N}(\lambda_{i}-\bar{\lambda})(\lambda_{i}-\bar{\lambda})'$$
denote the sample covariance matrix of $\lambda_i$, where $\bar{\lambda}=\frac{1}{N}\sum_{i=1}^{N}\lambda_i$ and let $\mathbb E[\mathbf{\Psi}_{N}] = \mathbf{\Psi}$. Similarly, define
$$\mathbf{S}_N=\frac{1}{N-1}\sum_{i=1}^{N}(y_i-\bar{y})(y_i-\bar{y})'$$
as the sample covariance matrix of $y_i$, where $\bar{y}=\frac{1}{N}\sum_{i=1}^{N}y_i$.
The expectation of the sample covariance matrix of $y_i$ is then given by
where $\theta^0=(\alpha,\mathbf{F},\mathbf{\Psi},\mathbf{D})$ is the true vector of structural parameters. The matrix $\mathbf{\Sigma}$ is a $T\times T$ symmetric positive definite (PD) matrix, $\mathbf{\Psi}$ is an $\bar{r}\times\bar{r}$ symmetric PD matrix, $\mathbf{F}\mathbf{\Psi} \mathbf{F}'$ is a $T\times T$ symmetric positive semi-definite (PSD) matrix of rank $\bar{r}$, and $\mathbf{D}$ is a $T\times T$ diagonal PD matrix. $\mathbf{\Omega}=\mathbf{F}\mathbf{\Psi} \mathbf{F}' + \mathbf{D}$ is a $T\times T$ symmetric PD matrix corresponding to the expected value of the sample covariance matrix of $y_i$ in a basic model without an autoregressive term $\alpha=0$ anderson1956statistical.
This paper aims to establish that the parameters of the model are globally identifiable from the Gaussian quasi log-likelihood perspective, which is equivalent to demonstrating global identification from the first two moments. Specifically, this means that the expectation and variance of $y_{i}$ uniquely determine the structural parameters of the model, as there exists a one-to-one mapping between them.
To formalize this argument, define $\delta^\dag=\mathbf{\Gamma} \delta$. The Gaussian quasi log-likelihood function is given by
$$-\frac N 2 \ln |\mathbf{\Sigma}(\theta)| -\frac 1 2 \sum_{i=1}^N (y_i -\delta^\dag)' \mathbf{\Sigma}(\theta)^{-1} (y_i-\delta^\dag)$$
Concentrating out $\delta^\dag$, which is estimated by $\bar y$, and dividing by $N$, we obtain\footnote{In the MLE, the definition of $\mathbf{S}_N $ needs a slight modification, namely, $1/(N-1)$ is replaced by $1/N$.}
$$\frac 1 N \ell_N(\theta) =-\frac 1 2 \ln |\mathbf{\Sigma}(\theta)| -\frac 1 2 \mathrm{tr}[ \mathbf{\Sigma}(\theta)^{-1} \mathbf{S}_N]$$
For a fixed $T$ and as $N\to\infty$, by the law of large numbers, $\mathbf{S}_N\overset{p}{\to} \mathbf{\Sigma}(\theta^0)$, we have
Furthermore, it is known arnold1981theory that this function has a unique maximizer at $\mathbf{\Sigma}(\theta)=\mathbf{\Sigma}(\theta^0)$. While this result ensures that $\theta^0$ is a global maximizer of the objective function, it does not immediately establish uniqueness—hence, global identification—because it does not preclude the possibility that other parameter values $\tilde{\theta}\neq\theta^0$ satisfy $\mathbf{\Sigma}(\tilde{\theta})=\mathbf{\Sigma}(\theta^0)$. In the following sections, we investigate whether any alternative parameter set $\tilde{\theta}\neq \theta^0$ satisfies $\mathbf{\Sigma}(\tilde{\theta})=\mathbf{\Sigma}(\theta^0)$. We demonstrate that, almost surely, no such alternative parametrization exists. Consequently, the quasi log-likelihood function is uniquely maximized at the true parameter value $\theta^0$, establishing global identification. This result, which has not been previously shown and was thought to be untrue hayakawa2023short, constitutes a novel contribution.
To prove this claim, consider an alternative parameterization
$$\mathbf{\tilde{\Gamma}} =
$$
By definition, $\mathbf{\tilde{F}} \mathbf{\tilde{\Psi}} \mathbf{\tilde{F}}'$ is a symmetric PSD $T\times T$ matrix of rank $\bar{r}$.
The model is globally identified at $\theta^0$ if whenever $\mathbf{\Sigma}(\tilde{\theta})=\mathbf{\Sigma}(\theta^0)$, it must hold that $\tilde{\alpha}=\alpha$, $\mathbf{\tilde{D}}=\mathbf{D}$, and $\mathbf{\tilde{F}\tilde{\Psi}\tilde{F}}’=\mathbf{F\Psi F’}$. Formally:
If and only if $\tilde{\alpha}=\alpha$, $\mathbf{\tilde{D}}=\mathbf{D}$, and $\mathbf{\tilde{F}\tilde{\Psi}\tilde{F}}’=\mathbf{F\Psi F’}$ (or simply $\tilde{\theta}=\theta^0$).
In order to proceed with our global identification proof, we need to make some assumptions. These assumptions are not restrictive and are rather standard in the factor analysis literature.
This assumption comes from the fact that factor models (and thus interactive effects) are identifiable up to a $\bar{r}\times\bar{r}$ matrix rotation, so $\bar{r}^2$ restrictions must be imposed for identification. A discussion of why this is the case can be found in Section 5 of anderson1956statistical. We work with this restriction for convenience, but any other normalization that imposes $\bar{r}^2$ restriction would also work.
Additionally, throughout the paper, we will reference several established results from linear algebra and probability theory. To ensure they are easily accessible to the reader, we present them as lemmas in this section.
where $\mathbf{A}_{-i,-j}$ is the minor obtained by removing row $i$ and column $j$, and $A_{i,j}$ represents the $(i,j)$-th element of the matrix $\mathbf{A}$.
In this paper, we will apply Lemma (ref) for cases in which $\mathbf{A}$ itself is a minor of a larger matrix. Let $\mathbf{H}$ be a $T\times T$ matrix, and let $\mathbf{M}^{\mathbf{H}}_{R,C}$ be a minor of $\mathbf{H}$ of dimension $k$ including rows $R=\{r_1,\dots,r_k\}$ and columns $C=\{c_1,\dots,c_k\}$ of the original matrix. Therefore, for any $r\in R$ we can denote $i\in\{1,\dots,k\}$ as the position of row $r$ in the minor $\mathbf{M}^{\mathbf{H}}_{R,C}$, and the Laplace decomposition of this minor is given by:
$$\det(\mathbf{M}^{\mathbf{H}}_{R,C}) = \sum_{j=1}^{k}(-1)^{i+j}H_{r,c_j}\det\left(\mathbf{M}^{\mathbf{H}}_{R-r,C-c_{j}}\right)$$
where $H_{r,c_j}$ is the $(r,c_j)$-th element of $\mathbf{H}$, and $\mathbf{M}^{\mathbf{H}}_{R-r,C-c_{j}}$ is the minor\footnote{Where $R-r$ means removal $r$ from the set $R$, and $C-c_j$ means removal of $c_j$ from the set $C$. They correspond to the usual notation $R\setminus \{r\}$ and $C\setminus \{c_j\}$. } of $\mathbf{H}$ that includes rows $R-r$, and columns $C-c_j$.
In our induction analysis to be used later, we will further need to indicate the dimension of the minor. The minor on the left hand side will be written as $\mathbf{M}^{\mathbf{H},k}_{R,C}$, and the minor on the right hand side will be written as $\mathbf{M}^{\mathbf{H},k-1}_{R-r,C-c_{j}}$
In this section, we demonstrate that, under Assumptions (ref)-(ref), the parameters of the dynamic panel data model with interactive effects described in (ref) are globally identified.
The proof we propose draws inspiration from Theorem 5.1 of anderson1956statistical. In that theorem, the authors establish a sufficient condition for global identification of a factor model that did not include an autoregressive component. In their simpler model, global identification requires
if and only if $\mathbf{\tilde{D}}=\mathbf{D}$ (a $T\times T$ diagonal PD matrix), and $\mathbf{\tilde{F}} \mathbf{\tilde{\Psi}} \mathbf{\tilde{F}}'=\mathbf{F}\mathbf{\Psi}\mathbf{F}'$ (a $T\times T$ symmetric PSD matrix of rank $\bar{r}$).
Their proof relies on the fact that the off-diagonal elements of $\mathbf{\tilde{F}} \mathbf{\tilde{\Psi}} \mathbf{\tilde{F}}'$ must match those of $\mathbf{F}\mathbf{\Psi} \mathbf{F}'$ for (ref) to hold, because both $\mathbf{D}\ \text{and} \ \mathbf{\tilde{D}}$ are diagonal matrices. Thus, the problem reduces to showing that the diagonal elements of these matrices are also equal. To establish this, they rewrite (ref) as
$$\mathbf{F}\mathbf{\Psi} \mathbf{F}' = \mathbf{\tilde{F}} \mathbf{\tilde{\Psi}} \mathbf{\tilde{F}}'+\mathbf{\tilde{D}}- \mathbf{D}$$
Then, they leverage the fact that all determinants of the $(\bar{r}+1)\times(\bar{r}+1)$ minors of $\mathbf{\tilde{F}} \mathbf{\tilde{\Psi}} \mathbf{\tilde{F}}'+\mathbf{\tilde{D}}- \mathbf{D}$ must be zero for (ref) to hold (Lemma (ref)). This follows as a necessary condition because, by assumption, $\mathbf{F}\mathbf{\Psi} \mathbf{F}'$ is of rank $\bar{r}$. This procedure enables them to pin down the diagonal terms of $\mathbf{\tilde{F}} \mathbf{\tilde{\Psi}} \mathbf{\tilde{F}}'$, thereby establishing that $\mathbf{\tilde{F}} \mathbf{\tilde{\Psi}} \mathbf{\tilde{F}}'=\mathbf{F}\mathbf{\Psi} \mathbf{F}'$. Once this is established, it follows immediately that $\mathbf{\tilde{D}}=\mathbf{D}$, completing the proof.
In their approach, the authors focus exclusively on the minors involving diagonal elements of the matrix and do not consider what we will refer to as diagonal exclusion minors (minors that do not include diagonal terms of the original matrix) as these do not contain relevant information for their analysis. In contrast, our proof takes a different approach: we rearrange the terms of our original problem—Equation (ref)—in a way that allows us to exploit the information contained in the diagonal exclusion minors of this transformed problem. This novel approach enables us to demonstrate that (ref) holds if and only if $\tilde{\alpha}=\alpha$, thereby establishing the global identification of $\alpha$.
Once we establish that $\alpha$ is globally identified, it follows that our identification condition can be rewritten in the form proposed by anderson1956statistical, ensuring that all other parameters can also be globally identified.
Thus, the primary focus of our analysis is to show that $\alpha$ is globally identified (or equivalently, that it must be the case that $\tilde{\alpha}=\alpha$ in order for (ref) to hold). This task requires some intermediate results, which we will articulate and prove throughout the paper. The main contribution of this work is to formally establish the following Theorem.
Once this is shown, it is immediate to conclude that all the parameters of the model are almost surely globally identified as well, which is formally stated in the following Theorem
{\bf Remark}: The qualifier \enquote*{almost surely} is with respect to the distribution of $\mathbf{F}$. For any fixed $\alpha$, identification holds for almost all realizations of $\mathbf{F}$. The practical implication is that as long as there is sufficient variation in $f_t$, both $\alpha$ and other parameters are globally identifiable. This is particularly true when $T$ is sufficiently large, ensuring a sufficient number of distinct realizations of $f_t$.
In order to prove Theorem (ref), we will rearrange the problem in a way that resembles the proof of anderson1956statistical and then establish some properties of this modified problem. Start from (ref). As $\mathbf{\tilde{\Gamma}}$ is a unit-diagonal lower triangular matrix, it is invertible so that we can pre-multiply and post-multiply by $\mathbf{\tilde{\Gamma}}^{-1}$ and $\mathbf{\tilde{\Gamma}}^{-1'}$ and re-write
$$\mathbf{\tilde{\Gamma}}^{-1}\mathbf{\Gamma}(\mathbf{F}\mathbf{\Psi}\mathbf{F}'+\mathbf{D})\mathbf{\Gamma}'\mathbf{\tilde{\Gamma}}^{-1'} = \mathbf{\tilde{F}} \mathbf{\tilde{\Psi}} \mathbf{\tilde{F}}'+\mathbf{\tilde{D}}$$
Now, define
$$\mathbf{Q} = \mathbf{\tilde{\Gamma}^{-1}}\mathbf{\Gamma} =
$$
Note that $\mathbf{Q}$ can be decomposed into
$$\mathbf{Q} = \mathbf{I}_T + (\alpha-\tilde{\alpha})\mathbf{L}$$
where $\mathbf{L}$ is a strictly lower-triangular Toeplitz matrix given by
$$\mathbf{L} =
$$
Then, subtract $\mathbf{\tilde{D}}$ on both sides of the equation
To simplify notation, denote the left-hand side of (ref) by
We next establish some preliminary results.
Furthermore, we can show that $\mathbf{O}$ can be decomposed in the following convenient way
where $\mathbf{\Omega}$ is defined as in (ref) and $\mathbf{J}(\tilde{\alpha}, \theta^0)$ is given by:
With this result in hand, we can proceed to derive the following corollaries.
By definition, $R$ and $C$ must satisfy
The minors in the definition are squared matrices, thus $|R|=|C|=k$. They do not contain diagonal elements, thus $R \cap C= \emptyset$. It follows that $|R|=|C|=k \leq \frac{T}{2}$.
Now, we are ready to state and show the following Propositions.
In mathematical terms, what we show is that given any diagonal exclusion minor of $\mathbf{O}$ of dimension $k$, say $\mathbf{M}^{\mathbf{O},k}_{R,C}$, we can write its determinant as the sum of the determinant of the diagonal exclusion minor of $\mathbf{\Omega}$ associated with the same rows and columns, say $\mathbf{M}^{\mathbf{\Omega},k}_{R,C}$, and a term that can be written as $(\alpha-\tilde{\alpha})\tilde{J}_{R,C}\left(\tilde{\alpha},\theta^0\right)$
where $\tilde{J}_{R,C}\left(\tilde{\alpha},\theta^0\right)$ is a scalar that is a function of the true structural parameters $\theta^0$ and $\tilde{\alpha}$. Note that the value of $\tilde{J}_{R,C}\left(\tilde{\alpha},\theta^0\right)$ will depend on which columns and rows were taken from the original matrix so in general, $\tilde{J}_{R,C}\left(\tilde{\alpha},\theta^0\right) \neq \tilde{J}_{R',C'}\left(\tilde{\alpha},\theta^0\right)$ for any other valid set of rows and columns $R'$ and $C'$. To keep notation as simple as possible, we sometimes use $\tilde{J}_{R,C}$ instead of $\tilde{J}_{R,C}\left(\tilde{\alpha},\theta^0\right)$ to represent the same object.
Propositions (ref) and (ref) provide the essential results for proving Theorem (ref). However, before we proceed with the proof, we need to establish one additional result. Consider the following claim:
From this result, we can derive the following corollary
In this subsection, we build on the intermediate results derived in subsection (ref) to prove Theorem (ref). Specifically, we demonstrate that (ref) can only hold if $\tilde{\alpha}=\alpha$, except for a set of factor realizations that occur with probability zero. We begin by outlining the structure of the proof, then proceed with a detailed argument for the case of a single factor to illustrate the reasoning. Finally, we generalize the proof to the case of $\bar{r}$ factors.
\bf Proof of Theorem 1: Pick any diagonal exclusion minor of $\mathbf{O}$ of dimension $\bar{r}+1$, say $\mathbf{M}^{\mathbf{O},\bar{r}+1}_{R,C}$. By Claim (ref) the determinant of this minor must satisfy
$$\det\left(\mathbf{M}^{\mathbf{O},\bar{r}+1}_{R,C}\right)=0$$
By Proposition (ref) we know that
$$\det\left(\mathbf{M}^{\mathbf{O},\bar{r}+1}_{R,C}\right) =\det\left(\mathbf{M}^{\mathbf{\Omega},\bar{r}+1}_{R,C}\right) +(\alpha-\tilde{\alpha})\tilde{J}_{R,C}\left(\tilde{\alpha},\theta^0\right)=0$$
Moreover, by Corollary (ref) we have that $\det\left(\mathbf{M}^{\mathbf{\Omega},\bar{r}+1}_{R,C}\right)=0$. Then:
By Assumption (ref) we know that $\mathbf{O}$ has at least $2$ different diagonal exclusion minors.\footnote{In particular, when $T=2(\bar{r}+1)$ and $\bar{r}>0$ there are at least $\frac{\binom{2(\bar{r}+1)}{\bar{r}+1}}{2}\geq 2$ diagonal exclusion minors} Then, consider any other diagonal exclusion minor of $\mathbf{O}$, say $\mathbf{M}^{\mathbf{O},\bar{r}+1}_{R',C'}$. Using the same logic, we can see that:
$$\det\left(\mathbf{M}^{\mathbf{O},\bar{r}+1}_{R',C'}\right)= (\alpha-\tilde{\alpha})\tilde{J}_{R',C'}\left(\tilde{\alpha},\theta^0\right)=0$$
Then, it must be the case that either $\tilde{\alpha}=\alpha$ or $\tilde{J}_{R,C}\left(\tilde{\alpha},\theta^0\right)=\tilde{J}_{R',C'}\left(\tilde{\alpha},\theta^0\right)=0$ in order for (ref) to hold (Claim (ref)). However, we show that $\tilde{J}_{R,C}\left(\tilde{\alpha},\theta^0\right)=\tilde{J}_{R',C'}\left(\tilde{\alpha},\theta^0\right)=0$ is only possible in a set of probability zero.\\
Consider the case of a single factor $\bar{r}=1$
Consider the following $2\times 2$ diagonal exclusion minor of $\mathbf{O}$:
$$\mathbf{M}_{(1,2),(3,4)}^{\mathbf{O},2} =
$$
See Appendix (ref) for the explicit functional form of each $O_{i,j}$. Combining Proposition (ref), Corollary (ref), and Claim (ref) we get that:
where in this case $\tilde{J}_{(1,2),(3,4)}\left(\tilde{\alpha},\theta^0\right)$ is a random polynomial in $\tilde{\alpha}$ of degree 1. Note that one of the following four alternatives must be true in order for (ref) to be satisfied:
Cases (ref) and (ref) correspond to the case in which $\tilde{J}_{(1,2),(3,4)}\left(\tilde{\alpha},\theta^0\right)$ is the zero polynomial, so that all the coefficients are trivially equal to 0. Case (ref) corresponds to the case in which we can find a root $\tilde{\alpha}^{\star}$ that solves $\tilde{J}_{(1,2),(3,4)}(\tilde{\alpha}^{\star},\theta^0)=0$. However, once $\tilde{\alpha}$ is pinned-down in this way, we have no other free parameter, so the same root must satisfy all the zero-determinant conditions (Claim (ref)) associated with the remaining diagonal exclusion minors of $\mathbf{O}$. Case (ref) corresponds to the case where global identification is achieved.
We show that Cases (ref)-(ref) are only possible in a set of factor realizations that occur with probability zero. Thus, with probability 1, Case (ref) must hold so $\alpha$ is almost surely globally identified.
Case (ref): $\Psi=0$
This case is obvious. We assumed $\Psi>0$, so we can instantly rule this out.\\
Case (ref): $f_2=\frac{d_2}{d_1 \alpha}$
From Assumptions (ref) and (ref), we know that $d_1,d_2,\alpha$ are fixed population parameters, whereas $f_2$ is drawn from a continuous probability distribution. Thus, we can invoke Lemma (ref) and ensure that $P\left(f_2 = \frac{d_2}{d_1 \alpha}\right) = 0$. Then, Case 2 can only occur in a set of probability 0.\\
Case (ref): $\tilde{\alpha}=\frac{f_4}{f_3}$
Setting $\tilde{\alpha}^{\star}=\frac{f_4}{f_3}$ so that $\tilde{J}_{(1,2),(3,4)}(\tilde{\alpha}^{\star},\theta^0)=0$ leaves us with no other free parameter. Therefore, in order for Case (ref) to be possible, it must be the case that this same $\tilde{\alpha}^{\star}$ simultaneously make all the $\tilde{J}_{R',C'}(\tilde{\alpha}^{\star},\theta^0)$ equal to 0 for any other valid set of rows and columns (in particular, it requires $\tilde{J}_{(2,3),(1,4)}(\tilde \alpha^{\star},\theta^0)=0$). To see why this will not be the case, consider the following $2\times 2$ diagonal exclusion minor of $\mathbf{O}$:
$$\mathbf{M}_{(2,3),(1,4)}^{\mathbf{O},2} =
$$ Following the same procedure as for (ref),
$$\det\left(\mathbf{M}_{(2,3),(1,4)}^{\mathbf{O},2}\right)=(\alpha-\tilde{\alpha})\tilde{J}_{(2,3),(1,4)}\left(\tilde{\alpha},\theta^0\right)=0$$
Algebra calculation shows that $\tilde{J}_{(2,3),(1,4)}\left(\tilde{\alpha},\theta^0\right)$ is a polynomial of degree 2 in $\tilde{\alpha}$:
with the coefficients given by:
Recall that $\tilde{J}_{(2,3),(1,4)}\left(\tilde{\alpha},\theta^0\right)=0$ must hold in order for Case (ref) to be plausible. There are two ways in which $\tilde{J}_{(2,3),(1,4)}\left(\tilde{\alpha},\theta^0\right)$ can be 0: either all the coefficients are equal to 0, so that $\tilde{J}_{(2,3),(1,4)}\left(\tilde{\alpha},\theta^0\right)$ is the zero polynomial; or the coefficients of the polynomial are such that real roots exist and at least one of them coincide with $\tilde{\alpha}^{\star}=\frac{f_4}{f_3}$.
We can see that Assumption (ref) implies that $a=b=c=0$ can only occur on a set with probability 0, as it involves the factor realizations lying on a particular low-dimensional manifold (see Lemma (ref)). That is, the first case has zero probability of occurring.
We next show that the second case also has zero probability of occurring. To see this, note that $\tilde{J}_{(2,3),(1,4)}(\tilde{\alpha}, \theta^0)$ is a polynomial of degree 2 in $\tilde{\alpha}$. Its roots are:
$$\tilde{\alpha} =\frac{-b \pm \sqrt{b^2 - 4ac}}{2a}$$
Conditional on the roots being real, it must be the case that either
$$\frac{f_4}{f_3} = \frac{-b + \sqrt{b^2 - 4ac}}{2a} \quad \text{or} \quad \frac{f_4}{f_3} = \frac{-b - \sqrt{b^2 - 4ac}}{2a}$$
Examining the expressions of $a$, $b$, and $c$ and making use of Lemma (ref) and Assumption (ref), we conclude that these events can never occur with positive probability (the factors $f_2,f_3,f_4$ must lie exactly in a lower dimensional manifold). Hence $\tilde{\alpha}=\frac{f_4}{f_3}$ is only possible on a set of zero probability.
Having ruled out Cases (ref)-(ref), we are left with Case (ref), which is the only option that can happen with positive probability when $\bar{r}=1$, so it must be the case that $\tilde{\alpha}=\alpha$ almost surely in order for (ref) to be satisfied. We have thus established global identification for $\bar{r}=1$. $\square$\\
Consider the general case with $\bar{r}$ factors
Start from (ref)
$$\det\left(\mathbf{M}^{\mathbf{O},\bar{r}+1}_{R,C}\right)= (\alpha-\tilde{\alpha})\tilde{J}_{R,C}\left(\tilde{\alpha},\theta^0\right)=0$$
Our goal is to show that $\tilde{J}_{R,C}\left(\tilde{\alpha},\theta^0\right)=0$ for all possible valid combinations of $R$ and $C$ can only occur with probability zero, implying that $\tilde{\alpha}=\alpha$ must hold almost surely. Recall from Proposition (ref) that $\tilde{J}_{R,C}\left(\tilde{\alpha},\theta^0\right)$ is a random polynomial in $\tilde{\alpha}$ of degree in between $1$ and $2\bar{r}+1$, given by:
$$\tilde{J}_{R,C} = \sum_{j=1}^{\bar{r}+1}(-1)^{\bar{r}+1+j}\left[J_{r_{\bar{r}+1},c_j}\det\left(\mathbf{M}^{\mathbf{\Omega},\bar{r}}_{R-r_{\bar{r}+1},C-c_j}\right)+\Omega_{r_{\bar{r}+1},c_j}\tilde{J}_{R-r_{\bar{r}+1},C-c_j} +(\alpha-\tilde{\alpha})J_{r_{\bar{r}+1},c_j}\tilde{J}_{R-r_{\bar{r}+1},C-c_j}\right]$$
By Assumptions (ref) and (ref), $\alpha,\mathbf{D} \ \text{and} \ \mathbf{\Psi}$ are deterministic population parameters, and $\mathbf{F}$ is drawn from a continuous distribution. The condition $\tilde{J}_{R,C}\left(\tilde{\alpha},\theta^0\right)=0$ can occur in two possible ways: either all of the coefficients of this random polynomial are equal to 0, or by setting $\tilde{\alpha}$ equal to one of the roots of this polynomial $\tilde{J}_{R,C}\left(\tilde{\alpha}^{\star},\theta^0\right)=0$ (by the fundamental theorem of algebra we know that $\tilde{J}_{R,C}\left(\tilde{\alpha},\theta^0\right)$ has at most $2\bar{r}+1$ real roots).
The first case is automatically ruled-out by Proposition (ref), as it states that $\tilde{J}_{R,C}(\tilde{\alpha},\theta^0)$ is a random polynomial in $\tilde{\alpha}$ of degree in between $1$ and $2\bar{r}+1$ almost surely. This result guarantees that with probability one $\tilde{J}_{R,C}(\tilde{\alpha},\theta^0)$ is not the zero polynomial.
Now consider the second case. If $\tilde{J}_{R,C}\left(\tilde{\alpha},\theta^0\right)$ has no real roots, then it is clear that it must be that $\tilde{\alpha}=\alpha$ in order for (ref) to hold, so the proof would be over. Now, assume $\tilde{J}_{R,C}\left(\tilde{\alpha},\theta^0\right)$ has at least one real root. Pick any of these roots, say $\tilde{\alpha}^{\star}_{R,C}(\theta^0)$. By definition, $\tilde{\alpha}^{\star}_{R,C}(\theta^0)$ solves:
$$\tilde{J}_{R,C}\left(\tilde{\alpha}^{\star}_{R,C}(\theta^0),\theta^0\right)=0$$
Note that $\tilde{\alpha}^{\star}_{R,C}(\theta^0)$ is a function of the other true data-generating process parameters and can be obtained from the above condition. Moreover, it also depends on the rows $R$ and columns $C$ chosen as in general the random polynomials $\tilde{J}_{R,C}\left(\tilde{\alpha},\theta^0\right)$ and $\tilde{J}_{R',C'}\left(\tilde{\alpha},\theta^0\right)$ are different (except for a set of factor realizations of probability zero).
As we used the diagonal exclusion minor associated with rows and columns $R,C$ to pin down $\tilde{\alpha}$ we have no other free parameter, so this same $\tilde{\alpha}^{\star}_{R,C}\left(\theta^0\right)$ must satisfy:
$$\tilde{J}_{R',C'}(\tilde{\alpha}^{\star}_{R,C}\left(\theta^0\right),\theta^0)=0$$
for any other valid set of $R',C'$. This can only occur in two different ways: either $\tilde{J}_{R',C'}\left(\tilde{\alpha},\theta^0\right)$ is the zero polynomial, or the random polynomial $\tilde{J}_{R',C'}\left(\tilde{\alpha},\theta^0\right)$ shares at least one root with $\tilde{J}_{R,C}\left(\tilde{\alpha},\theta^0\right)$.
As we have previously discussed, the case in which $\tilde{J}_{R',C'}\left(\tilde{\alpha},\theta^0\right)$ is the zero polynomial is an event that occurs with probability zero. Therefore, consider the second case. We know that the coefficients of $\tilde{J}_{R,C}\left(\tilde{\alpha},\theta^0\right)$ and $\tilde{J}_{R',C'}\left(\tilde{\alpha},\theta^0\right)$ can be exactly equal only in a set of probability zero, so the polynomials are almost surely different. Then, the case in which they share the root $\tilde{\alpha}^{\star}_{R,C}\left(\theta^0\right)$ can be expressed as:
Note that the above condition depends only on the structural parameters of the model. Moreover, it defines a particular low-dimensional manifold in which the random factors $\mathbf{F}$ must lie in order for both polynomials to share the root $\tilde{\alpha}^{\star}_{R,C}\left(\theta^0\right)$. However, by Assumption (ref) and Lemma (ref), the event in which the random factors lie in this particular low-dimensional manifold occurs with probability zero.
Then, the case $\tilde{J}_{R,C}\left(\tilde{\alpha},\theta^0\right)=\tilde{J}_{R',C'}\left(\tilde{\alpha},\theta^0\right)=0$ can only occur in a set of probability 0. Thus, it must be the case that $\tilde{\alpha}=\alpha$ in order for (ref) to be satisfied. This implies that $\alpha$ is almost surely globally identified. $\square$
We have shown that $\alpha$ is almost surely globally identified, meaning that (ref) is satisfied if and only if $\tilde{\alpha}=\alpha$ with probability 1. We can then conclude that all the parameters in our model are globally identified. This is stated in Theorem (ref). We next provide a proof.
\bf Proof of Theorem 2: As $\alpha$ is almost surely globally identified, then $\tilde{\alpha}=\alpha$ with probability 1 and $\mathbf{Q}=\mathbf{I}_{T}$. Then, we are left to show:
if and only if $\mathbf{\tilde{D}}=\mathbf{D}$, and $\mathbf{\tilde{F}} \mathbf{\tilde{\Psi}} \mathbf{\tilde{F}}'=\mathbf{F}\mathbf{\Psi}\mathbf{F}'$, where $\mathbf{\tilde{F}} \mathbf{\tilde{\Psi}} \mathbf{\tilde{F}}'$ must be a $T\times T$ symmetric PSD matrix of rank $\bar{r}$ and $\mathbf{\tilde{D}}$ a $T\times T$ PD diagonal matrix.
Under the Assumptions of our model, this is the case shown by anderson1956statistical, so all the rest of the parameters are almost surely globally identified as well. $\square$
Consider a panel data model with time, individual and interactive effects:
where $\gamma_i$ denotes the individual fixed effect for observation $i$, and the rest of the parameters are defined as in (ref). It is a well-known result that individual fixed-effects can be absorbed into the factor structure by treating $\gamma_i$ as the factor loadings of a special factor, which takes the value of 1 for all $t$. As a result, the factor matrix does not satisfy the conditions of Assumption (ref) because the factor associated with the individual fixed-effects follows a degenerate distribution.
Therefore, the results derived in Theorem (ref) do not apply in this case. We can overcome this issue in two ways. The first approach consists of showing that our proof can be adapted to allow for one factor with a degenerate distribution. In practice, global identification requires that there is enough variation in some of the factors, but there could be some other factors following degenerate distributions. It is important to note that in this case we do not estimate $\gamma_i$ (doing so would cause incidental parameter problems), but we absorb it in the factor structure and identify the variance of $\gamma_i$ and the covariance with the rest of the factor loadings through $\mathbf{\Psi}$.
The second approach consists of differencing the data to get rid of the individual fixed effects as in hayakawa2023short. While hayakawa2023short claims that global identification of $\alpha$ cannot be attained in the differenced data model we show that their conjecture is incorrect and that we can use a slightly modified version of our proof to derive global identification when we work in differences.
In practice, we advocate for working with the data in levels. Differencing the data eliminates important information about the process contained when the model is studied in levels. Moreover, it introduces time dependence between the errors as a mechanical result from differencing the idiosyncratic error. It also requires additional periods for identification as a result of both losing one observation by differencing and the temporal dependence between the errors.
We start from the model (ref) and absorb the individual fixed effects into the factor structure:
where everything is defined as in (ref) but $\eta_i=[\gamma_i,\lambda_i]$ and $F_t=[1,f_t]$. Using the same procedure as in Section (ref), we project the first observation \( y_{i1} \) onto the vector \([1, \gamma_i, \lambda_i]\), yielding the following expression: \( y_{i1} = \delta_1^{*} + \gamma_i f_{\gamma}^{*} + \lambda_i' f_{1}^{*} + \varepsilon_{i1}^{*} \), where \( (\delta_1^{*}, f_{\gamma}^{*}, f_1^{*}) \) are the projection coefficients and \( \varepsilon_{i1}^{*} \) is the projection residual. For notational simplicity, we drop the asterisks from the coefficients and from the residual, as these are free parameters, and we do not require the error terms \( \varepsilon_{it} \) to have identical distributions across \( t \). Therefore, for \( t = 1 \), we can rewrite the equation as:
\[ y_{i1} = \delta_1 + \gamma_i f_{\gamma} + \lambda_i' f_1 + \varepsilon_{i1}.\] Combining this with equation ((ref)) and stacking the observations over \( t \), we obtain: $$\mathbf{B}y_{i} = \delta + \mathbf{F}\eta_{i} + \varepsilon_{i}$$
where $\mathbf{B}$, $y_i$, $\delta$, and $\varepsilon_{i}$ are all defined earlier, but
$$\mathbf{F} =
, \quad and \eta_{i} =
$$
Note that the first element in the first column of \( \mathbf{F} \) is \( f_\gamma \), not 1. Importantly, the first column of \( \mathbf{F} \) consists of many 1’s, which are fixed constants and follow degenerate distributions. As a result, the factor matrix does not satisfy the conditions of Assumption (ref). However, since this assumption is a sufficient but not necessary condition, we show that the model remains globally identified. The key insight is that, as long as there is sufficient variation in some of the factors, identification can still be achieved.
For simplicity, we provide a proof for the global identification of $\alpha$ in the presence of individual fixed-effects and a single unrestricted factor ($dim(f_t)=1$), and $f_t$ is drawn from a continuous distribution. In this case, $\bar{r}=2$. However, the proof can be extended to more general cases encompassing multiple factors.
Note that the presence of individual fixed-effects imposes structure on $\mathbf{F}$, so we require fewer restrictions to eliminate rotational indeterminacy. We only need to impose two restrictions on $\mathbf {F}$ to prevent indeterminacy (see Appendix (ref) for detailed explanation). We shall impose the normalization in the following way, referred to as tail normalization:
This normalization restricts the last two realizations of the factor $f$ to be 0 and 1 respectively. This is purely for convenience in terms of exposition. Normalizing other entries (except the first row) has the same effect. Under this normalization, $\mathbf F \mathbf{A}= \mathbf F $, if and only if $\mathbf{A}=\mathbf{I}_2$, thus eliminating rotational indeterminacy.
Before showing that $\alpha$ can be almost surely identified in this context, we shall explicitly state the identification assumptions:
Assumptions (ref)-(ref) are an adaptation of the ones outlined in Section (ref) for the particular case we are studying. While we conjecture that $f_{\gamma}$ drawn from a continuous distribution is not strictly needed to attain identification, it substantially simplifies the proof exposition.
Assumption (ref) is imposed because our strategy is not able to show identification in that particular case (it will become clearer why during the proof). We argue that this does not undermine the generality of our results for two main reasons. First, the case in which $d_{t} = \alpha d_{t-1}$ with $\alpha\neq1$ is a restriction on the values of the parameters that have Lebesgue measure zero in the parameter space. Second, and more importantly, it is hard to think any true data generating process such that the variance of the idiosyncratic errors should evolve as $d_{t} = \alpha d_{t-1}$ when $\alpha\neq1$, so this is a case unlikely to be of practical relevance. Note that the case in which $\alpha=1$ and $d_{t} = d_{t-1}$ is identifiable. Thus, when for $\alpha=1$, the model is identifiable under both homoskedasticity and heterokedasticity.
Under these assumptions, we can show that $\alpha$ is almost surely globally identified:
Given global identification of $\alpha$, it becomes clear that the rest of the parameters are almost surely globally identified as well. This can be formalized in the following Theorem.
Another alternative to dealing with the presence of individual fixed effects is to difference the data as in hayakawa2023short. While the authors claim that global identification is not necessarily attainable when we consider the data in differences, we show that our proof and results can be adapted to achieve global identification of $\alpha$ in this context. To see this, consider a re-labeled version of the DGP studied in the previous subsection:\footnote{The DGP is exactly equivalent to the one considered in the previous subsection and in hayakawa2023short. The choice of notation is arbitrary and done for convenience.}
where everything is defined exactly as in (ref), except that we assume that $u_{it}$ is homoskedastic with variance $\sigma^2<\infty$.\footnote{hayakawa2023short assumes that $u_{it}$ is homoskedastic in their proof, so we follow this assumption in order to simplify the comparison and exposition. However, the proof would work in the same way if $u_{it}$ had heteroskedastic variance $V(u_{it})=\sigma_{t}^{2}$ as in the previous sections of the paper.} Now, consider the model for the differenced data:
Let:
And write the model for the differenced data as:
We can see that this model looks exactly like the one we have considered throughout the paper derived from (ref) with the only difference that $\varepsilon_{it}$ are no longer independent over $t$. The reason is that, as $u_{it}$ was assumed to be a white noise process, then its first difference follows a non-invertible MA process. Aside from that, the models are equal. Thus, we can project the initial observation of the differenced process $y_{i1}$ onto $[1,\lambda_i]$ and stack observations over $t$ the same way we did in Section (ref) and write:
We are in a situation very similar to the one we considered in Section (ref), with the only difference that the covariance matrix of the idiosyncratic errors differs. In the case considered in Section (ref):
$$\mathbf{D} = \mathbb E[\varepsilon_{i}\varepsilon_{i}'] =
$$
whereas for the differenced data the covariance matrix is no longer diagonal and it equals:
$$\mathbf{\dot{D}} = \mathbb E[\varepsilon_{i}\varepsilon_{i}'] =
$$
a tridiagonal matrix. The parameters $\sigma_{1}^{2}$ and $\sigma_{c}$ capture the fact that the initial idiosyncratic error $\varepsilon_{i1}$ may follow a different distribution from the subsequent errors. Specifically, \[ \sigma_{1}^{2} = \mathbb E[\varepsilon_{i1}^{2}], \qquad \sigma_{c} = \mathbb E[\varepsilon_{i1}\varepsilon_{i2}] = \mathbb E[\varepsilon_{i1}\Delta u_{i2}] = -\,\mathbb E[\varepsilon_{i1}u_{i1}]. \] For any $t \ge 3$, we have $\mathbb E[\varepsilon_{i1}\varepsilon_{it}] = \mathbb E[\varepsilon_{i1}\Delta u_{it}] = 0$, implying that the projection of the initial observation affects only the leading $2\times2$ block of $\mathbf{\dot{D}}$.
Then, the expected value of the sample variance of the differenced data is equal to:
which is equal to the one derived in Section (ref) changing $\mathbf{D}$ for $\mathbf{\dot{D}}$. Thus, showing global identification in the differenced model is equivalent to showing:
if and only if $\mathbf{\tilde{\dot{D}}}=\mathbf{\dot{D}}$, $\tilde{\alpha}=\alpha$, and $\mathbf{\tilde{F}} \mathbf{\tilde{\Psi}} \mathbf{\tilde{F}}'=\mathbf{F}\mathbf{\Psi}\mathbf{F}'$.
While the proof derived in previous sections does not directly apply to this case, it can be adapted very easily. For instance, throughout the paper, we used the fact that $\mathbf{D}$ was diagonal to isolate the identification problem of $\alpha$ from the one of the rest of the parameters. In the case of the differenced data, the counterpart of the matrix $\mathbf{D}$, $\mathbf{\dot{D}}$, is not diagonal, so we cannot use the same argument. However, it is easy to note that the matrix:
is such that the elements outside the main diagonal and the adjacent diagonals do not depend on $\mathbf{\tilde{\dot{D}}}$. In particular, the tridiagonal exclusion minors of $\mathbf{\dot{O}}$ only depend on the true DGP parameters $\theta^0$ and $\tilde{\alpha}$.
These restrictions ensure that no entry of the tridiagonal band (main diagonal, first super-, or subdiagonal) is included. The properties of tridiagonal exclusion minors of $\mathbf{\dot{O}}$ are the same as those derived for diagonal exclusion minors of $\mathbf{O}$ in Section (ref). The change of $\mathbf{D}$ to $\mathbf{\dot{D}}$ does not affect the essence of the result, it simply requires re-defining some of the objects. Thus, our proof follows smoothly for tridiagonal exclusion minors if we impose similar assumptions on the parameters as we did before. Let $T$ denote the dimension of the data in levels, so that we have $T-1$ periods when working in differences. The assumptions required for almost sure identification in the differenced data are:
The assumptions are basically the same we imposed in Section (ref) but applied to the differenced processes. The only major difference appears in the time periods required for identification (Assumption (ref)). We need three more periods to achieve identification in this case than in the baseline case of the paper. There are two reasons for this. First, differencing the data mechanically reduces the number of observations by 1, so in order to have $T$ observations for the differenced model there must be $T+1$ observations in levels.
The second reason is that in order to have $(\bar{r}+1)\times(\bar{r}+1)$ tridiagonal exclusion minors the original matrix should at least be of dimension $2(\bar{r}+1)+1$. However, when $T=2(\bar{r}+1)+1$ we can only have one tridiagonal exclusion minor because $\mathbf{\dot{O}}$ is symmetric (see Appendix (ref) for a proof), and our identification strategy requires having at least two different tridiagonal exclusion minors. Thus, we need to have at least one additional observation to be able to build many tridiagonal exclusion minors. For example, if there is a single factor $\bar{r}=1$ and we want to build $2\times2$ tridiagonal exclusion minors of $\mathbf{\dot{O}}$, we should have that $dim(\mathbf{\dot{O}})\geq 6$. In our original proof we only needed $dim(\mathbf{O})\geq 4$. In general, this adds two additional observations. Once this is taken into account, we can use the techniques employed in the previous sections to show that the autoregressive coefficient $\alpha$ is almost surely identified when we consider the differenced data:
This result contradicts the claim stated in hayakawa2023short regarding the impossibility of attaining global identification in dynamic panel models additive time and fixed effects, as well as interactive effects. Moreover, williams2020identification extends anderson1956statistical to more general structures of dependence between the errors, not requiring their covariance matrix to be diagonal. We can use their results to show that the rest of the parameters are globally identified as well.
Consider a dynamic panel with individual fixed effects as in the canonical work of arellano1991some:
where $\varepsilon_{it}$ are i.i.d.\ across $i$, independent of $\gamma_i$, satisfy $\mathbb E[\varepsilon_{it}]=0$, and have $\mathbb{V}(\varepsilon_{it})=d_t$.
In arellano1991some, the restriction $|\alpha|<1$ is imposed. Subsequent work argues that $\alpha$ is not identified in the unit-root case based on their estimation method (see blundell1998initial and sentana2024finite). In this subsection we show that, in fact, $\alpha$ is almost surely identified for any value, including $\alpha=1$.
To proceed, we augment (ref) with time effects:
As in previous sections, we treat the first observation separately and project $y_{i1}$ onto the span of $[1,\gamma_i]$:
$$y_{i1} \;=\; \delta_1^{*} \;+\; f_\gamma\, \gamma_i \;+\; \varepsilon^{*}_{i1}$$
where $(\delta_1^{*},f_\gamma)$ are projection coefficients and $\varepsilon^{*}_{i1}$ is orthogonal to $\gamma_i$ by construction. Since we place no restrictions on the distribution of $\varepsilon_{i1}$, we simply relabel $(\delta_1^{*},\varepsilon_{i1}^{*})$ as $(\delta_1,\varepsilon_{i1})$ without loss of generality.
We include time effects $\delta_t$ for generality, even if our method does not rely on time effects to identify the model. If (ref) is the true DGP, one may impose $\delta_t=0$ for $t\ge2$ (and retain a free $\delta_1$ arising from the projection at $t=1$) to improve efficiency. In any case, $\delta_t$ can be concentrated out as in Section (ref), and our identification arguments are based on second-moment conditions that are invariant to this step.
As we have outlined in subsection (ref), we can write a dynamic panel model with individual fixed effects as an interactive fixed-effects model with a unique factor given by $\mathbf{F}=[f_\gamma,1,\dots,1]'$ and factor loadings given by $\gamma_i$ (this is a particular case of the model discussed in subsection (ref) when $\bar{r}=1$ and the only factor corresponds to the individual fixed-effects). Then, we can stack observations over $t$ as before and write:
$$y_{i} = \mathbf{\Gamma}\delta + \mathbf{\Gamma}\mathbf{F}\gamma_i+\mathbf{\Gamma}\varepsilon_{i}$$
where everything is defined as before except that now $\mathbf{F}=[f_\gamma,1,\dots,1]'$ and $\gamma_i$ is a scalar containing the individual fixed effect of observation $i$. Also note that the matrix $\mathbf{\Psi}$ is now a scalar $\Psi$ corresponding to the population variance of the individual fixed effects $\mathbb{V}(\gamma_i)$. The assumptions required for identification are:
Remark: Assumption (ref) is just a technical device to simplify exposition. Another way of seeing this, is that the parameters of the model are globally identified except for a set of projection coefficients of the initial observation that have Lebesgue measure zero in the support of $f_\gamma$. This means that unless very particular and strict conditions are imposed on the initial observation, any value of $\alpha$ can be identified in a dynamic panel with individual effects, including $\alpha=1$. Moreover, we also conjecture that this assumption is not strictly needed to attain identification, but it substantially simplifies the proof exposition without having major consequences in terms of generalization.
Under these assumptions $\alpha$ can be almost surely globally identified:
Once $\alpha$ is identified the rest of the parameters are identified as well:
The results derived in this paper can be used to establish identification in a wide range of dynamic models featuring time, individual, or interactive effects. The techniques of the proof we have introduced can be used to accommodate models with more sophisticated structures in the error term than a heteroskedastic white noise and provide a novel framework for addressing identification in these settings.
For example, going back to the baseline model of the paper (ref), an immediate implication of the results derived in the previous subsections is that if the error $\varepsilon_{it}$ is an MA(q) process then almost sure global identification of $\alpha$ can be derived using minors that exclude all entries within the band formed by the main diagonal and the first $q$ super- and subdiagonals of $\mathbf{O}$ provided that the time dimension is big enough.
More generally, we believe that some modifications of our original proof can be used to show identification in a wide range of dynamic models where the covariance between the error terms feature some level of sparsity (i.e. the matrix $\mathbf{D}$ has \enquote*{enough} zeros).
In this paper, we demonstrate that the parameters of a dynamic panel data model with interactive effects can be globally identified. We extend the seminal result of anderson1956statistical to dynamic panels with interactive effects, a problem that was previously thought to be non-globally identifiable.
The significance of our result lies in the fact that, despite the high degree of nonlinearity in the problem, when the data-generating process corresponds to the specified model, there exists a one-to-one correspondence between the population moments and the model parameters.
Our proof shows that global identification is achievable when \( T \geq 2(\bar{r} + 1) \). An interesting open question is whether global identification can be attained under weaker conditions, such as when \( T \) satisfies the Ledermann bound plus one. The requirement for an additional period, beyond the Ledermann bound, arises from the dynamic nature of the model, which includes an extra parameter compared to the static case.