Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
111,269 characters · 14 sections · 82 citation commands
When it counts - Econometric identification of the basic factor model based on GLT structures
{\em Keywords:} Identifiability; sparsity; rank deficiency; rotational invariance; variance identification
\centerline{JEL classification: C11, C38, C63}
Ever since the pioneering work of thu:vec,thu:mul, factor analysis has been a popular method to model the covariance matrix ${\mathbf \Omega}$ of correlated, multivariate observations ${\mathbf y}_t$ of dimension $m$, see e.g.\ and:int for a comprehensive review. Assuming $r$ uncorrelated factors, the basic factor model yields the representation ${\mathbf \Omega}= \boldsymbol{\Lambda} \boldsymbol{\Lambda} ^{\top} + {\mathbf \Sigma}_0 $, with a $m\times r$ factor loading matrix $\boldsymbol{\Lambda}$ and a diagonal matrix ${\mathbf \Sigma}_0 $. The considerable reduction of the number of parameters compared to the $m(m+1)/2$ elements of an unconstrained covariance matrix ${\mathbf \Omega}$ is the main motivation for applying factor models to covariance estimation, especially if $m$ is large; see, among many others, fan-etal:hig_je in finance and for-etal:ope in economics. In addition, shrinkage estimation has been shown to lead to very efficient covariance estimation, see, for example, kas:spa in Bayesian factor analysis and led-wol:pow in a non-Bayesian context.
In numerous applications, factor analysis reaches beyond covariance modelling. From the very beginning, the goal of factor analysis has been to extract the underlying loading matrix $\boldsymbol{\Lambda}$ to understand the driving forces behind the observed correlation between the features, see e.g.\ owe-wan:bic for a recent review. However, also in this setting, the only source of information is the observed covariance of the data, making the decomposition of the covariance matrix ${\mathbf \Omega}$ into the cross-covariance matrix $\boldsymbol{\Lambda} \boldsymbol{\Lambda}^{\top}$ and the variance $ {\mathbf \Sigma}_0 $ of the idiosyncratic errors more challenging than estimating only ${\mathbf \Omega}$ itself.
A huge literature, dating back to koo-rei:ide and rei:ide, has addressed this problem of identification which can be resolved only by imposing additional structure on the factor model. and-rub:sta considered identification as a two-step procedure, namely identification of ${\mathbf \Sigma}_0$ from ${\mathbf \Omega}$ (variance identification) and subsequent identification of $\boldsymbol{\Lambda}$ from $\boldsymbol{\Lambda} \boldsymbol{\Lambda} ^{\top}$ (solving rotational invariance). The most popular constraint in econometrics, statistics and machine learning for solving rotational invariance is to consider positive lower triangular loading matrices, see e.g.\ gew-zho:mea,wes:bay_fac,lop-wes:bay, albeit other strategies have been put forward, see e.g.\ neu:ide, bai-ng:pri, ass-etal:bay, cha-etal:inv, and wil:ide. Only a few papers have addressed variance identification bek:ide and to the best of our knowledge so far no structure has been put forward that simultaneously addresses both identification problems.
In this work, we discuss a new identification strategy based on generalized lower triangular (GLT) structures, see Figure (ref) for illustration. This concept was originally introduced as part of an MCMC sampler for sparse Bayesian factor analysis where the number of factors is unknown in the (unpublished) work of fru-lop:spa. In the present paper, GLT structures are given a full and comprehensive mathematical treatment and are applied in fru-etal:spa to develop an efficient reversible jump MCMC (RJMCMC) sampler for sparse Bayesian factor analysis under very general shrinkage priors. It will be proven that GLT structures simultaneously address rotational invariance and variance identification in factor models. Variance identification relies on a counting rule for the number of non-zero elements in the loading matrix $\boldsymbol{\Lambda}$, which is a sufficient condition that extends previous work by sat:stu.
In addition, we will show that GLT structures are useful in exploratory factor analysis where the factor dimension $r$ is unknown. Identification of the number of factors in applied factor analysis is a notoriously difficult problem, with considerable ambiguity which method works best, be it BIC-type criteria bai-ng:det2002, marginal likelihoods lop-wes:bay, techniques from Bayesian non-parametrics involving infinite-dimensional factor models bha-dun:spa,roc-geo:fas,leg-etal:bay or more heuristic procedures kau-sch:bay. Imposing an unordered GLT structure in exploratory factor analysis allows to identify the true loading matrix $\boldsymbol{\Lambda}$ and the matrix ${\mathbf \Sigma}_0 $ and to easily spot all spurious columns in a possibly overfitting model. This strategy underlies the RJMCMC sampler of fru-etal:spa to estimate the number of factors.
The paper is structured as follows. Section (ref) reviews the role of identification in factor analysis using illustrative examples. Section (ref) introduces GLT structures, proves identification for sparse GLT structures and shows that any unconstrained loading matrix has a unique representation as a GLT matrix. Section (ref) addresses variance identification under GLT structures. Section (ref) discusses exploratory factor analysis under unordered GLT structures, while Section (ref) presents an illustrative application. Section (ref) concludes.
Let ${\mathbf y}_t=(y_{1t}, \ldots, y_{m t})^{\top}$ be an observation vector of $m$ measurements, which is assumed to arise from a multivariate normal distribution, ${\mathbf y}_t \sim \mathcal{N} _{m}\left({\mathbf{0}},{\mathbf \Omega}\right)$, with zero mean and covariance matrix ${\mathbf \Omega}$. In factor analysis, the correlation among the observations is assumed to be driven by a latent $r$-variate random variable ${\mathbf f}_t=(f_{1t}, \ldots, f_{r t})^{\top}$, the so-called common factors, through the following observation equation:
where the $m\times r$ matrix $\boldsymbol{\Lambda}$ containing the factor loadings $\Lambda_{ij}$ is of full column rank, $\mbox{\rm rk}\,(\boldsymbol{\Lambda})=r$, equal to the factor dimension $r$. In the present paper, we focus on the so-called basic factor model where the vector $\boldsymbol{\epsilon}_t =(\epsilon_{1t}, \ldots, \epsilon_{m t})^{\top}$ accounts for independent, idiosyncratic variation of each measurement and is distributed as $\boldsymbol{\epsilon}_t \sim \mathcal{N} _{m}\left({\mathbf{0}},{\mathbf \Sigma}_0\right) $, with ${\mathbf \Sigma}_0=\mbox{\rm Diag}\!\left(\sigma^2_{1},\ldots,\sigma^2_{m}\right)$ being a positive definite diagonal matrix. The common factors are orthogonal, meaning that ${\mathbf f}_t \sim \mathcal{N} _{r}\left({\mathbf{0}},{{\mathbf I}}_{r}\right),$ and independent of $\boldsymbol{\epsilon}_t$. In this case, the observation equation ((ref)) implies the following covariance matrix ${\mathbf \Omega}$, when we integrate w.r.t.\ the latent common factors ${\mathbf f}_t$:
Hence, all dependence among the measurements in ${\mathbf y}_t$ is explained through the latent common factors and the off-diagonal elements of $\boldsymbol{\Lambda} \boldsymbol{\Lambda} ^{\top}$ define the marginal covariance between any two measurements $y_{i_1,t}$ and $y_{i_2, t}$:
where $\boldsymbol{\Lambda} _{i,\bullet}$ is the $i$th row of $\boldsymbol{\Lambda}$. Consequently, we will refer to $\boldsymbol{\Lambda} \boldsymbol{\Lambda} ^{\top}$ as the cross-covariance matrix. Since the number of factors, $r$, is often considerably smaller than the number of measurements, $m$, ((ref)) can be seen as a parsimonious representation of the dependence between the measurements, often with considerably fewer parameters in $\boldsymbol{\Lambda}$ than the $m(m -1)/2$ off-diagonal elements in an unconstrained covariance matrix ${\mathbf \Omega}$.
Since the factors ${\mathbf f}_t$ are unobserved, the only information available to estimate $\boldsymbol{\Lambda}$ and ${\mathbf \Sigma}_0$ is the covariance matrix ${\mathbf \Omega}$. A rigorous approach toward identification of factor models was first offered by rei:ide and and-rub:sta. Identification in the context of a basic factor model means the following. For any pair $(\boldsymbol{\beta} ,{\mathbf \Sigma})$, where $\boldsymbol{\beta}$ is an $m \times r$ matrix and ${\mathbf \Sigma}$ is a positive definite diagonal matrix, that satisfies ((ref)), i.e.:
it follows that $\boldsymbol{\beta}=\boldsymbol{\Lambda}$ and ${\mathbf \Sigma} ={\mathbf \Sigma}_0$. Note that both parameter pairs imply the same Gaussian distribution ${\mathbf y}_t \sim \mathcal{N} _{m}\left({\mathbf{0}},{\mathbf \Omega}\right)$ for every possible realisation ${\mathbf y}_t$.
and-rub:sta considered identification as a two-step procedure. The first step is identification of the variance decomposition, i.e.\ identification of ${\mathbf \Sigma}_0$ from ((ref)), which implies identification of $\boldsymbol{\Lambda} \boldsymbol{\Lambda} ^{\top}$. The second step is subsequent identification of $\boldsymbol{\Lambda}$ from $\boldsymbol{\Lambda} \boldsymbol{\Lambda} ^{\top}$, also know as solving the rotational invariance problem. The literature on factor analysis often reduces identification of factor models to the second problem, however as we will argue in the present paper, variance identification is equally important.
\paragraph*{Rotational invariance.} Let us assume for the moment that $\boldsymbol{\Lambda} \boldsymbol{\Lambda} ^{\top}$ is identified. Consider, for further illustration, the following factor loading matrix $\boldsymbol{\Lambda} $ and a loading matrix $\boldsymbol{\beta} = \boldsymbol{\Lambda} {\mathbf P}_{\alpha b} $ defined as a rotation of $\boldsymbol{\Lambda} $:
For any $\alpha \in [0, 2\pi)$ and $b\in\{0,1\}$, the factor loading matrix $\boldsymbol{\beta} $ yields the same cross-covariance matrix for ${\mathbf y}_t$ as $\boldsymbol{\Lambda}$, as is easily verified:
The rotational invariance apparent in ((ref)) holds more generally for any basic factor model ((ref)). Take any arbitrary $r\times r$ rotation matrix ${\mathbf P}$ (i.e.\ ${\mathbf P} {\mathbf P}^{\top}= {{\mathbf I}}_{r}$) and define the basic factor model
where $\boldsymbol{\beta} = \boldsymbol{\Lambda} {\mathbf P}$ and ${\mathbf f}_t^{\star} = {\mathbf P}^{\top} {\mathbf f}_t$. Then both models imply the same covariance ${\mathbf \Omega}$, given by ((ref)). Hence, without imposing further constraints, $\boldsymbol{\Lambda}$ is in general not identified from the cross-covariance matrix $ \boldsymbol{\Lambda} \boldsymbol{\Lambda}^{\top} $. If interest lies in interpreting the factors through the factor loading matrix $\boldsymbol{\Lambda}$, rotational invariance has to be resolved. The usual way of dealing with rotational invariance is to constrain $\boldsymbol{\Lambda}$ in such a way that the only possible rotation is the identity ${\mathbf P}={{\mathbf I}}_{r}$. For orthogonal factors at least $r(r-1)/2$ restrictions on the elements of $\boldsymbol{\Lambda}$ are needed to eliminate rotational indeterminacy and-rub:sta.
The most popular constraints are positive lower triangular (PLT) loading matrices, where the upper triangular part is constrained to be zero and the main diagonal elements $\Lambda_{11},\ldots, \Lambda_{r r}$ of $\boldsymbol{\Lambda}$ are strictly positive, see Figure (ref) for illustration. Despite its popularity, the PLT structure is restrictive, as outlined already by joe:gen. Let $\boldsymbol{\beta} \boldsymbol{\beta}^{\top} $ be an arbitrary cross-covariance matrix with factor loading matrix $\boldsymbol{\beta}$. A PLT representation of $ \boldsymbol{\beta} \boldsymbol{\beta}^{\top} $ is possible iff a rotation matrix $ {\mathbf P}$ exists such that $\boldsymbol{\beta} $ can be rotated into a PLT matrix $\boldsymbol{\Lambda} = \boldsymbol{\beta} {\mathbf P} $. However, as example ((ref)) illustrates this is not necessarily the case. Obviously, $ \boldsymbol{\Lambda}$ is not a PLT matrix, since $\Lambda_{22}=0$. Any of the possible rotations $\boldsymbol{\beta} = \boldsymbol{\Lambda} {\mathbf P}_{\alpha b}$ have {\em non-zero} elements above the main diagonal and are not PLT matrices either. This example demonstrates that the PLT representation is restrictive. To circumvent this problem in example ((ref)), one could reorder the measurements in an appropriate manner. However, in applied factor analysis, such an appropriate ordering is typically not known in advance and the choice of the first $r$ measurements is an important modeling decision under PLT constraints, see e.g.\ lop-wes:bay and car-etal:hig.
We discuss in Section (ref) a new identification strategy to resolve rotational invariance in factor models based on the concept of generalized lower triangular (GLT) structures. Loosely speaking, GLT structures generalize PLT structures by freeing the position of the first non-zero factor loading in each column, see the loading matrix $ \boldsymbol{\Lambda}$ in ((ref)) and Figure (ref) for an example. We show in Section (ref) that a unique GLT structure $\boldsymbol{\Lambda}$ can be identified for any cross-covariance matrix $\boldsymbol{\beta} \boldsymbol{\beta}^{\top}$, provided that variance identification holds and, consequently, $\boldsymbol{\beta} \boldsymbol{\beta}^{\top}$ itself is identified. Even if $ \boldsymbol{\beta} \boldsymbol{\beta}^{\top} $ is obtained from a loading matrix $\boldsymbol{\beta}$ that does not take the form of a GLT structure, such as the matrix $\boldsymbol{\beta}$ in ((ref)), we show in Section (ref) that a {\em unique} orthogonal matrix ${\mathbf G}$ exists which represents $\boldsymbol{\beta}$ as a rotation of a unique GLT structure $\boldsymbol{\Lambda}$:
which we call rotation into GLT. Hence, the GLT representation is unrestrictive in the sense of joe:gen and is, indeed, a new and generic way to resolve rotational invariance for any factor loading matrix.
\paragraph*{Sparse factor loading matrices.} The factor loading matrix $\boldsymbol{\Lambda}$ given in ((ref)) is an example of a sparse loading matrix. While only a single zero loading would be needed to resolve rotational invariance, six zeros are present and each factor loads only on dedicated measurements. Such sparse loading matrices are generated by a binary indicator matrix $\boldsymbol{\delta}$ of 0s and 1s of the same dimension as $\boldsymbol{\Lambda}$, where $\Lambda_{ij}=0$ iff $\delta_{ij}=0$, and $\Lambda_{ij} \in \mathbb{R}$ is unconstrained otherwise. The binary matrix $\boldsymbol{\delta}=\mathbb{I} (\boldsymbol{\Lambda} \neq 0)$, where the indicator function is applied element-wise, is called the sparsity matrix corresponding to $\boldsymbol{\Lambda}$. The sparsity matrix $\boldsymbol{\delta}$ contains a lot of information about the structure of $\boldsymbol{\Lambda}$, see Figure (ref) for illustration. The indicator matrix on the right hand side tells us that $\boldsymbol{\Lambda}$ obeys the PLT constraint. The fifth row of the left and center matrices contains only zeros, which tells us that observation $y_{5t}$ is uncorrelated with the remaining observations, since $\mbox{\rm Cov}(y_{it}, y_{5t}) = 0$ for all $i\neq 5$.
\paragraph*{Variance identification.} Constraints that resolve rotational invariance typically take variance identification, i.e.\ identification of $\boldsymbol{\Lambda} \boldsymbol{\Lambda} ^{\top}$, for granted, see e.g.\ gew-zho:mea. Variance identification refers to the problem that the idiosyncratic variances $\sigma^2_1, \ldots, \sigma^2_m$ in ${\mathbf \Sigma}_0$ are identified only from the diagonal elements of ${\mathbf \Omega}$, as all other elements are independent of the $\sigma^2_i$s; see again (ref). To achieve variance identification of $\sigma^2_i$ from $\Omega_{ii}= \boldsymbol{\Lambda} _{i,\bullet} \boldsymbol{\Lambda} _{i,\bullet}^{\top} + \sigma^2_i$, all factor loadings have to be identified solely from the off-diagonal elements of ${\mathbf \Omega}$. Variance identification, however, is easily violated, as the following considerations illustrate.
Let us return to the factor model defined in ((ref)). The corresponding covariance matrix ${\mathbf \Omega}$ is given by: {
} Let us assume that the sparsity pattern $\boldsymbol{\delta}=\mathbb{I} (\boldsymbol{\Lambda} \neq 0)$ of $\boldsymbol{\Lambda}$ is known, but the specific values of the unconstrained loadings $(\lambda_{11}, \ldots, \lambda_{62})$ are unknown. An interesting question is the following. Knowing ${\mathbf \Omega}$, can the unconstrained loadings $\lambda_{11}, \ldots, \lambda_{62}$ and the variances $\sigma^2_1, \ldots, \sigma^2_m$ be identified uniquely? Given ${\mathbf \Omega}$, the three nonzero covariances $\mbox{\rm Cov}(y_{1t}, y_{2t})= \lambda_{11} \lambda_{21}$, $\mbox{\rm Cov}(y_{1t}, y_{3t})=\lambda_{11} \lambda_{31}$ and $\mbox{\rm Cov}(y_{2t}, y_{3t})=\lambda_{21} \lambda_{31}$ are available to identify the three factor loadings $(\lambda_{11}, \lambda_{21}, \lambda_{31})$. Similarly, the nonzero covariances $\mbox{\rm Cov}(y_{4t}, y_{5t})= \lambda_{42} \lambda_{52}$, $\mbox{\rm Cov}(y_{4t}, y_{6t})=\lambda_{42} \lambda_{62}$ and $\mbox{\rm Cov}(y_{5t}, y_{6t})=\lambda_{52} \lambda_{62}$ are available to identify the factor loadings $(\lambda_{42}, \lambda_{52}, \lambda_{62})$, hence variance identification is given. However, if we remove the last measurement from the loading factor matrix defined in ((ref)), we obtain
and the corresponding covariance matrix reads:
While the three factor loadings $(\lambda_{11}, \lambda_{21}, \lambda_{31})$ are still identified from the off-diagonal elements of ${\mathbf \Omega}$ as before, variance identification of $\sigma^2_4$ and $\sigma^2_5$ fails. Since $\mbox{\rm Cov}(y_{4t}, y_{5t})= \lambda_{42} \lambda_{52}$ is the only non-zero element that depends on the loadings $\lambda_{42}$ and $\lambda_{52}$, infinitely many different parameters $(\lambda_{42},\lambda_{52}, \sigma^2_4,\sigma^2_5)$ imply the same covariance matrix ${\mathbf \Omega}$. From these considerations it is evident that a minimum of three non-zero loadings is necessary in each column to achieve variance identification, a condition which has been noted as early as and-rub:sta. At the same time, this condition is not sufficient, as it is satisfied by the loading matrix $\boldsymbol{\beta}$ in (ref), although variance identification does not hold. In general, variance identification is not straightforward to verify. We will introduce in Section (ref) a new and convenient way to verify variance identification for GLT structures.
\paragraph*{The row deletion property.} As explained above, we need to verify uniqueness of the variance decomposition, i.e.\ the identification of the idiosyncratic variances $\sigma^2_1, \ldots, \sigma^2_m$ in ${\mathbf \Sigma}_0$ from the covariance matrix ${\mathbf \Omega}$ given in ((ref)). The identification of ${\mathbf \Sigma}_0$ guarantees that $\boldsymbol{\Lambda} \boldsymbol{\Lambda} ^{\top}$ is identified. The second step of identification is then to ensure uniqueness of the factor loadings, i.e.\ unique identification of $\boldsymbol{\Lambda}$ from $\boldsymbol{\Lambda} \boldsymbol{\Lambda} ^{\top}$. To verify variance identification, we rely in the present paper on a condition known as {\em row-deletion property}.
and-rub:sta prove that the row-deletion property is a sufficient condition for the identification of $\boldsymbol{\Lambda} \boldsymbol{\Lambda}^{\top} $ and ${\mathbf \Sigma}_0$ from the marginal covariance matrix ${\mathbf \Omega}$ given in ((ref)). For any (not necessarily GLT) factor loading matrix $\boldsymbol{\Lambda}$, the row deletion property \bf AR\ can be trivially tested by a step-by-step analysis, where each single row of $\boldsymbol{\Lambda}$ is sequentially deleted and the two distinct submatrices are determined from examining the remaining matrix, as suggested e.g.\ by hay-mar:exa. However, this procedure is inefficient and challenging in higher dimensions.
Hence, it is helpful to have more structural conditions for verifying variance identification under the row deletion property \bf AR . The literature provides several necessary conditions for the row deletion property \bf AR\ that are based on counting the number of non-zero factor loadings in $\boldsymbol{\Lambda}$. and-rub:sta, for instance, prove the following necessary conditions for \bf AR : for every nonsingular $r$-dimensional square matrix ${\mathbf G}$, the matrix $\boldsymbol{\beta}=\boldsymbol{\Lambda} {\mathbf G}$ contains in each column at least 3 and in each pair of columns at least 5 nonzero factor loadings. sat:stu extends these {\em necessary} conditions in the following way: every subset of $1\leq q \leq r$ columns of $\boldsymbol{\beta}=\boldsymbol{\Lambda} {\mathbf G}$ contains at least $2q+1$ nonzero factor loadings for every nonsingular matrix ${\mathbf G}$. We call this the $3579$ counting rule for obvious reasons.
For illustration, let us return to the examples in ((ref)) and ((ref)). First, apply the $3579$ counting rule to the unrestricted matrix $\boldsymbol{\beta}$ in ((ref)). Although the variance decomposition ${\mathbf \Omega}= \boldsymbol{\beta} \boldsymbol{\beta} ^{\top} + {\mathbf \Sigma} = \boldsymbol{\Lambda} \boldsymbol{\Lambda} ^{\top} + {\mathbf \Sigma}_0,$ is not unique, the counting rules are not violated, since $\boldsymbol{\beta}$ has five non-zero rows except for the cases $(\alpha,b) \in \{0, \frac{\pi}{2}, \pi, \frac{3\pi}{2}\}\times \{0,1\}$. Only for these eight specific cases, which correspond to the trivial rotations
we find immediately that the counting rules are violated, since one of the two columns has only two non-zero elements. This example shows the need to check such counting rules not only for a single loading matrix $\boldsymbol{\beta}$, but also for all rotations $\boldsymbol{\beta} {\mathbf P}$ admissible under the chosen strategy toward rotational invariance. On the other hand, if we apply the $3579$ counting rule to the unrestricted matrix $\boldsymbol{\beta}$ in ((ref)), we find that the {\em necessary} counting rules are satisfied for all rotations $\boldsymbol{\beta} {\mathbf P}_{\alpha b}$. For this specific example, we have already verified explicitly that variance identification holds and one might wonder if, in general, the $3579$ counting rule can lead to a {\em sufficient} criterion for variance identification under \bf AR .
Sufficient conditions for variance identification are hardly investigated in the literature. One exception is the popular factor analysis model where $\boldsymbol{\Lambda}$ takes the form of a {\em dense} PLT matrix, where all factor loadings on and below the main diagonal are left unrestricted and can take any in value in $\mathbb{R}$. For this model, condition \bf AR\ and hence variance identification holds, except for a set of measure 0, if the condition $m \geq 2r +1$ is satisfied. con-etal:bay_exp investigate identification of a dedicated factor model, where equation ((ref)) is combined with correlated (oblique) factors, ${\mathbf f}_t \sim \mathcal{N} _{r}\left({\mathbf{0}},\mathbf{R}\right)$, and the factor loading matrix $\boldsymbol{\Lambda}$ has a perfect simple structure, i.e.\ each observation loads on at most one factor, as in ((ref)) and ((ref)); however, the exact position of the non-zero elements is unknown. They prove necessary and sufficient conditions that imply uniqueness of the variance decomposition as well as uniqueness of the factor loading matrix, namely: the correlation matrix $\mathbf{R}$ is of full rank ($\mbox{\rm rk}\,(\mathbf{R})=r$) and each column of $\boldsymbol{\Lambda}$ contains at least three nonzero loadings.
In the present paper, we build on and extend this previous work. We provide sufficient conditions for variance identification of a GLT structure $\boldsymbol{\Lambda}$. These conditions are formulated as counting rules for the $m \times r$ sparsity matrix $\boldsymbol{\delta} = \mathbb{I} (\boldsymbol{\beta}\neq 0)$ of $\boldsymbol{\beta}$ and are equivalent to the $3579$ counting rules of sat:stu. More specifically, if the $3579$ counting rule holds for the sparsity matrix $\boldsymbol{\delta}$ of a GLT matrix $\boldsymbol{\Lambda}$, then this is a sufficient condition for the row deletion property \bf AR\ and consequently for variance identification, except for a set of measure 0.
\paragraph*{Identification of the number of factors.}
Identification of the number of factors is a notoriously difficult problem and analysing this problem from the view point of variance identification is helpful in understanding some fundamental difficulties. Assume that ${\mathbf \Omega}$ has a representation as in ((ref)) with $r$ factors which is variance identified. Then, on the one hand, no equivalent representation exists with $r' < r$ number of factors. On the other hand, as shown in rei:ide, any such structure $(\boldsymbol{\Lambda},{\mathbf \Sigma}_0)$ creates solutions $(\boldsymbol{\beta}_k, {\mathbf \Sigma}_k)$ with $m \times k$ loading matrices $\boldsymbol{\beta}_k$ of dimension $k= r+1, r+2, \ldots, m$ bigger than $r$ and ${\mathbf \Sigma}_k$ being a positive definite matrix different from ${\mathbf \Sigma}_0$ which imply the same covariance matrix ${\mathbf \Omega}$ as $(\boldsymbol{\Lambda},{\mathbf \Sigma}_0)$, i.e.:
Furthermore, for any fixed $k > r$, infinitely many such solutions $(\boldsymbol{\beta}_k, {\mathbf \Sigma}_k) $ can be created that satisfy the decomposition ((ref)) which, consequently, no longer is variance identified. This problem is prevalent regardless of the chosen strategy toward rotational invariance. For illustration, we return to example ((ref)) and construct an equivalent solution for $k=3$. While the first two columns of $\boldsymbol{\beta}_3$ are equal to $\boldsymbol{\Lambda}$, the third column is a so-called {\em spurious} factor with a single non-zero loading and ${\mathbf \Sigma}_3 $ is defined as follows:
We can place the spurious factor loading $\beta_{i3}$ in any row $i$ and $\beta_{i3}$ can take any value satisfying $0 < \beta_{i3}^2 < \sigma^2_i$. It is easy to verify that any such pair $(\boldsymbol{\beta}_3,{\mathbf \Sigma}_3)$ indeed implies the same covariance matrix ${\mathbf \Omega}$ as in ((ref)).
This ambiguity in an overfitting model renders the estimation of true number of factors $r$ a challenging problem and leads to considerable uncertainty how to choose the number of factors in applied factor analysis. In Section (ref), we follow up on this problem in more detail. An important necessary condition for $k$ to be the true number of factors is that variance identification of ${\mathbf \Sigma}_k$ in ((ref)) holds. Therefore, the counting rules that we introduce in this paper will also be useful in cases where the true number of factors $r$ is unknown.
\paragraph*{Overfitting GLT structures.} Finally, we investigate in Section (ref) the class of potentially overfitting GLT structures where the matrix $\boldsymbol{\beta}_k $ in ((ref)) is constrained to be an unordered GLT structure. We apply results by tum-sat:ide to this class and show how easily spurious factors and the underlying true factor loading matrix $\boldsymbol{\Lambda}$ are identified under GLT structures, even if the model is overfitting. Our strategy relies on the concept of extended variance identification and the extended row deletion property introduced by tum-sat:ide, where more than one row is deleted from the loading matrix. An extended counting rule will be introduced for the sparsity matrix of a GLT loading matrices $\boldsymbol{\beta}_k$ in Section (ref) which is useful in this context.
In this work, we introduce a new identification strategy to resolve rotational invariance based on the concept of generalized lower triangular (GLT) structures. First, we introduce the notion of {\em pivot rows} of a factor loading matrix $\boldsymbol{\Lambda}$.
For PLT factor loading matrices the pivot rows lie on the main diagonal, i.e.\ $(l_1, \ldots, l_r)=(1,\ldots,r)$, and the leading factor loadings $\Lambda_{jj} > 0$ are positive for all columns $j=1,\ldots,r$. GLT structures generalize the PLT constraint by freeing the pivot rows of a factor loading matrix $\boldsymbol{\Lambda}$ and allowing them to take arbitrary positions $(l_1, \ldots, l_r)$, the only constraint being that the pivot rows are pairwise distinct. GLT structures contain PLT matrices as the special case where $l_j= j$ for $j=1,\ldots,r$. Our generalization is particularly useful if the ordering of the measurements $y_{it}$ is in conflict with the PLT assumption. Since $\Lambda_{jj}$ is allowed to be 0, measurements different from the first $r$ ones may lead the factors. For each factor $j$, the leading variable is the response variable $y_{l_j,t}$ corresponding to the pivot row $l_j$.
We will distinguish between two types of GLT structures, namely ordered and unordered GLT structures. The following definition introduces ordered GLT matrices. Unordered GLT structures will be motivated and defined below. Examples of ordered and unordered GLT matrices are displayed in Figure (ref) for a model with $r=6$ factors.
Evidently, imposing an ordered GLT structure resolves rotational invariance if the pivot rows are known. For any two ordered GLT matrices $\boldsymbol{\beta}$ and $ \boldsymbol{\Lambda} $ with {\em identical} pivot rows $l_1 , \ldots , l_r$, the identity $\boldsymbol{\beta} = \boldsymbol{\Lambda} {\mathbf P}$ evidently holds iff ${\mathbf P}={{\mathbf I}}_{r}$. In practice, the pivot rows $l_1 , \ldots , l_r$ of a GLT structure are unknown and need to be identified from the marginal covariance matrix ${\mathbf \Omega}$ for a given number of factors $r$. Given variance identification, i.e.\ assuming that the cross-covariance matrix $\boldsymbol{\Lambda} \boldsymbol{\Lambda}^{\top}$ is identified, a particularly important issue for the identification of a GLT factor model is whether $\boldsymbol{\Lambda} $ is uniquely identified from $\boldsymbol{\Lambda} \boldsymbol{\Lambda}^{\top}$ if the pivot rows $l_1 , \ldots , l_r$ are {\em unknown}. Non-trivial rotations $\boldsymbol{\beta} = \boldsymbol{\Lambda} {\mathbf P}$ of a loading matrix $\boldsymbol{\Lambda}$ with pivot rows $l_1 , \ldots , l_r$ might exist such that $\boldsymbol{\beta} \boldsymbol{\beta}^{\top} = \boldsymbol{\Lambda} \boldsymbol{\Lambda}^{\top} $, while the pivot rows $\tilde{l}_1 , \ldots , \tilde{l}_r$ of $\boldsymbol{\beta} $ are different from the pivot rows of $\boldsymbol{\Lambda}$. Very assuringly, Theorem (ref) shows that this is not the case: not only the pivot rows, but the entire loading matrices $\boldsymbol{\Lambda}$ and $\boldsymbol{\beta}$ are identical, if $\boldsymbol{\Lambda} \boldsymbol{\Lambda}^{\top} = \boldsymbol{\beta} \boldsymbol{\beta}^{\top}$ (see Appendix (ref) for a proof).
Definition (ref) introduces, as an extension of Definition (ref), unordered GLT structures under which $\boldsymbol{\Lambda}$ is identified from $\boldsymbol{\Lambda} \boldsymbol{\Lambda} ^{\top}$ only up to signed permutations. A signed permutation permutes the columns of the factor loading matrix $\boldsymbol{\Lambda}$ and switches the sign of all factor loadings in any specific column. This leads to a trivial case of rotational invariance. For $r=2$, for instance, the eight signed permutations of the loading matrix $\boldsymbol{\Lambda} $ defined in ((ref)) are depicted in ((ref)). More formally, $ \boldsymbol{\beta} $ is a signed permutation of $ \boldsymbol{\Lambda} $, iff
where the permutation matrix ${\mathbf P}_{\rho}$ corresponds to one of the $r$! permutations of the $r$ columns of $\boldsymbol{\Lambda}$ and the reflection matrix ${\mathbf P}_{\pm}=\mbox{\rm Diag}\!\left(\pm 1, \ldots, \pm 1\right)$ corresponds to one of the $2^r$ ways to switch the signs of the $r$ columns of $\boldsymbol{\Lambda}$. Often, it is convenient to employ identification rules that guarantee identification of $\boldsymbol{\Lambda} $ only up to such column and sign switching, see e.g.\ con-etal:bay_exp. Any structure $\boldsymbol{\Lambda}$ obeying such an identification rule represents a whole equivalence class of matrices given by all $2^r r !$ signed permutation $\boldsymbol{\beta} = \boldsymbol{\Lambda} {\mathbf P} _{\pm} {\mathbf P} _{\rho}$ of $\boldsymbol{\Lambda}$. This trivial form of the rotational invariance does not impose any additional mathematical challenges and is often convenient from a computational viewpoint, in particular for Bayesian inference, see for e.g.\ con-etal:bay_exp and fru-etal:spa.
It is easy to verify how identification up to trivial rotational invariance can be achieved for GLT structures and motivates the following definition of unordered GLT structures as loadings matrices $ \boldsymbol{\beta}$ where the pivot rows $l_1, \ldots , l_r$ simply occupy $r$ different rows. In Definition (ref), no order constraint is imposed on the pivot rows and no sign constraint is imposed on the leading factor loadings. This very general structure allows to design highly efficient sampling schemes for sparse Bayesian factor analysis under GLT structures, see fru-etal:spa.
Theorem (ref) is easily extended to unordered GLT structures. Any signed permutation $ \boldsymbol{\beta} = \boldsymbol{\Lambda} {\mathbf P} _{\rho} {\mathbf P} _{\pm}$ of $\boldsymbol{\Lambda} $ is uniquely identified from $\boldsymbol{\beta} \boldsymbol{\beta}^{\top}= \boldsymbol{\Lambda} \boldsymbol{\Lambda} ^{\top}$, provided that $\boldsymbol{\Lambda} \boldsymbol{\Lambda} ^{\top}$ is identified. Hence, under unordered GLT structures the factor loading matrix $\boldsymbol{\Lambda}$ is uniquely identified up to signed permutations. Full identification can easily be obtained from unordered GLT structures $\boldsymbol{\beta}$. Any unordered GLT structure $\boldsymbol{\beta}$ has unordered pivot rows $l_1 , \ldots , l_{r}$, occupying different rows. The corresponding ordered GLT structure $\boldsymbol{\Lambda} $ is recovered from $\boldsymbol{\beta}$ by sorting the columns in ascending order according to the pivot rows. In other words, the pivot rows of $\boldsymbol{\Lambda} $ are equal to the order statistics $l_{(1)} , \ldots , l_{(r)}$ of the pivot rows $l_1, \ldots , l_ r$ of $\boldsymbol{\beta}$, see again Figure (ref). This procedure resolves rotational invariance, since the pivot rows $l_1, \ldots , l_ r$ in the unordered GLT structure are distinct. Furthermore, imposing the condition $\Lambda_{l_j,j}>0$ in each column $j$ resolves sign switching: if $\Lambda_{l_j,j}<0$, then the sign of all factor loadings $\Lambda_{ij}$ in column $j$ is switched.
In Definition (ref) and (ref), \lq\lq structural\rq\rq\ zeros are introduced for a GLT structure for all factor loading above the pivot row $l_j$, while the factor loading $\Lambda_{l_j,j}$ in the pivot row is non-zero by definition. We call $\boldsymbol{\Lambda}$ a {\em dense} GLT structure if all loadings below the pivot row are unconstrained and can take any value in $\mathbb{R}$.
A sparse GLT structure results if factor loadings at unspecified places below the pivot rows are zero and only the remaining loadings are unconstrained. A sparse loading matrix $\boldsymbol{\Lambda}$ can be characterized by the so-called sparsity matrix, defined as a binary indicator matrix $\boldsymbol{\delta}$ of 0/1s of the same size as $\boldsymbol{\Lambda}$, where $\delta_{ij}=\mathbb{I} (\Lambda_{ij} \neq 0)$. Let $\boldsymbol{\delta} ^{\Lambda}$ be the sparsity matrix of a GLT matrix $\boldsymbol{\Lambda}$. The sparsity matrix $ \boldsymbol{\delta}$ corresponding to the signed permutation $\boldsymbol{\beta} = \boldsymbol{\Lambda} {\mathbf P} _{\rho} {\mathbf P} _{\pm}$ is equal to $ \boldsymbol{\delta} = \boldsymbol{\delta}^{\Lambda} {\mathbf P} _{\rho}$ and is invariant to sign switching. Hence, for any sparse unordered GLT matrix $\boldsymbol{\beta}$, the corresponding sparsity matrix $\boldsymbol{\delta}$ obeys an unordered GLT structure with the same pivot rows as $\boldsymbol{\beta}$, see Figure (ref) for illustration.
In sparse factor analysis, single factor loadings take zero-values with positive probability and the corresponding sparsity matrix $\boldsymbol{\delta}$ is a binary matrix that has to be identified from the data. Identification in sparse factor analysis has to provide conditions under which the entire 0/1 pattern in $\boldsymbol{\delta}$ can be identified from the covariance matrix ${\mathbf \Omega}$ if $\boldsymbol{\delta}$ is unknown. Whether this is possible hinges on variance identification, i.e.\ whether the decomposition of ${\mathbf \Omega}$ into $\boldsymbol{\Lambda} \boldsymbol{\Lambda}^{\top}$ and ${\mathbf \Sigma}_0$ is unique. How variance identification can be verified for (sparse) GLT structures is investigated in detail in Section (ref). Let us assume at this point that variance identification holds, i.e. the cross-covariance matrix $\boldsymbol{\Lambda} \boldsymbol{\Lambda}^{\top}$ is identified. Then an important step toward the identification of a sparse factor model is to verify whether the 0/1 pattern of $ \boldsymbol{\Lambda}$, characterized by $\boldsymbol{\delta}$, is uniquely identified from $\boldsymbol{\Lambda} \boldsymbol{\Lambda}^{\top}$. Very importantly, if $\boldsymbol{\Lambda} $ is assumed to be a GLT structure, then the entire GLT structure $\boldsymbol{\Lambda}$ and hence the indicator matrix $\boldsymbol{\delta}$ is uniquely identified from $\boldsymbol{\Lambda} \boldsymbol{\Lambda} ^{\top}$, as follows immediately from Theorem (ref), since $\delta_{i j}= 0$, iff $\Lambda_{i j}= 0 $ for all $i,j$. By identifying the 0/1 pattern in $\boldsymbol{\delta}$ we can uniquely identify the pivot rows of $\boldsymbol{\Lambda}$ and the sparsity pattern below.
We would like to emphasize that in sparse factor analysis with unconstrained loading matrices $\boldsymbol{\Lambda}$ this is not necessarily the case. The indicator matrix $\boldsymbol{\delta}$ is in general {\em not} uniquely identified from $\boldsymbol{\Lambda} \boldsymbol{\Lambda}^{\top}$, because (non-trivial) rotations ${\mathbf P}$ change the zero pattern in $\boldsymbol{\beta} = \boldsymbol{\Lambda} {\mathbf P}$, while $\boldsymbol{\beta} \boldsymbol{\beta}^{\top} = \boldsymbol{\Lambda} \boldsymbol{\Lambda} ^{\top}$. For illustration, let us return to the example in ((ref)) where we showed that $\boldsymbol{\Lambda} \boldsymbol{\Lambda}^{\top}$ is uniquely identified if the true sparsity matrix $\boldsymbol{\delta} ^{\Lambda}$ is known. Now assume that $\boldsymbol{\delta} ^{\Lambda}$ is unknown and allow the loading matrix $\boldsymbol{\beta} = \boldsymbol{\Lambda} {\mathbf P}$ to be any rotation of $\boldsymbol{\Lambda}$. It is then evident that the corresponding sparsity matrix $\boldsymbol{\delta}$ is not unique and two solutions exists. For all rotations where $(\alpha,b) \in \{0, \frac{\pi}{2}, \pi, \frac{3\pi}{2}\}\times\{0,1\}$, $\boldsymbol{\beta}$ correspond to one of the eight signed permutation of $\boldsymbol{\Lambda}$ given in (ref) and the sparsity matrix $\boldsymbol{\delta}$ is equal to $\boldsymbol{\delta}^{\boldsymbol{\Lambda}}$ up to this signed permutation. For all other rotations, all elements of $\boldsymbol{\beta}$ are different from zero and $\boldsymbol{\delta}$ is simply a matrix of ones.
As discussed above, GLT structures generalize the PLT constraint, but one might wonder how restrictive this structure still is. We will show in this section that for a basic factor model with unconstrained loading matrix $\boldsymbol{\beta}$ there exists an equivalent representation involving a unique GLT structure $\boldsymbol{\Lambda}$ which is related to $\boldsymbol{\beta}$ by an orthogonal transformation, provided that uniqueness of the variance decomposition holds.
The proof of this result uses a relationship between a matrix with GLT structure and the so-called reduced row echelon form in linear algebra that results from the Gauss-Jordan elimination for solving linear systems, see e.g.\ ant-ror:ele. Any transposed GLT loading matrix $\boldsymbol{\Lambda}^{\top}$ has a row echelon form which can be turned into a reduced row echelon form (RREF) ${\mathbf B} = {\mathbf A}^{\top} \boldsymbol{\Lambda}^{\top} $ with the help of an $r\times r$ matrix ${\mathbf A}$ which is constructed from the pivot rows $l_1,\ldots, l_r$ of $\boldsymbol{\Lambda}$ and invertible by definition:
Since the RREF of any matrix is unique, see e.g.\ yus:red, we find that the pivot columns of ${\mathbf B} $ coincide with the pivot rows $l_1,\ldots, l_r$ of $\boldsymbol{\Lambda}$. Hence, for a basic factor model
with an arbitrary, unstructured loading matrix $\boldsymbol{\beta}$ with full column rank $r$, we prove in Theorem (ref) that the RREF of $\boldsymbol{\beta}^{\top}$ can be used to represent $\boldsymbol{\beta}$ as a unique GLT structure $\boldsymbol{\Lambda}$, where the pivot rows $l_1,\ldots, l_r$ of $\boldsymbol{\Lambda}$ coincide with the pivot columns of the RREF of $\boldsymbol{\beta}^{\top}$ (see Appendix (ref) for a proof).
Would it be possible to obtain a similar results with the factor loading matrix $\boldsymbol{\Lambda}$ being constrained to be a PLT structure? The answer is definitely no, as has already been established in Section (ref) for example ((ref)). As mentioned above, GLT structures encompass PLT structures as a special case. Hence, if a PLT representation $\boldsymbol{\Lambda}$ exists for a loading matrix $\boldsymbol{\beta} =\boldsymbol{\Lambda} {\mathbf P} $, then the GLT representation in ((ref)) {\em automatically} reduces to the PLT structure $\boldsymbol{\Lambda}$, since $\mathbf{R} = \boldsymbol{\beta}_1^{\top}$ is obtained from the first $r$ rows of $\boldsymbol{\beta}$ and the \lq\lq rotation into GLT\rq\rq\ is equal to the identity, ${\mathbf Q}={{\mathbf I}}_{r}$. On the other hand, if the GLT representation $\boldsymbol{\Lambda}$ differs from a PLT structure, then no equivalent PLT representation exists. Hence, forcing a PLT structure in the representation ((ref)) may introduce a systematic bias in estimating the marginal covariance matrix ${\mathbf \Omega}$.
As mentioned in the previous sections, constraints imposed on the structure of a factor loading matrix $\boldsymbol{\Lambda}$ will resolve rotational invariance only if uniqueness of the variance decomposition holds and the cross-covariance matrix $\boldsymbol{\Lambda} \boldsymbol{\Lambda} ^{\top}$ is identified. However, rotational constraints alone do not necessarily guarantee uniqueness of the variance decomposition. Consider, for instance, a sparse PLT loading matrix where in some column $j$ in addition to the diagonal element $\Lambda_{jj}$ (which is nonzero by definition) only a single further factor loading $\Lambda_{n_j,j}$ in some row $n_j>j$ is nonzero. Such a loading matrix obviously violates the necessary condition for variance identification that each column contains at least three nonzero elements. Similarly, while GLT structures resolve rotational invariance, they do not guarantee uniqueness of the variance decomposition either.
In Section (ref), we derive sufficient conditions for variance identification of GLT structures based on the 3579\ counting rule of sat:stu. In Section (ref), we discuss how to verify variance identification for sparse GLT structures in practice.
We will show how to verify from the 0/1 pattern $\boldsymbol{\delta}$ of an unordered GLT structure $\boldsymbol{\beta} $, whether the row deletion property \bf AR\ holds for $\boldsymbol{\beta} $ and all its signed permutations. Our condition is a structural counting rule expressed solely in terms of the sparsity matrix $\boldsymbol{\delta}$ underlying $\boldsymbol{\beta} $ and does not involve the values of the unconstrained factor loadings in $\boldsymbol{\beta}$, which can take any value in $\mathbb{R}$. For any factor model, variance identification is invariant to signed permutations. If we can verify variance identification for a single signed permutation $ \boldsymbol{\beta} = \boldsymbol{\Lambda} {\mathbf P} _{\pm} {\mathbf P} _{\rho} $ of $\boldsymbol{\Lambda}$, as defined in ((ref)), then variance identification of $\boldsymbol{\Lambda}$ holds, since $ \boldsymbol{\beta}$ and $\boldsymbol{\Lambda} $ imply the same cross-covariance matrix $\boldsymbol{\Lambda} \boldsymbol{\Lambda} ^{\top}$. Hence, we focus in this section on variance identification of unordered GLT structures.
In Definition (ref), we recall the so-called {\em extended row deletion property}, introduced by tum-sat:ide.
The row-deletion property of and-rub:sta results as a special case where $s={1}$. As will be shown in Section (ref), the extended row deletion properties $\mbox{RD} (r,s)$ for $s >{1}$ are useful in exploratory factor analysis, when the factor dimension $r$ is unknown. In Definition (ref), we introduce a counting rule for binary matrices.
Note that the counting rule $\mbox{CR} (r,s)$, like the extended row deletion property $\mbox{RD} (r,s)$, is invariant to signed permutations. Lemma (ref) in Appendix (ref) summarizes further useful properties of $\mbox{CR} (r,s)$.
For a given binary matrix $\boldsymbol{\delta}$ of dimension $m \times r$, let $\Theta_{\delta}$ be the space generated by the non-zero elements of all unordered GLT structure $\boldsymbol{\beta}$ with sparsity matrix $\boldsymbol{\delta}$ and all their $2^r r ! -1 $ trivial rotations $\boldsymbol{\beta} {\mathbf P} _{\pm} {\mathbf P} _{\rho}$. We prove in Theorem (ref) that for GLT structures the counting rule $\mbox{CR} (r,s)$ and the extended row deletion property $\mbox{RD} (r,s)$ are equivalent conditions for all loading matrices in $\Theta_{\delta}$, except for a set of measure 0.
See Appendix (ref) for a proof. The special case $s=1$ is relevant for verifying the row deletion property \bf AR . It proves that for unordered GLT structures the 3579\ counting rule of sat:stu is not only a necessary, but also a {\em sufficient} condition for \bf AR\ to hold. In addition, this means that the counting rule needs to be verified only for the sparsity matrix $\boldsymbol{\delta}$ of a {\em single trivial rotation} $ \boldsymbol{\beta} =\boldsymbol{\Lambda} {\mathbf P} _{\pm} {\mathbf P} _{\rho}$ rather than for every nonsingular matrix ${\mathbf G}$. This result is summarized in Corollary (ref).
A few comments are in order. If $\boldsymbol{\delta}$ satisfies $\mbox{CR} (r,1)$, then $\mbox{\bf AR}$ holds for all $\boldsymbol{\beta} \in \Theta_{\delta}$ and a sufficient condition for variance identification is satisfied. As shown by and-rub:sta, $\mbox{\bf AR}$ is a {\em necessary} condition for variance identification only for $r=1$ and $r=2$. tum-sat:ide show the same for $r=3$, provided that $m\geq 7$. It follows that $\mbox{CR} (r,1)$ is a {\em necessary and sufficient} condition for variance identification for the models summarized in (c). In all other cases, variance identification may hold for loading matrices $\boldsymbol{\beta} \in \Theta_{\delta}$, even if $\boldsymbol{\delta}$ violates $\mbox{CR} (r,1)$.
The definition of unordered GLT structures given in Section (ref) imposes no constraint on the pivot rows $l_1, \ldots , l_{r}$ beyond the assumption that they are distinct. This flexibility can lead to GLT structures that can never satisfy the 3579\ rule, even if all elements below the pivot rows are non-zero. Consider, for instance, a GLT matrix with the pivot row in column $r$ being equal to $l_r= m -1$. The loading matrix has at most two nonzero elements in column $r$ and violates the necessary condition for variance identification. This example shows that there is an upper bound for the pivot elements beyond which the 3579\ rule can never hold. This insight is formalized in Definition (ref).
Evidently, an ordered GLT structure $\boldsymbol{\Lambda}$ fulfills condition \bf GLT-AR\ if the pivot rows $l_1, \ldots , l_{r}$ of $\boldsymbol{\Lambda}$ satisfy the constraint $ l_j \leq m - 2(r - j+1)$. For the special case of a PLT structure where $l_j=j$, this constraint reduces to $ m \geq 2r +1 $ which is equivalent to a well-known upper bound for the number of factors. For dense unordered GLT structures with $m$ (non-zero) rows, condition \bf GLT-AR\ is a sufficient condition for \bf AR . For sparse GLT structures \bf GLT-AR\ is only a necessary condition for \bf AR\ and the 3579\ rule has to be verified explicitly, as shown by the example discussed above. Very conveniently for verifying variance identification in sparse factor analysis based on GLT structures, Theorem (ref) and Corollary (ref) operate solely on the sparsity matrix $\boldsymbol{\delta}$ corresponding to $\boldsymbol{\beta}$.
To verify $\mbox{CR} (r,s)$ in practice, all submatrices of $q$ columns have to be extracted from the sparsity matrix $\boldsymbol{\delta}$ to verify if at least $2q+1$ rows of this submatrix are non-zero. For $q=1,2, r -1, r$, this condition is easily verified from simple functionals of $\boldsymbol{\delta}$, see Corollary (ref) which follows immediately from Theorem (ref) (see Appendix (ref) for details).
Using Corollary (ref) for $s=1$, one can efficiently verify, if the 3579\ counting rule and hence the row deletion property \bf AR\ holds for unordered GLT factor models with up to $r \leq 4$ factors. For models with more than four factors ($r > 4$), a more elaborated strategy is needed. After checking the conditions of Corollary (ref), $\mbox{CR} (r,s)$ could be verified for a given binary matrix $\boldsymbol{\delta}$ by iterating over all remaining $\footnotesize{ r ! /(q ! (r -q ) !) }$ subsets of $q=3, \ldots, r-2$ columns of $\boldsymbol{\delta}$. While this is a finite task, such a na\"ive approach may need to visit $2^r-1$ matrices in order to make a decision and the combinatorial explosion quickly becomes an issue in practice as $r$ increases. Recent work by hos-fru:cov establishes the applicability of this framework for large models.
In this section, we discuss how the concept of GLT structures is helpful for addressing identification problems in exploratory factor analysis (EFA). Consider data $\{{\mathbf y}_1, \ldots,{\mathbf y}_T\}$ from a multivariate Gaussian distribution, ${\mathbf y}_t \sim \mathcal{N} _{m}\left({\mathbf{0}},{\mathbf \Omega}\right)$, where an investigator wants to perform factor analysis since she expects that the covariances of the measurements $y_{it}$ are driven by common factors. In practice, the number of factors is typically unknown and often it is not obvious, whether all $m$ measurements in ${\mathbf y}_t$ are actually correlated. It is then common to employ EFA by fitting a basic factor model to the entire collection of measurements in ${\mathbf y}_t$, i.e.\ assuming the model
with an assumed number of factors $k$, a $m\times k$ loading matrix $\boldsymbol{\beta}_k$ with elements $\beta_{ij}$ and a diagonal matrix ${\mathbf \Sigma}_k$ with strictly positive entries. The EFA model ((ref)) is potentially overfitting in two ways. First, the true number of factors $r$ is possibly smaller than $k$, i.e.\ $\boldsymbol{\beta}_k$ has too many columns. Second, some measurements in ${\mathbf y}_t$ are possibly irrelevant, which means that $\boldsymbol{\beta}_k$ allows for too many non-zero rows. The goal is then to determine the true number of factors and to identify irrelevant measurements from the EFA model ((ref)).
We will address identification under the assumption that the data are generated by a basic factor model with loading matrix $\boldsymbol{\beta}_0$ with $r$ factors which implies the following covariance matrix ${\mathbf \Omega}$:
Instead of ((ref)), for a given $k$, the EFA model ((ref)) yields the alternative representation of ${\mathbf \Omega}$:
The question is then under which conditions can the true loading matrix $\boldsymbol{\beta}_0$ be recovered from ((ref)). Let us assume for the moment that no constraint that resolves rotational invariance is imposed on $\boldsymbol{\beta}_0$ or $\boldsymbol{\beta}_k$.
\paragraph*{\lq\lq Revealing the truth\rq\rq\ in an overfitting EFA model.}
A fundamental problem in factor analysis is the following. If the EFA model is overfitting, i.e.\ $k > r$, could we nevertheless recover the true loading matrix $\boldsymbol{\beta}_0$ directly from $\boldsymbol{\beta}_k$? We will show how this can be achieved mathematically by combining the important work by tum-sat:ide with the framework of GLT structures. We have demonstrated in Section (ref) using example ((ref)) that solutions in an overfitting model can be constructed by adding spurious columns rei:ide,gew-sin:int. Additional solutions are obtained as rotations of such solutions. For instance, one of the following solutions may result:
both with the same ${\mathbf \Sigma}_3$ as in ((ref)). The first case is a signed permutation of $\boldsymbol{\beta}_{3}$, while the second case combines a signed permutation of $\boldsymbol{\beta}_{3}$ with a rotation of the spurious and $\boldsymbol{\Lambda} $'s first column involving ${\mathbf P}_{\alpha b}$. In the first case, despite the rotation, both the spurious column and the columns of $\boldsymbol{\Lambda} $ are clearly visible, while in the second case the presence of a spurious column is by no means obvious and the columns of $\boldsymbol{\Lambda} $ are disguised.
In general, for an EFA model that is overfitting by a single column, i.e.\ $k = r+1$, and $\boldsymbol{\beta}_{k}$ is left unconstrained, infinitely many representations $(\boldsymbol{\beta}_k,{\mathbf \Sigma}_k)$ with covariance matrix ${\mathbf \Omega}= \boldsymbol{\beta} _k \boldsymbol{\beta}_k^{\top} + {\mathbf \Sigma}_k$ can be constructed in the following way. Let the first $r$ columns of $\boldsymbol{\beta}_k$ be equal to $\boldsymbol{\beta}_0$ and append an extra column to its right. In this extra column, which will be called a spurious column, add a single non-zero loading $\beta_{l_k, k}$ in any row $1 \leq l_k \leq m $ taking any value that satisfies $0< \beta_{l_k, k}^2< \sigma^2_{l_k}$; then reduce the idiosyncratic variance in row $l_k$ to $\sigma^2_{l_k} - \beta_{l_k, k}^2$; and finally apply an arbitrary rotation ${\mathbf P}$:
Interesting questions are then the following: under which conditions is ((ref)) an exhaustive representation of all possible solutions $ \boldsymbol{\beta}_{k}$ in an EFA model where the {\em degree of overfitting} defined as $s=k-r$ is equal to one? How can all solutions $ \boldsymbol{\beta}_{k}$ be represented if $s>1$?
Such identifiability problems in overfitting EFA models have been analyzed in depth by tum-sat:ide. They show that a stronger condition than $\mbox{RD} (r,1)$ is needed for $\boldsymbol{\beta}_0$ in the underlying variance decomposition ((ref)) to ensure that only spurious and no additional common factors are added in the overfitting representation ((ref)). In addition, tum-sat:ide provide a general representation of the factor loading matrix $\boldsymbol{\beta}_k $ in overfitting representation ((ref)) with $k > r$.
The $m\times s$-matrix ${\mathbf M}_s $ is a so-called spurious factor loading matrix that does not contribute to explaining the covariance in ${\mathbf y}_t$, since
While this theorem is an important result, without imposing further structure on the factor loading matrix $\boldsymbol{\beta}_k$ in the EFA model it cannot be applied immediately to \lq\lq recover the truth\rq\rq , as the separation of $\boldsymbol{\beta}_k$ into the true factor loading matrix $\boldsymbol{\beta}_0$ and the spurious factor loading matrix ${\mathbf M}_s$ is possible only up to a rotation ${\mathbf T}_k$ of $\boldsymbol{\beta}_k$. However, the truth\rq\rq\ in an overfitting EFA model can be recovered, if tum-sat:ide is applied within the class of unordered GLT structures introduced in this paper. If we assume that $\boldsymbol{\Lambda}$ is a GLT structure which satisfies the extended row deletion property $\mbox{RD} (r,1+S)$, we prove in Theorem (ref) the following result. If $\boldsymbol{\beta}_k$ in an overfitting EFA model is an unordered GLT structure, then $\boldsymbol{\beta}_k$ has a representation, where the rotation in ((ref)) is a signed permutation ${\mathbf T}_k = {\mathbf P} _{\pm} {\mathbf P} _{\rho}$. Hence, spurious factors in $\boldsymbol{\beta}_k$ are easily spotted and $\boldsymbol{\Lambda}$ can be recovered immediately from $\boldsymbol{\beta}_k$.
See Appendix (ref) for a proof.
\paragraph*{Identifying irrelevant variables.}
In applied factor analysis, the assumption that each measurement $y_{it}$ is correlated with at least one other measurement is too restrictive, because irrelevant measurements might be present that are uncorrelated with all the other measurements. As argued by boi-ng:are, it is useful to identify such variables. Within the framework of sparse factor analysis, irrelevant variables are identified in kau-sch:ide by exploring the sparsity matrix $\boldsymbol{\delta}$ of a factor loading matrix $\boldsymbol{\beta}_0$ with respect to zero rows. Since $\mbox{\rm Cov}(y_{it}, y_{lt}) = 0$ for all $l \neq i$, if the entire $i$th row of $ \boldsymbol{\beta}_0$ is zero (see also ((ref))), the presence of $m_0$ irrelevant measurements causes the corresponding $m_0$ rows of $ \boldsymbol{\beta}_0$ and $\boldsymbol{\delta}$ to be zero. As before, we assume that the variance decomposition ((ref)) of the underlying basic factor model is variance identified.
Let us first investigate identification of the zero rows in $ \boldsymbol{\beta}_0$ and the corresponding sparsity matrix $\boldsymbol{\delta}$ for the case that the assumed and the true number of factors in the EFA model ((ref)) are identical, i.e.\ $k = r$. Since variance identification of ((ref)) in the underlying model holds, we obtain that ${\mathbf \Sigma}_0={\mathbf \Sigma}_r$, $\boldsymbol{\beta}_0 \boldsymbol{\beta}_0^{\top}=\boldsymbol{\beta}_r \boldsymbol{\beta}_r^{\top}$ and $\boldsymbol{\beta}_r = \boldsymbol{\beta}_0 {\mathbf P} $ is a rotation of $\boldsymbol{\beta}_0$. Therefore, the position of the zero rows both in $\boldsymbol{\beta}_0$ and $\boldsymbol{\beta}_r $ are identical and all irrelevant variables can be identified from $\boldsymbol{\beta}_r$ or the corresponding sparsity matrix $\boldsymbol{\delta}$, regardless of the strategy toward rotational invariance.
What makes this task challenging in applied factor analysis is that in practice only the total number $m$ of observations is known, whereas the investigator is ignorant both about the number of factors $r$ and the number of irrelevant measurements $m_0$. In such a situation, variance identification of ${\mathbf \Sigma}_k$ for an EFA model with $k$ assumed factors is easily lost if too many irrelevant variables are included in relation to $k$. These considerations have important implication for exploratory factor analysis. While the investigator can choose $ k$, she is ignorant about the number of irrelevant variables and the recovered model might not be variance identified. For this reason, it is relevant to verify in any case that the solution $\boldsymbol{\beta}_k$ obtained from any EFA model satisfies variance identification.
Under \bf AR\ this means that the loading matrix of the correlated measurements, i.e.\ the non-zero rows of $ \boldsymbol{\beta}_0$, satisfies $\mbox{RD} (r,1)$. If variance identification relies on \bf AR , then a minimum requirement for $\boldsymbol{\beta}_k$ to satisfy $\mbox{RD} (k,1)$ is that $ 2 k + 1 \leq m-m_0 $. If no irrelevant measurement are present, then the well-known upper bound $ k \leq \frac{m-1}{2}$ results. However, if irrelevant measurements are present, then there is a trade-off between $m_0$ and $k$: the more irrelevant measurements are included, the smaller the maximum number of assumed factors $k$ has to be. Hence, the presence of $m_0$ zero rows in $ \boldsymbol{\beta}_0$, while $ \boldsymbol{\beta}_k$ in the EFA model is allowed to have $m$ potentially non-zero rows requires stronger conditions for variance identification than for an EFA model where the underlying loading matrix $ \boldsymbol{\beta}_0$ contains only non-zero rows. More specifically, for a given number $m_0 \in \mathbb{N}$ of irrelevant measurements, variance identification necessitates the more stringent upper bound $ k \leq \frac{m-m_0-1}{2} $, where $m-m_0$ is the number of non-zero rows. On the other hand, for a given number of factors $k$ in an EFA model, the maximum number of irrelevant measurements that can be included is given by $ m_0 \leq m -(2k + 1)$.
\paragraph*{Identifying the number of factors through an EFA model.}
Let us assume that the variance decomposition ((ref)) of the unknown underlying basic factor model is identified. As shown by rei:ide, the true number of factors $r$ is equal to the smallest value $k$ that satisfies ((ref)). However, in practice, it is not obvious how to solve this \lq\lq minimization\rq\rq\ problem. As the following considerations show, verifying variance identification for $\boldsymbol{\beta}_k$ in an EFA model can be helpful in this regard.
If $r$ is unknown, then we need to find a decomposition of ${\mathbf \Omega}$ as in ((ref)) where $ {\mathbf \Sigma}_k $ is variance identified. Since the true underlying decomposition ((ref)) is variance identified, any solution where $ {\mathbf \Sigma}_k $ is not variance identified can be rejected. As has been discussed above, any overfitting EFA model, where $k > r$, has infinitely many decompositions of ${\mathbf \Omega}$ and therefore is never variance identified. Hence, if any solution $ {\mathbf \Sigma}_k $ of an EFA model with $k$ assumed factors is not variance identified, then we can deduce that $k$ is bigger than $r$. On the other hand, if variance identification holds for $ {\mathbf \Sigma}_k $, then the decompositions ((ref)) and ((ref)) are equivalent and we can conclude that $r=k$, ${\mathbf \Sigma}_0= {\mathbf \Sigma}_k$ and therefore $ \boldsymbol{\beta}_0 \boldsymbol{\beta}_0^{\top}=\boldsymbol{\beta}_k \boldsymbol{\beta}^{\top}_k$. As a consequence, we can identify the true loading matrix $ \boldsymbol{\beta}_0=\boldsymbol{\beta}_k {\mathbf P}$ from $\boldsymbol{\beta}_k$ mathematically up to a rotation ${\mathbf P}$ and-rub:sta.
This insight shows that verifying variance identification is relevant beyond resolving rotational invariance and is essential for recovering the true number of factors. This has important implications for applied factor analysis. Most importantly, the {\em rank} or the {\em number of non-zero columns} of a factor loading matrix $\boldsymbol{\beta}_k $ recovered from an EFA model with {\em assumed} number $k$ of factors might overfit the true number of factors $r$, if variance identification for ${\mathbf \Sigma}_k$ is not satisfied and the variance decomposition is not unique. Hence, extracting the number of factors from an EFA model makes only sense in connection with ensuring that variance identification holds.
A common goal of Bayesian factor analysis is to identify the unknown factor dimension $r$ of a factor loading matrix from the overfitting factor model ((ref)) with potentially $k>r$ factors, see, among many others, roc-geo:fas, fru-lop:spa, and ohn-kim:pos. Often, spike-and slab priors are employed, where the elements $\beta_{ij}$ of the loading matrix $\boldsymbol{\beta}_k$ apriori are allowed to be exactly zero with positive probability. This is achieved through a prior on the corresponding $m\times k$ sparsity matrix $\boldsymbol{\delta}_k$. In each column $j$, the indicators $\delta_{i j}$ are active apriori with a column-specific probability $\tau_j$, i.e. $\mbox{\rm Pr} (\delta_{i j}=1|\tau_j)=\tau_j$ for $i=1, \ldots,m$, where the slab probabilities $\tau_1, \ldots, \tau_k$ arise from an exchangeable shrinkage prior:
If $\gamma$ is unknown, then ((ref)) is called a two-parameter-beta (2PB) prior. If $\gamma=1$, then ((ref)) is called a one-parameter-beta (1PB) prior and takes the form:
Prior ((ref)) converges to the Indian buffet process prior teh-etal:sti for $k \rightarrow \infty$. As recently shown by fru:gen, prior ((ref)) has a representation as a cumulative shrinkage process (CUSP) prior leg-etal:bay.
This specification leads to a Dirac-spike-and-slab prior for the factor loadings,
where the columns of the loading matrix are increasingly pulled toward 0 as the column index increases. In ((ref)), a Gaussian slab distribution is assumed with a random global shrinkage parameter $\kappa$, although other slab distributions are possible, see e.g.\ zha-etal:bay_gro and fru-etal:spa.
The hyperparameters $\alpha$ and $\gamma$ are instrumental in controlling prior sparsity. Choosing $\alpha=k$ and $\gamma=1$ leads to a uniform distribution for $\tau_j$, with the {\em smallest} slab probability $\tau_{(1)}= \min_{j=1,\ldots,k} \tau_j$ also being uniform, while the largest slab probability $\tau_{(k)}= \max_{j=1,\ldots,k} \tau_j \sim \mathcal{B}\left(k,1\right)$, see fru:gen. Such a prior is likely to overfit the number of factors, regardless of all other assumptions. A prior with $\alpha<k$ and $\gamma=1$ induces sparsity, since the {\em largest} slab probability $\tau_{(k)} \sim \mathcal{B}\left(\alpha,1\right) $, while the smallest slab probability $\tau_{(1)} \sim \mathcal{B}\left(\alpha/k,1\right)$. To control the small probabilities, which are important in identifying the true number of factors, $\alpha$ is assumed to be a random parameter and learnt from the data under the prior $\alpha \sim \mathcal{G}\left(a^\alpha, b^\alpha\right)$. $\gamma$ controls the prior information in ((ref)). Priors with $\gamma >1$ and $\gamma < 1$, respectively, decrease and increase the difference between $\tau_{(1)}$ and $\tau_{(k)}$. Typically, $\gamma$ is unknown and is estimated from the data using the prior $\gamma \sim \mathcal{G}\left(a^\gamma, b^\gamma\right)$.
\paragraph*{MCMC estimation.} For a given choice of hyperparameters, Markov chain Monte Carlo (MCMC) methods are applied to sample from the posterior distribution $p(\boldsymbol{\beta}_k, {\mathbf \Sigma}_k, \boldsymbol{\delta}_k|{\mathbf y})$, given $T$ multivariate observations ${\mathbf y}=({\mathbf y}_1, \ldots, {\mathbf y}_T)$, see e.g.\ kau-sch:bay among many others. In fru-etal:spa, such a sampler is developed for GLT factor models. To move between factor models of different factor dimension, fru-etal:spa exploit Theorem (ref) to add and delete spurious columns through a reversible jump MCMC (RJMCMC) sampler. For each posterior draw $\boldsymbol{\beta}_k$, the active columns $\boldsymbol{\beta}_r$ (i.e.\ all columns with at least 2 non-zero elements) and the corresponding sparsity matrix $\boldsymbol{\delta}_r$ are determined. If $\boldsymbol{\delta}_r$ satisfies the counting rule $\mbox{CR} (r,1)$, then $\boldsymbol{\beta}_r$ is a signed permutation of $\boldsymbol{\Lambda}$ with the corresponding covariance matrix ${\mathbf \Sigma}_r = {\mathbf \Sigma}_k + {\mathbf M} ^{\Lambda}_s ({\mathbf M} ^{\Lambda}_s)^{\top}$, where ${\mathbf M} ^{\Lambda}_s$ contains the spurious columns of $\boldsymbol{\beta}_k$. These variance identified draws are kept for further inference and the number of columns of $\boldsymbol{\beta}_r$ is considered a posterior draw of the unknown factor dimension $r$. This algorithm is easily extended to EFA models without any constraints.
For illustration, we perform a simulation study and consider three different data scenarios with $m=30$ and $T=150$. In all three scenarios, $r_{\mbox{\tiny \rm true}}=5$ factors are assumed, however, the zero/non-zero pattern is quite different. The first setting is a {\em dedicated} factor model, where the first 6 variables load on factor 1, the next 6 variables load on factor 2, and so forth, and the final 6 variables load on factor 5. A dedicated factor model has a GLT structure by definition. The second scenario is a {\em block} factor model, where the first 15 observations load only on factor 1 and 2, while the remaining 15 observations only load on factor 3, 4 and 5 and the covariance matrix has a block-diagonal structure. All loadings within a block are non-zero. The third scenario is a {\em dense} factor loading matrix without any zero loadings and the corresponding GLT representation has a PLT structure. For all three scenarios, non-zero factor loadings are drawn as $\lambda_{ij}=(-1)^{b_{ij}} (1 + 0.1 \mathcal{N}\left(0,1\right))$, where the exponent $b_{ij}$ is a binary variable with $\mbox{\rm Pr} (b_{ij}=1)=0.2$. In all three scenarios, ${\mathbf \Sigma}_0={\mathbf I}$. 21 data sets are sampled under these three scenarios from the Gaussian factor model ((ref)).
A sparse overfitting factor model is fitted to each simulated data set with the maximum number of factors $k= 14$ being equal to the upper bound. Regarding the structure, we compare a model where the non-zero columns of $\boldsymbol{\beta}_k$ are left unconstrained with a model where a GLT structure is imposed. Inference is based on the Bayesian approach described in Section (ref) with two different shrinkage priors on the sparsity matrix $\boldsymbol{\delta}_k$: the 1PB prior ((ref)) with random hyperparameter $\alpha \sim \mathcal{G}\left(6,2\right)$ and the 2PB prior ((ref)) with random hyperparameters $\alpha \sim \mathcal{G}\left(6,2\right)$ and $ \gamma \sim \mathcal{G}\left(6,6\right)$. MCMC estimation is run for 3000 iterations after a burn-in of 2000 using the RJMCMC algorithm of fru-etal:spa.
For each of the 21 simulated data sets, we evaluate all 12 combinations of data scenarios, structural constraints (GLT versus unconstrained) and priors on the sparsity matrix (1PB versus 2PB) through Monte Carlo estimates of following statistics: to assess the performance in estimating the true number $r_{\mbox{\tiny \rm true}}$ of factors, we consider the mode $\hat{r}$ of the posterior distribution $p(r|{\mathbf y}) $ and the magnitude of the posterior ordinate $p(\hat{r}=r_{\mbox{\tiny \rm true}}|{\mathbf y})$. To assess the accuracy in estimating the covariance matrix ${\mathbf \Omega}$ of the data, we consider the mean squared error (MSE) defined by
which accounts both for posterior variance and bias of the estimated covariance matrix ${\mathbf \Omega}_r= \boldsymbol{\beta}_r \boldsymbol{\beta}_r ^\top + {\mathbf \Sigma}_r$ in comparison to the true matrix. Table (ref) reports, for all 12 combinations the median, the 5% and the 95% quantile of these statistics across all simulated data sets. For inference under GLT structures, posterior draws which are not variance identified have been removed. The fraction of variance identified draws is also reported in the table and is in general pretty high. As common for sparse Bayesian factor analysis with unstructured loading matrices, the posterior draws are not screened for variance identification and inference is based on all draws.
Some interesting conclusions can be drawn from Table (ref). First of all, sparse Bayesian factor analysis under the GLT constraint successfully recovers the true number of factors in all three scenarios. For most of the simulated data sets, the posterior ordinate $p(\hat{r}=r_{\mbox{\tiny \rm true}}|{\mathbf y})$ is larger than 0.9. Sparse Bayesian factor analysis with unstructured loading matrices is also quite successful in recovering $r_{\mbox{\tiny \rm true}}$, but with less confidence. Both over- and underfitting can be observed and the posterior ordinate $p(\hat{r}=r_{\mbox{\tiny \rm true}}|{\mathbf y})$ is much smaller than under a GLT structure. For both structures, the 2PB prior yields higher posterior ordinates than the 1PB prior.
Recently, hos-fru:cov proved that the counting rule $\mbox{CR} (r,1)$ can also be applied to verify variance identification for unconstrained loading matrices. As is evident from Table (ref), the fraction of variance identified draws is however, much smaller than under GLT structures. Nevertheless, inference w.r.t.\ to the number of factors can be improved also for an unconstrained EFA model by rejecting all draws that do not obey the counting rule $\mbox{CR} (r,1)$.
It should be emphasized that the ability of Bayesian factor analysis to recover the number of factors from an overfitting model is closely tied to choosing a suitable shrinkage prior on the sparsity matrix $\boldsymbol{\delta}_k$. For illustration, we also consider a uniform prior for $\tau_j$ and report the corresponding statistics in Table (ref). As expected from the considerations in Section (ref), considerable overfitting is observed for all simulated data sets, regardless of the chosen structure.
We have given a full and comprehensive mathematical treatment to generalized lower triangular (GLT) structures, a new identification strategy that improves on the popular positive lower triangular (PLT) assumption for factor loadings matrices. We have proven that GLT retains PLT's good properties: uniqueness and rotational invariance. At the same time and unlike PLT, GLT exists for any factor loadings matrix; i.e.\ it is not a restrictive assumption. Furthermore, we have shown that verifying variance identification under GLT structures is simple and is based purely on the zero-nonzero pattern of the factor loadings matrix. Additionally, we have embedded the GLT model class into exploratory factor analysis with unknown factor dimension and discussed how easily spurious factors and irrelevant variables are recognized in that setup. At the end, we demonstrated the power of the framework in a simulation study.