The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
111,248 characters
When it counts - Econometric identification of the basic factor model based on GLT structures
\title{When it counts - Econometric identification of the basic factor model based on GLT structures}
\author{Sylvia Fr\"uhwirth-Schnatter\footnote{Department of Finance, Accounting, and Statistics, WU Vienna University of Economics and Business, Austria. Email: {\tt [email removed]}} \and
Darjus Hosszejni\footnote{Department of Finance, Accounting, and Statistics, WU Vienna University of Economics and Business, Austria. Email: {\tt [email removed]}}
\and Hedibert Freitas Lopes\footnote{School of Mathematical and Statistical Sciences, Arizona State University, Tempe, USA \& Insper Institute of Education and Research, S\~ao Paulo, Brazil. Email: {\tt [email removed]}}, \,
}
\maketitle
\begin{abstract}
Despite the popularity of factor models with sparse loading matrices, little attention has been given to formally address identifiability of these models beyond standard rotation-based identification such as the positive lower triangular (PLT) constraint.
To fill this gap, we review the advantages of variance identification in sparse factor analysis and introduce the generalized lower triangular (GLT) structures.
We show that the GLT assumption is an improvement over PLT without compromise: GLT is also unique but, unlike PLT, a non-restrictive assumption.
Furthermore, we provide a simple counting rule for variance identification under GLT structures, and we demonstrate that within this model class the unknown number of common factors can be recovered in an exploratory factor analysis.
Our methodology is illustrated for simulated data in the context of post-processing posterior draws in Bayesian sparse factor analysis.
\end{abstract}
\vspace{0.5cm}
{\em Keywords:}
Identifiability; sparsity; rank deficiency; rotational invariance; variance identification
\vspace{0.5cm}
\centerline{JEL classification: C11, C38, C63}
\section{Introduction}
Ever since the pioneering work of \citet{thu:vec,thu:mul}, factor analysis has been a popular method to model the covariance matrix ${\mathbf \Omega}$ of correlated, multivariate observations ${\mathbf y}_t$ of dimension $m$, see e.g.\ \citet{and:int} for a comprehensive review.
Assuming $r$ uncorrelated factors, the basic factor model yields the representation ${\mathbf \Omega}= \boldsymbol{\Lambda} \boldsymbol{\Lambda} ^{\top} + {\mathbf \Sigma}_0 $, with a $m\times r$ factor loading matrix $\boldsymbol{\Lambda}$ and a diagonal matrix ${\mathbf \Sigma}_0 $.
The considerable reduction of the number of parameters compared to the $m(m+1)/2$ elements of an unconstrained covariance matrix ${\mathbf \Omega}$ is the main motivation for applying factor models to covariance estimation,
especially if $m$ is large; see, among many others, \citet{fan-etal:hig_je} in finance and \citet{for-etal:ope} in economics. In addition, shrinkage estimation has been shown to lead to very efficient covariance estimation,
see, for example, \citet{kas:spa} in Bayesian factor analysis and \citet{led-wol:pow} in a non-Bayesian context.
In numerous applications, factor analysis reaches beyond covariance modelling. From the very beginning, the goal
of factor analysis has been to extract the underlying loading matrix $\boldsymbol{\Lambda}$ to understand the driving forces behind the observed correlation between the features, see e.g.\ \citet{owe-wan:bic} for a recent review.
However, also in this setting, the only source of information is the observed covariance of the data,
making the decomposition of the covariance matrix ${\mathbf \Omega}$ into the cross-covariance matrix $\boldsymbol{\Lambda} \boldsymbol{\Lambda}^{\top}$ and the variance $ {\mathbf \Sigma}_0 $ of the idiosyncratic errors more challenging than estimating only ${\mathbf \Omega}$ itself.
A huge literature, dating back to \citet{koo-rei:ide} and \citet{rei:ide}, has addressed this problem of identification
which can be resolved only by imposing additional structure on the factor model. \citet{and-rub:sta} considered identification as a two-step procedure, namely identification of ${\mathbf \Sigma}_0$ from ${\mathbf \Omega}$ (variance identification) and subsequent identification of $\boldsymbol{\Lambda}$ from $\boldsymbol{\Lambda} \boldsymbol{\Lambda} ^{\top}$ (solving rotational invariance). The most popular constraint in econometrics, statistics and machine learning for solving rotational invariance is to consider positive lower triangular loading matrices, see e.g.\ \citet{gew-zho:mea,wes:bay_fac,lop-wes:bay}, albeit other strategies have been put forward, see e.g.\ \citet{neu:ide}, \citet{bai-ng:pri}, \citet{ass-etal:bay}, \citet{cha-etal:inv}, and \citet{wil:ide}. Only a few papers have addressed variance identification \citep[e.g.][]{bek:ide} and to the best of our knowledge so far no structure has been put forward that simultaneously addresses both identification problems.
In this work, we discuss a new identification strategy based on generalized lower triangular (GLT) structures, see Figure~\ref{figglt} for illustration. This concept was originally introduced as part of an MCMC sampler for sparse Bayesian factor analysis where the number of factors is unknown in the (unpublished) work of \citet{fru-lop:spa}. In the present paper, GLT structures are given a full and comprehensive mathematical treatment and are applied in \citet{fru-etal:spa} to develop an efficient reversible jump MCMC (RJMCMC) sampler for sparse Bayesian factor analysis under very general shrinkage priors.
It will be proven that GLT structures simultaneously address rotational invariance and variance identification in factor models. Variance identification relies on a counting rule for the number of non-zero elements in the loading matrix $\boldsymbol{\Lambda}$, which is a sufficient condition that extends previous work by \citet{sat:stu}.
In addition, we will show that GLT structures are useful in exploratory factor analysis where the factor dimension
$r$ is unknown. Identification of the number of factors in applied factor analysis is a notoriously difficult problem,
with considerable ambiguity which method works best, be it BIC-type criteria
\citep{bai-ng:det2002}, marginal likelihoods \citep{lop-wes:bay}, techniques from Bayesian non-parametrics involving
infinite-dimensional factor models \citep{bha-dun:spa,roc-geo:fas,leg-etal:bay} or more heuristic procedures \citep{kau-sch:bay}.
Imposing an unordered GLT structure in exploratory factor analysis allows to identify the true loading matrix
$\boldsymbol{\Lambda}$ and the matrix ${\mathbf \Sigma}_0 $ and to easily spot all spurious columns in a possibly overfitting model.
This strategy underlies the RJMCMC sampler of \citet{fru-etal:spa} to estimate the number of factors.
\begin{figure}[t]
\centering
\begin{tikzpicture} [x=3mm, y=3mm, yscale=-1]
\node[leading] at (1,1) {};
\node[nonzero] at (1,3) {};
\node[nonzero] at (1,4) {};
\node[nonzero] at (1,6) {};
\node[nonzero] at (1,9) {};
\node[nonzero] at (1,10) {};
\node[nonzero] at (1,11) {};
\node[nonzero] at (1,16) {};
\node[nonzero] at (1,17) {};
\node[nonzero] at (1,19) {};
\node[leading] at (2,3) {};
\node[nonzero] at (2,4) {};
\node[nonzero] at (2,6) {};
\node[nonzero] at (2,10) {};
\node[nonzero] at (2,11) {};
\node[nonzero] at (2,13) {};
\node[nonzero] at (2,14) {};
\node[nonzero] at (2,17) {};
\node[nonzero] at (2,21) {};
\node[nonzero] at (2,22) {};
\node[leading] at (3,10) {};
\node[nonzero] at (3,11) {};
\node[nonzero] at (3,13) {};
\node[nonzero] at (3,14) {};
\node[nonzero] at (3,16) {};
\node[nonzero] at (3,17) {};
\node[nonzero] at (3,19) {};
\node[nonzero] at (3,21) {};
\node[nonzero] at (3,22) {};
\node[leading] at (4,11) {};
\node[nonzero] at (4,14) {};
\node[nonzero] at (4,16) {};
\node[nonzero] at (4,18) {};
\node[nonzero] at (4,21) {};
\node[leading] at (5,14) {};
\node[nonzero] at (5,16) {};
\node[nonzero] at (5,17) {};
\node[nonzero] at (5,18) {};
\node[nonzero] at (5,19) {};
\node[nonzero] at (5,21) {};
\node[nonzero] at (5,22) {};
\node[leading] at (6,17) {};
\node[nonzero] at (6,18) {};
\node[nonzero] at (6,19) {};
\node[nonzero] at (6,21) {};
\draw[matrix box] (0,0) rectangle (7,23);
\begin{scope}[every node/.style={axis label, label distance=-1.5mm}]
\node (v1) [label=left:{$0$}] at (0,0) {};
\node (v2) [label=left:{$5$}] at (0,5) {};
\node (v3) [label=left:{$10$}] at (0,10) {};
\node (v4) [label=left:{$15$}] at (0,15) {};
\node (v5) [label=left:{$20$}] at (0,20) {};
\node (h1) [label=below:{$0$}] at (0,23) {};
\node (h2) [label=below:{$2$}] at (2,23) {};
\node (h3) [label=below:{$4$}] at (4,23) {};
\node (h4) [label=below:{$6$}] at (6,23) {};
\end{scope}
\begin{scope}[matrix box]
\draw (v2.center) -- +(0.3,0);
\draw (v3.center) -- +(0.3,0);
\draw (v4.center) -- +(0.3,0);
\draw (v5.center) -- +(0.3,0);
\draw (h2.center) -- +(0,-0.3);
\draw (h3.center) -- +(0,-0.3);
\draw (h4.center) -- +(0,-0.3);
\end{scope}
\begin{scope}[xshift=5cm]
\node[leading] at (1,3) {};
\node[nonzero] at (1,4) {};
\node[nonzero] at (1,6) {};
\node[nonzero] at (1,10) {};
\node[nonzero] at (1,11) {};
\node[nonzero] at (1,13) {};
\node[nonzero] at (1,14) {};
\node[nonzero] at (1,17) {};
\node[nonzero] at (1,21) {};
\node[nonzero] at (1,22) {};
\node[leading] at (2,10) {};
\node[nonzero] at (2,11) {};
\node[nonzero] at (2,13) {};
\node[nonzero] at (2,14) {};
\node[nonzero] at (2,16) {};
\node[nonzero] at (2,17) {};
\node[nonzero] at (2,19) {};
\node[nonzero] at (2,21) {};
\node[nonzero] at (2,22) {};
\node[leading] at (3,1) {};
\node[nonzero] at (3,3) {};
\node[nonzero] at (3,4) {};
\node[nonzero] at (3,6) {};
\node[nonzero] at (3,9) {};
\node[nonzero] at (3,10) {};
\node[nonzero] at (3,11) {};
\node[nonzero] at (3,16) {};
\node[nonzero] at (3,17) {};
\node[nonzero] at (3,19) {};
\node[leading] at (4,11) {};
\node[nonzero] at (4,14) {};
\node[nonzero] at (4,16) {};
\node[nonzero] at (4,18) {};
\node[nonzero] at (4,21) {};
\node[leading] at (5,14) {};
\node[nonzero] at (5,16) {};
\node[nonzero] at (5,17) {};
\node[nonzero] at (5,18) {};
\node[nonzero] at (5,19) {};
\node[nonzero] at (5,21) {};
\node[nonzero] at (5,22) {};
\node[leading] at (6,17) {};
\node[nonzero] at (6,18) {};
\node[nonzero] at (6,19) {};
\node[nonzero] at (6,21) {};
\draw[matrix box] (0,0) rectangle (7,23);
\begin{scope}[every node/.style={axis label, label distance=-1.5mm}]
\node (v1) [label=left:{$0$}] at (0,0) {};
\node (v2) [label=left:{$5$}] at (0,5) {};
\node (v3) [label=left:{$10$}] at (0,10) {};
\node (v4) [label=left:{$15$}] at (0,15) {};
\node (v5) [label=left:{$20$}] at (0,20) {};
\node (h1) [label=below:{$0$}] at (0,23) {};
\node (h2) [label=below:{$2$}] at (2,23) {};
\node (h3) [label=below:{$4$}] at (4,23) {};
\node (h4) [label=below:{$6$}] at (6,23) {};
\end{scope}
\begin{scope}[matrix box]
\draw (v2.center) -- +(0.3,0);
\draw (v3.center) -- +(0.3,0);
\draw (v4.center) -- +(0.3,0);
\draw (v5.center) -- +(0.3,0);
\draw (h2.center) -- +(0,-0.3);
\draw (h3.center) -- +(0,-0.3);
\draw (h4.center) -- +(0,-0.3);
\end{scope}
\end{scope}
\begin{scope}[xshift=10cm]
\node[leading] at (1,1) {};
\node[nonzero] at (1,3) {};
\node[nonzero] at (1,4) {};
\node[nonzero] at (1,6) {};
\node[nonzero] at (1,9) {};
\node[nonzero] at (1,10) {};
\node[nonzero] at (1,11) {};
\node[nonzero] at (1,16) {};
\node[nonzero] at (1,17) {};
\node[nonzero] at (1,19) {};
\node[leading] at (2,2) {};
\node[nonzero] at (2,3) {};
\node[nonzero] at (2,4) {};
\node[nonzero] at (2,6) {};
\node[nonzero] at (2,10) {};
\node[nonzero] at (2,11) {};
\node[nonzero] at (2,13) {};
\node[nonzero] at (2,14) {};
\node[nonzero] at (2,17) {};
\node[nonzero] at (2,21) {};
\node[nonzero] at (2,22) {};
\node[leading] at (3,3) {};
\node[nonzero] at (3,10) {};
\node[nonzero] at (3,11) {};
\node[nonzero] at (3,13) {};
\node[nonzero] at (3,14) {};
\node[nonzero] at (3,16) {};
\node[nonzero] at (3,17) {};
\node[nonzero] at (3,19) {};
\node[nonzero] at (3,21) {};
\node[nonzero] at (3,22) {};
\node[leading] at (4,4) {};
\node[nonzero] at (4,11) {};
\node[nonzero] at (4,14) {};
\node[nonzero] at (4,16) {};
\node[nonzero] at (4,18) {};
\node[nonzero] at (4,21) {};
\node[leading] at (5,5) {};
\node[nonzero] at (5,14) {};
\node[nonzero] at (5,16) {};
\node[nonzero] at (5,17) {};
\node[nonzero] at (5,18) {};
\node[nonzero] at (5,19) {};
\node[nonzero] at (5,21) {};
\node[nonzero] at (5,22) {};
\node[leading] at (6,6) {};
\node[nonzero] at (6,17) {};
\node[nonzero] at (6,18) {};
\node[nonzero] at (6,19) {};
\node[nonzero] at (6,21) {};
\draw[matrix box] (0,0) rectangle (7,23);
\begin{scope}[every node/.style={axis label, label distance=-1.5mm}]
\node (v1) [label=left:{$0$}] at (0,0) {};
\node (v2) [label=left:{$5$}] at (0,5) {};
\node (v3) [label=left:{$10$}] at (0,10) {};
\node (v4) [label=left:{$15$}] at (0,15) {};
\node (v5) [label=left:{$20$}] at (0,20) {};
\node (h1) [label=below:{$0$}] at (0,23) {};
\node (h2) [label=below:{$2$}] at (2,23) {};
\node (h3) [label=below:{$4$}] at (4,23) {};
\node (h4) [label=below:{$6$}] at (6,23) {};
\end{scope}
\begin{scope}[matrix box]
\draw (v2.center) -- +(0.3,0);
\draw (v3.center) -- +(0.3,0);
\draw (v4.center) -- +(0.3,0);
\draw (v5.center) -- +(0.3,0);
\draw (h2.center) -- +(0,-0.3);
\draw (h3.center) -- +(0,-0.3);
\draw (h4.center) -- +(0,-0.3);
\end{scope}
\end{scope}
\end{tikzpicture}
\caption{Left: ordered sparse GLT matrix with six factors.
Center: one of the $2^6 \cdot 6$! corresponding unordered sparse GLT matrices.
Right: a corresponding sparse PLT matrix, i.e.\ enforced non-zeros on the main diagonal.
The pivot rows $(l_1, \ldots, l_6) =(1,3,10,11,14,17)$ are marked by triangles.
Non-zero loadings are marked by circles, zero loadings are left blank.}
\label{figglt}
\end{figure}
The paper is structured as follows. Section~\ref{secmotivate} reviews the role of identification in factor analysis
using illustrative examples. Section~\ref{uniqload} introduces GLT structures, proves identification for
sparse GLT structures and shows that any unconstrained loading matrix has a unique representation as a GLT matrix.
Section~\ref{varidesp} addresses variance identification under GLT structures.
Section~\ref{secEFA} discusses exploratory factor analysis under
unordered GLT structures, while Section~\ref{secapll} presents an illustrative application. Section~\ref{secconcluse} concludes.
\section{The role of identification in factor analysis} \label{secmotivate}
Let ${\mathbf y}_t=(y_{1t}, \ldots, y_{m t})^{\top}$ be an observation vector of $m$ measurements, which is assumed to arise from a multivariate normal distribution, ${\mathbf y}_t \sim \mathcal{N} _{m}\left({\mathbf{0}},{\mathbf \Omega}\right)$, with zero mean and covariance matrix ${\mathbf \Omega}$.
In factor analysis, the correlation among the observations is assumed to be driven by a latent $r$-variate random variable ${\mathbf f}_t=(f_{1t}, \ldots, f_{r t})^{\top}$, the so-called common factors, through the following observation equation:
\begin{eqnarray} \label{fac1}
{\mathbf y}_t = \boldsymbol{\Lambda} {\mathbf f}_t + \boldsymbol{\epsilon}_t,
\end{eqnarray}
where the $m\times r$ matrix $\boldsymbol{\Lambda}$ containing the
factor loadings $\Lambda_{ij}$ is of full column rank, $\mbox{\rm rk}\,(\boldsymbol{\Lambda})=r$,
equal to the factor dimension $r$.
In the present paper, we
focus on the so-called basic factor model where
the vector $\boldsymbol{\epsilon}_t =(\epsilon_{1t}, \ldots, \epsilon_{m t})^{\top}$
accounts for independent, idiosyncratic variation of each measurement
and is distributed as $\boldsymbol{\epsilon}_t
\sim \mathcal{N} _{m}\left({\mathbf{0}},{\mathbf \Sigma}_0\right) $, with
${\mathbf \Sigma}_0=\mbox{\rm Diag}\!\left(\sigma^2_{1},\ldots,\sigma^2_{m}\right)$
being a positive definite diagonal matrix.
The common factors are orthogonal, meaning that
${\mathbf f}_t \sim \mathcal{N} _{r}\left({\mathbf{0}},{{\mathbf I}}_{r}\right),$
and independent of
$\boldsymbol{\epsilon}_t$.
In this case, the observation equation (\ref{fac1})
implies the following covariance matrix ${\mathbf \Omega}$,
when we integrate w.r.t.\ the latent common factors ${\mathbf f}_t$:
\begin{eqnarray}
{\mathbf \Omega}= \boldsymbol{\Lambda} \boldsymbol{\Lambda} ^{\top} + {\mathbf \Sigma}_0 . \label{fac4}
\end{eqnarray}
Hence, all dependence among the measurements in ${\mathbf y}_t$
is explained through the latent common factors and
the off-diagonal elements
of $\boldsymbol{\Lambda} \boldsymbol{\Lambda} ^{\top}$ define
the marginal covariance between any two measurements $y_{i_1,t}$ and $y_{i_2, t}$:
\begin{eqnarray}
\mbox{\rm Cov}(y_{i_1,t}, y_{i_2, t}) = \boldsymbol{\Lambda} _{i_1,\bullet} \boldsymbol{\Lambda} _{i_2,\bullet}^{\top},
\label{fac5}
\end{eqnarray}
where $\boldsymbol{\Lambda} _{i,\bullet}$ is the $i$th row of $\boldsymbol{\Lambda}$. Consequently,
we will refer to $\boldsymbol{\Lambda} \boldsymbol{\Lambda} ^{\top}$ as the cross-covariance matrix.
Since the number of factors, $r$, is often considerably
smaller than the number of measurements, $m$, (\ref{fac4}) can be seen as a
parsimonious representation of the dependence between the measurements, often
with considerably fewer parameters in $\boldsymbol{\Lambda}$ than the $m(m -1)/2$
off-diagonal elements in an unconstrained covariance matrix ${\mathbf \Omega}$.
Since the factors ${\mathbf f}_t$ are unobserved, the only information
available to estimate $\boldsymbol{\Lambda}$ and ${\mathbf \Sigma}_0$
is the covariance matrix ${\mathbf \Omega}$. A rigorous approach toward identification of factor
models was first offered by \citet{rei:ide} and \citet{and-rub:sta}.
Identification in the context of a basic factor model means the following. For any pair
$(\boldsymbol{\beta} ,{\mathbf \Sigma})$, where $\boldsymbol{\beta}$ is an $m \times r$ matrix and
${\mathbf \Sigma}$ is a positive definite diagonal matrix,
that satisfies (\ref{fac4}), i.e.:
\begin{eqnarray}
{\mathbf \Omega} = \boldsymbol{\beta} \boldsymbol{\beta}^{\top} + {\mathbf \Sigma}
= \boldsymbol{\Lambda} \boldsymbol{\Lambda} ^{\top} + {\mathbf \Sigma}_0, \label{facide1}
\end{eqnarray}
it follows that $\boldsymbol{\beta}=\boldsymbol{\Lambda}$ and
${\mathbf \Sigma} ={\mathbf \Sigma}_0$.
Note that both parameter pairs imply the same Gaussian distribution ${\mathbf y}_t \sim \mathcal{N} _{m}\left({\mathbf{0}},{\mathbf \Omega}\right)$
for every possible realisation ${\mathbf y}_t$.
\citet{and-rub:sta} considered identification as a two-step procedure. The first step is
identification of the variance decomposition, i.e.\ identification of ${\mathbf \Sigma}_0$ from (\ref{fac4}),
which implies identification of $\boldsymbol{\Lambda} \boldsymbol{\Lambda} ^{\top}$.
The second step is subsequent identification of $\boldsymbol{\Lambda}$ from
$\boldsymbol{\Lambda} \boldsymbol{\Lambda} ^{\top}$, also know as solving the rotational invariance problem.
The literature on factor analysis often
reduces identification of factor models to the second problem,
however as we will argue in the present paper, variance identification is equally important.
\paragraph*{Rotational invariance.}
Let us assume for the moment
that $\boldsymbol{\Lambda} \boldsymbol{\Lambda} ^{\top}$ is identified.
Consider, for further illustration, the following factor loading matrix
$\boldsymbol{\Lambda} $ and a loading matrix $\boldsymbol{\beta} = \boldsymbol{\Lambda} {\mathbf P}_{\alpha b} $ defined as
a rotation of $\boldsymbol{\Lambda} $:
\begin{eqnarray} \label{example1}
\boldsymbol{\Lambda} = \left(\begin{array}{cc}
\lambda_{11} & 0 \\
\lambda _{21} & 0 \\
\lambda_{31} & 0 \\
0 & \lambda_{42} \\
0 & \lambda_{52}\\
0 & \lambda_{62}
\end{array}\right), \quad {\mathbf P}_{\alpha b} = \left(
\begin{array}{rr}
\cos \alpha & (-1)^b\sin \alpha \\
-\sin \alpha & (-1)^b\cos \alpha \\
\end{array}
\right),
\quad
\boldsymbol{\beta}= \boldsymbol{\Lambda} {\mathbf P}_{\alpha b} = \left(\begin{array}{cc}
\beta_{11} & \beta_{21} \\
\beta _{21} & \beta _{22} \\
\beta_{51} & \beta_{52} \\
\beta _{41} & \beta_{42} \\
\beta _{51} & \beta_{52}\\
\beta _{61} & \beta_{62}
\end{array}\right).
\end{eqnarray}
For any
$\alpha \in [0, 2\pi)$ and $b\in\{0,1\}$,
the factor loading matrix
$\boldsymbol{\beta} $
yields the same cross-covariance matrix for ${\mathbf y}_t$ as $\boldsymbol{\Lambda}$, as is easily verified:
\begin{eqnarray}
\boldsymbol{\beta} \boldsymbol{\beta}^{\top} =
\boldsymbol{\Lambda} {\mathbf P}_{\alpha b} \mathbf P^{\top}_{\alpha b} \boldsymbol{\Lambda} ^{\top} = \boldsymbol{\Lambda} \boldsymbol{\Lambda} ^{\top}.
\label{rotation}
\end{eqnarray}
The rotational invariance apparent in (\ref{rotation}) holds
more generally for
any basic factor model (\ref{fac1}). Take any
arbitrary $r\times r$ rotation matrix
${\mathbf P}$ (i.e.\ ${\mathbf P} {\mathbf P}^{\top}= {{\mathbf I}}_{r}$)
and define the basic factor model
\begin{eqnarray} \label{fac1rot}
{\mathbf f}_t^{\star} \sim \mathcal{N} _{r}\left({\mathbf{0}},{{\mathbf I}}_{r}\right), \quad
{\mathbf y}_t = \boldsymbol{\beta} {\mathbf f}^{\star}_t + \boldsymbol{\epsilon}_t, \quad \boldsymbol{\epsilon}_t
\sim \mathcal{N} _{m}\left({\mathbf{0}},{\mathbf \Sigma}_0\right),
\end{eqnarray}
where $\boldsymbol{\beta} = \boldsymbol{\Lambda} {\mathbf P}$ and
${\mathbf f}_t^{\star} = {\mathbf P}^{\top} {\mathbf f}_t$. Then both models
imply the
same covariance ${\mathbf \Omega}$, given by (\ref{fac4}).
Hence, without imposing further constraints, $\boldsymbol{\Lambda}$ is in general
not identified from the cross-covariance matrix $ \boldsymbol{\Lambda} \boldsymbol{\Lambda}^{\top} $.
If interest lies in interpreting the factors through the factor loading matrix $\boldsymbol{\Lambda}$,
rotational invariance has to be resolved.
The usual way of dealing with rotational invariance is to constrain $\boldsymbol{\Lambda}$ in such a
way that the only possible rotation is the identity ${\mathbf P}={{\mathbf I}}_{r}$.
For orthogonal factors at least $r(r-1)/2$ restrictions on the elements of
$\boldsymbol{\Lambda}$ are needed to eliminate rotational indeterminacy \citep{and-rub:sta}.
The most popular constraints are positive lower triangular (PLT) loading matrices, where the upper triangular part is constrained to be zero and the main diagonal elements $\Lambda_{11},\ldots, \Lambda_{r r}$ of $\boldsymbol{\Lambda}$ are strictly positive, see Figure~\ref{figglt} for illustration.
Despite its popularity, the PLT structure is restrictive, as outlined already by \citet{joe:gen}.
Let $\boldsymbol{\beta} \boldsymbol{\beta}^{\top} $ be an arbitrary cross-covariance matrix with factor loading matrix $\boldsymbol{\beta}$. A
PLT representation of $ \boldsymbol{\beta} \boldsymbol{\beta}^{\top} $
is possible iff a rotation matrix $ {\mathbf P}$
exists such that $\boldsymbol{\beta} $ can be rotated into a PLT matrix $\boldsymbol{\Lambda} = \boldsymbol{\beta} {\mathbf P} $.
However, as example (\ref{example1}) illustrates this is not necessarily the case.
Obviously, $ \boldsymbol{\Lambda}$ is not a PLT matrix, since
$\Lambda_{22}=0$. Any of the possible rotations
$\boldsymbol{\beta} = \boldsymbol{\Lambda} {\mathbf P}_{\alpha b}$ have {\em non-zero}
elements above the main diagonal and are not PLT matrices either.
This example demonstrates that the
PLT representation is restrictive.
To circumvent this problem in example (\ref{example1}), one could reorder the measurements
in an appropriate manner. However, in applied
factor analysis, such an appropriate ordering is typically not known in advance and
the choice of the first
$r$ measurements
is an important modeling decision under PLT constraints, see e.g.\ \citet{lop-wes:bay} and~\citet{car-etal:hig}.
We discuss in Section~\ref{uniqload} a new identification strategy to resolve rotational invariance in factor models
based on the concept of generalized lower triangular (GLT) structures.
Loosely speaking, GLT structures generalize PLT structures by
freeing the position of the first non-zero factor loading in each
column, see the loading matrix $ \boldsymbol{\Lambda}$ in (\ref{example1})
and Figure~\ref{figglt} for an example.
We show in Section~\ref{secGLT} that
a unique GLT structure $\boldsymbol{\Lambda}$ can be identified for
any cross-covariance matrix $\boldsymbol{\beta} \boldsymbol{\beta}^{\top}$, provided that
variance identification holds and, consequently,
$\boldsymbol{\beta} \boldsymbol{\beta}^{\top}$ itself is identified.
Even if $ \boldsymbol{\beta} \boldsymbol{\beta}^{\top} $ is obtained
from a loading matrix $\boldsymbol{\beta}$ that does not take the form of a
GLT structure,
such as the matrix $\boldsymbol{\beta}$ in (\ref{example1}),
we show in Section~\ref{secRREF}
that a {\em unique} orthogonal matrix ${\mathbf G}$ exists which represents
$\boldsymbol{\beta}$ as a rotation of a unique GLT structure $\boldsymbol{\Lambda}$:
\begin{eqnarray} \label{rotRREFint}
\boldsymbol{\Lambda}= \boldsymbol{\beta} {\mathbf G},
\end{eqnarray}
which we call
rotation into GLT. Hence,
the GLT representation is unrestrictive in the sense of \citet{joe:gen}
and is, indeed, a new and generic way to resolve rotational invariance
for any factor loading matrix.
\paragraph*{Sparse factor loading matrices.}
The factor loading matrix $\boldsymbol{\Lambda}$ given in (\ref{example1}) is an
example of a sparse loading matrix. While only a single zero loading would
be needed to resolve rotational invariance,
six zeros are present and each factor loads only on dedicated
measurements. Such sparse loading matrices
are generated by a binary indicator matrix $\boldsymbol{\delta}$
of 0s and 1s of the same dimension as $\boldsymbol{\Lambda}$,
where $\Lambda_{ij}=0$ iff $\delta_{ij}=0$, and
$\Lambda_{ij} \in \mathbb{R}$ is unconstrained otherwise.
The binary matrix $\boldsymbol{\delta}=\mathbb{I} (\boldsymbol{\Lambda} \neq 0)$, where the indicator function
is applied element-wise, is called the sparsity matrix corresponding to $\boldsymbol{\Lambda}$.
The sparsity matrix $\boldsymbol{\delta}$ contains a lot of information about the structure of
$\boldsymbol{\Lambda}$, see Figure~\ref{figglt} for illustration.
The indicator matrix on the right hand side tells us that $\boldsymbol{\Lambda}$ obeys
the PLT constraint.
The fifth row of the left and center matrices contains only zeros, which tells us that observation $y_{5t}$ is uncorrelated with the remaining observations,
since $\mbox{\rm Cov}(y_{it}, y_{5t}) = 0$ for all $i\neq 5$.
\paragraph*{Variance identification.}
Constraints that resolve rotational invariance
typically take variance identification, i.e.\ identification of
$\boldsymbol{\Lambda} \boldsymbol{\Lambda} ^{\top}$, for granted, see e.g.\ \citet{gew-zho:mea}.
Variance identification refers to the problem that
the idiosyncratic variances $\sigma^2_1, \ldots, \sigma^2_m$ in ${\mathbf \Sigma}_0$
are identified only from the diagonal elements
of ${\mathbf \Omega}$, as all other elements are independent of the $\sigma^2_i$s; see again \eqref{fac5}.
To achieve variance identification of $\sigma^2_i$ from $\Omega_{ii}=
\boldsymbol{\Lambda} _{i,\bullet} \boldsymbol{\Lambda} _{i,\bullet}^{\top} + \sigma^2_i$,
all factor loadings
have to be identified solely from the off-diagonal elements of ${\mathbf \Omega}$.
Variance identification,
however, is easily violated,
as the following considerations illustrate.
Let us return to the factor model defined in (\ref{example1}).
The corresponding covariance matrix ${\mathbf \Omega}$ is given by:
{\begin{eqnarray} \small
{\mathbf \Omega} = \left(
\begin{array}{cccccc}
{ \lambda_{11}^2 + {\sigma^2_1} } & { \lambda_{11} \lambda_{21}} & { \lambda_{11} \lambda_{31}} & && \\
{ \lambda_{11} \lambda_{21}} & { \lambda_{21}^2 + {\sigma^2_2} } & { \lambda_{21} \lambda_{31}} & &
{\mathbf{0}} &\\
{ \lambda_{11} \lambda_{31}} & { \lambda_{21} \lambda_{31}} & { \lambda_{31}^2 + {\sigma^2_3} } & && \\
& && \lambda_{42}^2 + {\sigma^2_4} & \lambda_{42} \lambda_{52} & \lambda_{42} \lambda_{62} \\
& {\mathbf{0}} & & \lambda_{42} \lambda_{52} & \lambda_{52}^2 + {\sigma^2_5} & \lambda_{52} \lambda_{62} \\
& & & \lambda_{42} \lambda_{62} & \lambda_{52} \lambda_{62}& \lambda_{62}^2 + {\sigma^2_6} \\
\end{array}\right). \label{Omegatrue}
\end{eqnarray}}
Let us assume that the sparsity pattern $\boldsymbol{\delta}=\mathbb{I} (\boldsymbol{\Lambda} \neq 0)$
of $\boldsymbol{\Lambda}$ is known, but the specific values of the
unconstrained loadings $(\lambda_{11}, \ldots, \lambda_{62})$ are
unknown.
An interesting question is the following. Knowing ${\mathbf \Omega}$,
can the unconstrained loadings $\lambda_{11}, \ldots,
\lambda_{62}$ and the variances $\sigma^2_1, \ldots, \sigma^2_m$ be identified uniquely?
Given ${\mathbf \Omega}$, the three nonzero covariances $\mbox{\rm Cov}(y_{1t}, y_{2t})= \lambda_{11} \lambda_{21}$,
$\mbox{\rm Cov}(y_{1t}, y_{3t})=\lambda_{11} \lambda_{31}$
and $\mbox{\rm Cov}(y_{2t}, y_{3t})=\lambda_{21} \lambda_{31}$
are available to identify the three factor loadings $(\lambda_{11}, \lambda_{21}, \lambda_{31})$.
Similarly,
the nonzero covariances $\mbox{\rm Cov}(y_{4t}, y_{5t})= \lambda_{42} \lambda_{52}$,
$\mbox{\rm Cov}(y_{4t}, y_{6t})=\lambda_{42} \lambda_{62}$
and $\mbox{\rm Cov}(y_{5t}, y_{6t})=\lambda_{52} \lambda_{62}$
are available to identify the factor loadings $(\lambda_{42}, \lambda_{52}, \lambda_{62})$,
hence variance identification is given.
However, if we remove the last measurement from the loading
factor matrix defined in (\ref{example1}), we obtain
\begin{eqnarray} \label{example2}
\boldsymbol{\Lambda} = \left(\begin{array}{cc}
\lambda_{11} & 0 \\
\lambda _{21} & 0 \\
\lambda_{31} & 0 \\
0 & \lambda_{42} \\
0 & \lambda_{52}\\
\end{array}\right),
\quad
\boldsymbol{\beta}= \boldsymbol{\Lambda} {\mathbf P}_{\alpha b}= \left(\begin{array}{cc}
\beta_{11} & \beta_{21} \\
\beta _{21} & \beta _{22} \\
\beta_{51} & \beta_{52} \\
\beta _{41} & \beta_{42} \\
\beta _{51} & \beta_{52}\\
\end{array}\right),
\end{eqnarray}
and the corresponding covariance matrix reads:
\begin{eqnarray*} {\small
{\mathbf \Omega} = \left(
\begin{array}{ccccc}
{ \lambda_{11}^2 + {\sigma^2_1} } & { \lambda_{11} \lambda_{21}} & { \lambda_{11} \lambda_{31}} & & \\
{ \lambda_{11} \lambda_{21}} & { \lambda_{21}^2 + {\sigma^2_2} } & { \lambda_{21} \lambda_{31}} &
{\mathbf{0}} &\\
{ \lambda_{11} \lambda_{31}} & { \lambda_{21} \lambda_{31}} & { \lambda_{31}^2 + {\sigma^2_3} } & & \\
& && \lambda_{42}^2 + {\sigma^2_4} & \lambda_{42} \lambda_{52} \\
& {\mathbf{0}} & & \lambda_{42} \lambda_{52} & \lambda_{52}^2 + {\sigma^2_5} \large
\end{array}\right)}.
\end{eqnarray*}
While the three factor loadings $(\lambda_{11}, \lambda_{21}, \lambda_{31})$ are still identified from
the off-diagonal elements of ${\mathbf \Omega}$ as before,
variance identification of $\sigma^2_4$ and $\sigma^2_5$ fails.
Since $\mbox{\rm Cov}(y_{4t}, y_{5t})= \lambda_{42} \lambda_{52}$ is the only non-zero
element
that depends on
the loadings $\lambda_{42}$ and $\lambda_{52}$,
infinitely many different parameters $(\lambda_{42},\lambda_{52},
\sigma^2_4,\sigma^2_5)$ imply the same covariance matrix ${\mathbf \Omega}$.
From these considerations it is evident that a minimum of three
non-zero loadings is necessary
in each column to achieve variance identification, a condition
which has been noted as early as \citet{and-rub:sta}. At the same time,
this condition
is not sufficient, as it is satisfied by the loading matrix $\boldsymbol{\beta}$ in
\eqref{example2}, although variance identification does
not hold.
In general, variance identification is not straightforward
to verify. We will introduce in Section~\ref{sec:3579} a new and convenient way to
verify variance identification for GLT structures.
\paragraph*{The row deletion property.}
As explained above, we need to verify
uniqueness of the variance decomposition, i.e.\ the identification of the idiosyncratic
variances $\sigma^2_1, \ldots, \sigma^2_m$ in
${\mathbf \Sigma}_0$ from the covariance matrix ${\mathbf \Omega}$ given in (\ref{fac4}).
The identification of ${\mathbf \Sigma}_0$ guarantees that $\boldsymbol{\Lambda} \boldsymbol{\Lambda} ^{\top}$ is
identified.
The second step of identification is then
to ensure uniqueness of the factor loadings, i.e.\ unique identification of $\boldsymbol{\Lambda}$
from $\boldsymbol{\Lambda} \boldsymbol{\Lambda} ^{\top}$.
To verify variance identification,
we rely in the present paper on a condition
known as {\em row-deletion property}.
\begin{defn}[Row deletion property \mbox{\bf AR}\ \citep{and-rub:sta}] \label{ARdef}
An $m \times r$ factor loading matrix $\boldsymbol{\Lambda}$
satisfies the row-deletion property if the following condition is satisfied:
whenever an arbitrary row is deleted from $\boldsymbol{\Lambda}$, two disjoint submatrices of rank $r$ remain.
\end{defn}
\noindent \citet[Theorem~5.1]{and-rub:sta} prove that the row-deletion property
is a sufficient condition
for
the identification of $\boldsymbol{\Lambda} \boldsymbol{\Lambda}^{\top} $ and ${\mathbf \Sigma}_0$
from the marginal covariance matrix ${\mathbf \Omega}$ given in (\ref{fac4}).
For any (not necessarily GLT) factor loading matrix $\boldsymbol{\Lambda}$, the
row deletion property \mbox{\bf AR}\
can be trivially tested by a step-by-step analysis, where each single row of
$\boldsymbol{\Lambda}$ is sequentially deleted and the two distinct submatrices are determined
from examining the remaining matrix, as suggested e.g.\ by \citet{hay-mar:exa}.
However, this procedure is inefficient and challenging in higher dimensions.
Hence, it is helpful to have more structural conditions for verifying variance identification
under the row deletion property \mbox{\bf AR} .
The literature provides several necessary
conditions for the row deletion property \mbox{\bf AR}\ that are based on
counting the number of non-zero factor loadings in $\boldsymbol{\Lambda}$.
\citet{and-rub:sta}, for instance, prove
the following necessary conditions for \mbox{\bf AR} : for every nonsingular $r$-dimensional
square matrix ${\mathbf G}$,
the matrix $\boldsymbol{\beta}=\boldsymbol{\Lambda} {\mathbf G}$ contains in each column \emph{at least 3}
and in each pair of columns \emph{at least 5} nonzero factor loadings.
\citet[Theorem~3.3]{sat:stu} extends these {\em necessary}
conditions in the following way: every subset of
$1\leq q \leq r$ columns
of $\boldsymbol{\beta}=\boldsymbol{\Lambda} {\mathbf G}$ contains \emph{at least $2q+1$} nonzero factor loadings for every
nonsingular matrix ${\mathbf G}$. We call this the $3579$ counting rule for obvious reasons.
For illustration, let us return to the examples in (\ref{example1}) and (\ref{example2}).
First, apply the $3579$ counting rule to the unrestricted matrix $\boldsymbol{\beta}$ in (\ref{example2}).
Although the variance decomposition
${\mathbf \Omega}= \boldsymbol{\beta} \boldsymbol{\beta} ^{\top} + {\mathbf \Sigma}
= \boldsymbol{\Lambda} \boldsymbol{\Lambda} ^{\top} + {\mathbf \Sigma}_0,$
is not unique, the counting rules are not violated,
since $\boldsymbol{\beta}$ has five non-zero rows
except for the cases $(\alpha,b) \in \{0, \frac{\pi}{2}, \pi, \frac{3\pi}{2}\}\times \{0,1\}$.
Only for these eight specific cases, which correspond to the trivial rotations
\begin{eqnarray} \label{example3}
&&
\left(\begin{array}{cc}
\lambda_{11} & 0 \\
\lambda _{21} & 0 \\
\lambda_{31} & 0 \\
0 & \lambda_{42} \\
0 & \lambda_{52}\\
\end{array}\right) \quad
\left(\begin{array}{cc}
\lambda_{11} & 0 \\
\lambda _{21} & 0 \\
\lambda_{31} & 0 \\
0 & -\lambda_{42} \\
0 & -\lambda_{52}\\
\end{array}\right) \quad
\left(\begin{array}{cc}
-\lambda_{11} & 0 \\
-\lambda _{21} & 0 \\
-\lambda_{31} & 0 \\
0 & \lambda_{42} \\
0 & \lambda_{52}\\
\end{array}\right) \quad
\left(\begin{array}{cc}
-\lambda_{11} & 0 \\
-\lambda _{21} & 0 \\
-\lambda_{31} & 0 \\
0 & -\lambda_{42} \\
0 & -\lambda_{52}\\
\end{array}\right)\\
&& \nonumber \left(\begin{array}{cc}
0 & \lambda_{11} \\
0 & \lambda _{21} \\
0 & \lambda_{31} \\
\lambda_{42} & 0\\
\lambda_{52} & 0\\
\end{array}\right) \quad
\left(\begin{array}{cc}
0 & -\lambda_{11} \\
0 & - \lambda _{21} \\
0 & -\lambda_{31} \\
\lambda_{42} & 0\\
\lambda_{52} & 0\\
\end{array}\right) \quad
\left(\begin{array}{cc}
0 & \lambda_{11} \\
0 & \lambda _{21} \\
0 & \lambda_{31} \\
-\lambda_{42} & 0\\
-\lambda_{52} & 0\\
\end{array}\right) \quad
\left(\begin{array}{cc}
0 & -\lambda_{11} \\
0 & -\lambda _{21} \\
0 & -\lambda_{31} \\
-\lambda_{42} & 0\\
-\lambda_{52} & 0\\
\end{array}\right),
\end{eqnarray}
we find immediately that the counting rules are violated, since
one of the two columns has only two non-zero elements.
This example shows the need to check such counting rules not only
for a single loading matrix $\boldsymbol{\beta}$, but also for all rotations $\boldsymbol{\beta} {\mathbf P}$
admissible under the chosen strategy toward rotational invariance.
On the other hand, if we apply the $3579$ counting rule to
the unrestricted matrix $\boldsymbol{\beta}$ in (\ref{example1}),
we find that the {\em necessary} counting rules are satisfied for all rotations $\boldsymbol{\beta} {\mathbf P}_{\alpha b}$.
For this specific example, we have already verified explicitly that variance identification holds
and one might wonder if, in general, the $3579$ counting rule can lead to
a {\em sufficient} criterion for variance identification under \mbox{\bf AR} .
Sufficient conditions for variance identification are hardly investigated in the literature.
One exception is the popular factor analysis model where $\boldsymbol{\Lambda}$
takes the form of a {\em dense} PLT matrix, where
all factor loadings
on and below the main diagonal are left unrestricted and can take any in value in $\mathbb{R}$.
For this model, condition \mbox{\bf AR}\ and hence variance identification
holds, except for a set of measure 0, if the condition
$m \geq 2r +1$ is satisfied.
\citet{con-etal:bay_exp} investigate identification of a dedicated factor model, where equation
(\ref{fac1}) is
combined with correlated (oblique) factors, ${\mathbf f}_t \sim \mathcal{N} _{r}\left({\mathbf{0}},\mathbf{R}\right)$,
and the factor loading matrix $\boldsymbol{\Lambda}$ has a perfect simple structure, i.e.\ each observation
loads on at most one factor, as in (\ref{example1}) and (\ref{example2}); however, the exact
position of the non-zero elements is unknown.
They prove
necessary and sufficient conditions that imply uniqueness of the variance decomposition
as well as uniqueness of the
factor loading matrix,
namely:
the correlation matrix $\mathbf{R}$ is of full rank ($\mbox{\rm rk}\,(\mathbf{R})=r$) and
each column of $\boldsymbol{\Lambda}$ contains at least three nonzero loadings.
In the present paper, we build on and extend this previous work.
We provide
sufficient conditions for variance identification of a GLT structure $\boldsymbol{\Lambda}$.
These conditions are formulated as counting rules
for the $m \times r$
sparsity matrix $\boldsymbol{\delta} = \mathbb{I} (\boldsymbol{\beta}\neq 0)$ of $\boldsymbol{\beta}$
and are
equivalent to the $3579$ counting rules of \citet[Theorem~3.3]{sat:stu}.
More specifically, if
the $3579$ counting rule holds for the sparsity matrix $\boldsymbol{\delta}$ of a GLT matrix
$\boldsymbol{\Lambda}$, then this is a sufficient condition
for the row deletion property \mbox{\bf AR}\ and consequently for variance identification, except for a set of measure 0.
\paragraph*{Identification of the number of factors.}
Identification of the number of factors is a notoriously difficult problem and
analysing this problem from the view point of variance identification is helpful
in understanding some fundamental difficulties.
Assume that ${\mathbf \Omega}$ has a representation as in (\ref{fac4}) with $r$ factors
which is variance identified.
Then, on the one hand, no equivalent representation exists with $r' < r$ number of factors.
On the other hand, as shown in
\citet[Theorem~3.3]{rei:ide}, any such structure $(\boldsymbol{\Lambda},{\mathbf \Sigma}_0)$
creates solutions $(\boldsymbol{\beta}_k, {\mathbf \Sigma}_k)$ with $m \times k$ loading matrices
$\boldsymbol{\beta}_k$ of dimension $k= r+1, r+2, \ldots, m$ bigger than $r$
and ${\mathbf \Sigma}_k$ being a positive definite matrix different from ${\mathbf \Sigma}_0$
which imply the same covariance matrix ${\mathbf \Omega}$ as $(\boldsymbol{\Lambda},{\mathbf \Sigma}_0)$,
i.e.:
\begin{eqnarray}
{\mathbf \Omega}= \boldsymbol{\beta}_k \boldsymbol{\beta}^{\top}_k + {\mathbf \Sigma}_k . \label{facover}
\end{eqnarray}
Furthermore, for any fixed $k > r$,
infinitely many such solutions $(\boldsymbol{\beta}_k, {\mathbf \Sigma}_k) $ can be created
that satisfy the decomposition (\ref{facover}) which, consequently,
no longer is variance identified.
This problem is prevalent regardless
of the chosen strategy toward rotational invariance.
For illustration, we return to example (\ref{example1}) and
construct an equivalent solution for $k=3$.
While the first two columns of $\boldsymbol{\beta}_3$ are equal to $\boldsymbol{\Lambda}$,
the third column is a so-called {\em spurious} factor with a single non-zero loading
and ${\mathbf \Sigma}_3 $ is defined as follows:
\begin{eqnarray} \label{example4}
\boldsymbol{\beta}_3 = \left(\begin{array}{ccc}
\lambda_{11} & 0 & 0\\
\lambda _{21} & 0 & \beta_{23} \\
\lambda_{31} & 0 & 0 \\
0 & \lambda_{42} & 0 \\
0 & \lambda_{52} & 0\\
0 & \lambda_{62} & 0
\end{array}\right), \quad
{\mathbf \Sigma}_3 = \mbox{\rm Diag}\!\left(\sigma^2_1,\sigma^2_2 - \beta_{23}^2 ,\sigma^2_3 , \sigma^2_4,\sigma^2_5,\sigma^2_6\right).
\end{eqnarray}
We can place the spurious factor loading $\beta_{i3}$ in any row $i$
and $\beta_{i3}$ can take any value
satisfying $0 < \beta_{i3}^2 < \sigma^2_i$.
It is easy to verify that any such pair
$(\boldsymbol{\beta}_3,{\mathbf \Sigma}_3)$ indeed implies the same covariance matrix ${\mathbf \Omega}$ as in (\ref{Omegatrue}).
This ambiguity in an overfitting model
renders the estimation of true number of factors
$r$ a challenging problem and
leads to considerable
uncertainty how to choose the number of factors in applied factor analysis.
In Section~\ref{secEFA}, we follow up on this problem in more detail.
An important necessary condition for $k$ to be the true number of factors
is that variance identification of
${\mathbf \Sigma}_k$ in (\ref{facover}) holds. Therefore, the
counting rules that we introduce in this paper will also be useful
in cases where the true number of factors $r$ is unknown.
\paragraph*{Overfitting GLT structures.}
Finally, we investigate in Section~\ref{secEFA} the class of potentially overfitting
GLT structures where the matrix $\boldsymbol{\beta}_k $ in (\ref{facover}) is constrained
to be an unordered GLT structure. We apply results by \citet{tum-sat:ide}
to this class and show how easily
spurious factors and the underlying true factor loading matrix $\boldsymbol{\Lambda}$
are identified under GLT structures, even if the model is overfitting. Our strategy relies on
the concept of extended variance identification
and the extended row deletion property introduced by \citet{tum-sat:ide},
where more than one row is deleted from the loading
matrix. An extended counting rule will be introduced
for the sparsity matrix of a GLT loading matrices $\boldsymbol{\beta}_k$ in Section~\ref{varidesp}
which is useful in this context.
\section{Solving rotational invariance through GLT structures} \label{uniqload}
\subsection{Ordered and unordered GLT structures} \label{secGLT}
In this work, we introduce a new identification strategy to resolve rotational invariance
based on the concept of generalized lower triangular (GLT) structures.
First, we introduce the notion of {\em pivot rows} of a factor loading matrix $\boldsymbol{\Lambda}$.
\begin{defn}[\bf Pivot rows] \label{Pivotdef}
Consider an $m \times r$ factor loading matrix $\boldsymbol{\Lambda}$
with $r$ non-zero columns.
For each column $j=1, \ldots, r$ of
$\boldsymbol{\Lambda}$,
the pivot row $l_j$ is defined as the row index of the first
non-zero factor loading in column $j$, i.e.\ $ \Lambda_{ij}=0, \forall \, i<l_j$ and
$\Lambda_{l_j,j} \neq 0$.
The factor loading $\Lambda_{l_j,j}$ is called the leading factor loading of column $j$.
\end{defn}
\noindent
For PLT factor loading matrices the pivot rows lie on the main diagonal,
i.e.\
$(l_1, \ldots, l_r)=(1,\ldots,r)$,
and the leading factor loadings $\Lambda_{jj} > 0$ are positive
for all columns $j=1,\ldots,r$.
GLT structures generalize the PLT constraint
by freeing the pivot rows of a factor loading matrix $\boldsymbol{\Lambda}$ and
allowing them to take arbitrary positions $(l_1, \ldots, l_r)$,
the
only constraint being that the pivot rows are pairwise distinct.
GLT structures contain PLT matrices as the special case where $l_j= j$
for $j=1,\ldots,r$.
Our generalization is particularly useful if the ordering
of the measurements $y_{it}$ is in conflict with the PLT assumption. Since
$\Lambda_{jj}$ is allowed to be 0, measurements different from the first
$r$ ones may lead the factors. For each factor $j$, the leading variable is the
response variable $y_{l_j,t}$ corresponding to the pivot row $l_j$.
We will distinguish between two types of GLT structures, namely ordered and unordered GLT structures.
The following definition introduces ordered GLT matrices.
Unordered GLT structures will be motivated and defined below.
Examples of ordered and unordered
GLT matrices
are displayed in Figure~\ref{figglt} for a model
with $r=6$ factors.
\begin{defn}[\bf Ordered GLT structures] \label{GLTdef}
An $m \times r$ factor loading matrix $\boldsymbol{\Lambda}$ with full column rank $r$
has an ordered GLT structure if the pivot rows $l_1 , \ldots , l_r$ of $\boldsymbol{\Lambda}$ are ordered, i.e.\
$l_1 < \ldots < l_r$, and the leading factor loadings are positive, i.e.\
$\Lambda_{l_j,j} > 0$ for $j=1,\ldots,r$.
\end{defn}
\noindent Evidently, imposing an ordered GLT structure resolves rotational invariance
if the pivot rows are known. For any two ordered GLT matrices
$\boldsymbol{\beta}$ and $ \boldsymbol{\Lambda} $ with {\em identical} pivot rows
$l_1 , \ldots , l_r$,
the identity $\boldsymbol{\beta} = \boldsymbol{\Lambda} {\mathbf P}$ evidently
holds iff ${\mathbf P}={{\mathbf I}}_{r}$.
In practice, the pivot rows $l_1 , \ldots , l_r$ of a GLT structure are unknown and need to be
identified from the marginal covariance matrix ${\mathbf \Omega}$ for a given number of factors $r$.
Given variance identification, i.e.\ assuming that the cross-covariance matrix
$\boldsymbol{\Lambda} \boldsymbol{\Lambda}^{\top}$ is identified,
a particularly important issue for the identification of a GLT factor model
is whether $\boldsymbol{\Lambda} $ is uniquely identified from $\boldsymbol{\Lambda} \boldsymbol{\Lambda}^{\top}$
if the pivot rows $l_1 , \ldots , l_r$ are {\em unknown}.
Non-trivial rotations $\boldsymbol{\beta} = \boldsymbol{\Lambda} {\mathbf P}$ of a loading matrix $\boldsymbol{\Lambda}$ with pivot rows $l_1 , \ldots , l_r$ might exist such that $\boldsymbol{\beta} \boldsymbol{\beta}^{\top} = \boldsymbol{\Lambda} \boldsymbol{\Lambda}^{\top} $,
while the pivot rows $\tilde{l}_1 , \ldots , \tilde{l}_r$ of $\boldsymbol{\beta} $ are different
from the pivot rows of $\boldsymbol{\Lambda}$. Very assuringly,
Theorem~\ref{theGLT} shows that this is not the case: not only the pivot rows, but the entire loading matrices
$\boldsymbol{\Lambda}$ and $\boldsymbol{\beta}$ are identical, if $\boldsymbol{\Lambda} \boldsymbol{\Lambda}^{\top} = \boldsymbol{\beta} \boldsymbol{\beta}^{\top}$ (see Appendix~\ref{app:proof} for a proof).
\begin{thm}\label{theGLT}
An ordered GLT structure is uniquely identified, provided that uniqueness of the variance decomposition holds, i.e.:
if $\boldsymbol{\Lambda}$ and $\boldsymbol{\beta}$ are GLT matrices, respectively, with
pivot rows $l_1 < \ldots < l_r$ and $\tilde{l}_1 < \ldots < \tilde{l}_r$
that satisfy
$\boldsymbol{\beta} \boldsymbol{\beta}^{\top} =
\boldsymbol{\Lambda} \boldsymbol{\Lambda} ^{\top}$, then $\boldsymbol{\beta} = \boldsymbol{\Lambda}$ and
consequently $(\tilde{l}_1 , \ldots , \tilde{l}_r)=(l_1,\ldots , l_r)$.
\end{thm}
Definition~\ref{UGLTdef} introduces, as an extension of Definition~\ref{GLTdef},
unordered GLT structures
under which $\boldsymbol{\Lambda}$
is identified from $\boldsymbol{\Lambda} \boldsymbol{\Lambda} ^{\top}$ only up to signed permutations.
A signed permutation permutes the columns of the factor loading matrix $\boldsymbol{\Lambda}$ and switches the
sign of all factor loadings in any specific column.
This leads to a trivial case of rotational
invariance.
For $r=2$, for instance, the eight signed permutations of the loading matrix $\boldsymbol{\Lambda} $ defined
in (\ref{example2}) are depicted in (\ref{example3}).
More formally, $ \boldsymbol{\beta} $ is
a signed permutation of $ \boldsymbol{\Lambda} $, iff
\begin{align}\label{eq:Ralpha}
\boldsymbol{\beta} = \boldsymbol{\Lambda} {\mathbf P} _{\pm} {\mathbf P} _{\rho} ,
\end{align}
where the permutation matrix ${\mathbf P}_{\rho}$ corresponds to one of the $r$! permutations of the $r$ columns of $\boldsymbol{\Lambda}$
and the reflection matrix
${\mathbf P}_{\pm}=\mbox{\rm Diag}\!\left(\pm 1, \ldots, \pm 1\right)$ corresponds
to one of the $2^r$ ways to switch the signs of the $r$ columns of $\boldsymbol{\Lambda}$.
Often, it is convenient to employ identification rules
that guarantee identification of $\boldsymbol{\Lambda} $
only up to such column and sign switching,
see e.g.\ \citet{con-etal:bay_exp}. Any structure $\boldsymbol{\Lambda}$ obeying such an
identification rule represents a whole equivalence class of matrices
given by all $2^r r !$ signed permutation
$\boldsymbol{\beta} = \boldsymbol{\Lambda} {\mathbf P} _{\pm} {\mathbf P} _{\rho}$ of $\boldsymbol{\Lambda}$.
This trivial form of the rotational invariance does not
impose any additional mathematical challenges
and is often convenient from a computational viewpoint,
in particular for Bayesian inference,
see for e.g.\ \citet{con-etal:bay_exp} and \citet{fru-etal:spa}.
It is easy to verify how identification up to trivial rotational invariance can
be achieved for GLT structures
and
motivates the following definition of unordered GLT structures
as loadings matrices $ \boldsymbol{\beta}$ where
the pivot rows $l_1, \ldots , l_r$ simply occupy $r$ different
rows. In Definition~\ref{UGLTdef}, no order constraint is imposed
on the pivot rows and no sign constraint is imposed on the leading
factor loadings. This very general structure allows to
design highly efficient sampling schemes for sparse Bayesian factor analysis under GLT structures,
see \cite{fru-etal:spa}.
\begin{defn}[\bf Unordered GLT structures] \label{UGLTdef}
An $m \times r$ factor loading matrix $\boldsymbol{\beta}$ with full column rank $r$
has an unordered GLT structure if the pivot rows $l_1, \ldots , l_r$ of $\boldsymbol{\beta}$
are pairwise distinct.
\end{defn}
\noindent Theorem~\ref{theGLT} is easily extended to unordered GLT structures.
Any signed permutation $ \boldsymbol{\beta} =
\boldsymbol{\Lambda} {\mathbf P} _{\rho} {\mathbf P} _{\pm}$ of $\boldsymbol{\Lambda} $
is uniquely identified from $\boldsymbol{\beta} \boldsymbol{\beta}^{\top}= \boldsymbol{\Lambda} \boldsymbol{\Lambda} ^{\top}$, provided that $\boldsymbol{\Lambda} \boldsymbol{\Lambda} ^{\top}$ is identified.
Hence, under unordered GLT structures the factor loading matrix $\boldsymbol{\Lambda}$
is uniquely identified up to signed permutations.
Full identification can easily be obtained from unordered GLT structures $\boldsymbol{\beta}$.
Any unordered GLT structure $\boldsymbol{\beta}$ has unordered pivot rows
$l_1 , \ldots , l_{r}$, occupying different rows.
The corresponding ordered GLT structure $\boldsymbol{\Lambda} $ is recovered from $\boldsymbol{\beta}$
by sorting the columns in ascending order according to the pivot rows.
In other words, the pivot rows of $\boldsymbol{\Lambda} $ are equal to the order statistics
$l_{(1)} , \ldots , l_{(r)}$ of
the pivot rows $l_1, \ldots , l_ r$ of $\boldsymbol{\beta}$,
see again Figure~\ref{figglt}.
This procedure resolves rotational invariance, since the pivot rows
$l_1, \ldots , l_ r$ in the unordered GLT structure are distinct.
Furthermore, imposing the condition $\Lambda_{l_j,j}>0$ in each column $j$
resolves sign switching:
if $\Lambda_{l_j,j}<0$, then
the sign of all factor loadings $\Lambda_{ij}$ in column $j$ is switched.
\subsection{Sparse GLT structures} \label{sec:UGLT}
In Definition~\ref{GLTdef} and \ref{UGLTdef}, \lq\lq structural\rq\rq\ zeros are introduced
for a GLT structure for all
factor loading above the pivot row $l_j$, while the factor loading $\Lambda_{l_j,j}$ in the
pivot row is non-zero by definition. We call $\boldsymbol{\Lambda}$ a {\em dense} GLT structure
if all loadings
below the pivot row are unconstrained and
can take any value in $\mathbb{R}$.
A \emph{sparse} GLT structure results if factor loadings
at unspecified places below the pivot rows
are zero and only the remaining
loadings are unconstrained.
A sparse loading matrix $\boldsymbol{\Lambda}$ can be characterized by
the so-called sparsity matrix, defined as a
binary indicator matrix
$\boldsymbol{\delta}$ of 0/1s of the same size
as $\boldsymbol{\Lambda}$, where $\delta_{ij}=\mathbb{I} (\Lambda_{ij} \neq 0)$.
Let $\boldsymbol{\delta} ^{\Lambda}$ be
the sparsity matrix of a GLT matrix $\boldsymbol{\Lambda}$.
The sparsity matrix $ \boldsymbol{\delta}$ corresponding to the signed permutation
$\boldsymbol{\beta} =
\boldsymbol{\Lambda} {\mathbf P} _{\rho} {\mathbf P} _{\pm}$ is equal to
$ \boldsymbol{\delta} = \boldsymbol{\delta}^{\Lambda} {\mathbf P} _{\rho}$ and is
invariant to sign switching.
Hence, for any sparse unordered GLT matrix $\boldsymbol{\beta}$,
the corresponding sparsity matrix $\boldsymbol{\delta}$ obeys an unordered
GLT structure with the same pivot rows as $\boldsymbol{\beta}$,
see Figure~\ref{figglt} for illustration.
In sparse factor analysis,
single factor loadings take zero-values with positive probability and
the corresponding sparsity matrix $\boldsymbol{\delta}$ is a binary matrix that has to be identified
from the data.
Identification in sparse factor analysis
has to provide conditions under which
the entire 0/1 pattern in $\boldsymbol{\delta}$ can be identified from the
covariance matrix ${\mathbf \Omega}$ if $\boldsymbol{\delta}$ is unknown.
Whether this is possible hinges on variance identification, i.e.\ whether
the decomposition of ${\mathbf \Omega}$ into $\boldsymbol{\Lambda} \boldsymbol{\Lambda}^{\top}$ and ${\mathbf \Sigma}_0$ is unique.
How variance identification can be verified for (sparse) GLT structures is
investigated in detail in Section~\ref{varidesp}.
Let us assume at this point that variance identification holds, i.e.
the cross-covariance matrix $\boldsymbol{\Lambda} \boldsymbol{\Lambda}^{\top}$ is identified.
Then an important step toward the identification of a sparse factor model
is to verify whether the 0/1 pattern of $ \boldsymbol{\Lambda}$, characterized by $\boldsymbol{\delta}$,
is uniquely identified from $\boldsymbol{\Lambda} \boldsymbol{\Lambda}^{\top}$.
Very importantly, if $\boldsymbol{\Lambda} $ is assumed to be
a GLT structure, then the entire GLT structure $\boldsymbol{\Lambda}$ and hence
the indicator matrix $\boldsymbol{\delta}$ is uniquely identified from $\boldsymbol{\Lambda} \boldsymbol{\Lambda} ^{\top}$,
as follows immediately from Theorem~\ref{theGLT},
since $\delta_{i j}= 0$, iff $\Lambda_{i j}= 0 $ for all $i,j$.
By identifying the 0/1 pattern in $\boldsymbol{\delta}$ we can uniquely identify the pivot rows of
$\boldsymbol{\Lambda}$ and the sparsity
pattern below.
We would like to emphasize that in sparse factor analysis with unconstrained loading matrices
$\boldsymbol{\Lambda}$ this is not necessarily the case. The indicator matrix $\boldsymbol{\delta}$ is
in general {\em not} uniquely identified
from $\boldsymbol{\Lambda} \boldsymbol{\Lambda}^{\top}$, because
(non-trivial) rotations ${\mathbf P}$
change the zero pattern in
$\boldsymbol{\beta} = \boldsymbol{\Lambda} {\mathbf P}$, while $\boldsymbol{\beta} \boldsymbol{\beta}^{\top} =
\boldsymbol{\Lambda} \boldsymbol{\Lambda} ^{\top}$.
For illustration, let us return to the example in (\ref{example1}) where we showed that
$\boldsymbol{\Lambda} \boldsymbol{\Lambda}^{\top}$ is uniquely identified if the true
sparsity matrix $\boldsymbol{\delta} ^{\Lambda}$ is known. Now assume
that $\boldsymbol{\delta} ^{\Lambda}$ is unknown
and allow the loading matrix
$\boldsymbol{\beta} = \boldsymbol{\Lambda} {\mathbf P}$ to be any rotation of $\boldsymbol{\Lambda}$.
It is then evident that the corresponding sparsity matrix $\boldsymbol{\delta}$ is not unique
and two solutions exists. For all rotations where
$(\alpha,b) \in \{0, \frac{\pi}{2}, \pi, \frac{3\pi}{2}\}\times\{0,1\}$,
$\boldsymbol{\beta}$ correspond to one of the eight signed permutation of
$\boldsymbol{\Lambda}$
given in
\eqref{example3} and the sparsity matrix $\boldsymbol{\delta}$
is equal to $\boldsymbol{\delta}^{\boldsymbol{\Lambda}}$ up to this signed permutation.
For all other rotations, all elements of $\boldsymbol{\beta}$ are different
from zero and $\boldsymbol{\delta}$ is simply a matrix of ones.
\subsection{Rotation into GLT} \label{secRREF}
As discussed above, GLT structures generalize the PLT constraint, but
one might wonder how restrictive
this structure
still is. We will show in this section that for a basic factor model with unconstrained loading matrix
$\boldsymbol{\beta}$ there exists an equivalent representation involving a unique GLT structure
$\boldsymbol{\Lambda}$ which is related to $\boldsymbol{\beta}$ by an orthogonal transformation,
provided that uniqueness of the variance decomposition holds.
The proof of this
result uses a relationship between a matrix with GLT structure and the so-called
reduced row echelon form in linear algebra that results from the Gauss-Jordan elimination
for solving linear systems, see e.g.\ \citet{ant-ror:ele}.
Any
transposed GLT
loading matrix $\boldsymbol{\Lambda}^{\top}$ has a row echelon form which can be turned into
a reduced row echelon form (RREF) ${\mathbf B} = {\mathbf A}^{\top} \boldsymbol{\Lambda}^{\top} $
with the help of an
$r\times r$ matrix ${\mathbf A}$ which is constructed from the pivot rows
$l_1,\ldots, l_r$ of $\boldsymbol{\Lambda}$ and
invertible by definition:
\begin{eqnarray*}
{\mathbf A} ^ {-1} = \left(
\begin{array}{c}
\boldsymbol{\Lambda}_{l_1,\cdot} \\
\vdots \\
\boldsymbol{\Lambda}_{l_r,\cdot}
\end{array}
\right).
\end{eqnarray*}
Since the RREF of any matrix is unique, see e.g.\ \citet{yus:red}, we find that the pivot columns of
${\mathbf B} $ coincide with the pivot rows $l_1,\ldots, l_r$ of $\boldsymbol{\Lambda}$.
Hence, for a basic factor model
\begin{eqnarray*}
{\mathbf f}_t \sim \mathcal{N} _{r}\left({\mathbf{0}},{{\mathbf I}}_{r}\right), \qquad {\mathbf y}_t = \boldsymbol{\beta} {\mathbf f}_t + \boldsymbol{\epsilon}_t,
\end{eqnarray*}
with an arbitrary, unstructured loading matrix $\boldsymbol{\beta}$ with full column rank $r$,
we prove in Theorem~\ref{therref} that the RREF
of $\boldsymbol{\beta}^{\top}$ can be used to represent
$\boldsymbol{\beta}$ as a unique GLT structure $\boldsymbol{\Lambda}$, where
the pivot rows $l_1,\ldots, l_r$ of $\boldsymbol{\Lambda}$
coincide with the pivot columns of the RREF of $\boldsymbol{\beta}^{\top}$
(see Appendix~\ref{app:proof} for a proof).
\begin{thm}[Rotation into GLT]\label{therref}
Let $\boldsymbol{\beta}$ be an arbitrary loading matrix with full column rank $r$. Then the following
holds:
\begin{itemize}
\item[(a)] There exists an equivalent representation of $\boldsymbol{\beta}$ involving a
unique GLT structure $\boldsymbol{\Lambda}$,
\begin{eqnarray} \label{vardecRREF}
\boldsymbol{\beta} = \boldsymbol{\Lambda} {\mathbf G}^{\top},
\end{eqnarray}
where ${\mathbf G}$ is a unique orthogonal matrix.
$\boldsymbol{\Lambda}$ is called
the {\em GLT representation} of
$\boldsymbol{\beta}$.
\item[(b)]
Let $l_1 < \ldots < l_r$ be the pivot columns of the RREF ${\mathbf B} $
of $\boldsymbol{\beta}^{\top}$ and let
$\boldsymbol{\beta}_1$ be the
$r \times r$ submatrix
of $\boldsymbol{\beta}$ containing the
corresponding rows $l_1, \ldots, l_r$.
The GLT representation $\boldsymbol{\Lambda} = \boldsymbol{\beta} {\mathbf G}$
of $\boldsymbol{\beta}$ has pivot rows $l_1 , \ldots , l_r$ and is obtained
through rotation into GLT
with a rotation matrix
\begin{eqnarray} \label{vardectilde}
{\mathbf G} = {\mathbf Q} ,
\end{eqnarray}
which results from the QR decomposition ${\mathbf Q} \mathbf{R} =\boldsymbol{\beta}_1^{\top}$ of $\boldsymbol{\beta}_1^{\top}$.
\end{itemize}
\end{thm}
Would it be possible to obtain a similar
results
with the factor loading matrix $\boldsymbol{\Lambda}$ being constrained to be a PLT structure?
The answer is definitely no,
as has already been established in Section~\ref{secmotivate} for example (\ref{example1}).
As mentioned above, GLT
structures encompass PLT structures as a special case. Hence,
if a PLT representation $\boldsymbol{\Lambda}$
exists for a loading matrix $\boldsymbol{\beta} =\boldsymbol{\Lambda} {\mathbf P} $, then the GLT representation
in (\ref{vardectilde}) {\em automatically} reduces to
the PLT structure $\boldsymbol{\Lambda}$, since
$\mathbf{R} = \boldsymbol{\beta}_1^{\top}$ is obtained from the first $r$ rows of $\boldsymbol{\beta}$ and
the \lq\lq rotation into GLT\rq\rq\
is equal to the identity, ${\mathbf Q}={{\mathbf I}}_{r}$.
On the other hand, if the GLT representation $\boldsymbol{\Lambda}$
differs from a PLT structure, then no equivalent
PLT representation exists. Hence, forcing a PLT structure
in the representation (\ref{fac1}) may introduce a systematic bias in estimating the marginal covariance matrix ${\mathbf \Omega}$.
\section{Variance identification and GLT structures} \label{varidesp}
As mentioned in the previous sections, constraints imposed on the structure of a factor loading matrix $\boldsymbol{\Lambda}$ will resolve rotational invariance only if uniqueness of the variance
decomposition holds and the cross-covariance matrix $\boldsymbol{\Lambda} \boldsymbol{\Lambda} ^{\top}$
is identified. However, rotational constraints alone do not necessarily guarantee
uniqueness of the variance decomposition.
Consider, for instance, a sparse PLT loading matrix where in some column $j$ in addition
to the diagonal element $\Lambda_{jj}$ (which is nonzero by definition) only
a single further factor loading $\Lambda_{n_j,j}$ in some row $n_j>j$
is nonzero.
Such a loading matrix obviously violates the necessary condition for variance identification that
each column contains at least three nonzero elements.
Similarly, while GLT structures resolve rotational invariance, they do not
guarantee uniqueness of the variance decomposition either.
In Section~\ref{sec:3579}, we derive sufficient conditions for variance identification
of GLT structures based on the 3579\ counting rule
of \citet[Theorem~3.3]{sat:stu}.
In Section~\ref{varbeyGLT}, we discuss how to verify variance
identification for sparse GLT structures in practice.
\subsection{Counting rules for variance identification} \label{sec:3579}
We will show how to verify from the 0/1 pattern $\boldsymbol{\delta}$ of an unordered GLT structure
$\boldsymbol{\beta} $,
whether the row deletion property \mbox{\bf AR}\ holds for
$\boldsymbol{\beta} $ and all its signed permutations. Our condition is
a structural counting rule expressed
solely in terms of
the sparsity matrix $\boldsymbol{\delta}$ underlying $\boldsymbol{\beta} $
and does not involve the values of the unconstrained
factor loadings in $\boldsymbol{\beta}$, which can take any value in $\mathbb{R}$.
For any factor model, variance identification is invariant to signed permutations. If
we can verify
variance identification for a single signed permutation
$ \boldsymbol{\beta} = \boldsymbol{\Lambda} {\mathbf P} _{\pm} {\mathbf P} _{\rho} $
of $\boldsymbol{\Lambda}$, as defined in (\ref{eq:Ralpha}), then variance identification
of $\boldsymbol{\Lambda}$ holds, since $ \boldsymbol{\beta}$ and $\boldsymbol{\Lambda} $ imply the same
cross-covariance matrix $\boldsymbol{\Lambda} \boldsymbol{\Lambda} ^{\top}$.
Hence, we focus in this section on variance identification of unordered GLT structures.
\noindent In Definition~\ref{ARdefext}, we recall the
so-called {\em extended row deletion property}, introduced by
\citet{tum-sat:ide}.
\begin{defn}[Extended row deletion property $\mbox{RD} (r,s)$] \label{ARdefext}
A $m \times r$ factor loading matrix $\boldsymbol{\beta}$ satisfies the row-deletion property
$\mbox{RD} (r,s)$, if the following condition is satisfied: whenever
$s \in \mathbb{N}_0$
rows are deleted from $\boldsymbol{\beta}$,
then two disjoint submatrices of rank $r$ remain.
\end{defn}
\noindent
The row-deletion property of \citet{and-rub:sta}
results as a special case where $s={1}$.
As will be shown in Section~\ref{secEFA},
the extended row deletion properties $\mbox{RD} (r,s)$ for $s >{1}$ are
useful in exploratory factor analysis,
when the factor dimension $r$ is unknown.
In Definition~\ref{CRdef}, we introduce a counting rule for binary matrices.
\begin{defn}[Counting rule $\mbox{CR} (r,s)$] \label{CRdef}
Let $\boldsymbol{\delta}$ be an $m \times r$ binary matrix.
For each $q =1,\ldots,r$, consider all submatrices $\boldsymbol{\delta}_{q,\ell}$, $\ell=1, \ldots, {\small \left(\begin{array}{c} r \\[-3mm] q \end{array}\right)}$,
built from $q$ columns of $\boldsymbol{\delta}$.
$\boldsymbol{\delta}$ is said to satisfy the
$\mbox{CR} (r,s)$ counting rule for $s \in \mathbb{N}_0$
if the matrix
$\boldsymbol{\delta}_{q,\ell}$ has at least $2\ell+s $ nonzero rows for all $(q,\ell)$.
\end{defn}
\noindent
Note that the
counting rule $\mbox{CR} (r,s)$, like the extended row deletion property
$\mbox{RD} (r,s)$, is invariant to
signed permutations. Lemma~\ref{lemCRrules} in Appendix~\ref{app:proof} summarizes
further useful properties of $\mbox{CR} (r,s)$.
For a given binary matrix $\boldsymbol{\delta}$ of dimension $m \times r$,
let $\Theta_{\delta}$ be the space generated by
the non-zero elements of
all unordered GLT structure
$\boldsymbol{\beta}$ with sparsity matrix $\boldsymbol{\delta}$
and all their $2^r r ! -1 $ trivial rotations
$\boldsymbol{\beta} {\mathbf P} _{\pm} {\mathbf P} _{\rho}$.
We prove in Theorem~\ref{Crule357} that for GLT structures the counting rule
$\mbox{CR} (r,s)$ and the extended row deletion property $\mbox{RD} (r,s)$
are equivalent conditions
for all loading matrices in $\Theta_{\delta}$, except for a set of measure 0.
\begin{thm}\label{Crule357}
Let $\boldsymbol{\delta}$ be a binary $m \times r$ matrix with unordered GLT structure.
Then the following holds:
\begin{enumerate}
\item[(a)] If $\boldsymbol{\delta}$ violates the counting rule $ \mbox{CR} (r,s)$,
then the extended row deletion property $\mbox{RD} (r,s)$ is violated
for all $\boldsymbol{\beta} \in \Theta_{\delta}$ generated by $\boldsymbol{\delta}$.
\item[(b)] If $\boldsymbol{\delta}$ satisfies the counting rule $ \mbox{CR} (r,s)$,
then the extended row deletion property $\mbox{RD} (r,s)$ holds
for all $\boldsymbol{\beta} \in \Theta_{\delta}$
except for a set of measure 0.
\end{enumerate}
\end{thm}
\noindent See Appendix~\ref{app:proof} for a proof. The special case
$s=1$ is relevant for verifying the row deletion property \mbox{\bf AR} .
It proves that for unordered GLT structures
the 3579\ counting rule of \citet{sat:stu} is not only a necessary,
but also a {\em sufficient} condition for \mbox{\bf AR}\ to hold.
In addition, this means
that the counting rule needs to be verified
only for the sparsity matrix $\boldsymbol{\delta}$ of a
{\em single trivial rotation} $ \boldsymbol{\beta} =\boldsymbol{\Lambda} {\mathbf P} _{\pm} {\mathbf P} _{\rho}$
rather than for every nonsingular matrix ${\mathbf G}$.
This result is summarized in Corollary~\ref{rule357}.
\begin{cor}[Variance identification rule for GLT structures]\label{rule357}
For any unordered $m \times r$ GLT structure $\boldsymbol{\beta} $,
the following holds:
\begin{enumerate}
\item[(a)] If $\boldsymbol{\delta}$ satisfies the
3579\ counting rule, i.e.\ every column of $\boldsymbol{\delta}$ has at least 3 non-zero elements,
every pair of columns at least 5 and, more generally,
every possible combination of $q=3, \ldots, r $ columns has at least $2q+1$ non-zero elements,
then variance identification is given for all $\boldsymbol{\beta} \in \Theta_{\delta}$
except for a set of measure 0;
i.e.\ for any other factor decomposition of the marginal covariance matrix ${\mathbf \Omega} =
\boldsymbol{\beta} \boldsymbol{\beta}^{\top} + {\mathbf \Sigma}
= \tilde{\boldsymbol{\beta}} \tilde{\boldsymbol{\beta}}^{\top} + \tilde{{\mathbf \Sigma}}$,
where $\tilde{\boldsymbol{\beta}}$ is an unordered GLT matrix,
it follows that $\tilde{{\mathbf \Sigma}}= {\mathbf \Sigma}$, i.e.
$\tilde{\boldsymbol{\beta}} \tilde{\boldsymbol{\beta}}^{\top}=\boldsymbol{\beta} \boldsymbol{\beta}^{\top}$,
and $\tilde{\boldsymbol{\beta}} = \boldsymbol{\beta} {\mathbf P} _{\pm} {\mathbf P} _{\rho}$.
\item[(b)] If $\boldsymbol{\delta}$ violates the
3579\ counting rule, then for all $\boldsymbol{\beta} \in \Theta_{\delta}$
the row deletion property \mbox{\bf AR}\ does not hold.
\item[(c)] For $r=1$, $r=2$, and $r=3$,
condition $\mbox{CR} (r,1)$ is both sufficient and necessary for variance identification.
\end{enumerate}
\end{cor}
\noindent A few comments are in order.
If $\boldsymbol{\delta}$ satisfies $\mbox{CR} (r,1)$,
then $\mbox{\bf AR}$ holds for all $\boldsymbol{\beta} \in \Theta_{\delta}$
and a sufficient condition for variance identification is satisfied.
As shown by \citet{and-rub:sta},
$\mbox{\bf AR}$ is a {\em necessary} condition for variance identification only for
$r=1$ and $r=2$. \citet[Theorem~3]{tum-sat:ide} show the same for
$r=3$, provided that $m\geq 7$.
It follows that $\mbox{CR} (r,1)$ is a {\em necessary and sufficient} condition for variance identification
for the models summarized in (c). In all other cases, variance identification may hold
for loading matrices $\boldsymbol{\beta} \in \Theta_{\delta}$,
even if $\boldsymbol{\delta}$ violates $\mbox{CR} (r,1)$.
The definition of unordered GLT structures given in Section~\ref{uniqload} imposes no constraint on the pivot
rows $l_1, \ldots , l_{r}$ beyond the assumption that they are distinct. This flexibility
can lead to GLT structures that can never satisfy the 3579\ rule, even if all elements below
the pivot rows are non-zero. Consider, for instance, a GLT matrix with
the pivot row in column $r$ being equal to $l_r= m -1$.
The loading matrix has at most two nonzero elements in column $r$ and violates the necessary
condition for variance identification.
This example shows that there is an upper bound for the pivot elements beyond which the 3579\ rule
can never hold. This insight is formalized in Definition~\ref{GLTAR}.
\begin{defn}\label{GLTAR} An unordered GLT structure $\boldsymbol{\beta}$
fulfills condition \mbox{\bf GLT-AR}\ if the following constraint on the
pivot rows $l_1, \ldots , l_{r}$ of $\boldsymbol{\beta}$ is satisfied,
where $z_j$ is the rank of $l_j$ in the ordered sequence $ l_{(1)} <
\ldots < l_{(r)}$:
\begin{eqnarray} \label{condlj}
l_j \leq m - 2(r - z_j +1).
\end{eqnarray}
\end{defn}
\noindent Evidently, an ordered GLT structure $\boldsymbol{\Lambda}$
fulfills condition \mbox{\bf GLT-AR}\ if the
pivot rows $l_1, \ldots , l_{r}$ of $\boldsymbol{\Lambda}$ satisfy the constraint
$ l_j \leq m - 2(r - j+1)$.
For the special case of a PLT structure where $l_j=j$, this constraint reduces to
$ m \geq 2r +1 $ which is equivalent to
a well-known upper bound for the number of factors.
For dense unordered GLT structures with $m$ (non-zero) rows,
condition \mbox{\bf GLT-AR}\ is a sufficient condition for \mbox{\bf AR} .
For sparse GLT structures
\mbox{\bf GLT-AR}\ is only a necessary condition for \mbox{\bf AR}\ and the 3579\ rule
has to be verified explicitly, as shown by the example discussed above.
Very conveniently for verifying variance identification in sparse factor analysis based
on GLT structures, Theorem~\ref{Crule357} and Corollary~\ref{rule357} operate
solely on the sparsity matrix $\boldsymbol{\delta}$ corresponding to $\boldsymbol{\beta}$.
\subsection{Variance identification in practice} \label{varbeyGLT}
To verify $\mbox{CR} (r,s)$ in practice, all submatrices of $q$ columns have to be
extracted from
the sparsity matrix $\boldsymbol{\delta}$ to verify if at least $2q+1$ rows of this
submatrix are non-zero.
For $q=1,2, r -1, r$,
this condition is easily verified from simple functionals of $\boldsymbol{\delta}$,
see Corollary~\ref{Lemma1} which follows immediately from Theorem~\ref{Crule357}
(see Appendix~\ref{app:proof} for details).
\begin{cor}[{\bf Simple counting rules for $\mbox{CR} (r,s)$}]\label{Lemma1}
Let $\boldsymbol{\delta}$ be a $m \times r $ unordered GLT sparsity matrix.
The following conditions on $\boldsymbol{\delta}$ are necessary for $\mbox{CR} (r,s)$ to hold:
\begin{eqnarray}
&{\mathbf{1}}_{r \times m} \cdot \boldsymbol{\delta} + \boldsymbol{\delta}^{\top} ( {\mathbf{1}}_{ m \times r}
- \boldsymbol{\delta} ) \geq 4 + s - 2 {{\mathbf I}}_{r} , & \label{checkNC2} \\
&{\mathbf{1}}_{1 \times m} \cdot \mathbb{I} ( \boldsymbol{\delta}^\star >0) \geq 2r + s, \quad \boldsymbol{\delta}^\star= \boldsymbol{\delta} \cdot {\mathbf{1}}_{r \times 1}, &
\label{checkNC1} \\
& {\mathbf{1}}_{1 \times m} \cdot \mathbb{I} (\boldsymbol{\delta}^\star >0) \geq 2(r -1) + s , \quad \boldsymbol{\delta}^\star= \boldsymbol{\delta} ( {\mathbf{1}}_{ m \times m} - {{\mathbf I}}_{m}),
& \label{checkNC3}
\end{eqnarray}
where the indicator function $ \mathbb{I} (\boldsymbol{\delta}^\star >0)$ is applied element-wise
and ${\mathbf{1}}_{ n \times k}$ denotes a $ n \times k$ matrix of ones.
For $r \leq 4$, these conditions are also sufficient for
$\mbox{CR} (r,s)$ to hold for $\boldsymbol{\delta}$.
\end{cor}
\noindent Using Corollary~\ref{Lemma1} for $s=1$,
one can efficiently verify, if the 3579\
counting rule and hence the row deletion property \mbox{\bf AR}\ holds for unordered GLT factor models
with up to $r \leq 4$ factors.
For models with more than four factors ($r > 4$), a more elaborated strategy is needed.
After checking the conditions of Corollary~\ref{Lemma1},
$\mbox{CR} (r,s)$
could be verified for a given binary matrix $\boldsymbol{\delta}$ by
iterating over all remaining $\footnotesize{ r ! /(q ! (r -q ) !) }$
subsets of $q=3, \ldots, r-2$ columns of $\boldsymbol{\delta}$.
While this
is a finite task, such a na\"ive approach may need to visit $2^r-1$ matrices in order to make a decision and the combinatorial explosion quickly becomes an issue in practice as $r$ increases.
Recent work by \citet{hos-fru:cov} establishes the applicability of this framework for large models.
\section{Identification in exploratory factor analysis} \label{secEFA}
In this section, we discuss how the concept of GLT structures is helpful
for addressing identification problems in exploratory factor analysis (EFA).
Consider data $\{{\mathbf y}_1, \ldots,{\mathbf y}_T\}$
from a multivariate Gaussian distribution,
${\mathbf y}_t \sim \mathcal{N} _{m}\left({\mathbf{0}},{\mathbf \Omega}\right)$, where an investigator wants
to perform factor analysis
since she expects that the covariances of the
measurements $y_{it}$ are driven by common factors.
In practice, the number of factors
is typically unknown
and often it is not obvious,
whether all $m$ measurements in ${\mathbf y}_t$ are actually correlated.
It is then common to employ EFA by fitting
a basic factor model to the entire
collection of measurements in ${\mathbf y}_t$, i.e.\ assuming the model
\begin{eqnarray} \label{fac1reg}
{\mathbf y}_t = \boldsymbol{\beta}_k {\mathbf f}_t + \boldsymbol{\epsilon}_t, \qquad \boldsymbol{\epsilon}_t \sim \mathcal{N} _{m}\left({\mathbf{0}},{\mathbf \Sigma}_k\right) ,
\end{eqnarray}
with an \emph{assumed} number of factors $k$,
a $m\times k$ loading matrix $\boldsymbol{\beta}_k$
with elements $\beta_{ij}$
and a diagonal matrix ${\mathbf \Sigma}_k$ with strictly positive entries.
The EFA model (\ref{fac1reg}) is potentially overfitting
in two ways.
First, the true number of factors $r$ is
possibly smaller than $k$, i.e.\ $\boldsymbol{\beta}_k$ has too many columns.
Second, some measurements in ${\mathbf y}_t$ are possibly irrelevant, which means that $\boldsymbol{\beta}_k$
allows for too many non-zero rows.
The goal is then to determine the true number of factors and to identify
irrelevant measurements from the EFA model (\ref{fac1reg}).
We will address identification under the assumption that the data
are generated by a basic factor model with loading matrix $\boldsymbol{\beta}_0$
with $r$ factors which implies the following
covariance matrix ${\mathbf \Omega}$:
\begin{eqnarray}
{\mathbf \Omega}= \boldsymbol{\beta}_0 \boldsymbol{\beta}_0^{\top} + {\mathbf \Sigma}_0. \label{fac4over}
\end{eqnarray}
Instead of (\ref{fac4over}), for a given $k$, the EFA model (\ref{fac1reg}) yields the alternative
representation of ${\mathbf \Omega}$:
\begin{eqnarray}
{\mathbf \Omega}= \boldsymbol{\beta}_k \boldsymbol{\beta}_k^{\top} + {\mathbf \Sigma}_k . \label{fac4beta}
\end{eqnarray}
The question is then under which conditions
can the true loading matrix $\boldsymbol{\beta}_0$
be recovered from (\ref{fac4beta}).
Let us assume for the moment that
no constraint that resolves rotational invariance is imposed on
$\boldsymbol{\beta}_0$ or $\boldsymbol{\beta}_k$.
\paragraph*{\lq\lq Revealing the truth\rq\rq\ in an overfitting EFA model.}
A fundamental problem in factor analysis
is the following. If the EFA model is overfitting, i.e.\ $k > r$,
could we nevertheless recover
the true loading matrix $\boldsymbol{\beta}_0$ directly from
$\boldsymbol{\beta}_k$? We will show how this can be achieved mathematically by combining
the important work by \citet{tum-sat:ide} with the framework of GLT structures.
We have demonstrated in Section~\ref{secmotivate} using example (\ref{example4})
that solutions in an overfitting model can be constructed by adding spurious columns
\citep{rei:ide,gew-sin:int}.
Additional solutions are obtained as rotations
of such solutions. For instance,
one of the following solutions may result:
\begin{eqnarray*}
\tilde{\boldsymbol{\beta}}_3 = \left(\begin{array}{ccc}
0 & \lambda_{11} & 0 \\
\beta_{23} & \lambda _{21} & 0 \\
0 & \lambda_{31} & 0 \\
0 & 0 & \lambda_{42} \\
0 & 0 & \lambda_{52} \\
0 & 0 & \lambda_{62}
\end{array}\right), \quad
\tilde{\boldsymbol{\beta}}_3 = \left(\begin{array}{ccc}
-\lambda_{11} \sin \alpha & 0 & \lambda_{11} \cos \alpha \\
\beta_{23} \cos \alpha -\lambda_{21} \sin \alpha & 0 & \lambda _{21} \cos \alpha \\
-\lambda_{31} \sin \alpha & 0 & \lambda_{31} \cos \alpha\\
0 & \lambda_{42} & 0 \\
0 & \lambda_{52} & 0 \\
0 & \lambda_{62} & 0\\
\end{array}\right),
\end{eqnarray*}
both with the same ${\mathbf \Sigma}_3$ as in (\ref{example4}). The first case is a signed permutation
of $\boldsymbol{\beta}_{3}$, while the second case
combines a signed permutation of $\boldsymbol{\beta}_{3}$ with a
rotation
of the spurious and $\boldsymbol{\Lambda} $'s
first column involving ${\mathbf P}_{\alpha b}$.
In the first case, despite the rotation,
both the spurious column and the columns of $\boldsymbol{\Lambda} $ are clearly visible,
while in the second case the presence of a spurious column
is by no means obvious and the columns of $\boldsymbol{\Lambda} $ are disguised.
In general, for an EFA model
that is overfitting by a single column, i.e.\ $k = r+1$,
and $\boldsymbol{\beta}_{k}$ is left unconstrained,
infinitely many representations $(\boldsymbol{\beta}_k,{\mathbf \Sigma}_k)$ with covariance matrix
${\mathbf \Omega}= \boldsymbol{\beta} _k \boldsymbol{\beta}_k^{\top} + {\mathbf \Sigma}_k$ can be constructed in the following way.
Let the first $r$ columns of $\boldsymbol{\beta}_k$ be equal to $\boldsymbol{\beta}_0$ and append an extra column to its right.
In this extra column, which will be called a spurious column, add a single non-zero loading
$\beta_{l_k, k}$ in any row
$1 \leq l_k \leq m $ taking any value that satisfies
$0< \beta_{l_k, k}^2< \sigma^2_{l_k}$;
then reduce the idiosyncratic variance in row $l_k$
to $\sigma^2_{l_k} - \beta_{l_k, k}^2$;
and finally apply an arbitrary rotation ${\mathbf P}$:
\begin{eqnarray} \label{adsp}
\boldsymbol{\beta}_{k}= \left(\begin{array}{cc}
\boldsymbol{\beta}_0 & \left|\begin{array}{c}
0 \\
{ \beta_{l_k, k}} \\
0
\end{array} \right.
\end{array} \right) {\mathbf P} , \qquad
{\mathbf \Sigma}_{k} = \mbox{\rm Diag}\!\left(\sigma^2_1,\ldots, \sigma^2_{l_k} -
{ \beta_{l_k, k}^2}, \ldots, \sigma^2_m\right).
\end{eqnarray}
Interesting questions are then the following: under which conditions
is (\ref{adsp}) an exhaustive representation of all possible solutions
$ \boldsymbol{\beta}_{k}$ in an EFA model where the {\em degree of overfitting}
defined as $s=k-r$ is equal to one?
How can all solutions $ \boldsymbol{\beta}_{k}$ be represented if $s>1$?
Such identifiability problems in overfitting EFA models have been analyzed
in depth by \citet{tum-sat:ide}.
They show that a stronger condition
than $\mbox{RD} (r,1)$ is needed for $\boldsymbol{\beta}_0$ in the
underlying variance decomposition (\ref{fac4over}) to
ensure that only spurious and no additional common factors are added in
the overfitting representation (\ref{fac4beta}).
In addition, \citet{tum-sat:ide} provide a
general representation of the factor loading matrix $\boldsymbol{\beta}_k $ in overfitting
representation
(\ref{fac4beta}) with $k > r$.
\begin{thm}\emph{\citep[Theorem~1]{tum-sat:ide}}
Suppose that ${\mathbf \Omega}$ has a decomposition as in (\ref{fac4over})
with $r$ factors and that
for some $S \in \mathbb{N}$ with
$m \geq 2r + S + 1 $
the extended row deletion property $\mbox{RD} (r,1+S)$ holds for $\boldsymbol{\beta}_0$.
If $ {\mathbf \Omega} $ has another decomposition such that
${\mathbf \Omega}= \boldsymbol{\beta}_k \boldsymbol{\beta}_k^{\top} + {\mathbf \Sigma}_k$ where
$\boldsymbol{\beta}_k$ is a $m\times (r+s)$-matrix of rank
$k=r+s$ with $ 1 \leq s \leq S$, then there exists an orthogonal matrix
${\mathbf T}_k$ of rank $k$ such that
\begin{eqnarray} \label{decover}
\boldsymbol{\beta}_k {\mathbf T}_k = \left(\begin{array}{cc}
\boldsymbol{\beta}_0 & {\mathbf M}_s
\end{array} \right) , \qquad {\mathbf \Sigma}_k = {\mathbf \Sigma}_0 - {\mathbf M}_s {\mathbf M}_s^{\top},
\end{eqnarray}
where the off-diagonal elements of ${\mathbf M}_s {\mathbf M}_s^{\top}$ are zero.
\end{thm}
\noindent The $m\times s$-matrix ${\mathbf M}_s $
is a so-called \emph{spurious factor loading matrix} that does not
contribute to explaining the covariance in ${\mathbf y}_t$, since
\begin{eqnarray*}
\boldsymbol{\beta}_k \boldsymbol{\beta}_k ^{\top} + {\mathbf \Sigma}_k = \boldsymbol{\beta}_k {\mathbf T}_k {\mathbf T}_k^{\top}
\boldsymbol{\beta}_k^{\top} + {\mathbf \Sigma}_k =
\boldsymbol{\beta}_0 \boldsymbol{\beta}_0^{\top} + {\mathbf M}_s {\mathbf M}_s ^{\top} + ({\mathbf \Sigma}_0 - {\mathbf M}_s {\mathbf M}_s^{\top} ) =
\boldsymbol{\beta}_0 \boldsymbol{\beta}_0^{\top} + {\mathbf \Sigma}_0 ={\mathbf \Omega} . \label{fac4A}
\end{eqnarray*}
\noindent While this theorem is an important result,
without imposing further structure on the factor loading matrix
$\boldsymbol{\beta}_k$ in the EFA model it cannot be applied immediately
to \lq\lq recover the truth\rq\rq ,
as the separation of $\boldsymbol{\beta}_k$ into the true factor loading matrix $\boldsymbol{\beta}_0$ and
the spurious factor loading matrix ${\mathbf M}_s$ is possible only up to a rotation
${\mathbf T}_k$ of $\boldsymbol{\beta}_k$.
However,
the truth\rq\rq\ in an overfitting EFA model can be recovered, if \citet[Theorem~1]{tum-sat:ide}
is applied
within the class of unordered GLT structures introduced in this paper.
If we assume that $\boldsymbol{\Lambda}$ is a GLT structure
which satisfies the extended row deletion property $\mbox{RD} (r,1+S)$,
we prove in Theorem~\ref{theoverGLT} the following result.
If $\boldsymbol{\beta}_k$ in an overfitting EFA model is an unordered GLT structure,
then $\boldsymbol{\beta}_k$ has a representation,
where the rotation in (\ref{decover}) is a signed
permutation ${\mathbf T}_k = {\mathbf P} _{\pm} {\mathbf P} _{\rho}$.
Hence, spurious factors in $\boldsymbol{\beta}_k$
are easily spotted and $\boldsymbol{\Lambda}$ can be recovered immediately from
$\boldsymbol{\beta}_k$.
\begin{defn}[Unordered spurious GLT structure] \label{SpurM}
A $m \times s$ unordered GLT factor loading matrix ${\mathbf M}^\Lambda_s$
with pivots rows $\{ {n_1}, \ldots, {n_s} \}$ is an
unordered spurious GLT structure if
all columns are spurious columns with a single nonzero loading in the corresponding
pivot row.
\end{defn}
\begin{thm} \label{theoverGLT}
Let $\boldsymbol{\Lambda} $ be a $m \times r$ GLT factor loading matrix
with pivot rows $l_{1} < \ldots < l_{r}$ which
obeys the extended
row deletion property $\mbox{RD} (r,1+S)$ for some $S \in \mathbb{N}$.
Assume that the $m \times k$ matrix $\boldsymbol{\beta}_k$ in the EFA variance decomposition
${\mathbf \Omega}= \boldsymbol{\beta} _k \boldsymbol{\beta}_k^{\top} + {\mathbf \Sigma}_k$ is of rank $\mbox{\rm rk}\,(\boldsymbol{\beta} _k)=
k=r + s$, where
$1 \leq s \leq S $. If $\boldsymbol{\beta} _k$ is restricted
to be an unordered GLT matrix, then
(\ref{decover}) reduces to
\begin{eqnarray*}
\boldsymbol{\beta}_k {\mathbf P} _{\pm} {\mathbf P} _{\rho} = \left(\begin{array}{cc}
\boldsymbol{\Lambda} & {\mathbf M} ^{\Lambda}_s
\end{array} \right), \quad {\mathbf \Sigma}_k = {\mathbf \Sigma}_0 - {\mathbf M} ^{\Lambda}_s
({\mathbf M} ^{\Lambda}_s)^{\top},
\end{eqnarray*}
where ${\mathbf M} ^{\Lambda}_s$ is a spurious ordered GLT structure
with pivot rows ${n_1} < \ldots < {n_s}$
which are distinct from the $r$ pivot rows in $\boldsymbol{\Lambda}$.
Hence, $r$ columns of $\boldsymbol{\beta}_k$
are a signed permutation of the true loading matrix $\boldsymbol{\Lambda}$,
while the remaining $s$ columns of $\boldsymbol{\beta}_k$ are
an \textit{unordered spurious GLT structure} with pivots ${n_1}, \ldots , {n_s}$.
\end{thm}
\noindent See Appendix~\ref{app:proof} for a proof.
\paragraph*{Identifying irrelevant variables.} \label{irr}
In applied factor analysis, the assumption that each measurement $y_{it}$ is correlated with
at least one other measurement is too restrictive,
because irrelevant measurements might be present
that are uncorrelated with all the other measurements.
As argued by \citet{boi-ng:are}, it is useful to identify such variables.
Within the framework of sparse factor analysis, irrelevant variables are identified
in \citet{kau-sch:ide} by exploring the sparsity matrix $\boldsymbol{\delta}$ of a factor loading
matrix $\boldsymbol{\beta}_0$ with respect to zero rows.
Since $\mbox{\rm Cov}(y_{it}, y_{lt}) = 0$ for all $l \neq i$, if
the entire $i$th row
of $ \boldsymbol{\beta}_0$ is zero (see also (\ref{fac5})),
the presence of $m_0$ irrelevant measurements causes the corresponding $m_0$ rows of $ \boldsymbol{\beta}_0$
and $\boldsymbol{\delta}$ to be zero.
As before, we assume that the variance decomposition (\ref{fac4over})
of the underlying basic factor model is variance identified.
Let us first investigate identification of the zero rows in $ \boldsymbol{\beta}_0$ and
the corresponding sparsity matrix $\boldsymbol{\delta}$
for the case that the assumed and the true number of factors in
the EFA model (\ref{fac1reg}) are identical, i.e.\ $k = r$.
Since variance identification of (\ref{fac4over})
in the underlying model holds, we obtain that
${\mathbf \Sigma}_0={\mathbf \Sigma}_r$, $\boldsymbol{\beta}_0 \boldsymbol{\beta}_0^{\top}=\boldsymbol{\beta}_r
\boldsymbol{\beta}_r^{\top}$ and
$\boldsymbol{\beta}_r = \boldsymbol{\beta}_0 {\mathbf P} $ is a rotation of $\boldsymbol{\beta}_0$.
Therefore, the position of the zero rows both in $\boldsymbol{\beta}_0$ and $\boldsymbol{\beta}_r $
are identical and all irrelevant variables can be identified from $\boldsymbol{\beta}_r$
or the corresponding sparsity matrix $\boldsymbol{\delta}$,
regardless
of the strategy toward rotational invariance.
What makes this task challenging in applied factor analysis is that
in practice only
the total number $m$ of observations is known, whereas
the investigator is ignorant both about the
number of factors $r$ and the number of irrelevant measurements $m_0$.
In such a situation, variance identification of ${\mathbf \Sigma}_k$
for an EFA model with $k$ assumed factors is easily lost
if too many irrelevant variables are included in relation to $k$.
These considerations have important implication for exploratory factor analysis. While the investigator can
choose $ k$, she is ignorant about the number of irrelevant variables
and the recovered model might not be variance identified.
For this reason, it is relevant to verify in any case
that the solution $\boldsymbol{\beta}_k$ obtained
from any EFA model satisfies variance identification.
Under \mbox{\bf AR}\ this means that
the loading matrix of the correlated measurements, i.e.\
the non-zero rows of $ \boldsymbol{\beta}_0$, satisfies $\mbox{RD} (r,1)$.
If variance identification relies on \mbox{\bf AR} ,
then a minimum requirement for $\boldsymbol{\beta}_k$ to satisfy $\mbox{RD} (k,1)$ is that
$ 2 k + 1 \leq m-m_0 $.
If no irrelevant measurement are present, then the well-known upper bound
$ k \leq \frac{m-1}{2}$ results.
However, if irrelevant measurements are present, then
there is a trade-off between $m_0$ and $k$: the more irrelevant measurements
are included, the smaller the maximum number of assumed factors $k$ has to be.
Hence, the presence of $m_0$ zero rows in $ \boldsymbol{\beta}_0$, while $ \boldsymbol{\beta}_k$
in the EFA model is allowed to have $m$ potentially non-zero rows requires stronger
conditions for variance identification
than for an EFA model where the underlying loading matrix
$ \boldsymbol{\beta}_0$ contains only non-zero rows.
More specifically, for a given number $m_0 \in \mathbb{N}$ of irrelevant measurements,
variance identification
necessitates the more stringent upper bound
$ k \leq \frac{m-m_0-1}{2} $, where $m-m_0$ is the number of non-zero rows.
On the other hand, for a given number of factors
$k$ in an EFA model, the maximum number of irrelevant measurements
that can be included is given by $ m_0 \leq m -(2k + 1)$.
\paragraph*{Identifying the number of factors through an EFA model.}
Let us assume that the variance decomposition (\ref{fac4over})
of the unknown underlying basic factor model is identified.
As shown by \citet{rei:ide},
the true number of factors $r$ is equal
to the smallest value $k$ that satisfies (\ref{fac4beta}).
However, in practice, it is not
obvious how to solve this \lq\lq minimization\rq\rq\ problem.
As the following considerations show, verifying variance identification for
$\boldsymbol{\beta}_k$ in an EFA model
can be helpful in this regard.
If $r$ is unknown, then we need
to find a decomposition of ${\mathbf \Omega}$ as in
(\ref{fac4beta}) where $ {\mathbf \Sigma}_k $ is variance identified.
Since the true underlying decomposition (\ref{fac4over}) is variance identified,
any solution where $ {\mathbf \Sigma}_k $ is not variance identified can be rejected.
As has been discussed above, any overfitting EFA model,
where $k > r$, has infinitely many decompositions of ${\mathbf \Omega}$ and therefore is
never variance identified. Hence, if
any solution $ {\mathbf \Sigma}_k $ of an EFA model with $k$ assumed factors
is not variance identified, then we
can deduce that $k$ is bigger than $r$.
On the other hand, if variance identification holds for $ {\mathbf \Sigma}_k $,
then the decompositions (\ref{fac4over}) and (\ref{fac4beta}) are equivalent and
we can conclude that $r=k$, ${\mathbf \Sigma}_0= {\mathbf \Sigma}_k$ and
therefore $ \boldsymbol{\beta}_0 \boldsymbol{\beta}_0^{\top}=\boldsymbol{\beta}_k \boldsymbol{\beta}^{\top}_k$.
As a consequence, we can identify the true loading matrix
$ \boldsymbol{\beta}_0=\boldsymbol{\beta}_k {\mathbf P}$ from $\boldsymbol{\beta}_k$ mathematically up to a rotation
${\mathbf P}$ \citep[Lemma~5.1]{and-rub:sta}.
This insight shows that verifying variance identification
is relevant beyond resolving rotational invariance and is essential for recovering the
true number of factors. This has important implications for applied factor analysis.
Most importantly,
the {\em rank} or the {\em number of non-zero columns}
of a factor loading matrix $\boldsymbol{\beta}_k $ recovered
from an EFA model with {\em assumed} number $k$ of factors
might overfit the true number of factors $r$,
if
variance identification for ${\mathbf \Sigma}_k$ is not satisfied and the variance decomposition is not
unique.
Hence, extracting the number of factors from an EFA model
makes only sense in connection with ensuring that variance identification holds.
\section{Illustrative application} \label{secapll}
\subsection{Sparse Bayesian factor analysis} \label{secPBFA}
A common goal of Bayesian factor analysis is to identify the unknown factor dimension $r$
of a factor loading matrix from the overfitting factor model (\ref{fac1reg}) with potentially $k>r$ factors,
see, among many others, \cite{roc-geo:fas}, \cite{fru-lop:spa}, and \cite{ohn-kim:pos}.
Often, spike-and slab priors are employed,
where the elements $\beta_{ij}$ of the loading matrix $\boldsymbol{\beta}_k$
apriori are allowed to be exactly zero with positive
probability.
This is achieved through a prior on the
corresponding $m\times k$
sparsity matrix $\boldsymbol{\delta}_k$. In each column $j$, the indicators
$\delta_{i j}$ are active apriori with a column-specific probability $\tau_j$,
i.e.~$\mbox{\rm Pr} (\delta_{i j}=1|\tau_j)=\tau_j$ for $i=1, \ldots,m$, where
the slab probabilities $\tau_1, \ldots, \tau_k$ arise
from an exchangeable shrinkage prior:
\begin{eqnarray} \label{pri2P}
\tau_j| k \sim \mathcal{B}\left(\gamma \frac{\alpha}{k},\gamma\right), \quad j=1, \ldots,k.
\end{eqnarray}
If $\gamma$ is unknown, then
(\ref{pri2P}) is called a two-parameter-beta (2PB) prior.
If $\gamma=1$, then
(\ref{pri2P}) is called a one-parameter-beta (1PB) prior and takes the form:
\begin{eqnarray} \label{prialt}
\tau_j| k \sim \mathcal{B}\left(\frac{\alpha}{k},1\right), \quad j=1, \ldots,k.
\end{eqnarray}
Prior (\ref{prialt}) converges to the Indian buffet process prior \citep{teh-etal:sti}
for $k \rightarrow \infty$.
As recently shown by \citet{fru:gen}, prior (\ref{prialt})
has a representation as a cumulative shrinkage process
(CUSP) prior \citep{leg-etal:bay}.
This specification leads to a Dirac-spike-and-slab prior for the factor loadings,
\begin{eqnarray} \label{PriorLL}
& \beta_{i j} | \kappa, \sigma^2_i, \tau_{j}
\sim (1- \tau_{j}) \Delta_{0} + \tau_{j}
\mathcal{N}\left(0, \kappa \sigma^2_i \right),\\
& \nonumber \sigma^2_i \sim \mathcal{G}^{-1} \left(c^\sigma,b^\sigma\right), \quad \kappa \sim \mathcal{G}^{-1} \left(c^\kappa,b^\kappa\right), &
\end{eqnarray}
where the columns of the loading matrix are increasingly pulled toward 0 as the column index increases.
In (\ref{PriorLL}), a Gaussian slab distribution is assumed with a random global shrinkage parameter $\kappa$, although other slab distributions are possible, see e.g.\ \citet{zha-etal:bay_gro} and
\citet{fru-etal:spa}.
The hyperparameters $\alpha$ and $\gamma$ are instrumental in controlling prior sparsity. Choosing
$\alpha=k$ and $\gamma=1$ leads to a uniform distribution for $\tau_j$, with
the {\em smallest} slab probability $\tau_{(1)}= \min_{j=1,\ldots,k} \tau_j$ also being uniform, while
the largest slab probability $\tau_{(k)}= \max_{j=1,\ldots,k} \tau_j \sim \mathcal{B}\left(k,1\right)$,
see \cite{fru:gen}.
Such a prior is likely to overfit the number of factors, regardless of all other assumptions. A prior with $\alpha<k$ and $\gamma=1$ induces sparsity, since the {\em largest} slab probability $\tau_{(k)} \sim \mathcal{B}\left(\alpha,1\right) $, while the smallest slab probability
$\tau_{(1)} \sim \mathcal{B}\left(\alpha/k,1\right)$. To control the small probabilities,
which are important in identifying the true number of factors, $\alpha$ is assumed to be a random parameter and learnt from the data under
the prior $\alpha \sim \mathcal{G}\left(a^\alpha, b^\alpha\right)$.
$\gamma$ controls the prior information in (\ref{pri2P}). Priors with $\gamma >1$ and $\gamma < 1$, respectively,
decrease and increase the difference between $\tau_{(1)}$ and $\tau_{(k)}$. Typically, $\gamma$ is unknown and is estimated from the data using the prior
$\gamma \sim \mathcal{G}\left(a^\gamma, b^\gamma\right)$.
\paragraph*{MCMC estimation.} For a given choice of
hyperparameters, Markov chain Monte Carlo (MCMC) methods are
applied to sample from the posterior distribution
$p(\boldsymbol{\beta}_k, {\mathbf \Sigma}_k, \boldsymbol{\delta}_k|{\mathbf y})$, given $T$
multivariate observations ${\mathbf y}=({\mathbf y}_1, \ldots, {\mathbf y}_T)$,
see e.g.\ \citet{kau-sch:bay} among many others.
In \citet{fru-etal:spa}, such a sampler is developed for GLT factor models. To move between factor models of different factor dimension, \citet{fru-etal:spa} exploit Theorem~\ref{theoverGLT} to add and delete spurious columns through a reversible jump MCMC (RJMCMC) sampler.
For each posterior draw $\boldsymbol{\beta}_k$, the active columns $\boldsymbol{\beta}_r$ (i.e.\ all columns with at least 2 non-zero elements) and the corresponding sparsity matrix $\boldsymbol{\delta}_r$ are determined. If $\boldsymbol{\delta}_r$ satisfies the counting rule $\mbox{CR} (r,1)$, then $\boldsymbol{\beta}_r$ is a signed permutation of $\boldsymbol{\Lambda}$ with the corresponding covariance matrix
${\mathbf \Sigma}_r = {\mathbf \Sigma}_k + {\mathbf M} ^{\Lambda}_s
({\mathbf M} ^{\Lambda}_s)^{\top}$, where ${\mathbf M} ^{\Lambda}_s$ contains the spurious columns of
$\boldsymbol{\beta}_k$. These variance identified draws are kept for further inference and the number of columns of $\boldsymbol{\beta}_r$ is considered a posterior draw of the unknown factor dimension $r$.
This algorithm is easily extended to EFA models without any constraints.
\begin{table}[t!] \caption{
Sparse Bayesian factor analysis under GLT and unconstrained structures (EFA)
under
a 1PB prior ($\alpha \sim \mathcal{G}\left(6,2\right)$) and a 2PB prior ($\alpha \sim \mathcal{G}\left(6,2\right),\gamma \sim \mathcal{G}\left(6,6\right) $).
GLT and EFA-V
use only the variance identified draws ($M_V$ is the percentage of variance identified draws), EFA uses all posterior draws.
}\label{tab1}
\vspace*{2mm}
{
\begin{tabular}{lllcccc}
\hline
& & & $M_V$ & \multicolumn{1}{c}{$\hat{r}$} & \multicolumn{1}{c}{$p(\hat{r}=r_{\mbox{\tiny \rm true}}|{\mathbf y})$}
& \multicolumn{1}{c}{$\mbox{\rm MSE}_\Omega$} \\
Scenario & & Prior & Med(QR) & Med(QR) & Med(QR)
& Med(QR) \\
\hline Dedic & GLT & 1PB & 97.0 (91.5,98.3) & 5 (5,5) & 0.90 (0.94,0.99)
& 0.018 (0.014,0.030) \\
& & 2PB & 97.6 (87.7,98.9)& 5 (5,5) & 0.99 (0.83,1.00)
& 0.019 (0.016,0.027) \\
\cline{2-7}
& EFA & 1PB & - & 5 (5,6) & 0.66 (0.09,0.79)
& 0.020 (0.015,0.026) \\
& &2PB & - & 5 (5,6) & 0.69 (0.36,0.80)
& 0.019 (0.014,0.024)\\
\cline{2-7}
& EFA-V & 1PB & 80.3 (49.8,87.0) & 5 (5,6) & 0.81 (0.17,0.91)
& 0.020 (0.015,0.026) \\
& &2PB & 82.6 (63.4,87.9) & 5 (5,6) & 0.84 (0.53,0.92)
& 0.019 (0.014,0.024)\\
\hline
Block & GLT & 1PB & 96.5 (39.4,98.9) & 5 (5,5) & 0.99 (0.28,0.99)
& 0.12 (0.08,0.18) \\
& & 2PB & 98.7 (61.9,99.4) &5 (5,5) & 0.99 (0.54,1.00)
& 0.10 (0.08,0.14) \\
\cline{2-7}
& EFA & 1PB &-& 5 (4,5) & 0.78 (0.22,0.88)
& 0.14 (0.11,0.20) \\
& & 2PB &-& 5 (4,5) & 0.79 (0.08,0.89)
& 0.12 (0.08,0.24) \\
\cline{2-7}
& EFA-V & 1PB & 87.0 (55.0,91.5) & 5 (4,5) & 0.89 (0.09,0.96)
& 0.14 (0.11,0.20) \\
& & 2PB & 85.9 (28.3,90.4) & 5 (4,5) & 0.92 (0.03,0.97)
& 0.12 (0.08,0.24) \\
\hline Dense & GLT & 1PB & 95.7 (84.6,98.6) & 5 (5,5) & 0.98 (0.92,0.99)
& 0.67 (0.44,1.12) \\
& & 2PB & 99.4 (90.8,99.8) & 5 (5,5) & 0.99 (0.93,1.00)
& 0.68 (0.51,1.18) \\
\cline{2-7}
& EFA & 1PB &-& 5 (5,6) & 0.76 (0.43,0.85)
& 0.54 (0.39,0.76) \\
& & 2PB &-& 5 (5,5) & 0.80 (0.66,0.91)
& 0.59 (0.43,0.90) \\
\cline{2-7}
& EFA-V & 1PB & 84.4 (76.0,90.2) & 5 (5,6) & 0.89 (0.57,0.95)
& 0.54 (0.39,0.76) \\
& & 2PB & 89.7 (80.4,93.9) & 5 (5,5) & 0.93 (0.77,0.98)
& 0.59 (0.43,0.90)\\
\hline
\end{tabular}}
{\footnotesize
Med is the median and QR are
the 5\% and the 95\% quantile of the various statistics over the 21 simulated data sets.}
\end{table}
\subsection{An illustrative simulation study} \label{illapp}
For illustration, we perform a simulation study and consider three different data scenarios with
$m=30$ and $T=150$. In all three scenarios, $r_{\mbox{\tiny \rm true}}=5$ factors are assumed,
however, the zero/non-zero pattern is quite different. The first setting is a {\em dedicated} factor model,
where the first 6 variables load on factor 1, the next 6 variables load on factor 2, and so forth, and the final 6 variables
load on factor 5. A dedicated factor model has a GLT structure by definition. The second scenario is a {\em block} factor model,
where the first 15 observations load only on factor 1 and 2, while the remaining 15 observations only load on factor 3, 4 and 5 and the covariance matrix has a block-diagonal structure.
All loadings within a block are non-zero. The third scenario is a {\em dense} factor loading matrix
without any zero loadings and the corresponding GLT representation has a PLT structure.
For all three scenarios, non-zero factor loadings are drawn as
$\lambda_{ij}=(-1)^{b_{ij}} (1 + 0.1 \mathcal{N}\left(0,1\right))$, where the exponent $b_{ij}$ is a binary variable with $\mbox{\rm Pr} (b_{ij}=1)=0.2$.
In all three scenarios, ${\mathbf \Sigma}_0={\mathbf I}$.
21 data sets are sampled under these three scenarios from the Gaussian factor model
(\ref{fac1}).
\begin{table}[t!] \caption{
Bayesian factor analysis under GLT and unconstrained structures (EFA)
under
a uniform prior on $\tau_j$.
GLT and EFA-V
use only the variance identified draws ($M_V$ is the percentage of variance identified draws), EFA uses all posterior draws.
}\label{tab2}
\vspace*{2mm}
{
\begin{tabular}{llcccc}
\hline
& & $M_V$ & \multicolumn{1}{c}{$\hat{r}$} & \multicolumn{1}{c}{$p(\hat{r}=r_{\mbox{\tiny \rm true}}|{\mathbf y})$}
& \multicolumn{1}{c}{$\mbox{\rm MSE}_\Omega$} \\
Scenario & & Med(QR) & Med(QR) & Med(QR) & Med(QR)
\\
\hline
Dedic & GLT & 50.6 (32.5,62.2)
& 6 (5,7) & 0.38 (0.03,0.68) & 0.02
(0.01,0.03) \\
& EFA & - & 7 (6,8) & 0.06 (0,0.12) & 0.02 (0.02,0.03) \\
& EFA-V & 36.6 (24.4,44.8)
& 6 (5,7) & 0.17 (0,0.44) & 0.02 (0.02,0.03) \\
\hline
Block & GLT & 53.3 (29.3,71.3) & 5 (4,6)& 0.62 (0.18,0.85) & 0.11 (0.08,0.17) \\
& EFA &- &
6 (6,7) & 0.21 (0.00,0.35) & 0.13 (0.10,0.19) \\
& EFA-V & 43.3 (17.3,52.2) & 5 (5,7) & 0.47 (0.01,0.62) & 0.13 (0.11,0.19) \\
\hline
Dense & GLT & 62.4 (45.8,71.3)
& 5 (5,6) & 0.69 (0.05,0.84) &0.62 (0.44,1.31) \\
& EFA &-& 6 (6,7) & 0.12 (0.03,0.34) & 0.52 (0.42,0.74) \\ & EFA-V & 48.1 (30.1,56.7) & 5 (5,6) & 0.46 (0.10,0.63) & 0.52 (0.42,0.73) \\
\hline
\end{tabular}}
{\footnotesize
Med is the median and QR are
the 5\% and the 95\% quantile of the various statistics over the 21 simulated data sets.}
\end{table}
A sparse overfitting factor model is fitted to each simulated data set
with the maximum number of factors $k= 14$ being equal to the upper bound.
Regarding the structure, we compare a model where the non-zero columns of $\boldsymbol{\beta}_k$
are left unconstrained with a model where a GLT structure is imposed.
Inference is based on
the Bayesian approach described in Section~\ref{secPBFA}
with two different shrinkage priors on the sparsity matrix $\boldsymbol{\delta}_k$:
the 1PB prior (\ref{prialt}) with random hyperparameter $\alpha \sim \mathcal{G}\left(6,2\right)$ and
the 2PB prior (\ref{pri2P}) with random hyperparameters $\alpha \sim \mathcal{G}\left(6,2\right)$ and
$ \gamma \sim \mathcal{G}\left(6,6\right)$.
MCMC estimation is run for 3000 iterations after a burn-in of 2000 using the
RJMCMC algorithm of \citet{fru-etal:spa}.
For each
of the 21 simulated data sets, we evaluate all 12 combinations of data
scenarios, structural constraints (GLT versus unconstrained) and priors on the sparsity matrix (1PB versus 2PB)
through Monte Carlo estimates of following statistics:
to assess the performance in estimating
the true number $r_{\mbox{\tiny \rm true}}$ of factors,
we consider the mode $\hat{r}$ of the posterior distribution $p(r|{\mathbf y}) $
and the magnitude of the posterior ordinate $p(\hat{r}=r_{\mbox{\tiny \rm true}}|{\mathbf y})$. To assess the accuracy in
estimating the
covariance matrix ${\mathbf \Omega}$ of the data, we consider the mean squared error
(MSE) defined by
\begin{eqnarray*}
&& \mbox{\rm MSE}_\Omega=\sum_i\sum_{\ell \leq i}
\mbox{\rm E}(\left({\mathbf \Omega}_{r, i\ell}- {\mathbf \Omega}_{i\ell} \right)^2|{\mathbf y})/(m(m+1)/2),
\end{eqnarray*}
which accounts both for posterior variance and bias of
the estimated
covariance matrix ${\mathbf \Omega}_r= \boldsymbol{\beta}_r \boldsymbol{\beta}_r ^\top + {\mathbf \Sigma}_r$ in comparison to the true matrix.
Table~\ref{tab1} reports, for all 12 combinations the median, the 5\% and the 95\% quantile of these statistics across all simulated data sets.
For inference under GLT structures,
posterior draws which are not variance identified have been removed. The fraction of variance identified draws is also reported in the table and is in general pretty high.
As common for sparse Bayesian factor
analysis with unstructured loading matrices, the posterior draws are not screened for variance identification and inference is based on all draws.
Some interesting conclusions can be drawn from Table~\ref{tab1}. First of all, sparse Bayesian factor
analysis under the GLT constraint successfully recovers
the true number of factors in all three scenarios. For most of the simulated
data sets,
the posterior ordinate $p(\hat{r}=r_{\mbox{\tiny \rm true}}|{\mathbf y})$
is larger than 0.9.
Sparse Bayesian factor
analysis with unstructured loading matrices
is also quite successful in recovering
$r_{\mbox{\tiny \rm true}}$, but with less confidence. Both over- and underfitting can be observed and the posterior ordinate $p(\hat{r}=r_{\mbox{\tiny \rm true}}|{\mathbf y})$ is much smaller than under
a GLT structure.
For both structures, the 2PB prior yields higher posterior ordinates than the 1PB prior.
Recently, \citet{hos-fru:cov} proved that the counting
rule $\mbox{CR} (r,1)$ can also be applied to verify variance identification for unconstrained loading matrices. As is evident from Table~\ref{tab1},
the fraction of variance identified draws is however, much smaller than under GLT structures.
Nevertheless, inference w.r.t.\ to the number of factors can be improved also for an unconstrained EFA model by rejecting all draws that do not obey
the counting rule $\mbox{CR} (r,1)$.
It should be emphasized that the ability of Bayesian factor analysis to recover the number of factors from an overfitting model is closely tied to choosing a suitable shrinkage prior on the sparsity matrix $\boldsymbol{\delta}_k$.
For illustration, we also consider a uniform prior for $\tau_j$ and report the corresponding statistics in
Table~\ref{tab2}.
As expected from the considerations in Section~\ref{secPBFA}, considerable overfitting is observed for all simulated data sets, regardless of the chosen structure.
\section{Concluding remarks} \label{secconcluse}
We have given a full and comprehensive mathematical treatment to generalized lower triangular (GLT) structures, a new identification strategy that improves on the popular positive lower triangular (PLT) assumption for factor loadings matrices.
We have proven that GLT retains PLT's good properties: uniqueness and rotational invariance.
At the same time and unlike PLT, GLT exists for any factor loadings matrix; i.e.\ it is not a restrictive assumption.
Furthermore, we have shown that verifying variance identification under GLT structures is simple and is based purely on the zero-nonzero pattern of the factor loadings matrix.
Additionally, we have embedded the GLT model class into exploratory factor analysis with unknown factor dimension and discussed how easily spurious factors and irrelevant variables are recognized in that setup.
At the end, we demonstrated the power of the framework in a simulation study.
\bibliographystyle{chicago}
\bibliography{references}
\newpage