EconBase
← Back to paper

Panel Data Models with Time-Varying Latent Group Structures

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

126,833 characters

Panel Data Models with Time-Varying Latent Group Structures



\title{Panel Data Models with Time-Varying Latent Group Structures\thanks{
Phillips acknowledges support from a Lee Kong Chian Fellowship at SMU, the
Kelly Fund at the University of Auckland, and the NSF under Grant No. SES
18-50860. Su gratefully acknowledges the support from the National Natural
Science Foundation of China under Grant No. 72133002. Address correspondence
to: Liangjun Su, School of Economics and Management, Tsinghua University,
Beijing, 100084, China. E-mail: [email removed]. }}
\author{Yiren Wang$^{a}$, Peter C B Phillips$^{b}$ and Liangjun Su$^{c}$ \\
$^{a}$School of Economics, Singapore Management University, Singapore \\
$^{b}$Yale University, University of Auckland,
Singapore Management University\\
$^{c}$School of Economics and Management, Tsinghua University, China}
\maketitle

\begin{abstract}
This paper considers a linear panel model with interactive fixed effects and
unobserved individual and time heterogeneities that are captured by some
latent group structures and an unknown structural break, respectively. To enhance realism the model may have different numbers of groups and/or different group
memberships before and after the break. With the preliminary
nuclear-norm-regularized estimation followed by row- and column-wise linear
regressions, we estimate the break point based on the idea of binary
segmentation and the latent group structures together with the number of
groups before and after the break by sequential testing K-means algorithm
simultaneously. It is shown that the break point, the number of groups and the
group memberships can each be estimated correctly with probability approaching
one. Asymptotic distributions of the estimators of the slope coefficients
are established. Monte Carlo simulations demonstrate excellent finite sample
performance for the proposed estimation algorithm. An empirical application to real house price data across 377 Metropolitan Statistical Areas in the US from 1975 to 2014 suggests the presence both of structural breaks and of changes in group membership. \medskip

\noindent \textbf{Key words:} Interactive fixed effects, latent group
structure, structural break, nuclear norm regularization, sequential testing
K-means algorithm. \medskip

\noindent \textbf{JEL Classification:} C23, C33, C38, C51\medskip\
\end{abstract}

\ \ \ \ \pagebreak

\section{Introduction}

\label{sec:intro} Heterogeneous panel data models have been widely used in
empirical research in economics because they can capture a rich degree of
unobserved heterogeneity. But models with complete heterogeneity along
either the cross section or time dimension tend to possess too many
parameters to be identified, which results in slow convergence and
inefficient estimates. For this reason, researchers now frequently advocate
the use of panel data models with certain structures imposed along either
the cross section or time dimension. On the one hand, the recent burgeoning
of panels with latent group structures can be motivated from the observation
that different groups of individuals respond differently to exogenous
shocks. For instance, \cite{durlauf1995multiple}, \cite
{berthelemy1996economic}, and \cite{ben1998convergence} show economies in
different groups of income per capita and/or education level may converge to
different steady state equilibria. \cite{klapper2011impact}, \cite
{chu2012logistics}, and \cite{zhang2019threshold} show an exogenous shock
like policy implementation has different impacts on different individuals,
and \cite{long2012impact} argue that the influence of the 2008 financial
crisis on economic growth differs for emergent and developed
economies. On the other hand, the recent popularity of panels that evidence
structural change can be motivated by events such as financial crises,
the economic impact of technological progress, and more general economic transitions that occur during the time
periods covered by the data. See \cite{qian2016shrinkage} for a survey of
panel data models and research that consider estimation and inference concerning structural change.

In spite of the now large literature that studies separately individual heterogeneity or
time heterogeneity in the slope coefficients of panel models, few works consider both types of heterogeneity simultaneously. Exceptions include \cite{keane2020climate} and \cite{lu2023uniform} who consider linear
panel data models with two-dimensional unobserved heterogeneity in the slope
coefficients that are modelled via the usual additive structure, and \cite
{chernozhukov2019inference} and \cite{wang2022low-rank} who model the slope
coefficients via the use of low-rank matrices for conditional mean and
quantile regressions, respectively. In addition, \cite{okui2021heterogeneous}
and \cite{lumsdaine2023estimation} consider both individual heterogeneity
and time heterogeneity by modeling them as a grouped pattern and as structural
breaks, respectively. Specifically, \cite{okui2021heterogeneous} develop a
new panel data model with latent groups where the number of groups and the
group memberships do not change over time but the coefficients within each
group can change over time and they may have different break-dates; \cite
{lumsdaine2023estimation} consider the panels with a grouped pattern of
heterogeneity when the latent group membership structure and/or the values
of slope coefficients change at a break point. Both papers provide
algorithms to recover the latent group structure based on linear panel
models with or without individual fixed effects, but cannot allow for the
presence of more complicated fixed effects such as interactive fixed
effects (IFEs) to capture strong cross-sectional dependence in the data.

This paper proposes a linear panel data model with IFEs that enable the
slope coefficients to exhibit two-way heterogeneity. Following the lead of \cite
{okui2021heterogeneous} and \cite{lumsdaine2023estimation} and to encourage
the parameter parsimony, we use a latent group structure to capture
individual heterogeneity and an unknown structural break to capture time
heterogeneity. The latent group structure of the model accommodates different
group numbers and different group memberships before and after the
break. Given this complicated structure, the approach proposed is to estimate
the break point, the number of groups before and after the break, the group
membership before and after the break, and the group-specific parameters in
multiple steps. The key insight that permits this degree of complication is that
the slope coefficients of each of the $p$ regressors in the model are permitted
to vary across both cross section and time dimensions by means of a factor structure with a fixed
number of factors so that they may be conveniently stacked into a low-rank matrix.

In the first step, the low-rank nature of the slope matrices is explored and
initial estimates are obtained by nuclear norm regularization
(NNR), a machine learning technique popular in computer science that is increasingly used in econometrics. Such
initial matrix estimates are consistent in terms of the Frobenius norm but do
not have pointwise or uniform convergence for their elements. Despite
this, by applying singular value decomposition (SVD) to these estimates, we
can obtain estimates of the associated factors and the factor loadings that are
also consistent in terms of the Frobenius norm. In the second step, we use the
first-step initial estimates of the factors and factor loadings to run the
row- and column-wise linear regressions to update the estimates of the
factors and factor loadings which now possess pointwise and uniform
consistency and can be used for subsequent analyses. In the third step, we
estimate the break point by using the celebrated idea of binary segmentation,
as commonly used for break point estimation in the time series literature.
Once the break point is estimated, the full-sample is naturally split into
two subsamples. In the fourth step, we follow the lead of \cite
{lin2012estimation} and \cite{Jin2022Optimal} to focus on each subsample
before and after the estimated break point and propose a sequential testing
K-means algorithm to recover the latent group structure and obtain the
number of groups simultaneously. In the last step, we use the estimated
group structure to estimate the group-specific parameters. Asymptotic
analyses show that the break point, the number of groups and the group
memberships can be consistently estimated in Steps 3-4, so that the final
step estimates for the group-specific coefficients can enjoy the oracle
property. This means they have the same asymptotic distributions as the ones
obtained by knowing the break point and the latent group structures before
and after the break points.


The present paper relates to two branches of literature. First, it contributes to the panel data literature on one-way heterogeneity, especially with either latent group structures or structural breaks. With respect to latent group structures, there are several popular ways to recover the
latent groups. The first approach is the K-means algorithm. \cite
{lin2012estimation} apply the K-means algorithm to linear panel data models
with grouped slope coefficients and propose an information criterion and a
sequential testing approach to estimate the true number of groups. \cite
{sarafidis2015partially} analyze the unknown grouped slopes in the large $N$
and fixed $T$ framework, and \cite{zhang2019quantile} provide an iterative
algorithm based on K-means clustering for a panel quantile regression model.
\cite{bonhomme2015grouped} and \cite{ando2016panel} consider panels with
grouped fixed effects. The second approach is the Classifier-Lasso (C-Lasso)
that has become a popular clustering method since \cite{su2016identifying}.
This method is extended by \cite{lu2017determining}, \cite{su2018identifying}
, \cite{su2019sieve}, \cite{wang2019heterogeneous}, and \cite
{huang2020identifying} to various contexts. In addition, both the clustering
algorithm in regression via a data-driven segmentation (CARDS) approach and
binary segmentation are also considered in \cite{ke2015homogeneity}, \cite
{wang2018homogeneity}, \cite{ke2016structure} and \cite{wang2021identifying}
, among others. As for panel models with structural breaks, binary
segmentation has become a common approach to estimate the break point. See
\cite{bai2010common}, \cite{hsu2011change}, \cite{kim2011estimating}, \cite
{kim2014common} and \cite{baltagi2017estimation}, among others.
These papers focus on the case of a single break point in the model. In contrast,
\cite{qian2016shrinkage} and \cite{li2016panel} allow for multiple breaks in
linear panel models with either classical fixed effects or IFEs, and propose an adaptive grouped fused lasso (AGFL) approach to estimate the break points. Compared to the existing panel literature on one-way
heterogeneity, our model allows for two-way heterogeneity. In
particular, not only are different membership structures in different
time blocks permitted but also changes in the number of groups over time. As a result,
our model is more flexible than all existing models that allow only for
latent group structures or structural breaks, but not both.

Second, this paper contributes to the recent burgeoning literature that
models two-way heterogeneity in the slope coefficients of a panel model.
As mentioned above, there are two approaches to model two-way
heterogeneity in the slope coefficients. One approach models them in an additive
structure so that both individual and time effects enter the slope
coefficients additively, as in \cite{keane2020climate} and \cite
{lu2023uniform}. The other approach imposes certain low-rank structures on the
slope coefficient matrices in which case one models each slope coefficient
via the use of IFEs to capture strong cross sectional dependence
in the panel. In view of the low-rank structures, we can resort to NNR estimation
which has attracted increasing attention recently in panel data analyses.
NNR has been used in recent econometric research -- see \cite
{bai2019rank}, \cite{moon2018nuclear}, \cite{chernozhukov2019inference},
\cite{belloni2023high}, \cite{miao2023high}, \cite{feng2023regularized}, and
\cite{hong2023profile}, among others. But none of these papers imposes any
latent group structures on the slope coefficients. With latent group
structures and structural breaks imposed, \cite{okui2021heterogeneous} allows
the slope coefficients within each group to have common breaks and the break
points to vary across different groups, and they propose to estimate the
latent group structures, the structural breaks, and the group-specific
regression parameters by the grouped adaptive group fused lasso (GAGFL).
But neither the number of groups nor the group memberships is allowed
to change over time in \cite{okui2021heterogeneous}. In a companion paper,
\cite{lumsdaine2023estimation} allows the latent group membership structure
and/or the values of slope coefficients to change at a break point, and
proposes an estimation algorithm similar to the K-means of \cite
{bonhomme2015grouped}. Both \cite{okui2021heterogeneous} and \cite
{lumsdaine2023estimation} allow for at most one-way heterogeneity
(individual fixed effects) in the intercept and neither allows for IFEs to
capture strong cross section dependence. In contrast, this paper proposes an
algorithm to detect the unknown break point and to recover the group
structure based on linear panel model with IFEs, which involves a more
general model. In addition, \cite{lumsdaine2023estimation} first assume the
number of groups is known in the estimation algorithm and then estimate the
number of groups via an information criterion but they do not establish consistency
for such an estimate. Instead, we estimate the number of
groups and group membership simultaneously by the sequential testing K-means
algorithm and establish the consistency of the number of groups estimator.

The rest of the paper is organized as follows. We first introduce the linear
panel model with time-varying latent group structures in Section \ref
{sec:model} and provide the estimation algorithm in Section \ref
{sec:estimation}. The asymptotic properties are given in Section \ref
{sec:theory}. In Section \ref{sec:extension}, we propose an alternative
approach to detect the break point, provide the test statistics for the null
that the slope coefficients exhibit no structure change against the
alternative with one break point, and discuss the estimation for the model
with multiple breaks. In Sections \ref{sec:simul} and \ref{sec:emp}, we show
the finite sample performance of our method by Monte Carlo simulations and
an empirical application, respectively. Section \ref{sec:concl} concludes.
All proofs are provided in the online supplement.

\textit{Notation.} Let $\left\Vert \cdot \right\Vert _{\max },$ $\left\Vert
\cdot \right\Vert _{op},$ $\left\Vert \cdot \right\Vert $, and $\left\Vert
\cdot \right\Vert _{\ast }$ denote the (elementwise) maximum norm, operator
norm, Frobenius norm, and nuclear norm, respectively. Let $\odot $ denote
the element-wise Hadamard product. $\lfloor \cdot \rfloor $ and $\lceil
\cdot \rceil $ denote the floor and ceiling functions, respectively. Let $
a\vee b=\max \left( a,b\right) $ and $a\wedge b=\min (a,b)$. $a_{n}\lesssim
b_{n}$ means $a_{n}/b_{n}=O_{p}\left( 1\right) $ and $a_{n}\gg b_{n}$ means $
b_{n}a_{n}^{-1}=o(1)$. Let $A=\{A_{it}\}$ be a matrix with its $(i,t)$-th
entry denoted as $A_{it}$. Let denote $\{A_{j}\}_{j\in \lbrack p]\cup \{0\}}$
be the collection of matrices $A_{j},$ $j\in \{0,1,\cdots ,p\}$. For a
specific $A\in \mathbb{R}^{m\times n}$ with rank $n$, let $P_{A}=A(A^{\prime
}A)^{-1}A^{\prime }$ and $M_{A}=I_{m}-P_{A}$. When $A$ is symmetric, $
\lambda _{\max }(A)$, $\lambda _{\min }(A)$ and $\lambda _{n}(A)$ denote its
largest, smallest and $n$-th largest eigenvalues, respectively. The
operators $\rightsquigarrow $ and $\overset{p}{\longrightarrow }$ denote
convergence in distribution and in probability, respectively. Let $[n]=$ $
\{1,\cdots ,n\}$ for any positive integer $n$, let $\mathbf{1}\{\cdot \}$ be
the usual indicator function, and w.p.a.1 and a.s. abbreviate
\textquotedblleft with probability approaching 1\textquotedblright\ and
\textquotedblleft almost surely\textquotedblright , respectively.

\section{Model Setup}

\label{sec:model} In this paper we consider the following linear panel
model with IFEs:
\begin{equation}
Y_{it}=\Theta _{0,it}^{0}+X_{it}^{\prime }\Theta _{it}^{0}+e_{it},
\label{eq2.1}
\end{equation}
where $i\in \lbrack N]$, $t\in \lbrack T]$, $Y_{it}$ is the dependent
variable, $X_{it}=(X_{1,it},\cdots ,X_{p,it})^{\prime }$ is a $p\times 1$
vector of regressors, $\Theta _{it}^{0}=(\Theta _{1,it}^{0},\cdots ,\Theta
_{p,it}^{0})^{\prime }$ is a $p\times 1$ vector of slope coefficients, $
\Theta _{0,it}^{0}=\lambda _{i}^{0\prime }f_{t}^{0}$ is an intercept term
that exhibits a factor structure with $r_{0}$ factors, and $e_{it}$ is the
error term. Here, we assume $r_{0}$ is a fixed integer that does not change
as $\left( N,T\right) \rightarrow \infty .$ Let $\Lambda ^{0}=\left( \lambda
_{1}^{0},\cdots ,\lambda _{N}^{0}\right) ^{\prime }$ and $F^{0}=\left(
f_{1}^{0},\cdots ,f_{T}^{0}\right) ^{\prime }$. Let $Y=\left\{
Y_{it}\right\} ,$ $X_{j}=\left\{ X_{j,it}\right\} ,$ $\Theta
_{j}^{0}=\{\Theta _{j,it}^{0}\}$ and $E=\left\{ e_{it}\right\} ,$ all of
which are $N\times T$ matrices. Then we can rewrite (\ref{eq2.1}) in matrix form as
\begin{equation}
Y=\Theta _{0}^{0}+\sum_{j=1}^{p}X_{j}\odot \Theta _{j}^{0}+E.  \label{eq2.2}
\end{equation}
We assume that the slope coefficients follow time-varying latent group
structures, viz.,
\begin{equation*}
\Theta _{it}^{0}=\sum_{k\in \lbrack K_{t}]}\alpha _{kt}\mathbf{1}\{i\in
G_{kt}\},
\end{equation*}
where $\left\{ G_{kt}\right\} _{k\in K_{t}}$ forms a partition of $[N]$ for
each specific time $t$ with $K_{t}$ being the number of groups at time $t$.
Moreover, we assume that the group-specific slope coefficients $\alpha _{kt}$
or the memberships change at an unknown time point $T_{1}$, i.e.,
\begin{eqnarray*}
\alpha _{kt} &=&\left\{ \begin{aligned}
&\alpha_{k}^{(1)},\quad\text{for}\quad t=1,\dots,T_{1},\\ &
\alpha_{k}^{(2)},\quad\text{for}\quad t=T_{1}+1,\dots,T,\\ \end{aligned}
\right.  \\
G_{kt} &=&\left\{ \begin{aligned} &G_{k}^{(1)},\quad\text{for}\quad
t=1,\dots,T_{1},~ k=1,\dots,K^{(1)},\\ & G_{k}^{(2)},\quad \text{for}\quad
t=T_{1}+1,\dots,T,~ k=1,\dots,K^{(2)},\\ \end{aligned}\right.
\end{eqnarray*}
with $K^{(1)}$ and $K^{(2)}$ being the number of latent groups before and
after the break point, respectively. Let $g_{i}^{(1)}$ and $g_{i}^{(2)}$
respectively denote the individual group indices before and after the break:
\begin{equation*}
g_{i}^{(1)}=\sum_{k\in K^{(1)}}k\mathbf{1}\{i\in G_{k}^{(1)}\}\quad \text{and
}\quad g_{i}^{(2)}=\sum_{k\in K^{(2)}}k\mathbf{1}\{i\in G_{k}^{(2)}\}.
\end{equation*}

Let $r_{j}$ be the rank of $\Theta _{j}^{0}$ for $j\in \lbrack p]\cup \{0\}$
. It is easy to see that $\Theta _{j}^{0}$ exhibits a low-rank structure for
all $j.$ By the SVD, we have
\begin{equation*}
\Theta _{j}^{0}=\sqrt{NT}\mathcal{U}_{j}^{0}\Sigma _{j}^{0}\mathcal{V}
_{j}^{0\prime }:=U_{j}^{0}V_{j}^{0\prime },\quad j\in \lbrack p]\cup \{0\},
\end{equation*}
where $\mathcal{U}_{j}^{0}\in \mathbb{R}^{N\times r_{j}}$, $\mathcal{V}
_{j}^{0}\in \mathbb{R}^{T\times r_{j}}$, $\Sigma _{j}^{0}=\text{diag}(\sigma
_{1,j},\cdots ,\sigma _{r_{j},j})$, $U_{j}^{0}=\sqrt{N}\mathcal{U}
_{j}^{0}\Sigma _{j}^{0}$ with each row being $u_{i,j}^{0\prime }$, and $
V_{j}^{0}=\sqrt{T}\mathcal{V}_{j}^{0}$ with each row being $v_{t,j}^{0\prime
}$.

Note that we allow $\left\{ \Theta _{it}^{0}\right\} _{i=1}^{N}$ to exhibit
latent group structures before and after the break. For a particular $j\in
\lbrack p]$, the $N\times T$ matrix $\Theta _{j}^{0}$ may have no group
structure before or after the break, or no break, or more or fewer groups
after the break. Let $K_{j}^{(1)}$ and $K_{j}^{(2)}$ denote the number of
groups before and after the break, respectively, for $\{\Theta
_{j,it}^{0}\}_{i=1}^{N}$. Let $\mathcal{G}_{j}^{(\ell )}=\{G_{1,j}^{(\ell
)},\cdots ,G_{K_{j}^{(\ell )},j}^{(\ell )}\},$ $\ell =1,2,$ denote the
associated latent group structures. Define $N_{k,j}^{(\ell
)}=|G_{k,j}^{(\ell )}|$ and $\pi _{k,j}^{(\ell )}=\frac{N_{k,j}^{(\ell )}}{N}
$ for $\ell =1,2,$ where $\left\vert A\right\vert $ denotes the cardinality
of set $A.$ Further define $\tau _{T}:=\frac{T_{1}}{T}$. We show that $
\Theta _{j}^{0}$ has a low-rank structure in all of the following cases:

\begin{description}
\item[Case 1:] $\Theta _{j}^{0}$ exhibits neither structural break nor group
structure.\newline
In this case, $K_{j}^{(1)}=K_{j}^{(2)}=1$, and $\Theta _{j,it}^{0}=\alpha
_{j}$ $\forall \left( i,t\right) \in \lbrack N]\times \lbrack T]$. Without
loss of generality, assume that $\alpha _{j}>0.$ Then by the SVD, we have
\begin{align*}
& \mathcal{U}_{j}=\frac{1}{\sqrt{N}}\iota _{N}\in \mathbb{R}^{N\times
1},\quad \Sigma _{j}=\alpha _{j},\quad \mathcal{V}_{j}=\frac{1}{\sqrt{T}}
\iota _{T}\in \mathbb{R}^{T\times 1}, \\
& U_{j}=\alpha _{j}\iota _{N}\in \mathbb{R}^{N\times 1},\quad V_{j}=\iota
_{T}\in \mathbb{R}^{T\times 1},
\end{align*}
where $\iota _{d}=\left( 1,\cdots ,1\right) ^{\prime }\in \mathbb{R}
^{d\times 1}$ for any natural number $d$. Obviously, $r_{j}=1$ in Case 1.

\item[Case 2:] $\Theta _{j}^{0}$ exhibits no structural break but a group
structure.\newline
In this case, $K_{j}^{(1)}=K_{j}^{(2)}=K_{j}$, $
G_{k,j}^{(1)}=G_{k,j}^{(2)}=G_{k,j}$, $N_{k,j}^{(1)}=N_{k,j}^{(2)}=N_{k,j}$,
$\pi _{k,j}^{(1)}=\pi _{k,j}^{(2)}=\pi _{k,j}~\forall k\in \left[ K_{j}
\right] $, and $\Theta _{j,it}^{0}=\operatornamewithlimits{\sum}\limits
_{k\in \lbrack K_{j}]}\alpha _{k,j}\mathbf{1}\left\{ i\in G_{k,j}\right\} $
for $t\in \lbrack T]$. Therefore, we have
\begin{align*}
& \mathcal{U}_{j,i}=\frac{\sum_{k\in \lbrack K_{j}]}\alpha _{k,j}\mathbf{1}
\left\{ i\in G_{k,j}\right\} }{\sqrt{\sum_{k\in \lbrack K_{j}]}N_{k,j}\left(
\alpha _{k,j}\right) ^{2}}},\quad \Sigma _{j}=\sqrt{\sum_{k\in \lbrack
K_{j}]}\pi _{k,j}\left( \alpha _{k,j}\right) ^{2}},\quad \mathcal{V}_{j}=
\frac{1}{\sqrt{T}}\iota _{T}, \\
& u_{i,j}=\sum_{k\in \lbrack K_{j}]}\alpha _{k,j}\mathbf{1}\left\{ i\in
G_{k,j}\right\} ,\quad V_{j}=\iota _{T},
\end{align*}
where $\mathcal{U}_{j,i}$ is the $i$-th element in $\mathcal{U}_{j}$.
Obviously, $r_{j}=1$ in this case.

\item[Case 3:] $\Theta _{j}^{0}$ exhibits both a structural break and a
group structure. \
\end{description}

\begin{itemize}
\item[\textbf{(i)}] $K_{j}^{(1)}\neq K_{j}^{(2)}$, where we have different
numbers of groups before and after the break;

\item[\textbf{(ii)}] $K_{j}^{(1)}=K_{j}^{(2)}=K_{j}$ and $G_{k,j}^{(1)}\neq
G_{k,j}^{(2)}$, where we have the same number of groups before and after the
break, but the group memberships change after the break point;

\item[\textbf{(iii)}] $K_{j}^{(1)}=K_{j}^{(2)}=K_{j}$, $
G_{k,j}^{(1)}=G_{k,j}^{(2)}=G_{k,j}$ for $\forall k\in \lbrack K_{j}]$, and $
\alpha _{k,j}^{(1)}\neq \alpha _{k,j}^{(2)}$ for at least one $k\in \lbrack
K_{j}]$, where even though neither the number of groups nor group membership
changes after the break, there exists at least one group whose slope
coefficients change.
\end{itemize}

For any positive integer $d$, we use $\mathbf{0}_{d}$ to denote a $d\times 1$
vector of zeros. The following lemma lays down the foundation for break point
detection in our model.

\begin{lemma}
{\small \label{Lem:idfct} } For any $j\in [p]$ such that $\Theta_{j}^{0}$
lies in Case 3 above, we have $rank(\Theta_{j}^{0})\leq 2$. When $
rank(\Theta_{j}^{0})=2$, we have

\begin{itemize}
\item[(i)] $\Theta_{j}^{0}=\mathcal{U}_{j}\Sigma_{j}\mathcal{V}
_{j}^{\prime}=U_{j}V_{j}^{\prime }$ where $U_{j}=\mathcal{U}_{j}\Sigma_{j}/
\sqrt{T},$ $V_{j}=\sqrt{T}\mathcal{V}_{j}=D_{j}R_{j}$$,$ $D_{j}=
\begin{bmatrix}
\frac{1}{\sqrt{\tau_{T}}}\iota_{T_{1}} & \mathbf{0}_{T_{1}} \\
\mathbf{0}_{T-T_{1}} & \frac{1}{\sqrt{1-\tau_{T}}}\iota_{T-T_{1}}
\end{bmatrix}
$ and $R_{j}^{\prime }R_{j}=I_{2};$

\item[(ii)] $\left\Vert \frac{v_{t,j}^{0}}{\left\Vert v_{t,j}^{0}\right\Vert
}-\frac{v_{t^{\ast },j}^{0}}{\left\Vert v_{t^{\ast },j}^{0}\right\Vert }
\right\Vert =\sqrt{2}$ for any $t\leq T_{1}$ and $t^{\ast }>T_{1}$.
\end{itemize}
\end{lemma}

By Lemma \ref{Lem:idfct} for Case 3 and the above analyses for Cases 1 and
2, we conclude that $\Theta _{j}^{0}$ is a low-rank matrix with rank equal
to or less than $2$. In view of the low-rank structure of the slope
matrices, we propose to adopt the NNR to obtain the preliminary estimates
below. Moreover, under Case 3, Lemma \ref{Lem:idfct}(ii) indicates that
singular vectors of the slope matrix with rank 2 contain the structural
break information.


\section{Estimation}

\label{sec:estimation} In this section we provide the estimation algorithm.
We first assume that the ranks $r_{j}$ for $j\in [p]\cup \{0\}$ are known,
and then propose a singular value thresholding (SVT) procedure to estimate
them. After we recover the break point and the latent group structures, we
propose consistent estimates of the group-specific parameters.

\subsection{Estimation Algorithm}

Given $r_{j},~\forall j\in \lbrack p]\cup \{0\}$, we propose the following
four-step procedure to estimate the break point and to recover the latent
group structures before and after the break.

\begin{description}
\item[\textbf{Step 1:}] \textbf{Nuclear Norm Regularization (NNR).} We run
the nuclear norm regularized regression and obtain the preliminary estimates
as follows:
\begin{equation}
\{\tilde{\Theta}_{j}\}_{j\in \lbrack p]\cup \{0\}}=\operatorname*{arg\,min}\limits_{\left\{
\Theta _{j}\right\} _{j=0}^{p}}\frac{1}{NT}\left\Vert
Y-\sum_{j=1}^{p}X_{j}\odot \Theta _{j}-\Theta _{0}\right\Vert
^{2}+\sum_{j=0}^{p}\nu _{j}\left\Vert \Theta _{j}\right\Vert _{\ast },
\label{obj}
\end{equation}
where $\nu _{j}$ is the tuning parameter for $j\in \lbrack p]\cup \{0\}$.
For each $j$, conduct the SVD: $\frac{1}{\sqrt{NT}}\tilde{\Theta}_{j}=\hat{
\tilde{\mathcal{U}}}_{j}\hat{\tilde{\Sigma}}_{j}\hat{\tilde{\mathcal{V}}}
_{j}^{\prime }$, where $\hat{\tilde{\Sigma}}_{j}$ is a diagonal matrix with
the diagonal elements being the descending singular values of $\tilde{\Theta}
_{j}$. Let $\tilde{\mathcal{V}}_{j}$ consist of the first $r_{j}$ columns of
$\hat{\tilde{\mathcal{V}}}_{j},$ and $\tilde{V}_{j}=\sqrt{T}\tilde{\mathcal{V
}}_{j}.$ Let $\tilde{v}_{t,j}^{\prime }$ denote the $t$-th row of $\tilde{V}
_{j}$ for $t\in \lbrack T]$.

\item[\textbf{Step 2:}] \textbf{Row- and Column-Wise Regressions.} First run
the row-wise regressions of $Y_{it}$ on $\left( \tilde{v}_{t,0},\{\tilde{v}
_{t,j}X_{j,it}\}_{j\in \lbrack p]}\right) $ to obtain $\{\dot{u}
_{i,j}\}_{j\in \lbrack p]\cup \{0\}}$ for $i\in \lbrack N]$. Then run the
column-wise regressions of $Y_{it}$ on $\left( \dot{u}_{i,0},\{\dot{u}
_{i,j}X_{j,it}\}_{j\in \lbrack p]}\right) $ to obtain $\{\dot{v}
_{t,j}\}_{j\in \lbrack p]\cup \{0\}}$ for $t\in \lbrack T].$ Let $\dot{\Theta
}_{j,it}=\dot{u}_{i,j}^{\prime }\dot{v}_{t,j}$ for $\left( i,t\right) \in
\lbrack N]\times \lbrack T]$ and $j\in \lbrack p]\cup \{0\}$. Specifically,
the row- and column-wise regressions are given by
\begin{align}
\left\{ \dot{u}_{i,j}\right\} _{j\in \lbrack p]\cup \{0\}}& =\operatorname*{arg\,min}
_{\{u_{i,j}\}_{j\in \lbrack p]\cup \{0\}}}\frac{1}{T}\sum_{t\in \lbrack
T]}\left( Y_{it}-u_{i,0}^{\prime }\tilde{v}_{t,0}-\sum_{j=1}^{p}u_{i,j}^{
\prime }\tilde{v}_{t,j}X_{j,it}\right) ^{2},\quad i\in \lbrack N],
\label{obj2} \\
\left\{ \dot{v}_{t,j}\right\} _{j\in \lbrack p]\cup \{0\}}& =\operatorname*{arg\,min}
_{\{v_{t,j}\}_{j\in \lbrack p]\cup \{0\}}}\frac{1}{N}\sum_{i\in \lbrack
N]}\left( Y_{it}-v_{t,0}^{\prime }\dot{u}_{i,0}-\sum_{j=1}^{p}v_{t,j}^{
\prime }\dot{u}_{i,j}X_{j,it}\right) ^{2},\quad t\in \lbrack T].
\end{align}

\item[\textbf{Step 3:}] \textbf{Break Point Estimation.} We estimate the
break point as follows:
\begin{equation}
\hat{T}_{1}=\operatorname*{arg\,min}_{s\in \{2,\cdots ,T-1\}}\frac{1}{pNT}\sum_{j\in \lbrack
p]}\sum_{i\in \lbrack N]}\left\{ \sum_{t=1}^{s}\left( \dot{\Theta}_{j,it}-
\bar{\dot{\Theta}}_{j,i}^{(1s)}\right) ^{2}+\sum_{t=s+1}^{T}\left( \dot{
\Theta}_{j,it}-\bar{\dot{\Theta}}_{j,i}^{(2s)}\right) ^{2}\right\} ,
\label{BP_estimates}
\end{equation}
where $\bar{\dot{\Theta}}_{j,i}^{(1s)}=\frac{1}{s}\sum_{t=1}^{s}{\dot{\Theta}
_{j,it}}$ and $\bar{\dot{\Theta}}_{j,i}^{(2s)}=\frac{1}{T-s}\sum_{t=s+1}^{T}{
\dot{\Theta}_{j,it}}$.

\item[\textbf{Step 4:}] \textbf{Sequential Testing K-means (STK).} In this
step, we estimate the number of groups and the group membership before and
after the break by using the STK algorithm. For each $j\in \lbrack p]$,
define $\dot{\Theta}_{j,i}^{(1)}=(\dot{\Theta}_{j,i1},\cdots ,\dot{\Theta}
_{j,i\hat{T}_{1}})^{\prime }$, $\dot{\Theta}_{j,i}^{(2)}=(\dot{\Theta}_{j,i,
\hat{T}_{1}+1},\cdots ,\dot{\Theta}_{j,iT})^{\prime }$, $\dot{\beta}
_{i}^{(1)}=\frac{1}{\sqrt{\hat{T}_{1}}}(\dot{\Theta}_{1,i}^{(1)\prime
},\cdots ,\dot{\Theta}_{p,i}^{(1)\prime })^{\prime },$ and $\dot{\beta}
_{i}^{(2)}=\frac{1}{\sqrt{\hat{T}_{2}}}(\dot{\Theta}_{1,i}^{(2)\prime
},\cdots ,\dot{\Theta}_{p,i}^{(2)\prime })^{\prime }$.
Let $z_{\varsigma }$ be some predetermined value which will be specified in
the next subsection. Given the subsample before and after the estimated
break point, initialize $m=1$ and classify each subsample into $m$ groups
by the K-means algorithm with group membership obtained as $\hat{\mathcal{G}
}_{m}^{(\ell )}:=\{\hat{G}_{k,m}^{(\ell )}\}_{k\in \lbrack m]}$. Next, we
construct a suitable test statistic $\hat{\Gamma}_{m}^{(\ell )}$, defined by \eqref{eq:gamma-m} in the next subsection, and compare it to its critical value $z_{\varsigma }$ at significance level $\varsigma$ under the null hypothesis of $m$ subgroups, setting $m=m+1$ and moving to the next iteration if $\hat{\Gamma}_{m}^{(\ell )}>z_{\varsigma }$ and stopping the STK algorithm otherwise. Lastly, define $\hat{K}^{(\ell )}=m$ and $\hat{\mathcal{G}}^{(\ell )}=\hat{
\mathcal{G}}_{m}^{(\ell )}$. The next subsection provides a detailed breakdown of these steps in the STK algorithm.
\begin{figure}[h]
\centering
\includegraphics[width=0.8\textwidth]{STK.png}
\caption{The flow chart of STK algorithm}
\label{fig:STK}
\end{figure}

\end{description}

Several remarks are in order. First, the ranks of the intercept
and slope matrices are assumed known in Step 1 but consistent estimates
for them are proposed in the SVT below. Second, we obtain preliminary estimates by
NNR based on the low-rank structure of the intercept and slope matrices in
the model. These estimates are consistent in terms of the Frobenius norm but pointwise or uniform convergence for their elements is not established. Nonetheless, SVD can be employed to obtain preliminary estimates of the
factors and factor loadings to be used subsequently. Third, row- and column-wise linear regressions are conducted to obtain updated estimates of the factors and factor loadings for which we can establish pointwise and
uniform convergence rates. Fourth, using the consistent estimates obtained in
the second step, we can estimate the break point in Step 3 consistently by
using a binary segmentation process. Fifth, the STK algorithm in Step
4 then yields the estimated number of groups and the group memberships together.

In the latent group literature, it is standard and popular to assume the
number of groups in the K-means algorithm is known and then to estimate the
number of groups by using certain information criteria. In this case, one
needs to consider not only under- and just-fitting cases, but also
over-fitting cases. It is well known that the major difficulty with this
approach is showing that the over-fitting case occurs with probability
approaching zero. The STK algorithm ensures a focus on the
under-and just-fitting cases, which helps to avoid the difficulty caused
by K-means classification with a larger than true
number of groups. In addition, although this sequential algorithm approach is adopted, the
error from the previous iteration does not accumulate in the following
iterations owing to fact that the classification in each iteration is new
and is not based on the K-means outcomes in previous iterations.

\subsection{The STK algorithm}

This subsection describes the K-means algorithm and the construction
of the test statistics $\hat{\Gamma}_{m}^{(\ell )}$ that are used in the STK algorithm for $
\ell \in \{1,2\}$.

First, we define the objective function for the K-means algorithm with $m$
clusters at each iteration. Let $a_{k,m}^{(\ell )}$ be a $p\hat{T}_{1}\times
1$ and $p(T-\hat{T}_{1})\times 1$ vector for $\ell =1,2$, respectively. We
obtain the group membership with $m$ groups by solving the following
minimization problem
\begin{equation}
\left\{ \dot{a}_{k,m}^{(\ell )}\right\} _{k\in \lbrack m]}=\operatorname*{arg\,min}_{\left\{
a_{k,m}^{(\ell )}\right\} _{k\in \lbrack m]}}\frac{1}{N}\sum_{i\in \lbrack
N]}\min_{k\in \lbrack m]}\left\Vert \dot{\beta}_{i}^{(\ell )}-a_{k,m}^{(\ell
)}\right\Vert ^{2},  \label{kmeans_obj}
\end{equation}
which yields the membership estimates for each individual at the $m$-th
iteration as
\begin{equation}
\hat{g}_{i,m}^{(\ell )}=\operatorname*{arg\,min}_{k\in \lbrack m]}\left\Vert \dot{\beta}
_{i}^{(\ell )}-\dot{a}_{k,m}^{(\ell )}\right\Vert \quad \forall i\in \lbrack
N].  \label{group estimates}
\end{equation}
Let $\hat{G}_{k,m}^{(\ell )}:=\{i\in \lbrack N]:\hat{g}_{i,m}^{(\ell )}=k\}$.

Second, we discuss the construction of the test statistic based on the idea
of homogeneity test for several subsamples. At iteration $m$, we have $m$
potential subgroups $(\hat{G}_{1,m}^{(\ell )},\cdots ,\hat{G}_{m,m}^{(\ell
)})$ after the K-means classification for $\ell =1$ and 2. Let $\mathcal{
\hat{T}}_{1}=[\hat{T}_{1}]$, $\mathcal{\hat{T}}_{2}=[T]\backslash \lbrack
\hat{T}_{1}]$, $\mathcal{\hat{T}}_{1,-1}=\mathcal{\hat{T}}_{1}\backslash \{
\hat{T}_{1}\}$, $\mathcal{\hat{T}}_{2,-1}=\mathcal{\hat{T}}_{2}\backslash
\{T\}$, $\mathcal{\hat{T}}_{1,j}=\{1+j,\cdots ,\hat{T}_{1}\}$, and $\mathcal{
\hat{T}}_{2,j}=\{\hat{T}_{1}+1+j,\cdots ,T\}$ for some specific $j\in \hat{
\mathcal{T}}_{\ell ,-1}$. Based on these estimated subgroups, we can obtain
the estimates of the coefficients, factors and factor loadings for each
subgroup in regime $\ell $ as follows:
\begin{equation*}
\left( \left\{ \hat{\theta}_{i,k,m}^{(\ell )}\right\} _{i\in \hat{G}
_{k,m}^{(\ell )}},\hat{\Lambda}_{k,m}^{(\ell )},\hat{F}_{k,m}^{(\ell
)}\right) =\operatorname*{arg\,min}_{\left\{ \theta _{i},\lambda _{i},f_{t}\right\} _{i\in
\hat{G}_{k,m}^{(\ell )},t\in \mathcal{\hat{T}}_{\ell }}}\sum_{i\in \hat{G}
_{k,m}^{(\ell )}}\sum_{t\in \mathcal{\hat{T}}_{\ell }}\left(
Y_{it}-X_{it}^{\prime }\theta _{i}-\lambda _{i}^{\prime }f_{t}\right) ^{2},
\end{equation*}
where $\hat{\Lambda}_{k,m}^{(\ell )}=\{\hat{\lambda}_{i,k,m}^{\left( \ell
\right) }\}_{i\in \hat{G}_{k,m}^{(\ell )}}$ and $\hat{F}_{k,m}^{(\ell )}=\{
\hat{f}_{t,k,m}^{\left( \ell \right) }\}_{t\in \hat{\mathcal{T}}_{\ell }}.$
For all $i\in \lbrack N]$ and $t\in \lbrack T]$, define the residuals
\begin{equation*}
\hat{e}_{it}=\sum_{\ell =1}^{2}\left( Y_{it}-\hat{f}_{t,k,m}^{(\ell )\prime }
\hat{\lambda}_{i,k,m}^{(\ell )}-X_{it}^{\prime }\hat{\theta}_{i,k,m}^{(\ell
)}\right) \mathbf{1}\{t\in \mathcal{\hat{T}}_{\ell }\}.
\end{equation*}
Let $\hat{X}_{i}^{(1)}=(X_{i1},\cdots ,X_{i\hat{T}_{1}})^{\prime }$, $\hat{X}
_{i}^{(2)}=(X_{i,\hat{T}_{1}+1},\cdots ,X_{iT})^{\prime }$, and $\hat{T}
_{2}=T-\hat{T}_{1}.$ Define
\begin{align*}
& \hat{\bar{\theta}}_{k,m}^{(\ell )}=\frac{1}{|\hat{G}_{k,m}^{(\ell )}|}
\sum_{i\in \hat{G}_{k,m}^{(\ell )}}\hat{\theta}_{i,k,m}^{(\ell )},\quad M_{
\hat{F}_{k,m}^{(\ell )}}=I_{\hat{T}_{\ell }}-\frac{1}{\hat{T}_{\ell }}\hat{F}
_{k,m}^{(\ell )}\hat{F}_{k,m}^{(\ell )\prime }, \\
& \hat{S}_{ii,k,m}^{(\ell )}=\frac{1}{\hat{T}_{\ell }}\hat{X}_{i}^{(\ell
)\prime }M_{\hat{F}_{k,m}^{(\ell )}}\hat{X}_{i}^{(\ell )},\quad \hat{a}
_{ii,k}^{(\ell )}=\hat{\lambda}_{i,k,m}^{(\ell )\prime }\left( |\hat{G}
_{k,m}^{(\ell )}|^{-1}\hat{\Lambda}_{k,m}^{(\ell )\prime }\hat{\Lambda}
_{k,m}^{(\ell )}\right) ^{-1}\hat{\lambda}_{i,k,m}^{(\ell )}.
\end{align*}
Let $\hat{z}_{it}^{(\ell )\prime }$ be the $t$-th row of $M_{\hat{F}
_{k,m}^{(\ell )}}\hat{X}_{i}^{(\ell )}.$ For each subgroup $\hat{G}
_{k,m}^{(\ell )}$ with $k\in \lbrack m]$, we follow the lead of \cite
{pesaran2008testing} and \cite{ando2015simple} and define the following test statistic components
\begin{equation} \label{eq:gamma-km}
\hat{\Gamma}_{k,m}^{(\ell )}=\sqrt{|\hat{G}_{k,m}^{(\ell )}|}\cdot \frac{
\frac{1}{|\hat{G}_{k,m}^{(\ell )}|}\sum_{i\in \hat{G}_{k,m}^{(\ell )}}\hat{
\mathbb{S}}_{i,k,m}^{(\ell )}-p}{\sqrt{2p}},
\end{equation}
where
\begin{align*}
& \hat{\mathbb{S}}_{i,k,m}^{(\ell )}=\hat{T}_{\ell }(\hat{\theta}
_{i,k,m}^{(\ell )}-\hat{\bar{\theta}}_{k,m}^{(\ell )})^{\prime }\hat{S}
_{ii,k,m}^{(\ell )}(\hat{\Omega}_{i,k,m}^{(\ell )})^{-1}\hat{S}
_{ii,k,m}^{(\ell )}(\hat{\theta}_{i,k}^{(\ell )}-\hat{\bar{\theta}}
_{k}^{(\ell )})\left( 1-\frac{\hat{a}_{ii,k}^{(\ell )}}{|\hat{G}
_{k,m}^{(\ell )}|}\right) ^{2}, \\
& \hat{\Omega}_{i,k,m}^{(\ell )}=\frac{1}{\hat{T}_{\ell }}\sum_{t\in \hat{
\mathcal{T}}_{\ell }}\hat{z}_{it}^{(\ell )}\hat{z}_{it}^{(\ell )\prime }\hat{
e}_{it}^{2}+\frac{1}{\hat{T}_{\ell }}\sum_{j\in \hat{\mathcal{T}}_{\ell
,-1}}k(j/S_{T})\sum_{t\in \mathcal{\hat{T}}_{\ell ,j}}[\hat{z}_{it}^{(\ell )}
\hat{z}_{i,t+j}^{(\ell )\prime }\hat{e}_{it}\hat{e}_{i,t+j}+\hat{z}
_{i,t-j}^{(\ell )}\hat{z}_{it}^{(\ell )\prime }\hat{e}_{i,t-j}\hat{e}_{i,t}],
\end{align*}
$k(\cdot )$ is a kernel function, $S_{T}$ is a bandwidth/truncation
parameter, and $\hat{\Omega}_{i,k,m}^{(\ell )}$ is a traditional HAC
estimator. Using the components \eqref{eq:gamma-km} we now define the test statistic
\begin{align} \label{eq:gamma-m}
\hat{\Gamma}_{m}^{(\ell )}=\max_{k\in \lbrack m]}(\hat{\Gamma
}_{k,m}^{(\ell )})^{2}.
\end{align}


We will show that $\hat{\Gamma}_{m}^{(\ell )}$ is asymptotically distributed
as the maximum of $m$ independent $\chi ^{2}(1)$ random variables under the
null hypothesis that the slope coefficients in each of the $m$ subsamples are
homogeneous, whereas it diverges to infinity under the alternative. Let $
z_{\varsigma }$ denote the critical value at significance level $\varsigma ,$
which is calculated from the maximum of $m$ independent $\chi ^{2}(1)$
random variables. We reject the null of $m$ subgroups in favor of more
groups at level $\varsigma $ if $\hat{\Gamma}_{m}^{(\ell )}>z_{\varsigma }.$

\subsection{Rank Estimation}

To obtain the rank estimator, we use the low-rank estimators from (\ref{obj}
) and estimate $r_{j}$ by the singular value thresholding (SVT) criterion
\begin{equation*}
\hat{r}_{j}=\sum_{i=1}^{N\wedge T}\mathbf{1}\left\{ \sigma _{i}\left( \tilde{
\Theta}_{j}\right) \geq 0.5\left( \nu _{j}\left\Vert \tilde{\Theta}
_{j}\right\Vert _{op}\right) ^{1/2}\right\} \quad \forall j\in \{0\}\cup
\lbrack p],
\end{equation*}
where $\sigma _{i}\left( A\right) $ denotes the $i$-th largest singular
value of $A$ and $N\wedge T=\min (N,T).$ By arguments as used in the proof
of Proposition D.1 in \cite{chernozhukov2019inference} and that of Theorem
3.2 in \cite{hong2023profile}, we can show that $\mathbb{P}(\hat{r}
_{j}=r_{j})\rightarrow 1$ for each $j$ as $\left( N,T\right) \rightarrow
\infty .$

\subsection{Parameter Estimation}

Once we obtain the estimated break point, the number of groups and the group
membership before and after the estimated break point, we can estimate the
group-specific slope coefficients $\{\alpha _{k}^{(\ell )}\}_{k\in \lbrack
\hat{K}^{(\ell )}]}$ along with the factors and factor loadings as follows
\begin{equation}
\left( \hat{\Lambda}^{\left( \ell \right) },\hat{F}^{\left( \ell \right)
},\left\{ \hat{\alpha}_{k}^{(\ell )}\right\} _{k\in \lbrack \hat{K}^{(\ell
)}]}\right) =\operatorname*{arg\,min}\mathbb{L}\left( \Lambda ,F,\left\{ a_{k}^{(\ell
)}\right\} _{k\in \lbrack \hat{K}^{(\ell )}]}\right) ,  \label{Post_estimate}
\end{equation}
where $\mathbb{L}\left( \Lambda ,F,\left\{ a_{k}^{(\ell )}\right\} _{k\in
\lbrack \hat{K}^{(\ell )}]}\right) =\frac{1}{N\hat{T}_{\ell }}\sum_{k=1}^{
\hat{K}^{(\ell )}}\sum_{i\in \hat{G}_{k}^{(\ell )}}\sum_{t\in \mathcal{\hat{T
}}_{\ell }}\left( Y_{it}-\lambda _{i}^{\prime }f_{t}-X_{it}^{\prime
}a_{k}^{(\ell )}\right) ^{2}.$ Here, we ignore the fact that the prior- and
post-break regimes share the same set of factor loadings and estimate the
group-specific parameters separately for the two regimes at the cost of
sacrificing some efficiency for the factor loading estimates. Alternatively,
we can pool the observations before and after the break to estimate the
parameters as
\begin{equation*}
\left( \hat{\Lambda},\hat{F},\left\{ \hat{\alpha}_{k}^{(1)}\right\} _{k\in
\lbrack \hat{K}^{(1)}]},\left\{ \hat{\alpha}_{k}^{(2)}\right\} _{k\in
\lbrack \hat{K}^{(2)}]}\right) =\operatorname*{arg\,min}\mathbb{L}\left( \Lambda ,F,\left\{
a_{k}^{(1)}\right\} _{k\in \lbrack \hat{K}^{(1)}]},\left\{
a_{k}^{(2)}\right\} _{k\in \lbrack \hat{K}^{(2)}]}\right)
\end{equation*}
where
\begin{equation}
\mathbb{L}\left( \Lambda ,F,\left\{ a_{k}^{(1)}\right\} _{k\in \lbrack \hat{K
}^{(1)}]},\left\{ a_{k}^{(2)}\right\} _{k\in \lbrack \hat{K}^{(2)}]}\right) =
\mathbb{L}\left( \Lambda ,F,\left\{ a_{k}^{(1)}\right\} _{k\in \lbrack \hat{K
}^{(1)}]}\right) +\mathbb{L}\left( \Lambda ,F,\left\{ a_{k}^{(2)}\right\}
_{k\in \lbrack \hat{K}^{(2)}]}\right) .
\end{equation}

In either case, as should be clear because of the presence of the group structures,
establishment of the asymptotic properties of the post-classification
estimators of the group-specific slope coefficients becomes much more
involved than in \cite{bai2009panel} and \cite{moon2017dynamic}. For
this reason, we will focus on the estimates defined in (\ref{Post_estimate}).


\section{Asymptotic Theory}

\label{sec:theory} This section develops the asymptotic properties of
the estimators introduced above.

\subsection{Basic Assumptions}

Define $e_{i}=\left( e_{i1},\cdots ,e_{iT}\right) ^{\prime }$ and $
X_{j,i}=\left( X_{j,i1},\cdots ,X_{j,iT}\right) ^{\prime }.$ Let $V_{j}^{0}$
be a $T\times r_{j}$ matrix with its $t$-th row being $v_{t,j}^{0\prime }$,
and $U_{j}^{0}$ be the $N\times r_{j}$ matrix with its $i$-th row being $
u_{i,j}^{0\prime }$. Throughout the paper, we treat the factors $
\{V_{j}^{0}\}_{j\in \lbrack p]\cup \{0\}}$ as random and their loadings $
\{U_{j}^{0}\}_{j\in \lbrack p]\cup \{0\}}$ as deterministic. Let $\mathscr{D}
:=\sigma (\{V_{j}^{0}\}_{j\in \lbrack p]\cup \{0\}})$ denote the minimum $
\sigma $-field generated by $\{V_{j}^{0}\}_{j\in \lbrack p]\cup \{0\}}.$
Similarly, let $\mathscr{G}_{t}:=\sigma (\mathscr{D},\left\{ X_{is}\right\}
_{i\in \lbrack N],s\leq t+1},\left\{ e_{is}\right\} _{i\in \lbrack N],s\leq
t})$. Let $\max_{i}=\max_{i\in \lbrack N]},$ $\max_{t}=\max_{t\in \lbrack
T]}\ $and $\max_{i,t}=\max_{i\in \lbrack N],t\in \lbrack T]}.$ Let $M$ and $
C $ be generic bounded positive constants which may vary across lines. \

\begin{ass}
\label{ass:1}

\begin{enumerate}
\item[(i)] $\left\{ e_{it},X_{it}\right\}_{t\in [T]}$ are conditionally
independent across $i$ given $\mathscr{D}$.

\item[(ii)] $\mathbb{E}(e_{it}|X_{it},\mathscr{D})=0$.

\item[(iii)] For each $i$, $\left\{ \left( e_{it},X_{it}\right) ,t\geq
1\right\} $ is strong mixing conditional on $\mathscr{D}$ with the mixing
coefficient $\alpha _{i}(\cdot )$ satisfying $\max_{i}\alpha _{i}(z)\leq
M\vartheta ^{z}$ for some constant $\vartheta \in \left( 0,1\right) $.

\item[(iv)] There exists a constant $C>0$ such that $\max_{i}\frac{1}{T}
\sum_{t\in \lbrack T]}\left\Vert \xi _{it}\right\Vert ^{2}\leq C~a.s.\quad $
and$\quad \max_{t}\frac{1}{N}\sum_{i\in \lbrack N]}\left\Vert \xi
_{it}\right\Vert ^{2}$ $\leq C~a.s.$ for $\xi _{it}=e_{it},$ $X_{it}$ and $
X_{it}e_{it}$.

\item[(v)] $\max_{i,t}\mathbb{E}[\left\Vert \xi _{it}\right\Vert ^{q}\big|
\mathscr{D}]\leq M~a.s.$ and $\max_{i,i^{\ast },t}\mathbb{E}[\left\Vert
X_{it}e_{i^{\ast }t}\right\Vert ^{q}\big|\mathscr{D}]\leq M~a.s.$ for some $
q>8$ and $\xi _{it}=e_{it},$ $X_{it}$ and $X_{it}e_{it}.$

\item[(vi)] As $(N,T)\rightarrow \infty $, $\sqrt{N}(\log
N)^{2}T^{-1}\rightarrow 0$ and $T(\log N)^{2}N^{-3/2}\rightarrow 0$.
\end{enumerate}
\end{ass}

\begin{assp}{\ref*{ass:1}{$^*$} }
\label{ass:1*}
(i), (iv) and (v) are same as Assumption \ref{ass:1}(i), (iv) and (v). In addition:

\begin{itemize}
\item[(ii)] $\mathbb{E}(e_{it}|\mathscr{G}_{t-1})=0$ $\forall (i,t)\in
\lbrack N]\times \lbrack T]$, and $\max_{i,t}\mathbb{E(}e_{it}^{2}\big|
\mathscr{G}_{t-1})\leq M~a.s.$.

\item[(iii)] $\left\{ e_{it}\right\} _{i\in \lbrack N]}$ is conditionally
independent across $t$ given $\mathscr{D}$.
\end{itemize}
\end{assp}


Assumption \ref{ass:1}(i) imposes conditional independence on $\left\{
e_{it},X_{it}\right\} _{t\in \lbrack T]}$ across the cross sectional units.
Assumption \ref{ass:1}(ii) is the conditional moment condition. Assumption \ref
{ass:1}(iii) imposes conditional strong mixing conditions along the time
dimension. See \cite{Prakasa_Rao2009} for the definition of conditional
strong mixing and \cite{su2013testing} for an application in the panel
setup. Assumptions \ref{ass:1}(iv)-(v) impose conditions that
restrict the tail behavior of $\xi _{it}$. Note that neither the regressors
nor the errors are constrained to be bounded. Assumption \ref{ass:1}
(vi) imposes restrictions on $N$ and $T$ but does not require $N$ and $
T$ to diverge at the the same rate. It is possible to allow $N$ to
diverge to infinity faster but not too much faster than $T$, and vice versa.

Assumption \ref{ass:1*} is used for the study of dynamic panel data models. To be
specific, Assumption \ref{ass:1*}(ii) requires that the error sequence $
\left\{ e_{it},t\geq 1\right\} $ be a martingale difference sequence
(m.d.s.) with respect to the filtration $\mathscr{G}_{t},$ which allows for
lagged dependent variables in $X_{it}$. Assumption \ref{ass:1*}(iii) imposes
conditional independence of the errors over $t$. The presence of
serially correlated errors in dynamic panels typically induces endogeneity,
which invalidates least-squares-based PCA estimation.

\begin{ass}
\label{ass:2} rank$(\Theta _{j}^{0})=r_{j}\leq \bar{r}$ for $j\in \lbrack
p]\cup \{0\}$ and some fixed $\bar{r},$ and $\max_{j\in \lbrack p]\cup
\{0\}}||\Theta _{j}^{0}||_{\max }\leq M$.
\end{ass}

\noindent Assumption \ref{ass:2} imposes low-rank conditions on the coefficient
matrices, which facilitate the use of NNR in obtaining preliminary
estimates in the first step. As discussed in the previous section, we see
that the low-rank assumption for the slope matrices is satisfied for the
model in Section \ref{sec:model}. Moreover, we follow \cite{ma2020detecting}
and assume the elements of the coefficient matrices are uniformly bounded to
simplify the proofs. The boundedness of the slope coefficients is reasonable
given that their cardinality does not grow with the sample size. The
boundedness assumption for the intercept coefficient can be relaxed at the
cost of more lengthy arguments.

\begin{ass}
\label{ass:3} Let $\sigma _{l,j}$ denote the $l$-th largest singular values
of $\Theta _{j}^{0}$ for $j\in \lbrack p]\cup \{0\}.$ There exist some
constants $C_{\sigma }$ and $c_{\sigma }$ such that
\begin{equation*}
\infty >C_{\sigma }\geq \lim \sup_{(N,T)\rightarrow \infty }\max_{j\in
\lbrack p]}\sigma _{1,j}\geq \lim \inf_{(N,T)\rightarrow \infty }\min_{j\in
\lbrack p]}\sigma _{r_{j},j}\geq c_{\sigma }>0.
\end{equation*}
\end{ass}

\noindent Assumption \ref{ass:3} imposes some conditions on the singular values of the
coefficient matrices. These ensure that only pervasive factors are allowed when
the matrices are written as a factor structure. The assumption can be
readily verified given the latent group structures of the slope coefficients.

Consider the SVD for $\Theta _{j}^{0}$: $\Theta _{j}^{0}=\mathcal{U}
_{j}\Sigma _{j}\mathcal{V}_{j}^{\prime }$ $\forall j\in \lbrack p]\cup \{0\}$
. Decompose $\mathcal{U}_{j}=\left( \mathcal{U}_{j,r},\mathcal{U}
_{j,0}\right) $ and $\mathcal{V}_{j}=\left( \mathcal{V}_{j,r},\mathcal{V}
_{j,0}\right) $ with $\left( \mathcal{U}_{j,r},\mathcal{V}_{j,r}\right) $
being the singular vectors corresponding to nonzero singular values and $
\left( \mathcal{U}_{j,0},\mathcal{V}_{j,0}\right) $ being the singular
vectors corresponding to zero singular values. Hence, for any matrix $W\in
\mathbb{R}^{N\times T}$, we define
\begin{equation*}
\mathcal{P}_{j}^{\bot }\left( W\right) =\mathcal{U}_{j,0}\mathcal{U}
_{j,0}^{\prime }W\mathcal{V}_{j,0}\mathcal{V}_{j,0}^{\prime },\quad \mathcal{
P}_{j}\left( W\right) =W-\mathcal{P}_{j}^{\bot }\left( W\right) ,
\end{equation*}
where $\mathcal{P}_{j}\left( W\right) $ can be seen as the linear projection
of matrix $W$ into the low-rank space with $\mathcal{P}_{j}^{\bot }\left(
W\right) $ being its orthogonal space. Let $\Delta _{\Theta _{j}}=\Theta
_{j}-\Theta _{j}^{0}$ for any $\Theta _{j}$. Based on the spaces constructed
above, with some positive constants $C_{1}$ and $C_{2}$, we define the
restricted set for full-sample parameters as follows:
\begin{align}
\mathcal{R}(C_{1},C_{2})& :=\bigg\{(\{\Delta _{\Theta _{j}}\}_{j\in \lbrack
p]\cup \{0\}}):\sum_{j\in \lbrack p]\cup \{0\}}\left\Vert \mathcal{P}
_{j}^{\bot }(\Delta _{\Theta _{j}})\right\Vert _{\ast }\leq C_{1}\sum_{j\in
\lbrack p]\cup \{0\}}\left\Vert \mathcal{P}_{j}(\Delta _{\Theta
_{j}})\right\Vert _{\ast },  \notag  \label{RS} \\
& \sum_{j\in \lbrack p]\cup \{0\}}\left\Vert \Theta _{j}\right\Vert ^{2}\geq
C_{2}\sqrt{NT}\bigg\}.
\end{align}

Lemma \ref{Lem:RS} in the online appendix shows that our nuclear norm
estimators are in a restricted set larger than (\ref{RS}), which derives from
the restriction on the Frobenius norm in the definition of $\mathcal{R}
\left( C_{1},C_{2}\right) $. Intuitively, the first restriction in (\ref{RS}
)\ means the projection onto the orthogonal low-rank space of the estimator
error can be controlled by its projection onto the low-rank space. Theorem
\ref{Thm1} below largely hinges on this property.

\begin{ass}
\label{ass:4} For any $C_{2}>0$, there are constants $C_{3}$ and $C_{4}$
such that for any $(\{\Delta _{\Theta _{j}}\}_{j\in \lbrack p]\cup \{0\}})$ $
\in \mathcal{R}(3,C_{2})$, we have
\begin{equation*}
\left\Vert \Delta _{\Theta _{0}}+\sum_{j=1}^{p}\Delta _{\Theta _{j}}\odot
X_{j}\right\Vert ^{2}\geq C_{3}\sum_{j\in \lbrack p]\cup \{0\}}\left\Vert
\Delta _{\Theta _{j}}\right\Vert ^{2}-C_{4}(N+T)\quad \text{w.p.a.1.}
\end{equation*}
\end{ass}

\noindent Assumption \ref{ass:4} imposes the restricted strong convexity (RSC)
condition, which is similar to Assumption 3.1 in \cite
{chernozhukov2019inference}. The latter authors also provide some sufficient
conditions to verify such an assumption.

Let $r=\sum_{j\in \lbrack p]\cup \{0\}}r_{j}.$ Define the following $r\times
r$ matrices:
\begin{equation*}
\Phi _{i}=\frac{1}{T}\sum_{t=1}^{T}\phi _{it}^{0}\phi _{it}^{0\prime
}~~\forall i\in \lbrack N]\text{ and }\Psi _{t}=\frac{1}{N}\sum_{i\in
\lbrack N]}\psi _{it}^{0}\psi _{it}^{0\prime }~~\forall t\in \lbrack T],
\end{equation*}
where $\phi _{it}^{0}=(v_{t,0}^{0\prime },v_{t,1}^{0\prime }X_{1,it},\cdots
,v_{t,p}^{0\prime }X_{p,it})^{\prime },$ and $\psi
_{it}^{0}=(u_{i,0}^{0\prime },u_{i,1}^{0\prime }X_{1,it},\cdots
,u_{i,p}^{0\prime }X_{p,it})^{\prime }.$

\begin{ass}
\label{ass:5} There exist constants $C_{\phi}$ and $c_{\phi}$ such that
\begin{align*}
&\infty>C_{\phi}\geq \limsup_{T}\max_{t\in[T]}\lambda_{\max}(\Psi_{t})\geq
\liminf_{T}\min_{t\in[T]}\lambda_{\min}(\Psi_{t})\geq c_{\phi}>0, \\
&\infty>C_{\phi}\geq \limsup_{N}\max_{i\in
[N]}\lambda_{\max}(\Phi_{i})\geq\liminf_{N}\min_{i\in
[N]}\lambda_{\min}(\Phi_{i})\geq c_{\phi}>0.
\end{align*}
\end{ass}
\noindent Assumption \ref{ass:5} is similar to Assumption 8 in \cite{ma2020detecting}
and it imposes some rank conditions.


\subsection{Asymptotics of NNR Estimators and Singular Vector
Estimators}

Let $\eta _{N,1}=\frac{\sqrt{\log T}}{\sqrt{N\wedge T}}$ and $\eta _{N,2}=
\frac{\sqrt{\log (N\vee T)}}{\sqrt{N\wedge T}}(NT)^{1/q}.$ Let $\tilde{\sigma
}_{k,j}$ denote the $k$-th largest singular value of $\tilde{\Theta}_{j}$
for $j\in \lbrack p]\cup \{0\}.$ Our first main result is about the
consistency of the first-stage NNR estimators and the second-stage singular
vector estimators.

\begin{theorem}
\label{Thm1} Suppose that Assumptions \ref{ass:1}--\ref{ass:4} hold. Then $
\forall j\in \lbrack p]\cup \{0\}$, we have

\begin{itemize}
\item[(i)] $\frac{1}{\sqrt{NT}}||\tilde{\Theta}_{j}-\Theta
_{j}^{0}||=O_{p}(\eta _{N,1})$, $\max_{k\in \lbrack r_{j}]}|\tilde{\sigma}
_{k,j}-\sigma _{k,j}|=O_{p}(\eta _{N,1})$, and $||V_{j}^{0}-\tilde{V}
_{j}O_{j}||=O_{p}(\sqrt{T}\eta _{N,1})$ where $O_{j}$ is an orthogonal
matrix defined in the proof.

If in addition Assumption \ref{ass:5} is also satisfied, then we have

\item[(ii)] $\operatornamewithlimits{\max}\limits_{i\in \lbrack N]}||\dot{u}
_{i,j}-O_{j}u_{i,j}^{0}||=O_{p}(\eta _{N,2}),$ $\operatornamewithlimits{\max}
\limits_{t\in \lbrack T]}\left\Vert \dot{v}_{t,j}-O_{j}v_{t,j}^{0}\right
\Vert _{2}=O_{p}(\eta _{N,2})$,

\item[(iii)] $\operatornamewithlimits{\max}\limits_{i\in \lbrack N],t\in
\lbrack T]}|\dot{\Theta}_{j,it}-\Theta _{j,it}^{0}|=O_{p}(\eta _{N,2})$.
\end{itemize}
\end{theorem}

Theorem \ref{Thm1}(i) reports the error bounds for $\tilde{\Theta}_{j},$ $
\tilde{\sigma}_{k,j},$ and $\tilde{V}_{j}$. The $\log T$ term in the
numerator of $\eta _{N,1}$ is due to the use of some exponential inequality
for (conditional)\ strong mixing processes. Theorem \ref{Thm1}(ii)--(iii)
report the uniform convergence rate of the factor and factor loading
estimators. The extra $(NT)^{1/q}$ term in the definition of $\eta _{N,2}$
is due to the nonboundedness of $X_{j,it}$ in Assumption \ref{ass:1}(v), and
it disappears when $X_{j,it}$ is assumed to be uniformly bounded.

\subsection{Consistency of the Break Point Estimate}

Recall that $g_{i}^{(1)}$ and $g_{i}^{(2)}$ denote the true group individual
$i$ belongs to before and after the break, respectively. To estimate the
break point consistently, we add the following condition.

\begin{ass}
\label{ass:6}

\begin{enumerate}
\item[(i)] $\sqrt{\frac{1}{N}\sum_{i\in \lbrack N]}||\alpha
_{g_{i}^{(1)}}-\alpha _{g_{i}^{(2)}}||^{2}}=C_{5}\zeta _{NT}$, where $C_{5}$
is a positive constant and $\zeta _{NT}\gg \eta _{N,2}$.

\item[(ii)] $\tau_{T}:=\frac{T_{1}}{T}\rightarrow \tau \in (0,1)$ as $
T\rightarrow \infty.$
\end{enumerate}
\end{ass}

Assumption \ref{ass:6}(i) imposes conditions on the break size in order to
identify the break point. Note that we allow the average break size to
shrink to zero at a rate slower than $\sqrt{\frac{\log (N\vee T)}{N\wedge T
}}(NT)^{1/q}$. This rate is of much bigger magnitude than the optimal $
\left( NT\right) ^{-1/2}$-rate that can be detected in the panel threshold
regressions (PTRs) for several reasons. First, in PTRs, the slope
coefficients are usually assumed to be homogeneous so that each individual
is subject to the same change in the slope coefficients and one can use the
cross-sectional information effectively. In contrast, we allow for
heterogeneous slope coefficients here and the change can occur only for a
subset of cross section units but not all. In addition, in the presence of
latent group structures, we not only allow the slope coefficients of some
specific groups to change with group membership fixed, but also allow the
slope coefficient to remain the same for some groups while the group
memberships change after the break. Second, our break point estimation
relies on the binary segmentation idea borrowed from the time series
literature where one can allow break sizes of bigger magnitude than $
T^{-1/2} $ in order to identify the break ratio consistently but not the
break point consistently. As is apparent, even though we require bigger break
sizes, we can estimate the break date consistently by using information from
both the cross section and time dimensions. Third, as mentioned above, the
additional term $\log (N\vee T)$ in the above rate is mainly due to the use
of an exponential inequality and the factor $(NT)^{1/q}$ is due to the fact
that we only assume the existence of $q$-th order moments for some random
variables.

The following theorem indicates that we can estimate the break date $T_{1}$
consistently.

\begin{theorem}
\label{Thm2} Suppose Assumptions \ref{ass:1}--\ref{ass:6} hold, with the
true break point being $T_{1}$ and the estimator defined in (\ref
{BP_estimates}). Then $\mathbb{P}(\hat{T}_{1}=T_{1})\rightarrow 1$ as $
\left( N,T\right) \rightarrow \infty $.
\end{theorem}

Theorem \ref{Thm2} shows that we can estimate the true break date
consistently w.p.a.1 despite the fact that we allow the break size to shrink
to zero at a certain rate.

\subsection{Consistency of the Estimates of the Number of Groups and the
Latent Group Structures}

To study the asymptotic properties of the estimates of the number of groups
and the recovery of the latent group structures, we first add the following
definition.

\begin{definition}
\label{def:1} Fix $K^{(\ell)}>1$ and $m\leq K^{(\ell)}$. The estimated group
structure $\hat{\mathcal{G}}_{m}^{(\ell)}$ satisfies the non-splitting
property (NSP) if for any pair of individuals in the same true group, the
estimated group labels are the same.
\end{definition}

\noindent Definition \ref{def:1} describes the non-splitting property introduced by
\cite{Jin2022Optimal}. The latter authors show that the STK algorithm yields
the estimated group structures enjoying the NSP.

To proceed, we add the following assumptions.

\begin{ass}
\label{ass:7}

\begin{enumerate}
\item[(i)] Let $k_{s}$ and $k_{s^{\ast }}$ be different group indices.
Assume that $\min_{1\leq k_{s}<k_{s^{\ast }}\leq K^{(\ell )}}\left\Vert
\alpha _{k_{s}}^{(\ell )}-\alpha _{k_{s^{\ast }}}^{(\ell )}\right\Vert _{2}$
$\geq C_{5}$ for$~\ell \in \{1,2\}.$

\item[(ii)] Let $N_{k}^{(\ell )}$ be the number of individuals in group $k\ $
for $k\in \lbrack K^{(\ell )}]$. Define $\pi _{k}^{(\ell )}=\frac{
N_{k}^{(\ell )}}{N}$ for $\ell =1,2.$ Assume $K^{(\ell )}$ is fixed and
\begin{equation*}
\infty >\bar{C}\geq \lim \sup_{N}\sup_{k\in \lbrack K^{(\ell )}]}\pi
_{k}^{(\ell )}\geq \lim \inf_{N}\inf_{k\in \lbrack K^{(\ell )}]}\pi
_{k}^{(\ell )}\geq \underline{c}>0,~\ell =1,2.
\end{equation*}

\item[(iii)] For any combination of the collection of $n$ true groups with $
n\geq 2$, we have
\begin{equation*}
\frac{T_{\ell }}{\sqrt{N}}\sum_{s=1}^{n}N_{k_{s}}^{(\ell )}\left\Vert
\sum_{s^{\ast }\in \left[ n\right] ,s^{\ast }\neq s}(\alpha _{k_{s^{\ast
}}}^{(\ell )}-\alpha _{k_{s}}^{(\ell )})\right\Vert ^{2}/(\log
N)^{1/2}\rightarrow \infty ,~\ell =1,2.
\end{equation*}
\end{enumerate}
\end{ass}

\noindent \textbf{Remark 3.} Assumption \ref{ass:7}(i)--(ii) is the standard
assumption for K-means algorithm, which is comparable to Assumption 4 in
\cite{su2020strong} and greatly facilitates the subsequent analyses.
Assumption \ref{ass:7}(i) assumes that the minimum distance of two distinct
groups is bounded away from 0, and Assumption \ref{ass:7}(ii) imposes that
each group has asymptotically non-negligible number of units. For Assumption
\ref{ass:7}(iii), it can be shown to hold under mild conditions. Below we
explain this assumption in detail. When $n=2$, it is clear that
\begin{align*}
\frac{T_{\ell }}{\sqrt{N}}\sum_{s=1}^{n}N_{k_{s}}^{(\ell )}\left\Vert
\sum_{s^{\ast }\in \left[ n\right] ,s^{\ast }\neq s}(\alpha _{k_{s^{\ast
}}}^{(\ell )}-\alpha _{k_{s}}^{(\ell )})\right\Vert ^{2}& =\frac{T_{\ell }}{
\sqrt{N}}\left( N_{k_{1}}^{(\ell )}\left\Vert \alpha _{k_{2}}^{(\ell
)}-\alpha _{k_{1}}^{(\ell )}\right\Vert ^{2}+N_{k_{2}}^{(\ell )}\left\Vert
\alpha _{k_{1}}^{(\ell )}-\alpha _{k_{2}}^{(\ell )}\right\Vert ^{2}\right) \\
& \geq \frac{C_{5}^{2}T_{\ell }(N_{k_{1}}^{(\ell )}+N_{k_{2}}^{(\ell )})}{
\sqrt{N}}=O(T\sqrt{N})
\end{align*}
by Assumptions \ref{ass:6}(ii) and \ref{ass:7}(i)--(ii). When $n>2$, we
consider a special case such that $S_{s}=:||\sum_{s^{\ast }\in \left[ n
\right] ,s^{\ast }\neq s}(\alpha _{k_{s^{\ast }}}^{(\ell )}-\alpha
_{k_{s}}^{(\ell )})||=0$ for some specific $s=s_{0}\in \lbrack n]$. Then it
is easy to see $S_{s}$ is non-zero for all $s\in \lbrack n]\backslash
\{s_{0}\}$ under Assumption \ref{ass:7}(i). Hence, if we assume $S_{s}$ is
lower bounded by a constant $c$ for any $s\in \lbrack n]\backslash \{s_{0}\}$
, Assumption \ref{ass:7}(iii) will holds naturally. Similar arguments hold
for the other general cases.

\begin{ass}
\label{ass:8} Let $\mathcal{T}_{1}=\left[ T_{1}\right] \ $and $\mathcal{T}
_{2}=[T]\backslash \left[ T_{1}\right] $.

\begin{itemize}
\item[(i)] $\frac{1}{T_{\ell }}\sum_{t\in \mathcal{T}_{\ell
}}f_{t}^{0}f_{t}^{0\prime }\overset{p}{\longrightarrow }\Sigma _{F}^{(\ell
)}>0$ as $T\rightarrow \infty $. $\frac{1}{N_{k}^{(\ell )}}\Lambda
_{k}^{0,\left( \ell \right) \prime }\Lambda _{k}^{0,\left( \ell \right) }
\overset{p}{\longrightarrow }\Sigma _{\Lambda ,k}^{\left( \ell \right) }>0$
as $N\rightarrow \infty $, where $\Lambda _{k}^{0,\left( \ell \right) }$ is
a stack of $\lambda _{i}^{0}$ for all individuals in group $k\ $and $k\in
\lbrack K^{(\ell )}]$.

\item[(ii)] There exists a constant $C>0$ such that $\max_{i}\frac{1}{
T_{\ell }}\sum_{t\in \mathcal{T}_{\ell }}\left\Vert \xi _{it}\right\Vert
^{2}\leq C~a.s.$ for $\xi _{it}=e_{it},$ $X_{it}$ and $X_{it}e_{it}.$
\end{itemize}
\end{ass}



\noindent Assumption \ref{ass:8}(i) imposes some standard assumptions on the factors
and factor loadings. Assumption \ref{ass:8}(ii) is similar as Assumption \ref
{ass:1}(iv), which strengthens Assumption \ref{ass:1}(iv) to hold for two
time regimes.

The next theorem reports the asymptotic properties of the STK estimators.

\begin{theorem}
\label{Thm3} Let $\varsigma =\varsigma _{N}\rightarrow 0$ at rate $N^{-c}$
for some $c>0$ as $N\rightarrow \infty .$ Suppose that Assumption \ref{ass:1*}
and Assumptions \ref{ass:2}--\ref{ass:8} hold. Then for $\ell \in
\{1,2\}$, we have

\begin{itemize}
\item[(i)] if $m=K^{(\ell )}$,

\begin{itemize}
\item[(a)] $\operatornamewithlimits{\max}\limits_{i\in \lbrack N]}\mathbf{1}
\{\hat{g}_{i,K^{(\ell )}}^{(\ell )}\neq g_{i}^{(\ell )}\}=0$ w.p.a.1,

\item[(b)] $\hat{\Gamma}_{K^{(\ell )}}^{(\ell )}$ is asymptotically
distributed as the maximum of $K^{(\ell )}$ independent $\chi ^{2}(1)$
random variables,

\item[(c)] $\mathbb{P}(\hat{K}^{(\ell )}\leq K^{(\ell )})\geq 1-\alpha +o(1)$
,
\end{itemize}

\item[(ii)] if $m<K^{(\ell )}$, $\hat{\Gamma}_{m}^{(\ell )}/\log
N\rightarrow \infty $ w.p.a.1. Thus $\mathbb{P}(\hat{K}^{(\ell )}\neq
K^{(\ell )})\leq \varsigma +o(1)$.
\end{itemize}
\end{theorem}

Theorem \ref{Thm3} studies the asymptotic properties of the STK algorithm.
Since we allow $\varsigma =\varsigma _{N}$ to shrink to zero at rate $N^{-c},
$ the critical value $z_{\varsigma }$ diverges to infinity at rate $\log N$
as $N\rightarrow \infty $ by virtue of the tail properties of $\chi ^{2}(1)$ random
variables. At iteration $m$ such that $m<K^{(\ell )}$, w.p.a.1, the test
statistics $\hat{\Gamma}_{m}^{(\ell )}$ diverges to infinity at a rate
faster than $\log N$, which means the iteration will continue at the $(m+1)$
-th iteration. At iteration $m$ with $m=K^{(\ell )}$, however, we can easily
find that $z_{\varsigma }\rightarrow \infty $ while the test statistic $
\hat{\Gamma}_{m}^{(\ell )}$ is stochastically bounded. As a result, the
iteration stops w.p.a.1 and we have $\mathbb{P}(\hat{K}^{(\ell )}=K^{(\ell
)})\rightarrow 1$. As aforementioned, Theorem \ref{Thm3} ensures the
application of K-means algorithm only for the under-fitting and just-fitting
cases and it avoids the theoretical challenge of handling the over-fitting
case in the classification.


For dynamic panels, we can focus on Assumption \ref{ass:1*}, where the error
term is an m.d.s. Under this assumption, the HAC estimator $\hat{\Omega}
_{i,k,m}^{(\ell )}$ degenerates to $\frac{1}{\hat{T}_{\ell }}\sum_{t\in \hat{
\mathcal{T}}_{\ell }}\hat{z}_{it}^{(\ell )}\hat{z}_{it}^{(\ell )\prime }\hat{
e}_{it}^{2}$. For static panels, we typically allow for serially correlated
errors and employ the HAC estimator, and the results in Theorem \ref{Thm3}
continue to hold with Assumption \ref{ass:1*} replaced by Assumption \ref
{ass:1}. For the kernel function and bandwidth, we can follow \cite
{andrews1991heteroskedasticity} and let $k\left( \cdot \right) $ belong to
the following class of kernels
\begin{align*}
\mathcal{K}=\bigg\{& k(\cdot ):\mathbb{R}\mapsto \lbrack
-1,1]~|~k(0)=1,~k(u)=k(-u),~\int \left\vert k\left( u\right) \right\vert
du<\infty , \\
& k(\cdot )\text{ is continuous at 0 and at all but a finite number of other
points}\bigg\}.
\end{align*}
See, e.g., \cite{andrews1991heteroskedasticity} and \cite
{white2014asymptotic} for details.

\subsection{Distribution Theory for the Group-specific Slope Estimators}

For $\ell \in \{1,2\}$, let $\{\hat{\alpha}_{k}^{\ast (\ell )}\}_{k\in
K^{(\ell )}}$ be the oracle estimators of the group-specific slope
coefficients before and after the break point by using the true break and
membership information for all individuals. To study the asymptotic
distribution theory for $\{\hat{\alpha}_{k}^{(\ell )}\}_{k\in K^{(\ell
)}},~\ell \in \{1,2\}$, we only need to show that for the oracle estimators $
\{\hat{\alpha}_{k}^{\ast (\ell )}\}_{k\in K^{(\ell )}}$ based on Theorems
\ref{Thm2} and \ref{Thm3} by extending the result of \cite{bai2009panel} and
\cite{moon2017dynamic}.

To proceed, we add some notation. For $\ell \in \{1,2\}$, we first define
the matrix notation for each subgroup. For $j\in \lbrack p]$, let $
X_{j,i}^{(1)}=\left( X_{j,i1},\cdots ,X_{j,iT_{1}}\right) ^{\prime }$, $
X_{j,i}^{(2)}=\left( X_{j,i(T_{1}+1)},\cdots ,X_{j,iT}\right) ^{\prime }$, $
e_{i}^{(1)}=\left( e_{i1},\cdots ,e_{iT_{1}}\right) ^{\prime }$ and $
e_{i}^{(2)}=\left( e_{i(T_{1}+1)},\cdots ,e_{iT}\right) ^{\prime }$. Then we
use $\mathbb{X}_{j,k}^{(\ell )}\in \mathbb{R}^{N_{k}^{(\ell )}\times T_{\ell
}}$ and $E_{k}^{(\ell )}\in \mathbb{R}^{N_{k}^{(\ell )}\times T_{\ell }}$ to
denote the regressor and error matrix for subgroup $k\in \lbrack K^{(\ell )}]
$ with each row being $X_{j,i}^{(\ell )}$ and $e_{i}^{(\ell )}$,
respectively. Let $\mathcal{X}_{j,k}^{(\ell )}=M_{\Lambda _{k}^{0,(\ell )}}
\mathbb{X}_{j,k}^{(\ell )}M_{F^{0,(\ell )}}\in \mathbb{R}^{N_{k}^{(\ell
)}\times T_{\ell }}$ with the $\left( i,t\right) $-th entry given by $
\mathcal{X}_{j,k,it}^{(\ell )}.$ Let $\mathcal{X}_{k,it}^{(\ell )}=(\mathcal{
X}_{1,k,it}^{(\ell )},\cdots ,\mathcal{X}_{p,k,it}^{(\ell )})^{\prime }$.
Further define
\begin{align*}
& \mathbb{B}_{NT,1,j,k}^{(\ell )}=\frac{1}{N_{k}^{(\ell )}}tr\left[
P_{F^{0,(\ell )}}\mathbb{E}\left( E_{k}^{(\ell )\prime }\mathbb{X}
_{j,k}^{(\ell )}\big|\mathscr{D}\right) \right] , \\
& \mathbb{B}_{NT,2,j,k}^{(\ell )}=\frac{1}{T_{\ell }}tr\left[ \mathbb{E}
\left( E_{k}^{(\ell )}E_{k}^{(\ell )\prime }\big|\mathscr{D}\right)
M_{\Lambda _{k}^{0,(\ell )}}\mathbb{X}_{j,k}^{(\ell )}F^{0,(\ell )}\left(
F^{0,(\ell )\prime }F^{0,(\ell )}\right) ^{-1}\left( \Lambda _{k}^{0,(\ell
)\prime }\Lambda _{k}^{0,(\ell )}\right) ^{-1}\Lambda _{k}^{0,(\ell )\prime }
\right] , \\
& \mathbb{B}_{NT,3,j,k}^{(\ell )}=\frac{1}{N_{k}^{(\ell )}}tr\left[ \mathbb{E
}\left( E_{k}^{(\ell )}E_{k}^{(\ell )\prime }\big|\mathscr{D}\right)
M_{F^{0,(\ell )}}\mathbb{X}_{j,k}^{(\ell )}\Lambda _{k}^{0,(\ell )}\left(
\Lambda _{k}^{0,(\ell )\prime }\Lambda _{k}^{0,(\ell )}\right) ^{-1}\left(
F^{0,(\ell )\prime }F^{0,(\ell )}\right) ^{-1}F^{0,(\ell )\prime }\right] ,
\\
& \mathbb{B}_{NT,m,k}^{(\ell )}=\left( \mathbb{B}_{NT,m,1,k}^{(\ell
)},\cdots ,\mathbb{B}_{NT,m,p,k}^{(\ell )}\right) ^{\prime },\quad \forall
m\in \{1,2,3\}, \\
& \Omega _{k}^{(\ell )}=\frac{1}{N_{k}^{(\ell )}T_{\ell }}\sum_{i\in
G_{k}^{(\ell )}}\sum_{t\in \mathcal{T}_{\ell }}e_{it}^{2}\mathcal{X}
_{k,it}^{(\ell )}\mathcal{X}_{k,it}^{(\ell )\prime }.
\end{align*}
Let $\mathbb{W}_{NT,k}^{(\ell )}$ be a $p\times p$ matrix with $(j_{1},j_{2})
$-th entry $\frac{1}{N_{k}^{(\ell )}T_{\ell }}tr\left( M_{F^{0,(\ell
)}}\mathbb{X}_{j_{1},k}^{(\ell )\prime }M_{\Lambda _{k}^{0,(\ell )}}\mathbb{X
}_{j_{2},k}^{(\ell )}\right) $. Then we define the overall bias term for
each subgroup as
\begin{equation*}
\mathbb{B}_{NT,k}^{(\ell )}=-\rho _{k}^{(\ell )}\mathbb{B}_{NT,1,k}^{(\ell
)}-(\rho _{k}^{(\ell )})^{-1}\mathbb{B}_{NT,2,k}^{(\ell )}-\rho _{k}^{(\ell
)}\mathbb{B}_{NT,3,k}^{(\ell )},
\end{equation*}
where $\rho _{k}^{(\ell )}=\sqrt{\frac{N_{k}^{(\ell )}}{T_{\ell }}}$. To
state the main result in this subsection, we add the following assumption.

\begin{ass}
\label{ass:10}

\begin{itemize}
\item[(i)] As $(N,T)\to\infty$, $T(\log T)N^{-4/3}\to 0$.

\item[(ii)] $\operatorname*{plim}_{(N,T)\rightarrow \infty }\frac{1}{N_{k}^{(\ell )}T_{\ell
}}\sum_{i\in G_{k}^{(\ell )}}\sum_{t\in \mathcal{T}_{\ell
}}X_{it}X_{it}^{\prime }>0$\ for $\ell \in \{1,2\}$ and $k\in \lbrack
K^{(\ell )}]$.

\item[(iii)] For $\ell\in\{1,2\}$ and $k\in[K^{(\ell)}]$, separate the $p$
regressors of each subgroups into $p_{1}$ ``low-rank regressors" $\mathbb{X}
_{j,k}^{(\ell)}$ such that $rank(\mathbb{X}_{j,k}^{(\ell)})=1,~\forall
j\in\{1,\cdots,p_{1}\},$ and ``high-rank regressors" $\mathbb{X}
_{j,k}^{(\ell)}$ such that $rank(\mathbb{X}_{j,k}^{(\ell)})>1,~\forall
j\in\{p_{1}+1,\cdots,p\}.$ Let $p_{2}:=p-p_{1}$. These two types of
regressors satisfy:

\begin{itemize}
\item[(iii.a)] Consider the linear combinations $b\cdot \mathbb{X}
_{high,k}^{(\ell )}:=\sum_{j=p_{1}+1}^{p}b_{j}\mathbb{X}_{j,k}^{(\ell )}$
for high-rank regressors with $p_{2}$-vectors $b$ such that $\left\Vert
b\right\Vert _{2}=1$ and $b=\left( b_{p_{1}+1},\cdots ,b_{p}\right) ^{\prime
}$. There exists a positive constant $C_{b}$ such that
\begin{equation*}
\operatornamewithlimits{\min}_{\left\{ \left\Vert b\right\Vert
_{2}=1\right\} }\sum_{n=2r_{0}+p_{1}+1}^{N}\lambda _{n}\left[ \frac{1}{
NT_{\ell }}\left( b\cdot \mathbb{X}_{high,k}^{(\ell )}\right) \left( b\cdot
\mathbb{X}_{high,k}^{(\ell )}\right) ^{\prime }\right] \geq C_{b}\quad
w.p.a.1.
\end{equation*}

\item[(iii.b)] For $j\in \lbrack p_{1}]$, write $\mathbb{X}_{j,k}^{(\ell
)}=w_{j,k}^{(\ell )}v_{j,k}^{(\ell )\prime }$ with $N_{k}^{(\ell )}$-vectors
$w_{j}^{(\ell )}$ and $T_{\ell }$-vectors $v_{j}^{(\ell )}$. Let $
w_{k}^{(\ell )}=(w_{1,k}^{(\ell )},\cdots ,w_{p_{1},k}^{(\ell )})\in \mathbb{
R}^{N\times p_{1}}$, $v^{(\ell )}=(v_{1}^{(\ell )},\cdots ,v_{p_{1}}^{(\ell
)})\in \mathbb{R}^{T_{\ell }\times p_{1}}$, $M_{w_{k}^{(\ell
)}}=I_{N_{k}^{(\ell )}}-w_{k}^{(\ell )}(w_{k}^{(\ell )\prime }w_{k}^{(\ell
)})^{-1}w_{k}^{(\ell )\prime }$ and $M_{v^{(\ell )}}=I_{T_{\ell }}-v^{(\ell
)}\left( v^{(\ell )\prime }v^{(\ell )}\right) ^{-1}v^{(\ell )\prime }$.
There exists a positive constant $C_{B}$ such that $(N_{k}^{(\ell
)})^{-1}\Lambda _{k}^{0,(\ell )\prime }M_{w_{k}^{(\ell )}}\Lambda
_{k}^{0,(\ell )}>C_{B}I_{r_{0}}$ and $T_{\ell }^{-1}F^{0,(\ell )\prime
}M_{w^{(\ell )}}F^{0,(\ell )}>C_{B}I_{r_{0}}$ w.p.a.1.
\end{itemize}

\item[(iv)] For $\forall j\in \lbrack p]$, $\ell \in \{1,2\}$, and $k\in
K^{(\ell )}$,
\begin{equation*}
\frac{1}{N_{k}^{(\ell )}\left( T_{\ell }\right) ^{2}}\sum_{i\in G_{k}^{(\ell
)}}\sum_{t_{1}\in \mathcal{T}_{\ell }}\sum_{t_{2}\in \mathcal{T}_{\ell
}}\sum_{s_{1}\in \mathcal{T}_{\ell }}\sum_{s_{2}\in \mathcal{T}_{\ell
}}\left\vert Cov\left( e_{it_{1}}\tilde{X}_{j,it_{2}},e_{is_{1}}\tilde{X}
_{j,is_{2}}\right) \right\vert =O_{p}(1),
\end{equation*}
where $\tilde{X}_{j,it}=X_{j,it}-\mathbb{E}\left( X_{j,it}|\mathscr{D}
\right) $.
\end{itemize}
\end{ass}

Assumption \ref{ass:10} imposes some conditions to help derive the asymptotic
distribution theory for the panel model with IFEs which allows for dynamics.
Assumption \ref{ass:10}(i) slightly strengthens Assumption \ref{ass:1}(vi).
Assumption \ref{ass:10}(ii) is the standard non-collinearity condition for
regressors, which is analogous to Assumption 4(i) in \cite{moon2017dynamic}.
Assumption \ref{ass:10}(iii) is the identification assumption which is
comparable to Assumption 4 in \cite{moon2017dynamic}. With the conditional
strong mixing condition in Assumption \ref{ass:1}(iii), we can verify
Assumption \ref{ass:10}(iv).

The following theorem establishes the asymptotic distribution of $\{\hat{
\alpha}_{k}^{(\ell )}\}_{k\in K^{(\ell )}}.$

\begin{theorem}
\label{Thm4} Suppose that Assumption \ref{ass:1} or \ref{ass:1*} and
Assumptions \ref{ass:2}--\ref{ass:10} hold. For $\ell \in \{1,2\}$, the
estimators $\{\hat{\alpha}_{k}^{(\ell )}\}_{k\in K^{(\ell )}}$ are
asymptotically equivalent to the oracle estimators $\{\hat{\alpha}_{k}^{\ast
(\ell )}\}_{k\in K^{(\ell )}}$, and we have
\begin{equation*}
\mathbb{W}_{NT}^{(\ell )}\mathbb{D}_{NT}^{(\ell )}
\begin{pmatrix}
\hat{\alpha}_{1}^{(\ell )}-\alpha _{1}^{(\ell )} \\
\vdots \\
\hat{\alpha}_{K^{(\ell )}}^{(\ell )}-\alpha _{K^{(\ell )}}^{(\ell )}
\end{pmatrix}
-\mathbb{B}_{NT}^{(\ell )}\rightsquigarrow \mathbb{N}\left( 0,\Omega ^{(\ell
)}\right) ,
\end{equation*}
such that $\mathbb{D}_{NT}^{(\ell )}=\text{diag}\left( \sqrt{N_{1}^{(\ell
)}T_{\ell }},\cdots ,\sqrt{N_{K^{(\ell )}}^{(\ell )}T_{\ell }}\right) $, $
\mathbb{W}_{NT}^{(\ell )}=\text{diag}\left( \mathbb{W}_{NT,1}^{(\ell
)},\cdots ,\mathbb{W}_{NT,K^{(\ell )}}^{(\ell )}\right) $, $\mathbb{B}
_{NT}^{(\ell )}=\text{diag}\left( \mathbb{B}_{NT,1}^{(\ell )},\cdots ,
\mathbb{B}_{NT,K^{(\ell )}}^{(\ell )}\right) $ and $\Omega ^{(\ell )}=\text{
diag}\left( \Omega _{1}^{(\ell )},\cdots ,\Omega _{K^{(\ell )}}^{(\ell
)}\right) $.
\end{theorem}

\noindent Theorem \ref{Thm4} establishes the asymptotic distribution for the
estimators of the group-specific slope coefficients before and after the
break. It shows that the parameter estimators from our algorithm enjoy the
oracle property given the results in Theorems \ref{Thm2} and \ref{Thm3}. The
proof of the above theorem can be done by following \cite{moon2017dynamic}
and \cite{lu2016shrinkage}.

\section{Alternatives and Extensions}

\label{sec:extension} This section first considers an alternative
method to estimate the break point and then discusses several possible
extensions.

\subsection{Alternative for Break Point Detection}

The algorithm proposed in Section \ref{sec:estimation} uses low-rank
estimates of $\Theta _{j}^{0}$ to find the break point estimates. However,
by Lemma \ref{Lem:idfct}(ii), we observe that the right singular vector
matrix of $\Theta _{j}^{0}$, i.e., $V_{j}^{0}$, contains the structural
break information when $r_{j}=2$. For this reason, we can propose an
alternative way to estimate the break point under the case where the maximum
rank of the slope matrix in the model is 2. Let $\dot{v}_{t,j}^{\ast }:=
\frac{\dot{v}_{t,j}}{\left\Vert \dot{v}_{t,j}\right\Vert }$ and $\dot{v}
_{t}^{\ast }:=\left( \dot{v}_{t,1}^{\ast \prime },\cdots ,\dot{v}
_{t,p}^{\ast \prime }\right) ^{\prime },$ with the true values being $
v_{t,j}^{\ast }:=\frac{O_{j}v_{t,j}^{0}}{\left\Vert
O_{j}v_{t,j}^{0}\right\Vert }$ and $v_{t}^{\ast }:=\left( v_{t,1}^{\ast
\prime },\cdots ,v_{t,p}^{\ast \prime }\right) ^{\prime }$, respectively.
Then Step 3 can be replaced by Step 3* below:

\begin{description}
\item[\textbf{Step 3*:}] \textbf{Break Point Estimation by Singular Vectors.}
We estimate the break point as follows:
\begin{equation}
\tilde{T}_{1}=\operatorname*{arg\,min}_{s\in \{2,\cdots ,T-1\}}\frac{1}{T}\left\{
\sum_{t=1}^{s}\left\Vert \dot{v}_{t}^{\ast }-\bar{\dot{v}}^{\ast
(1)s}\right\Vert ^{2}+\sum_{t=s+1}^{T}\left\Vert \dot{v}_{t}^{\ast }-\bar{
\dot{v}}^{\ast (2)s}\right\Vert ^{2}\right\} ,  \label{BP alternative}
\end{equation}
where $\bar{\dot{v}}^{\ast (1)s}=\frac{1}{s}\sum_{t=1}^{s}{\dot{v}_{t}}
^{\ast }$ and $\bar{\dot{v}}^{\ast (2)s}=\frac{1}{T-s}\sum_{t=s+1}^{T}{\dot{v
}}_{t}^{\ast }$.
\end{description}

The following two theorems establish the consistency of $\dot{v}_{t}^{\ast }$
and $\tilde{T}_{1},$ respectively.

\begin{theorem}
\label{Thm5} Suppose that Assumptions \ref{ass:1}--\ref{ass:5} hold. Then $
\max_{t}\left\Vert \dot{v}_{t}^{\ast }-v_{t}^{\ast }\right\Vert =O_{p}(\eta
_{N,2}).$
\end{theorem}

\begin{theorem}
\label{Thm6} Suppose that Assumptions \ref{ass:1}--\ref{ass:6} hold. Then $
\mathbb{P(}\tilde{T}_{1}=T_{1})\rightarrow 1$ as $\left( N,T\right)
\rightarrow \infty .$
\end{theorem}

\noindent Since the singular vectors of the slope matrices contain the structural
change information, Theorem \ref{Thm5} indicates that we can consistently
estimate the break point by using the factor estimates instead of the slope
matrix estimates in (\ref{BP_estimates}). Given Theorem \ref{Thm5} and Lemma
\ref{Lem:idfct}(iii), we can prove Theorem \ref{Thm6} with arguments
analogous to those used in the proof of Theorem \ref{Thm2}.

\subsection{Test for the Presence of a Structural Break}

\label{sec:extension_test} Section \ref{sec:model} considers
time-varying latent group structures with one break point. In this
subsection, we propose a test for the null that the slope coefficients are
time-invariant against the alternative that there's one structural break as
assumed in Section \ref{sec:model}.

Since various scenarios can occur once we allow for the presence of a
structural break in the latent group structures, the number of groups may
or may not change under the alternative and so may some of the
group-specific coefficients. First, information on the latent group structures may
be ignored and  the slope coefficients can be tested for possible time-variation. In this case, we can rewrite $\Theta
_{it}^{0}$ as
\begin{equation*}
\Theta _{it}^{0}=\Theta _{i}^{0}+c_{it},
\end{equation*}
where $\Theta _{i}^{0}:=\frac{1}{T}\sum_{t\in \lbrack T]}\Theta _{it}^{0}$.
The null and alternative hypotheses can then be specified as
\begin{eqnarray}
H_{0} &:&c_{it}=0~~\text{for all }i\in \lbrack N],\quad \text{and}  \notag
\label{BP_test} \\
H_{1} &:&c_{it}\neq 0~~\text{for some }i\in \lbrack N].
\end{eqnarray}
To construct the test statistics we follow \cite
{bai1998estimating} and consider a sup-$F$ test. Let $\mathcal{T}_{\epsilon
}:=\left\{ T_{1}:\epsilon T\leq T_{1}\leq (1-\epsilon )T\right\} ,$ where $
\epsilon >0$ is a tuning parameter that avoids breaks at the end of the
sample. Define
\begin{equation*}
F_{NT}(1|0):=\operatornamewithlimits{\max}\limits_{i\in \lbrack N]}
\operatornamewithlimits{\sup}\limits_{T_{1}\in \mathcal{T}_{\epsilon
}}F_{i}(T_{1}),
\end{equation*}
where
\begin{equation*}
F_{i}(T_{1})=\frac{T-2p}{p}\left[ \tilde{\beta}_{i}^{\left( 1\right)
}(T_{1})-\tilde{\beta}_{i}^{\left( 2\right) }\left( T_{1}\right) \right]
^{\prime }\left[ \hat{\Sigma}_{i}(T_{1})\right] ^{-1}\left[ \tilde{\beta}
_{i}^{\left( 1\right) }(T_{1})-\tilde{\beta}_{i}^{\left( 2\right) }\left(
T_{1}\right) \right] ,
\end{equation*}
$\tilde{\beta}_{i}^{\left( 1\right) }\left( T_{1}\right) $ and $\tilde{\beta}
_{i}^{\left( 2\right) }(T_{1})$ are the PCA slope estimators of $\Theta
_{i}^{0}$ in the linear panels with IFEs for each individual $i$ with the
prior-break observations $\left\{ (i,t):i\in \lbrack N],t\in \lbrack
T_{+}]\right\} $ and post-break observations $\left\{ (i,t):i\in \lbrack
N],t\in \lbrack T]\backslash \lbrack T_{+}]\right\} $, respectively,
\footnote{
See Section \ref{sec:panel_IFE} in the appendix for the detail of the PCA
estimation in linear panels with IFEs.} and $\hat{\Sigma}_{i}(T_{+})$ is the
consistent estimator for the asymptotic variance of $\tilde{\beta}
_{i}^{\left( 1\right) }(T_{1})-\tilde{\beta}_{i}^{\left( 2\right) }\left(
T_{1}\right) $. Following \cite{bai1998estimating}, we conjecture that the
asymptotic distribution of $\operatornamewithlimits{\sup}\limits_{T_{1}\in
\mathcal{T}_{\epsilon }}F_{i}(T_{1})$ depends on a $p$-vector of
Wiener processes on $\left[ 0,1\right] ,$ based on which the
corresponding distribution of $F_{NT}(1|0)$ can be obtained.

Alternatively, we can estimate the model with latent group structures by
assuming the presence of a break point at $T_{1}.$ Then we obtain the
estimates of the group-specific parameters $\{\alpha _{j}^{\left( 1\right)
}\left( T_{1}\right) \}_{j\in K^{\left( 1\right) }}$ prior to the potential
break point $T_{1}$ and those of the group-specific parameters $\{\alpha
_{j}^{\left( 2\right) }\left( T_{1}\right) \}_{j\in K^{\left( 2\right) }}$
after the potential break point $T_{1}.$ It is possible to construct a test
statistic based on the contrast of these two sets of estimates or the
corresponding residual sum of squares (RSS) and then take the supremum over $
T_{1}\in \mathcal{T}_{\epsilon }.$ Evidently, this approach is also
quite involved as it is necessary to determine the number of groups before and after
the break, $K^{\left( 1\right) }$ and $K^{\left( 1\right) },$ at each $
T_{1}. $ It is not clear how the estimation errors from these estimates and
those of the factors and factor loadings with slow convergence rates affect
the asymptotic properties of the estimators of the group-specific parameters.

Last, it is also possible to estimate the model with latent group structures
under the case of no structural changes to obtain the restricted residuals.
If there exists a structural change in the latent group structure, it should
be reflected in properties of the restricted residuals obtained under the null. We can then
consider the regression of the restricted residuals on the regressors
and construct an LM-type test statistic to check the goodness of fit of
such an auxiliary regression model as in \cite{su2013testing} and \cite
{su2020testing}. This analysis is left for future research.

\subsection{The Case of Multiple Breaks}

Section \ref{sec:model} considers only a one-time structural break in
the latent group structures. In practice it is possible to have multiple
breaks especially if $T$ is large. Here we generalize the model in Section
\ref{sec:model} to allow for multiple breaks. In this case, we have
\begin{equation*}
\alpha_{kt}=\left\{ \begin{aligned} &\alpha_{k}^{(1)}, \quad\text{for}\quad
t=1,\dots,T_{1},\\ & \alpha_{k}^{(2)}, \quad\text{for}\quad
t=T_{1}+1,\dots,T_{2},\\ &\vdots\\ & \alpha_{k}^{(b+1)},\quad\text{for}\quad
t=T_{b}+1,\dots,T,\\ \end{aligned}\right.
\end{equation*}
where $b\geq 1$ denotes the number of breaks.

To estimate the number of breaks and the break points $T_{1},\cdots ,T_{b}$,
in principle we can follow the sequential method proposed by \cite
{bai1998estimating}. First, using the full-sample data, we can construct $
F_{NT}(1|0)$ defined in the previous subsection and estimate the break point
as in \eqref{BP_estimates}. Second, for each regime before and after the
estimated break point, we test the hypothesis in \eqref{BP_test} and
estimate the break point for each regime separately. This sequential method is repeated
until the null can not be rejected for all subsamples, leading to the break point
estimates $\{\hat{T}_{a}\}_{a\in\lbrack \hat{b}]}$ where $\hat{b}$ is the estimated number of breaks. We
conjecture that we can establish the consistency of $\hat{b}$ and $\{\hat{T}
_{a}\}.$

After the estimated number of breaks and break points are obtained for each
subsample
\begin{equation*}
\left\{ (i,t):i\in [N],t\in \{\hat{T}_{a-1}+1,\cdots ,\hat{T}_{a}\}\right\},
\end{equation*}
where $a\in [\hat{b}+1]$ with $\hat{T}_{0}:=0$ and $\hat{T}_{\hat{b}+1}:=T$, we
can continue Step 4 in the estimation algorithm in Section \ref
{sec:estimation} to obtain the estimated group structure for each subsample.

\section{Monte Carlo Simulations}

\label{sec:simul} In this section we report simulation results for the
low-rank estimates, break point estimates, group membership estimates and
the group number estimates based on 1,000 replications, and tuning
parameter $\nu _{j}$ chosen by a procedure similar to that described in \cite
{chernozhukov2019inference}. We focus on the linear panel model $
Y_{it}=\lambda _{i}^{\prime }f_{t}+X_{it}^{\prime }\Theta _{it}+e_{it}$,
where $X_{it}=(X_{1,it},X_{2,it})^{\prime }$ and $\Theta _{it}=(\Theta
_{1,it},\Theta _{2,it})^{\prime }$.

\subsection{Data Generating Processes (DGPs)}

The following four main DGPs are employed.

\begin{description}
\item[\textbf{DGP 1:}] [\textbf{Static panel with homoskedasticity}] $
X_{1,it}\sim $ $i.i.d.$ $U(-2,2)$, $X_{2,it}\sim i.i.d.$ $U(-2,2)$, and errors
$e_{it}\sim i.i.d.$ $\mathbb{N}(0,1)$. For $\Theta _{1}$, we
randomly choose the break point $T_{1}$ from $0.4T$ to $0.6T$.

\item[\textbf{DGP 2:}] [\textbf{Static panel with heteroscedasticity}]
Compared to DGP 1, the errors $e_{it}\sim i.i.d.$ $\mathbb{N}
(0,\sigma _{it}^{2})$ with $\sigma _{it}^{2}\sim i.i.d.$ $U(0.5,1)$.
The settings for the regressors and break point are the same as those in DGP
1.

\item[\textbf{DGP 3:}] [\textbf{Serially correlated error}] Compared to
DGP 2, the errors $e_{it}=0.2e_{i,t-1}+\eta _{it},$ where $\eta
_{it}\sim i.i.d.$ $\mathbb{N}(0,1)$ and all other settings are the same as
in DGP 2.

\item[\textbf{DGP 4:}] [\textbf{Dynamic panel}] In this case, $
X_{1,it}=Y_{i,t-1}$ with $Y_{i,0}\sim i.i.d.~\mathbb{N}(0,1)$. $X_{2,it}\sim
i.i.d.~U(-2,2)$, and $e_{it}\sim i.i.d.~\mathbb{N}(0,0.5)$.
\end{description}

For each DGP above, we set $r_{0}=0$ and draw $\lambda _{i}$ and $f_{t}$
from $\mathbb{N}(0,1)$ independently. We define the slope coefficient based
on three subcases below.

\begin{description}
\item[\textbf{DGP X.1:}] In this case, the group membership and the number
of groups do not change after the break point and only the value of the slope
coefficient changes. We set the number of groups to be 2, the ratio of
individuals among the two groups is $N_{1}:N_{2}=0.5:0.5$, and the group
membership $G_{1}$ is obtained by a random draw from $[N]$ without
replacement. For DGPs 1.1, 2.1, and 3.1,
\begin{equation*}
\Theta _{1,it}=\Theta _{2,it}=\left\{ \begin{aligned} & 0.1, \quad i\in
G_{1},~t\in\{1,\cdots,T_{1} \},\\ & 0.9, \quad i\in
G_{2},~t\in\{1,\cdots,T_{1} \},\\ & 0.05, \quad i\in
G_{1},~t\in\{T_{1}+1,\cdots,T \},\\ & 0.45, \quad i\in
G_{2},~t\in\{T_{1}+1,\cdots,T \}. \end{aligned}\right.
\end{equation*}
For DGP 4.1, $\Theta _{2,it}$ is same as other DGPs X.1 for X$\in \{1,2,3\}$
, and
\begin{equation*}
\Theta _{1,it}=\left\{ \begin{aligned} & 0.1, \quad i\in
G_{1},~t\in\{1,\cdots,T_{1} \},\\ & 0.7, \quad i\in
G_{2},~t\in\{1,\cdots,T_{1} \},\\ & 0.05, \quad i\in
G_{1},~t\in\{T_{1}+1,\cdots,T \},\\ & 0.35, \quad i\in
G_{2},~t\in\{T_{1}+1,\cdots,T \}. \end{aligned}\right.
\end{equation*}

\item[\textbf{DGP X.2:}] Compared to DGP X.1, the values of the slope
coefficients for different groups do not change after the break point, but
the group membership changes. The number of groups is 2, the ratio of
individuals among the groups is still $N_{1}:N_{2}=0.5:0.5$.
Nevertheless, $\{G_{1}^{(1)},G_{2}^{(1)}\}$ is different from $
\{G_{1}^{(2)},G_{2}^{(2)}\}$ so that elements in both $G_{1}^{(1)}$ and $
G_{1}^{(2)}$ are independent draws from $[N]$ without replacement. In
addition, for DGPs 1.2, 2.2, and 3.2,
\begin{equation*}
\Theta _{1,it}=\Theta _{2,it}=\left\{ \begin{aligned} & 0.1, \quad i\in
G_{1}^{(1)},~t\in\{1,\cdots,T_{1} \},\\ & 0.9, \quad i\in
G_{2}^{(1)},~t\in\{1,\cdots,T_{1} \},\\ & 0.1, \quad i\in
G_{1}^{(2)},~t\in\{T_{1}+1,\cdots,T \},\\ & 0.9, \quad i\in
G_{2}^{(2)},~t\in\{T_{1}+1,\cdots,T \}. \end{aligned}\right.
\end{equation*}
For DGP 4.2, $\Theta _{2,it}$ is defined same as other DGPs X.2 for $X\in
\{1,2,3\}$, and
\begin{equation*}
\Theta _{1,it}=\left\{ \begin{aligned} & 0.1, \quad i\in
G_{1}^{(1)},~t\in\{1,\cdots,T_{1} \},\\ & 0.7, \quad i\in
G_{2}^{(1)},~t\in\{1,\cdots,T_{1} \},\\ & 0.1, \quad i\in
G_{1}^{(2)},~t\in\{T_{1}+1,\cdots,T \},\\ & 0.7, \quad i\in
G_{2}^{(2)},~t\in\{T_{1}+1,\cdots,T \}. \end{aligned}\right.
\end{equation*}

\item[\textbf{DGP X.3:}] Under this scenario, the number of groups changes
after the break. We set $N_{1}^{(1)}:N_{2}^{(1)}=0.5:0.5$ and $
N_{1}^{(2)}:N_{2}^{(2)}:N_{3}^{(2)}=0.4:0.3:0.3$ before and after the break,
respectively. Specifically, for DGPs 1.3, 2.3, and 3.3, we have
\begin{equation*}
\Theta _{1,it}=\Theta _{2,it}=\left\{ \begin{aligned} & 0.1, \quad i\in
G_{1}^{(1)},~t\in\{1,\cdots,T_{1} \},\\ & 0.9, \quad i\in
G_{2}^{(1)},~t\in\{1,\cdots,T_{1} \},\\ & 0.1, \quad i\in
G_{1}^{(2)},~t\in\{T_{1}+1,\cdots,T \},\\ & 0.5, \quad i\in
G_{2}^{(2)},~t\in\{T_{1}+1,\cdots,T \},\\ & 0.9, \quad i\in
G_{3}^{(2)},~t\in\{T_{1}+1,\cdots,T \}. \end{aligned}\right.
\end{equation*}
For DGP 4.3, $\Theta _{2,it}$ is defined as in DGP X.3 for X$\in \{1,2,3\}$,
and
\begin{equation*}
\Theta _{1,it}=\left\{ \begin{aligned} & 0.1, \quad i\in
G_{1}^{(1)},~t\in\{1,\cdots,T_{1} \},\\ & 0.7, \quad i\in
G_{2}^{(1)},~t\in\{1,\cdots,T_{1} \},\\ & 0.1, \quad i\in
G_{1}^{(2)},~t\in\{T_{1}+1,\cdots,T \},\\ & 0.4, \quad i\in
G_{2}^{(2)},~t\in\{T_{1}+1,\cdots,T \},\\ & 0.7, \quad i\in
G_{3}^{(2)},~t\in\{T_{1}+1,\cdots,T \}. \end{aligned}\right.
\end{equation*}
\end{description}

\subsection{Results}

Table \ref{tab:rank} reports the proportion of correct rank estimation for
the intercept (IFE) and slope matrices based on the SVT in Section 3.3. Note
that $r_{0}$ denotes the true rank of the intercept matrix and $r_{1}$ and $
r_{2}$ denote those of the two slope matrices. From Table \ref{tab:rank}, we
notice that the true ranks of both the intercept and slope matrices can be
almost perfectly estimated for the sample sizes under investigation.

\begin{table}[p]
\caption{Frequency of correct rank estimation} \label{tab:rank} \scriptsize \centering
\begin{tabular}{cccccccccccc}
\toprule\toprule\multicolumn{2}{c}{N} & \multicolumn{2}{c}{100} & \multicolumn{2}{c}{200} &
\multicolumn{2}{c}{N} & \multicolumn{2}{c}{100} & \multicolumn{2}{c}{200} \\
\multicolumn{2}{c}{T} & 100 & 200 & 100 & 200 & \multicolumn{2}{c}{T} & 100
& 200 & 100 & 200 \\
\midrule \multirow{3}[2]{*}{DGP 1.1} & $r_{0}=1$ & 1.00 & 1.00 & 1.00 & 1.00
& \multirow{3}[2]{*}{DGP 3.1} & $r_{0}=1$ & 1.00 & 1.00 & 1.00 & 1.00 \\
& $r_{1}=1$ & 1.00 & 1.00 & 1.00 & 1.00 &  & $r_{1}=1$ & 1.00 & 1.00 & 1.00
& 1.00 \\
& $r_{2}=1$ & 1.00 & 1.00 & 1.00 & 1.00 &  & $r_{2}=1$ & 1.00 & 1.00 & 1.00
& 1.00 \\
\midrule \multirow{3}[2]{*}{DGP 1.2} & $r_{0}=1$ & 1.00 & 1.00 & 1.00 & 1.00
& \multirow{3}[2]{*}{DGP 3.2} & $r_{0}=1$ & 1.00 & 1.00 & 1.00 & 1.00 \\
& $r_{1}=2$ & 1.00 & 1.00 & 1.00 & 1.00 &  & $r_{1}=2$ & 1.00 & 1.00 & 1.00
& 1.00 \\
& $r_{2}=2$ & 1.00 & 1.00 & 1.00 & 1.00 &  & $r_{2}=2$ & 1.00 & 1.00 & 1.00
& 1.00 \\
\midrule \multirow{3}[2]{*}{DGP 1.3} & $r_{0}=1$ & 1.00 & 1.00 & 1.00 & 1.00
& \multirow{3}[2]{*}{DGP 3.3} & $r_{0}=1$ & 1.00 & 1.00 & 1.00 & 1.00 \\
& $r_{1}=2$ & 1.00 & 1.00 & 1.00 & 1.00 &  & $r_{1}=2$ & 1.00 & 1.00 & 1.00
& 1.00 \\
& $r_{2}=2$ & 1.00 & 1.00 & 1.00 & 1.00 &  & $r_{2}=2$ & 0.998 & 1.00 & 1.00
& 1.00 \\
\midrule \multirow{3}[2]{*}{DGP 2.1} & $r_{0}=1$ & 1.00 & 1.00 & 1.00 & 1.00
& \multirow{3}[2]{*}{DGP 4.1} & $r_{0}=1$ & 1.00 & 1.00 & 1.00 & 1.00 \\
& $r_{1}=1$ & 1.00 & 1.00 & 1.00 & 1.00 &  & $r_{1}=1$ & 1.00 & 1.00 & 1.00
& 1.00 \\
& $r_{2}=1$ & 1.00 & 1.00 & 1.00 & 1.00 &  & $r_{2}=1$ & 1.00 & 1.00 & 1.00
& 1.00 \\
\midrule \multirow{3}[2]{*}{DGP 2.2} & $r_{0}=1$ & 1.00 & 1.00 & 1.00 & 1.00
& \multirow{3}[2]{*}{DGP 4.2} & $r_{0}=1$ & 1.00 & 1.00 & 1.00 & 1.00 \\
& $r_{1}=2$ & 1.00 & 1.00 & 1.00 & 1.00 &  & $r_{1}=1$ & 1.00 & 1.00 & 1.00
& 1.00 \\
& $r_{2}=2$ & 1.00 & 1.00 & 1.00 & 1.00 &  & $r_{2}=2$ & 1.00 & 1.00 & 1.00
& 1.00 \\
\midrule \multirow{3}[2]{*}{DGP 2.3} & $r_{0}=1$ & 1.00 & 1.00 & 1.00 & 1.00
& \multirow{3}[2]{*}{DGP 4.3} & $r_{0}=1$ & 1.00 & 1.00 & 1.00 & 1.00 \\
& $r_{1}=2$ & 1.00 & 1.00 & 1.00 & 1.00 &  & $r_{1}=1$ & 1.00 & 1.00 & 1.00
& 1.00 \\
& $r_{2}=2$ & 1.00 & 1.00 & 1.00 & 1.00 &  & $r_{2}=2$ & 1.00 & 1.00 & 1.00
& 1.00 \\
\bottomrule &  &  &  &  &  &  &  &  &  &  &
\end{tabular}
\end{table}

Table \ref{tab:BP_largeT} reports the results for the break point estimation
in Step 3 based on different $(N,T)$ combinations. We summarize some
important findings from Table \ref{tab:BP_largeT}. First, when the group
membership and the number of groups do not change as in DGP X.1 for X$\in
\left[ 3\right] ,$ the frequency of correct break point estimation may not
be 1 especially if $N$ is not large. This suggests that the binary
segmentation does not work perfectly in such a scenario. Second, the change
of group membership or the number of groups help to identify the break point
as reflected in the simulation results for DGP X.2 and X.3 for X$\in \left[ 4
\right] .$ In general, the binary segmentation works well in our setting.

\begin{table}[tbph]
\caption{Frequency of correct break point estimation}
\label{tab:BP_largeT}\scriptsize \centering
\begin{tabular}{cccccccccc}
\toprule\toprule N & \multicolumn{2}{c}{100} & \multicolumn{2}{c}{200} & N &
\multicolumn{2}{c}{100} & \multicolumn{2}{c}{200} \\
\midrule T & 100 & 200 & 100 & 200 & T & 100 & 200 & 100 & 200 \\
\midrule DGP 1.1 & 0.980 & 0.993 & 1.00 & 1.00 & DGP 3.1 & 0.985 & 0.972 &
1.00 & 0.999 \\
DGP 1.2 & 0.999 & 1.00 & 1.00 & 1.00 & DGP 3.2 & 1.00 & 1.00 & 1.00 & 1.00
\\
DGP 1.3 & 1.00 & 1.00 & 1.00 & 1.00 & DGP 3.3 & 1.00 & 1.00 & 1.00 & 1.00 \\
\midrule DGP 2.1 & 0.998 & 0.999 & 1.00 & 1.00 & DGP 4.1 & 1.00 & 1.00 & 1.00
& 1.00 \\
DGP 2.2 & 1.00 & 1.00 & 1.00 & 1.00 & DGP 4.2 & 1.00 & 1.00 & 1.00 & 1.00 \\
DGP 2.3 & 1.00 & 1.00 & 1.00 & 1.00 & DGP 4.3 & 1.00 & 1.00 & 1.00 & 1.00 \\
\bottomrule &  &  &  &  &  &  &  &  &
\end{tabular}
\end{table}

Table \ref{tab:group} reports the results for the group membership
estimation when the number of groups are either known (infeasible in
practice) or estimated from the data (feasible). With known number of
groups, the STK algorithm degenerates to the traditional K-means algorithm.
The \textquotedblleft Infeasible\textquotedblright\ part of Table \ref
{tab:group} reports the frequency of correct group membership estimation
before and after the estimated break point, $G_{B}$ and $G_{A}$, based on
the known true number of groups and K-means algorithm. Evidently, the
K-means classification exhibits excellent performance in this case.
Nevertheless, without prior information on the true number of groups, the STK
algorithm is able to estimate the group membership and the number of groups
simultaneously. In this case, the frequency of correct estimation of the
group membership and that of the number of groups are shown in the
\textquotedblleft Feasible\textquotedblright\ part in Table \ref{tab:group}
and in Table \ref{tab:K_1}, respectively. To implement the STK algorithm
with unknown number of groups, we set $\varsigma _{N}=N^{-2}$ to ensure the
consistency of the group number estimators. As expected, the performance of
the STK algorithm is slightly worse than that of the K-means algorithm with knowledge of the
true number of groups. But the performance improves when both $N$ and $T$
increase. Table \ref{tab:K_1} suggests that the number of groups can be
nearly perfectly estimated in DGPs 1.1, 1.2, 1.3 and 2.1. For the more
complicated DGPs (e.g., the dynamic case in DGPs 4.1, 4.2, and 4.3 or the
static panel with serially correlated errors in DGPs 3.1, 3.2, and 3.3), the
performance is not as good as that in the simple DGPs.

\begin{table}[p]
\caption{Frequency of correct group membership estimation}
\label{tab:group}\scriptsize \centering
\begin{tabular}{cccccccccccccc}
\toprule \toprule \multirow{26}[12]{*}{Infeasible} & \multicolumn{2}{c}{N} &
\multicolumn{2}{c}{100} & \multicolumn{2}{c}{200} &
\multirow{26}[12]{*}{Feasible} & \multicolumn{2}{c}{N} & \multicolumn{2}{c}{
100} & \multicolumn{2}{c}{200} \\
\cmidrule{2-7}\cmidrule{9-14} & \multicolumn{2}{c}{T} & 100 & 200 & 100 & 200
&  & \multicolumn{2}{c}{T} & 100 & 200 & 100 & 200 \\
\cmidrule{2-7}\cmidrule{9-14} & \multirow{2}[1]{*}{DGP 1.1} & $G_{B}$ & 1.00
& 1.00 & 1.00 & 1.00 &  & \multirow{2}[1]{*}{DGP 1.1} & $G_{B}$ & 1.00 & 1.00
& 1.00 & 1.00 \\
&  & $G_{A}$ & 1.00 & 1.00 & 1.00 & 1.00 &  &  & $G_{A}$ & 1.00 & 1.00 & 1.00
& 1.00 \\
& \multirow{2}[0]{*}{DGP 1.2} & $G_{B}$ & 1.00 & 1.00 & 1.00 & 1.00 &  &
\multirow{2}[0]{*}{DGP 1.2} & $G_{B}$ & 1.00 & 1.00 & 1.00 & 1.00 \\
&  & $G_{A}$ & 1.00 & 1.00 & 1.00 & 1.00 &  &  & $G_{A}$ & 1.00 & 1.00 & 1.00
& 1.00 \\
& \multirow{2}[1]{*}{DGP 1.3} & $G_{B}$ & 1.00 & 1.00 & 1.00 & 1.00 &  &
\multirow{2}[1]{*}{DGP 1.3} & $G_{B}$ & 1.00 & 1.00 & 1.00 & 1.00 \\
&  & $G_{A}$ & 0.989 & 0.999 & 0.978 & 0.999 &  &  & $G_{A}$ & 0.989 & 0.999
& 0.978 & 0.999 \\
\cmidrule{2-7}\cmidrule{9-14} & \multirow{2}[1]{*}{DGP 2.1} & $G_{B}$ & 1.00
& 1.00 & 1.00 & 1.00 &  & \multirow{2}[1]{*}{DGP 2.1} & $G_{B}$ & 1.00 & 1.00
& 1.00 & 1.00 \\
&  & $G_{A}$ & 1.00 & 1.00 & 1.00 & 1.00 &  &  & $G_{A}$ & 1.00 & 1.00 & 1.00
& 1.00 \\
& \multirow{2}[0]{*}{DGP 2.2} & $G_{B}$ & 1.00 & 1.00 & 1.00 & 1.00 &  &
\multirow{2}[0]{*}{DGP 2.2} & $G_{B}$ & 0.989 & 0.999 & 0.992 & 0.999 \\
&  & $G_{A}$ & 1.00 & 1.00 & 1.00 & 1.00 &  &  & $G_{A}$ & 0.992 & 0.999 &
0.977 & 0.998 \\
& \multirow{2}[1]{*}{DGP 2.3} & $G_{B}$ & 1.00 & 1.00 & 1.00 & 1.00 &  &
\multirow{2}[1]{*}{DGP 2.3} & $G_{B}$ & 0.992 & 0.999 & 0.961 & 0.999 \\
&  & $G_{A}$ & 0.998 & 1.00 & 0.999 & 1.00 &  &  & $G_{A}$ & 0.989 & 0.999 &
0.992 & 0.999 \\
\cmidrule{2-7}\cmidrule{9-14} & \multirow{2}[1]{*}{DGP 3.1} & $G_{B}$ & 1.00
& 1.00 & 1.00 & 1.00 &  & \multirow{2}[1]{*}{DGP 3.1} & $G_{B}$ & 0.981 &
0.999 & 0.949 & 0.999 \\
&  & $G_{A}$ & 1.00 & 1.00 & 1.00 & 1.00 &  &  & $G_{A}$ & 0.981 & 0.993 &
0.979 & 0.996 \\
& \multirow{2}[0]{*}{DGP 3.2} & $G_{B}$ & 1.00 & 1.00 & 1.00 & 1.00 &  &
\multirow{2}[0]{*}{DGP 3.2} & $G_{B}$ & 0.985 & 0.996 & 0.962 & 0.993 \\
&  & $G_{A}$ & 1.00 & 1.00 & 1.00 & 1.00 &  &  & $G_{A}$ & 0.985 & 0.994 &
0.973 & 0.998 \\
& \multirow{2}[1]{*}{DGP 3.3} & $G_{B}$ & 1.00 & 1.00 & 1.00 & 1.00 &  &
\multirow{2}[1]{*}{DGP 3.3} & $G_{B}$ & 0.985 & 0.998 & 0.973 & 0.995 \\
&  & $G_{A}$ & 0.981 & 0.997 & 0.982 & 0.999 &  &  & $G_{A}$ & 0.971 & 0.994
& 0.968 & 0.998 \\
\cmidrule{2-7}\cmidrule{9-14} & \multirow{2}[1]{*}{DGP 4.1} & $G_{B}$ & 1.00
& 1.00 & 1.00 & 1.00 &  & \multirow{2}[1]{*}{DGP 4.1} & $G_{B}$ & 0.975 &
0.999 & 0.984 & 0.999 \\
&  & $G_{A}$ & 1.00 & 1.00 & 1.00 & 1.00 &  &  & $G_{A}$ & 0.985 & 0.998 &
0.949 & 0.997 \\
& \multirow{2}[0]{*}{DGP 4.2} & $G_{B}$ & 1.00 & 1.00 & 1.00 & 1.00 &  &
\multirow{2}[0]{*}{DGP 4.2} & $G_{B}$ & 0.994 & 0.998 & 0.952 & 0.997 \\
&  & $G_{A}$ & 1.00 & 1.00 & 1.00 & 1.00 &  &  & $G_{A}$ & 0.977 & 0.999 &
0.985 & 0.999 \\
& \multirow{2}[1]{*}{DGP 4.3} & $G_{B}$ & 1.00 & 1.00 & 1.00 & 1.00 &  &
\multirow{2}[1]{*}{DGP 4.3} & $G_{B}$ & 0.983 & 0.998 & 0.948 & 0.999 \\
&  & $G_{A}$ & 1.00 & 1.00 & 1.00 & 1.00 &  &  & $G_{A}$ & 0.982 & 0.998 &
0.983 & 0.998 \\
\bottomrule &  &  &  &  &  &  &  &  &  &  &  &  &
\end{tabular}
\end{table}

\begin{table}[p]
\caption{Frequency of correct estimation of the number of groups}
\label{tab:K_1}\scriptsize \centering
\begin{tabular}{cccccccccccc}
	\toprule \toprule
\multicolumn{2}{c}{N} & \multicolumn{2}{c}{100} & \multicolumn{2}{c}{200} &
\multicolumn{2}{c}{N} & \multicolumn{2}{c}{100} & \multicolumn{2}{c}{200} \\
\multicolumn{2}{c}{T} & 100 & 200 & 100 & 200 & \multicolumn{2}{c}{T} & 100
& 200 & 100 & 200 \\
\midrule \multirow{2}[1]{*}{DGP 1.1} & $K^{(1)}=2$ & 1.00 & 1.00 & 0.999 &
1.00 & \multirow{2}[1]{*}{DGP 3.1} & $K^{(1)}=2$ & 0.880 & 0.993 & 0.675 &
0.985 \\
& $K^{(2)}=2$ & 1.00 & 1.00 & 1.00 & 1.00 &  & $K^{(2)}=2$ & 0.890 & 0.960 &
0.873 & 0.971 \\
\multirow{2}[0]{*}{DGP 1.2} & $K^{(1)}=2$ & 1.00 & 1.00 & 1.00 & 1.00 &
\multirow{2}[0]{*}{DGP 3.2} & $K^{(1)}=2$ & 0.868 & 0.985 & 0.759 & 0.940 \\
& $K^{(2)}=2$ & 1.00 & 1.00 & 1.00 & 0.999 &  & $K^{(2)}=2$ & 0.897 & 0.971
& 0.829 & 0.987 \\
\multirow{2}[1]{*}{DGP 1.3} & $K^{(1)}=2$ & 0.999 & 1.00 & 1.00 & 1.00 &
\multirow{2}[1]{*}{DGP 3.3} & $K^{(1)}=2$ & 0.889 & 0.988 & 0.802 & 0.965 \\
& $K^{(2)}=3$ & 1.00 & 0.999 & 1.00 & 1.00 &  & $K^{(2)}=3$ & 0.932 & 0.977
& 0.907 & 0.988 \\
\midrule \multirow{2}[1]{*}{DGP 2.1} & $K^{(1)}=2$ & 1.00 & 1.00 & 1.00 &
1.00 & \multirow{2}[1]{*}{DGP 4.1} & $K^{(1)}=2$ & 0.807 & 0.981 & 0.825 &
0.982 \\
& $K^{(2)}=2$ & 1.00 & 1.00 & 1.00 & 1.00 &  & $K^{(2)}=2$ & 0.919 & 0.988 &
0.714 & 0.980 \\
\multirow{2}[0]{*}{DGP 2.2} & $K^{(1)}=2$ & 0.919 & 0.995 & 0.940 & 0.994 &
\multirow{2}[0]{*}{DGP 4.2} & $K^{(1)}=2$ & 0.933 & 0.988 & 0.630 & 0.975 \\
& $K^{(2)}=2$ & 0.930 & 0.993 & 0.809 & 0.982 &  & $K^{(2)}=2$ & 0.758 &
0.988 & 0.870 & 0.989 \\
\multirow{2}[1]{*}{DGP 2.3} & $K^{(1)}=2$ & 0.940 & 0.989 & 0.724 & 0.990 &
\multirow{2}[1]{*}{DGP 4.3} & $K^{(1)}=2$ & 0.877 & 0.991 & 0.657 & 0.991 \\
& $K^{(2)}=3$ & 0.946 & 0.995 & 0.952 & 0.992 &  & $K^{(2)}=3$ & 0.900 &
0.987 & 0.874 & 0.980 \\
\bottomrule &  &  &  &  &  &  &  &  &  &  &
\end{tabular}
\end{table}

Table \ref{tab:K_2} presents more detailed results for the estimation of the
number of groups. For DGPs 1.X and DGP 2.X where we have static panels with
independent errors, the results show that the group membership and the
number of groups can be well estimated with nearly 100\% accuracy under
different $(N,T)$ combinations. For DGPs 3.X and 4.X where we have static
panels with serially correlated errors and dynamic panels, respectively, the
frequency of correct estimation of both the group membership and the number of
groups is not great when $T$ is small, but gradually
approaches unity as $T$ increases. One reason for this is that we need to use
HAC estimates of certain long-run variance objects in the STK algorithm and
it is well known that a relatively large value of $T$ is required in order
for the HAC estimates to be reasonably well behaved in finite samples.

\begin{table}[p]
\caption{Determination of the number of groups}
\label{tab:K_2}\scriptsize \centering
\begin{tabular}{ccccccccccc}
\toprule \toprule \multirow{2}[4]{*}{DGP } & \multirow{2}[4]{*}{N} &
\multirow{2}[4]{*}{T} & \multicolumn{4}{c}{$\hat{K}^{(1)}$} &  & $\hat{K}
^{(2)}$ &  &  \\
\cmidrule{4-11} &  &  & 2 & 3 & 4 & $\geq$5 & 2 & 3 & 4 & $\geq$5 \\
\midrule \multirow{4}[1]{*}{DGP 1.1} & \multirow{2}[1]{*}{100} & 100 &
\textbf{1.00} & 0 & 0 & 0 & \textbf{1.00} & 0 & 0 & 0 \\
&  & 200 & \textbf{1.00} & 0 & 0 & 0 & \textbf{1.00} & 0 & 0 & 0 \\
& \multirow{2}[0]{*}{200} & 100 & \textbf{0.999} & 0.001 & 0 & 0 & \textbf{
1.00} & 0 & 0 & 0 \\
&  & 200 & \textbf{1.00} & 0 & 0 & 0 & \textbf{1.00} & 0 & 0 & 0 \\
\multirow{4}[0]{*}{DGP 1.2} & \multirow{2}[0]{*}{100} & 100 & \textbf{1.00}
& 0.00 & 0.00 & 0.00 & \textbf{1.00} & 0.00 & 0.00 & 0.00 \\
&  & 200 & \textbf{1.00} & 0.00 & 0.00 & 0.00 & \textbf{1.00} & 0.00 & 0.00
& 0.00 \\
& \multirow{2}[0]{*}{200} & 100 & \textbf{1.00} & 0.00 & 0.00 & 0.00 &
\textbf{1.00} & 0.00 & 0.00 & 0.00 \\
&  & 200 & \textbf{1.00} & 0.00 & 0.00 & 0.00 & \textbf{1.00} & 0.00 & 0.00
& 0.00 \\
\multirow{4}[1]{*}{DGP 1.3} & \multirow{2}[0]{*}{100} & 100 & \textbf{0.999}
& 0.001 & 0.00 & 0.00 & 0.00 & \textbf{1.00} & 0.00 & 0.00 \\
&  & 200 & \textbf{1.00} & 0.00 & 0.00 & 0.00 & 0.00 & \textbf{0.999} & 0.001
& 0.00 \\
& \multirow{2}[1]{*}{200} & 100 & \textbf{1.00} & 0.00 & 0.00 & 0.00 & 0.00
& \textbf{1.00} & 0.00 & 0.00 \\
&  & 200 & \textbf{1.00} & 0.00 & 0.00 & 0.00 & 0.00 & \textbf{1.00} & 0.00
& 0.00 \\
\midrule \multirow{4}[1]{*}{DGP 2.1} & \multirow{2}[1]{*}{100} & 100 &
\textbf{0.933} & 0.058 & 0.009 & 0.00 & \textbf{0.936} & 0.060 & 0.003 &
0.001 \\
&  & 200 & \textbf{0.990} & 0.010 & 0.00 & 0.00 & \textbf{0.987} & 0.013 &
0.00 & 0.00 \\
& \multirow{2}[0]{*}{200} & 100 & \textbf{0.864} & 0.126 & 0.010 & 0.00 &
\textbf{0.901} & 0.090 & 0.009 & 0.00 \\
&  & 200 & \textbf{0.989} & 0.011 & 0.00 & 0.00 & \textbf{0.990} & 0.010 &
0.000 & 0.00 \\
\multirow{4}[0]{*}{DGP 2.2} & \multirow{2}[0]{*}{100} & 100 & \textbf{0.919}
& 0.074 & 0.007 & 0.00 & \textbf{0.930} & 0.067 & 0.003 & 0.00 \\
&  & 200 & \textbf{0.995} & 0.003 & 0.00 & 0.002 & \textbf{0.993} & 0.006 &
0.00 & 0.001 \\
& \multirow{2}[0]{*}{200} & 100 & \textbf{0.940} & 0.056 & 0.004 & 0.00 &
\textbf{0.809} & 0.164 & 0.027 & 0.00 \\
&  & 200 & \textbf{0.994} & 0.006 & 0.00 & 0.00 & \textbf{0.982} & 0.018 &
0.00 & 0.00 \\
\multirow{4}[1]{*}{DGP 2.3} & \multirow{2}[0]{*}{100} & 100 & \textbf{0.940}
& 0.055 & 0.005 & 0.00 & 0.00 & \textbf{0.946} & 0.039 & 0.015 \\
&  & 200 & \textbf{0.989} & 0.011 & 0.00 & 0.00 & 0.00 & \textbf{0.995} &
0.002 & 0.003 \\
& \multirow{2}[1]{*}{200} & 100 & \textbf{0.724} & 0.230 & 0.046 & 0.00 &
0.00 & \textbf{0.952} & 0.031 & 0.017 \\
&  & 200 & \textbf{0.990} & 0.010 & 0.00 & 0.00 & 0.00 & \textbf{0.992} &
0.006 & 0.002 \\
\midrule \multirow{4}[1]{*}{DGP 3.1} & \multirow{2}[1]{*}{100} & 100 &
\textbf{0.880} & 0.097 & 0.022 & 0.001 & \textbf{0.890} & 0.062 & 0.031 &
0.017 \\
&  & 200 & \textbf{0.993} & 0.007 & 0 & 0 & \textbf{0.960} & 0.019 & 0.012 &
0.009 \\
& \multirow{2}[0]{*}{200} & 100 & \textbf{0.675} & 0.224 & 0.099 & 0.002 &
\textbf{0.873} & 0.081 & 0.041 & 0.005 \\
&  & 200 & \textbf{0.985} & 0.015 & 0 & 0 & \textbf{0.971} & 0.023 & 0.005 &
0.001 \\
\multirow{4}[0]{*}{DGP 3.2} & \multirow{2}[0]{*}{100} & 100 & \textbf{0.868}
& 0.109 & 0.023 & 0.00 & \textbf{0.897} & 0.099 & 0.004 & 0.00 \\
&  & 200 & \textbf{0.985} & 0.008 & 0.003 & 0.004 & \textbf{0.971} & 0.021 &
0.006 & 0.002 \\
& \multirow{2}[0]{*}{200} & 100 & \textbf{0.759} & 0.198 & 0.042 & 0.001 &
\textbf{0.829} & 0.147 & 0.024 & 0.00 \\
&  & 200 & \textbf{0.940} & 0.055 & 0.005 & 0.00 & \textbf{0.987} & 0.013 &
0.000 & 0.00 \\
\multirow{4}[1]{*}{DGP 3.3} & \multirow{2}[0]{*}{100} & 100 & \textbf{0.889}
& 0.100 & 0.011 & 0.00 & 0.00 & \textbf{0.932} & 0.055 & 0.013 \\
&  & 200 & \textbf{0.988} & 0.009 & 0.003 & 0.00 & 0.00 & \textbf{0.977} &
0.013 & 0.010 \\
& \multirow{2}[1]{*}{200} & 100 & \textbf{0.802} & 0.175 & 0.023 & 0.00 &
0.00 & \textbf{0.907} & 0.073 & 0.020 \\
&  & 200 & \textbf{0.965} & 0.035 & 0.000 & 0.000 & 0.000 & \textbf{0.988} &
0.010 & 0.002 \\
\midrule \multirow{4}[1]{*}{DGP 4.1} & \multirow{2}[1]{*}{100} & 100 &
\textbf{0.807} & 0.084 & 0.089 & 0.02 & \textbf{0.919} & 0.041 & 0.019 &
0.021 \\
&  & 200 & \textbf{0.981} & 0.013 & 0.004 & 0.002 & \textbf{0.988} & 0.004 &
0.005 & 0.003 \\
& \multirow{2}[0]{*}{200} & 100 & \textbf{0.825} & 0.107 & 0.061 & 0.007 &
\textbf{0.714} & 0.118 & 0.084 & 0.084 \\
&  & 200 & \textbf{0.982} & 0.011 & 0.006 & 0.001 & \textbf{0.98} & 0.010 &
0.004 & 0.006 \\
\multirow{4}[0]{*}{DGP 4.2} & \multirow{2}[0]{*}{100} & 100 & \textbf{0.933}
& 0.051 & 0.012 & 0.004 & \textbf{0.758} & 0.141 & 0.089 & 0.012 \\
&  & 200 & \textbf{0.988} & 0.006 & 0.004 & 0.002 & \textbf{0.988} & 0.005 &
0.006 & 0.001 \\
& \multirow{2}[0]{*}{200} & 100 & \textbf{0.630} & 0.158 & 0.196 & 0.016 &
\textbf{0.870} & 0.080 & 0.048 & 0.002 \\
&  & 200 & \textbf{0.975} & 0.013 & 0.012 & 0.000 & \textbf{0.989} & 0.009 &
0.002 & 0.000 \\
\multirow{4}[1]{*}{DGP 4.3} & \multirow{2}[0]{*}{100} & 100 & \textbf{0.877}
& 0.076 & 0.042 & 0.005 & 0.000 & \textbf{0.900} & 0.055 & 0.045 \\
&  & 200 & \textbf{0.991} & 0.006 & 0.002 & 0.001 & 0.000 & \textbf{0.987} &
0.010 & 0.003 \\
& \multirow{2}[1]{*}{200} & 100 & \textbf{0.657} & 0.191 & 0.129 & 0.023 &
0.000 & \textbf{0.874} & 0.072 & 0.054 \\
&  & 200 & \textbf{0.991} & 0.005 & 0.004 & 0.000 & 0.000 & \textbf{0.980} &
0.012 & 0.008 \\
\bottomrule &  &  &  &  &  &  &  &  &  &
\end{tabular}
\end{table}

Table \ref{tab:infer} shows results for the post-classification estimator
for the first slope coefficient. We follow \cite{su2016identifying} to
define the evaluation criteria as bias and coverage. Specifically, we define
the bias to be the weighted versions of bias for slope estimator from all
estimated groups, i.e. $\text{Bias}=\sum_{k=1}^{K^{(1)}}\text{Bias}(\alpha
_{k,1}^{(\ell )})$ for $\ell \in \{1,2\}$. Similarly, we define the weighted
version of the coverage ratio of the 95\% confidence interval estimators. The
\textquotedblleft Infeasible" panel shows the result assuming the number of
groups information is known, and the \textquotedblleft Feasible" panel shows
the result without knowing the number of groups information by the STK
algorithm. From Table \ref{tab:infer}, we notice that the coverage ratio for
DGP 1 and 2 is close to 95\% under different combinations of $N$ and $T$ for
both the \textquotedblleft Infeasible" and \textquotedblleft Feasible"
panels, which is due to the higher correct classification ratio. For DGPs 3
and 4, by using the STK algorithm, the coverage ratio is a bit lower for $
T=100$, which is due to the inaccuracy of the group number and membership
estimators; but coverage approaches 95\% quickly when $T$ doubles.

\begin{table}[p]
\caption{Point estimation of $\protect\alpha _{\cdot ,1}^{(1)}$ and $\protect
\alpha _{\cdot ,1}^{(2)}$}
\label{tab:infer}\scriptsize \centering
\scalebox{0.9}{\begin{tabular}{ccc|cccc|cccc}
    \toprule
    \toprule
    \multirow{3}[6]{*}{DGP} & \multirow{3}[6]{*}{N} & \multicolumn{1}{c}{\multirow{3}[6]{*}{T}} & \multicolumn{4}{c}{Infeasible} & \multicolumn{4}{c}{Feasible} \\
\cmidrule{4-11}          &       & \multicolumn{1}{c}{} & \multicolumn{2}{c}{Before the break} & \multicolumn{2}{c}{After the break} & \multicolumn{2}{c}{Before the break} & \multicolumn{2}{c}{After the break} \\
\cmidrule{4-11}          &       & \multicolumn{1}{c}{} & Bias($\times 10^{-6}$) & Coverage & Bias($\times 10^{-6}$) & \multicolumn{1}{c}{Coverage} & Bias($\times 10^{-6}$) & Coverage & Bias($\times 10^{-6}$) & Coverage \\
    \midrule
    \multirow{4}[2]{*}{1.1} & \multirow{2}[1]{*}{100} & 100   & 2.585 & 0.951 & -2.869 & 0.946 & 2.585 & 0.951 & -2.869 & 0.946 \\
          &       & 200   & -1.944 & 0.944 & -8.920 & 0.945 & -1.958 & 0.944 & -8.920 & 0.945 \\
          & \multirow{2}[1]{*}{200} & 100   & -1.096 & 0.943 & 1.407 & 0.947 & -1.096 & 0.943 & 1.407 & 0.947 \\
          &       & 200   & -1.910 & 0.945 & 0.960 & 0.947 & -1.910 & 0.945 & 0.960 & 0.947 \\
    \midrule
    \multirow{4}[2]{*}{1.2} & \multirow{2}[1]{*}{100} & 100   & -1.050 & 0.949 & -27.398 & 0.941 & -1.050 & 0.949 & -27.398 & 0.941 \\
          &       & 200   & -5.449 & 0.930 & 7.616 & 0.953 & -5.449 & 0.930 & 7.655 & 0.953 \\
          & \multirow{2}[1]{*}{200} & 100   & 4.770 & 0.949 & 1.866 & 0.951 & 4.770 & 0.949 & 1.866 & 0.951 \\
          &       & 200   & 1.317 & 0.941 & 1.874 & 0.945 & 1.317 & 0.941 & 1.874 & 0.945 \\
    \midrule
    \multirow{4}[2]{*}{1.3} & \multirow{2}[1]{*}{100} & 100   & -0.961 & 0.943 & 11.417 & 0.944 & -1.050 & 0.949 & -27.398 & 0.941 \\
          &       & 200   & -4.213 & 0.951 & -5.002 & 0.941 & -5.449 & 0.930 & 7.655 & 0.953 \\
          & \multirow{2}[1]{*}{200} & 100   & -1.571 & 0.938 & -3.756 & 0.938 & 4.770 & 0.949 & 1.866 & 0.951 \\
          &       & 200   & 0.403 & 0.941 & -4.159 & 0.945 & 1.317 & 0.941 & 1.874 & 0.945 \\
    \midrule
    \multirow{4}[2]{*}{2.1} & \multirow{2}[1]{*}{100} & 100   & 14.840 & 0.944 & 9.410 & 0.950 & 14.816 & 0.943 & 9.406 & 0.950 \\
          &       & 200   & -7.222 & 0.951 & 1.795 & 0.951 & -7.222 & 0.951 & 1.795 & 0.951 \\
          & \multirow{2}[1]{*}{200} & 100   & 0.916 & 0.940 & 3.575 & 0.948 & 0.916 & 0.940 & 3.575 & 0.948 \\
          &       & 200   & 0.452 & 0.948 & -0.797 & 0.947 & 0.452 & 0.948 & -0.797 & 0.947 \\
    \midrule
    \multirow{4}[2]{*}{2.2} & \multirow{2}[1]{*}{100} & 100   & -21.379 & 0.946 & 0.234 & 0.937 & -21.379 & 0.946 & 0.234 & 0.937 \\
          &       & 200   & 0.264 & 0.942 & -15.542 & 0.953 & 0.264 & 0.942 & -15.542 & 0.953 \\
          & \multirow{2}[1]{*}{200} & 100   & -1.379 & 0.945 & -1.489 & 0.951 & -1.379 & 0.944 & -1.489 & 0.951 \\
          &       & 200   & -1.101 & 0.950 & 1.127 & 0.949 & -1.101 & 0.950 & 1.127 & 0.949 \\
    \midrule
    \multirow{4}[2]{*}{2.3} & \multirow{2}[1]{*}{100} & 100   & -8.610 & 0.945 & 5.254 & 0.952 & -8.610 & 0.945 & 5.261 & 0.952 \\
          &       & 200   & 0.927 & 0.949 & 5.840 & 0.949 & 0.927 & 0.949 & 5.840 & 0.949 \\
          & \multirow{2}[1]{*}{200} & 100   & -1.560 & 0.943 & -2.569 & 0.941 & -1.560 & 0.943 & -2.569 & 0.941 \\
          &       & 200   & -0.775 & 0.947 & 4.408 & 0.947 & -0.775 & 0.947 & 4.386 & 0.947 \\
    \midrule
    \multirow{4}[2]{*}{3.1} & \multirow{2}[1]{*}{100} & 100   & -20.928 & 0.955 & -73.947 & 0.945 & -26.250 & 0.927 & -77.613 & 0.920 \\
          &       & 200   & 3.066 & 0.949 & -12.443 & 0.937 & 2.884 & 0.940 & -13.116 & 0.934 \\
          & \multirow{2}[1]{*}{200} & 100   & -2.663 & 0.951 & -8.742 & 0.944 & -3.517 & 0.857 & -7.730 & 0.888 \\
          &       & 200   & -3.747 & 0.949 & -2.107 & 0.945 & -3.642 & 0.939 & -1.971 & 0.938 \\
    \midrule
    \multirow{4}[2]{*}{3.2} & \multirow{2}[1]{*}{100} & 100   & -55.980 & 0.952 & -10.846 & 0.943 & -58.714 & 0.926 & -15.109 & 0.863 \\
          &       & 200   & -2.774 & 0.950 & 4.690 & 0.946 & -3.218 & 0.945 & 4.913 & 0.942 \\
          & \multirow{2}[1]{*}{200} & 100   & 6.979 & 0.951 & 8.879 & 0.945 & 6.287 & 0.858 & 6.894 & 0.848 \\
          &       & 200   & -1.704 & 0.947 & 0.438 & 0.945 & -2.122 & 0.928 & 0.381 & 0.940 \\
    \midrule
    \multirow{4}[2]{*}{3.3} & \multirow{2}[1]{*}{100} & 100   & -25.340 & 0.950 & 37.639 & 0.907 & -29.905 & 0.924 & 37.016 & 0.890 \\
          &       & 200   & 2.042 & 0.947 & -6.431 & 0.960 & 1.667 & 0.940 & -6.245 & 0.960 \\
          & \multirow{2}[1]{*}{200} & 100   & -2.391 & 0.946 & 14.364 & 0.892 & -2.735 & 0.891 & 13.680 & 0.840 \\
          &       & 200   & 4.339 & 0.943 & 4.890 & 0.942 & 4.493 & 0.932 & 5.113 & 0.938 \\
    \midrule
    \multirow{4}[2]{*}{4.1} & \multirow{2}[1]{*}{100} & 100   & 800.620 & 0.930 & -466.590 & 0.929 & 777.650 & 0.928 & -454.980 & 0.924 \\
          &       & 200   & 126.760 & 0.931 & 550.210 & 0.942 & 126.220 & 0.942 & 548.160 & 0.943 \\
          & \multirow{2}[1]{*}{200} & 100   & -224.960 & 0.931 & -339.020 & 0.939 & -214.900 & 0.904 & -313.430 & 0.876 \\
          &       & 200   & 417.320 & 0.938 & 412.500 & 0.947 & 415.110 & 0.941 & 410.700 & 0.944 \\
    \midrule
    \multirow{4}[2]{*}{4.2} & \multirow{2}[1]{*}{100} & 100   & 1246.000 & 0.921 & 726.940 & 0.943 & 1205.600 & 0.918 & 709.260 & 0.903 \\
          &       & 200   & -440.880 & 0.943 & 83.433 & 0.944 & -436.670 & 0.943 & 81.585 & 0.951 \\
          & \multirow{2}[1]{*}{200} & 100   & -1600.500 & 0.930 & 1937.500 & 0.927 & -1538.300 & 0.901 & 1781.600 & 0.819 \\
          &       & 200   & -1513.300 & 0.950 & -272.500 & 0.946 & -1502.000 & 0.935 & -271.420 & 0.953 \\
    \midrule
    \multirow{4}[2]{*}{4.3} & \multirow{2}[1]{*}{100} & 100   & -2067.900 & 0.931 & -505.100 & 0.940 & -1951.700 & 0.866 & -491.200 & 0.929 \\
          &       & 200   & 317.360 & 0.946 & 411.770 & 0.945 & 316.550 & 0.951 & 407.820 & 0.935 \\
          & \multirow{2}[1]{*}{200} & 100   & 1279.900 & 0.930 & 3660.100 & 0.888 & 1246.900 & 0.906 & 3355.100 & 0.874 \\
          &       & 200   & -772.380 & 0.940 & -335.250 & 0.948 & -768.870 & 0.932 & -334.190 & 0.940 \\
    \bottomrule
    \end{tabular}}
\end{table}

\section{Empirical Study}

\label{sec:emp} The estimation methods were applied to
analyze the time-varying latent group structure of real house price changes
in Metropolitan Statistical Areas (MSAs) in the United States. Studies
of U.S. house price changes are plentiful in the literature. \cite
{malpezzi1999simple}, \cite{capozza2002determinants}, \cite{gallin2006long},
and \cite{ortalo2006housing} all show that the house price changes are closely
correlated with real income in the long run. \cite{su2023identifying}
consider a heterogenous spatial panel and show that real income growth
affects the U.S. house prices in different ways for different MSAs. In this
application, we examine whether there exist latent group structures for the
real income growth elasticity of house price changes and whether these structures change
over the time dimension.

\subsection{Model}

We consider the following panel data model with IFEs and two-way slope heterogeneity
\begin{equation}
\pi _{it}=\lambda _{i}^{\prime }f_{t}+\Theta _{1,it}ginc_{it}+\Theta
_{2,it}ginc_{i,t-1}+e_{it},  \label{emp:1}
\end{equation}
where the dependent variable $\pi _{it}$ measures the percentage of real
house price growth for MSA $i$ at time period $t$. The $\lambda _{i}$ and $f_{t}$
are the individual fixed effects and time fixed effects, the
covariate $ginc_{it}$ denotes the percentage of income growth for MSA $i$ at
time period $t$, and $ginc_{i,t-1}$ is the lagged value of $ginc_{it}$.
Unlike \cite{aquaro2021estimation} and \cite{su2023identifying} who consider
individual fixed effects and additive two-way fixed effects, respectively,
we allow the model to have IFEs. In the above model, we allow the slope
parameters $\left( \Theta _{1,it},\Theta _{2,it}\right) $ to exhibit
latent group structures along the cross-sectional dimension and an unknown
break along the time dimension.

\subsection{Data}

The dataset we use is obtained from \cite{aquaro2021estimation}, which is the
quarterly data for 377 MSAs over 1975 to 2014. To construct the growth rate
and the lagged term, we lose two periods of observations, which yields $
T=158 $. Similar to \cite{su2023identifying}, we deseasonalize the growth
rate of real house price and real income. We do not de-factor the variables
since our model contains IFEs to control the common shocks.

\subsection{Empirical Results}

We first apply the singular value thresholding to estimate the ranks of $
\Theta _{0}=\{\lambda _{i}^{\prime }f_{t}\},$ $\Theta _{1}=\{\Theta _{1,it}\}
$ and $\Theta _{2}=\{\Theta _{2,it}\}.$ The estimates are: $\hat{r}_{0}=1,$ $
\hat{r}_{1}=2,$ and $\hat{r}_{2}=2$. Before applying the proposed estimation
algorithm in Section \ref{sec:estimation}, we first test the presence of a
structural break as in Section \ref{sec:extension_test}. As given in
\cite{bai2003critical}, for each individual $i$, the critical value of test
statistic $\sup_{T_{1}\in \mathcal{T}_{\epsilon }}F_{i}(T_{1})$ is 15.37. We
then construct the sup-F test statistic for each MSA. Results show that $
\operatornamewithlimits{\text{min}}\limits_{i\in \lbrack N]}\sup_{T_{1}\in
\mathcal{T}_{\epsilon }}F_{i}(T_{1})=0.0195$ and the final test statistic is
$F_{NT}(1|0)=\operatornamewithlimits{\text{max}}\limits_{i\in \lbrack
N]}\sup_{T_{1}\in \mathcal{T}_{\epsilon }}F_{i}(T_{1})=2161.65$. Based on
this outcome, we reject the null that there is no structural break for slope
coefficients $\Theta _{1,it}$ and $\Theta _{2,it}$ at the 1\% significance level.

With the presence of a structural break, we apply the proposed multi-stage
estimation result in Section \ref{sec:estimation} to estimate the break date
and numbers of groups before and after the break. The estimated break date
is given by $\hat{T}_{1}=51,$ which suggests that the structural break
happens at the first quarter in 1988. We conjecture that this break may be related to
the catastrophic stock market crash that occurred on October 1987, which is
considered to be the first contemporary global financial crisis event.

\begin{figure}[p]
\caption{Group classification result 1975Q3-1987Q4}
\label{fig:before}\centering
\includegraphics[width=0.9\textwidth]{ABP_before.jpeg}
\includegraphics[width=0.9\textwidth]{ABP_after.jpeg}
\caption{Group classification result 1988Q1-2014Q4}
\label{fig:after}
\end{figure}

By setting $\varsigma _{N}=N^{-2}$ for the STK algorithm as in the
simulations, we obtain the estimated prior- and post-break numbers of groups
given by $\hat{K}^{(1)}=6$ and $\hat{K}^{(2)}=2,$ respectively. As for the
group structure, Figures \ref{fig:before} and \ref{fig:after} use six and
two colors to show the classification results for the 377 MSAs during 1975Q3
to 1987Q4 and 1988Q1 to 2014Q4, respectively. Table \ref{tab:emp_pool}
reports the pooled regression results for the full sample in column (1), the
subsample before the estimated break point in column (2), and the subsample
after the estimated break point in column (3). All the slope estimators are
bias-corrected. The pooled regression results in Table \ref{tab:emp_pool}
show that real income growth has positive and significant effect on
house prices. Comparing the two subsamples before and after the estimated
break, we observe that, with a 1 percentage increase in real income
growth, the real house price growth rate will increase 0.09 percentage
before 1988, which is 0.02 percentage points higher than that after 1988.
The slope estimates for the lagged term are similar for the two subsamples.

\begin{table}[h]
\caption{Results for the pooled regressions}
\label{tab:emp_pool}\centering
\begin{tabular}{cccc}
\toprule \toprule & Pooled (full sample) & Pooled ($1975Q3-1987Q4$) & Pooled
($1988Q1-2014Q4$) \\
& $(1)$ & $(2)$ & $(3)$ \\
\midrule $ginc_{it}$ & $0.1021^{***}$ & $0.0904^{***}$ & $0.0702^{***}$ \\
& (0.0067) & (0.0119) & (0.0065) \\
$ginc_{i,t-1}$ & $0.0590^{***}$ & $0.0392^{***}$ & $0.0401^{***}$ \\
& (0.0067) & (0.0117) & (0.0066) \\
\#individuals & 377 & 377 & 377 \\
 \bottomrule
\multicolumn{4}{p{16cm}}{$Note$: Column (1) reports the pooled regression
results for the full sample. Columns (2) and (3) report the pooled
regression results for the subsamples before and after the estimated break
point, respectively. Slope estimators are all bias-corrected. Values in
parentheses are standard errors and $^{***}$ indicates significance at $1\%$
level.}
\end{tabular}
\end{table}

To examine the difference for each of the 6 estimated groups before the
break, Table \ref{tab:emp_post_before} reports the post-classification
regression results for each estimated group before the estimated break. Even
though the effects of real income growth are positive for all estimated
groups, they differ vastly across groups. The effect of real income
growth for Group 2 is the highest, followed by Groups 5 and 3, and the
effects of real income growth on real house prices in all of these three
groups are higher than 0.15 percent. In contrast, the effects of real
income growth for the remaining three groups, viz., Groups 1, 4, and 6, are
less than 0.07 percent. Similarly, Table \ref{tab:emp_post_after} reports
the post-classification regression results for each estimated group after
the estimated break. The estimated slope coefficients for both groups are
statistically significant. Especially, during 1988Q1-2014Q1, the slope
estimator for the lagged term in the first group is much higher than that
for the second group.

We also applied the C-Lasso algorithm of \cite{su2016identifying} to estimate
the group structure before and after the estimated break point. The C-Lasso
approach in conjuction with IC detects 2 groups both before and after the
break. In view of the six groups detected by our present algorithm, we conjecture that the difference may
due to the smaller time periods before the break.

\begin{table}[h]
\caption{Results for the post-classification regressions before the break}
\label{tab:emp_post_before}\centering
\begin{tabular}{ccccccc}
\toprule \toprule & Group 1 & Group 2 & Group 3 & Group 4 & Group 5 & Group 6
\\
& $(1)$ & $(2)$ & $(3)$ & $(4)$ & $(5)$ & $(6)$ \\
\midrule $ginc_{it}$ & $0.0301$ & $0.3169^{***}$ & $0.1522^{***}$ & $0.0168$
& $0.1877^{**}$ & $0.0661^{***}$ \\
& (0.0345) & (0.0292) & (0.0408) & (0.0561) & (0.0775) & (0.0153) \\
$ginc_{i,t-1}$ & $0.1217^{***}$ & $-0.0191$ & $-0.0089$ & $-0.0298$ & $
-0.1269^{*}$ & $0.0331^{***}$ \\
& (0.0348) & (0.0288) & (0.0407) & (0.0560) & (0.0754) & (0.0151) \\
\#individuals & 60 & 92 & 35 & 12 & 36 & 142 \\
 \bottomrule
\multicolumn{7}{p{14cm}}{$Note$: Each column reports the regression results
for each estimated group during 1977Q3-1987Q4. Slope estimators are all
bias-corrected. Values in parentheses are standard errors. $^{***}$, $^{**}$
, and $^{*}$ indicate significance at $1\%$ level, $5\%$ level, $10\%$
level, respectively.}
\end{tabular}
\end{table}

\begin{table}[h]
\caption{Results for the post-classification regressions after the break}
\label{tab:emp_post_after}\centering
\begin{tabular}{ccc}
\toprule \toprule & Group 1 & Group 2 \\
& $(1)$ & $(2)$ \\
\midrule $ginc_{it}$ & $0.0670^{***}$ & $0.0714^{***}$ \\
& (0.0130) & (0.0079) \\
$ginc_{i,t-1}$ & $0.0870^{***}$ & $0.0275^{***}$ \\
& (0.0135) & (0.0079) \\
\#individuals & 103 & 274 \\
\bottomrule
\multicolumn{3}{p{8cm}}{$Note$: Each column reports the regression results
for each estimated group during 1988Q1-2014Q4. Slope estimators are all
bias-corrected. Values in parentheses are standard errors. $^{***}$
indicates significance at $1\%$ level.}
\end{tabular}
\end{table}








\section{Conclusion}

\label{sec:concl} This paper considers a linear panel data model with interactive fixed effects
and two-way heterogeneity such that the heterogeneity across individuals is
captured by latent group structures and the heterogeneity across time is
captured by an unknown structural break. We allow the model to have
different group numbers, or different group memberships, or just changes in
the slope coefficients for some specific groups before and after the break.
To estimate the unknown structural break, the number of groups and group
memberships before and after the break point, we propose an estimation
algorithm with initial nuclear-norm-regularized estimates, followed by row-
and column-wise linear regressions. Then, the break point estimator is
obtained by binary segmentation and the group structure together with the
number of groups are estimated simultaneously using a sequential testing K-means
algorithm. We show that the structural break estimator, the group number
estimators, and the group membership estimators before and after the break
point are all consistent, and the final post-classification slope
coefficient estimators enjoy the oracle property.

There are several interesting topics for further research. First, even
though we discuss a possible test for the existence of a break in the panel
data models with latent group structures, we have not fully worked out the
asymptotic theory, a challenge that deserves separate treatment. Second,
we assume the presence of a single break in the
data and it is interesting to extend our theory to allow for multiple
breaks. Third, the present treatment rules out both unit-root-type nonstationarity and
nonstochastic trending nonstationarity and it is interesting to extend our
theory to allow for nonstationarity. We will pursue these topics in future
research.

{\small {\
\bibliographystyle{apalike}
\bibliography{chapter2}
}} \newpage

{\small