EconBase
← Back to paper

Panel Data Models with Time-Varying Latent Group Structures

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

126,842 characters · 25 sections · 98 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Panel Data Models with Time-Varying Latent Group Structures

abstractThis paper considers a linear panel model with interactive fixed effects and unobserved individual and time heterogeneities that are captured by some latent group structures and an unknown structural break, respectively. To enhance realism the model may have different numbers of groups and/or different group memberships before and after the break. With the preliminary nuclear-norm-regularized estimation followed by row- and column-wise linear regressions, we estimate the break point based on the idea of binary segmentation and the latent group structures together with the number of groups before and after the break by sequential testing K-means algorithm simultaneously. It is shown that the break point, the number of groups and the group memberships can each be estimated correctly with probability approaching one. Asymptotic distributions of the estimators of the slope coefficients are established. Monte Carlo simulations demonstrate excellent finite sample performance for the proposed estimation algorithm. An empirical application to real house price data across 377 Metropolitan Statistical Areas in the US from 1975 to 2014 suggests the presence both of structural breaks and of changes in group membership. Key words: Interactive fixed effects, latent group structure, structural break, nuclear norm regularization, sequential testing K-means algorithm. JEL Classification: C23, C33, C38, C51\

\ \ \ \

Introduction

Heterogeneous panel data models have been widely used in empirical research in economics because they can capture a rich degree of unobserved heterogeneity. But models with complete heterogeneity along either the cross section or time dimension tend to possess too many parameters to be identified, which results in slow convergence and inefficient estimates. For this reason, researchers now frequently advocate the use of panel data models with certain structures imposed along either the cross section or time dimension. On the one hand, the recent burgeoning of panels with latent group structures can be motivated from the observation that different groups of individuals respond differently to exogenous shocks. For instance, durlauf1995multiple, berthelemy1996economic, and ben1998convergence show economies in different groups of income per capita and/or education level may converge to different steady state equilibria. klapper2011impact, chu2012logistics, and zhang2019threshold show an exogenous shock like policy implementation has different impacts on different individuals, and long2012impact argue that the influence of the 2008 financial crisis on economic growth differs for emergent and developed economies. On the other hand, the recent popularity of panels that evidence structural change can be motivated by events such as financial crises, the economic impact of technological progress, and more general economic transitions that occur during the time periods covered by the data. See qian2016shrinkage for a survey of panel data models and research that consider estimation and inference concerning structural change.

In spite of the now large literature that studies separately individual heterogeneity or time heterogeneity in the slope coefficients of panel models, few works consider both types of heterogeneity simultaneously. Exceptions include keane2020climate and lu2023uniform who consider linear panel data models with two-dimensional unobserved heterogeneity in the slope coefficients that are modelled via the usual additive structure, and chernozhukov2019inference and wang2022low-rank who model the slope coefficients via the use of low-rank matrices for conditional mean and quantile regressions, respectively. In addition, okui2021heterogeneous and lumsdaine2023estimation consider both individual heterogeneity and time heterogeneity by modeling them as a grouped pattern and as structural breaks, respectively. Specifically, okui2021heterogeneous develop a new panel data model with latent groups where the number of groups and the group memberships do not change over time but the coefficients within each group can change over time and they may have different break-dates; lumsdaine2023estimation consider the panels with a grouped pattern of heterogeneity when the latent group membership structure and/or the values of slope coefficients change at a break point. Both papers provide algorithms to recover the latent group structure based on linear panel models with or without individual fixed effects, but cannot allow for the presence of more complicated fixed effects such as interactive fixed effects (IFEs) to capture strong cross-sectional dependence in the data.

This paper proposes a linear panel data model with IFEs that enable the slope coefficients to exhibit two-way heterogeneity. Following the lead of okui2021heterogeneous and lumsdaine2023estimation and to encourage the parameter parsimony, we use a latent group structure to capture individual heterogeneity and an unknown structural break to capture time heterogeneity. The latent group structure of the model accommodates different group numbers and different group memberships before and after the break. Given this complicated structure, the approach proposed is to estimate the break point, the number of groups before and after the break, the group membership before and after the break, and the group-specific parameters in multiple steps. The key insight that permits this degree of complication is that the slope coefficients of each of the $p$ regressors in the model are permitted to vary across both cross section and time dimensions by means of a factor structure with a fixed number of factors so that they may be conveniently stacked into a low-rank matrix.

In the first step, the low-rank nature of the slope matrices is explored and initial estimates are obtained by nuclear norm regularization (NNR), a machine learning technique popular in computer science that is increasingly used in econometrics. Such initial matrix estimates are consistent in terms of the Frobenius norm but do not have pointwise or uniform convergence for their elements. Despite this, by applying singular value decomposition (SVD) to these estimates, we can obtain estimates of the associated factors and the factor loadings that are also consistent in terms of the Frobenius norm. In the second step, we use the first-step initial estimates of the factors and factor loadings to run the row- and column-wise linear regressions to update the estimates of the factors and factor loadings which now possess pointwise and uniform consistency and can be used for subsequent analyses. In the third step, we estimate the break point by using the celebrated idea of binary segmentation, as commonly used for break point estimation in the time series literature. Once the break point is estimated, the full-sample is naturally split into two subsamples. In the fourth step, we follow the lead of lin2012estimation and Jin2022Optimal to focus on each subsample before and after the estimated break point and propose a sequential testing K-means algorithm to recover the latent group structure and obtain the number of groups simultaneously. In the last step, we use the estimated group structure to estimate the group-specific parameters. Asymptotic analyses show that the break point, the number of groups and the group memberships can be consistently estimated in Steps 3-4, so that the final step estimates for the group-specific coefficients can enjoy the oracle property. This means they have the same asymptotic distributions as the ones obtained by knowing the break point and the latent group structures before and after the break points.

The present paper relates to two branches of literature. First, it contributes to the panel data literature on one-way heterogeneity, especially with either latent group structures or structural breaks. With respect to latent group structures, there are several popular ways to recover the latent groups. The first approach is the K-means algorithm. lin2012estimation apply the K-means algorithm to linear panel data models with grouped slope coefficients and propose an information criterion and a sequential testing approach to estimate the true number of groups. sarafidis2015partially analyze the unknown grouped slopes in the large $N$ and fixed $T$ framework, and zhang2019quantile provide an iterative algorithm based on K-means clustering for a panel quantile regression model. bonhomme2015grouped and ando2016panel consider panels with grouped fixed effects. The second approach is the Classifier-Lasso (C-Lasso) that has become a popular clustering method since su2016identifying. This method is extended by lu2017determining, su2018identifying , su2019sieve, wang2019heterogeneous, and huang2020identifying to various contexts. In addition, both the clustering algorithm in regression via a data-driven segmentation (CARDS) approach and binary segmentation are also considered in ke2015homogeneity, wang2018homogeneity, ke2016structure and wang2021identifying , among others. As for panel models with structural breaks, binary segmentation has become a common approach to estimate the break point. See bai2010common, hsu2011change, kim2011estimating, kim2014common and baltagi2017estimation, among others. These papers focus on the case of a single break point in the model. In contrast, qian2016shrinkage and li2016panel allow for multiple breaks in linear panel models with either classical fixed effects or IFEs, and propose an adaptive grouped fused lasso (AGFL) approach to estimate the break points. Compared to the existing panel literature on one-way heterogeneity, our model allows for two-way heterogeneity. In particular, not only are different membership structures in different time blocks permitted but also changes in the number of groups over time. As a result, our model is more flexible than all existing models that allow only for latent group structures or structural breaks, but not both.

Second, this paper contributes to the recent burgeoning literature that models two-way heterogeneity in the slope coefficients of a panel model. As mentioned above, there are two approaches to model two-way heterogeneity in the slope coefficients. One approach models them in an additive structure so that both individual and time effects enter the slope coefficients additively, as in keane2020climate and lu2023uniform. The other approach imposes certain low-rank structures on the slope coefficient matrices in which case one models each slope coefficient via the use of IFEs to capture strong cross sectional dependence in the panel. In view of the low-rank structures, we can resort to NNR estimation which has attracted increasing attention recently in panel data analyses. NNR has been used in recent econometric research -- see bai2019rank, moon2018nuclear, chernozhukov2019inference, belloni2023high, miao2023high, feng2023regularized, and hong2023profile, among others. But none of these papers imposes any latent group structures on the slope coefficients. With latent group structures and structural breaks imposed, okui2021heterogeneous allows the slope coefficients within each group to have common breaks and the break points to vary across different groups, and they propose to estimate the latent group structures, the structural breaks, and the group-specific regression parameters by the grouped adaptive group fused lasso (GAGFL). But neither the number of groups nor the group memberships is allowed to change over time in okui2021heterogeneous. In a companion paper, lumsdaine2023estimation allows the latent group membership structure and/or the values of slope coefficients to change at a break point, and proposes an estimation algorithm similar to the K-means of bonhomme2015grouped. Both okui2021heterogeneous and lumsdaine2023estimation allow for at most one-way heterogeneity (individual fixed effects) in the intercept and neither allows for IFEs to capture strong cross section dependence. In contrast, this paper proposes an algorithm to detect the unknown break point and to recover the group structure based on linear panel model with IFEs, which involves a more general model. In addition, lumsdaine2023estimation first assume the number of groups is known in the estimation algorithm and then estimate the number of groups via an information criterion but they do not establish consistency for such an estimate. Instead, we estimate the number of groups and group membership simultaneously by the sequential testing K-means algorithm and establish the consistency of the number of groups estimator.

The rest of the paper is organized as follows. We first introduce the linear panel model with time-varying latent group structures in Section (ref) and provide the estimation algorithm in Section (ref). The asymptotic properties are given in Section (ref). In Section (ref), we propose an alternative approach to detect the break point, provide the test statistics for the null that the slope coefficients exhibit no structure change against the alternative with one break point, and discuss the estimation for the model with multiple breaks. In Sections (ref) and (ref), we show the finite sample performance of our method by Monte Carlo simulations and an empirical application, respectively. Section (ref) concludes. All proofs are provided in the online supplement.

Notation. Let $\left\Vert \cdot \right\Vert _{\max },$ $\left\Vert \cdot \right\Vert _{op},$ $\left\Vert \cdot \right\Vert $, and $\left\Vert \cdot \right\Vert _{\ast }$ denote the (elementwise) maximum norm, operator norm, Frobenius norm, and nuclear norm, respectively. Let $\odot $ denote the element-wise Hadamard product. $\lfloor \cdot \rfloor $ and $\lceil \cdot \rceil $ denote the floor and ceiling functions, respectively. Let $ a\vee b=\max \left( a,b\right) $ and $a\wedge b=\min (a,b)$. $a_{n}\lesssim b_{n}$ means $a_{n}/b_{n}=O_{p}\left( 1\right) $ and $a_{n}\gg b_{n}$ means $ b_{n}a_{n}^{-1}=o(1)$. Let $A=\{A_{it}\}$ be a matrix with its $(i,t)$-th entry denoted as $A_{it}$. Let denote $\{A_{j}\}_{j\in \lbrack p]\cup \{0\}}$ be the collection of matrices $A_{j},$ $j\in \{0,1,\cdots ,p\}$. For a specific $A\in \mathbb{R}^{m\times n}$ with rank $n$, let $P_{A}=A(A^{\prime }A)^{-1}A^{\prime }$ and $M_{A}=I_{m}-P_{A}$. When $A$ is symmetric, $ \lambda _{\max }(A)$, $\lambda _{\min }(A)$ and $\lambda _{n}(A)$ denote its largest, smallest and $n$-th largest eigenvalues, respectively. The operators $\rightsquigarrow $ and $\overset{p}{\longrightarrow }$ denote convergence in distribution and in probability, respectively. Let $[n]=$ $ \{1,\cdots ,n\}$ for any positive integer $n$, let $\mathbf{1}\{\cdot \}$ be the usual indicator function, and w.p.a.1 and a.s. abbreviate \textquotedblleft with probability approaching 1\textquotedblright\ and \textquotedblleft almost surely\textquotedblright , respectively.

Model Setup

In this paper we consider the following linear panel model with IFEs:

equation[equation omitted — 96 chars of source]

where $i\in \lbrack N]$, $t\in \lbrack T]$, $Y_{it}$ is the dependent variable, $X_{it}=(X_{1,it},\cdots ,X_{p,it})^{\prime }$ is a $p\times 1$ vector of regressors, $\Theta _{it}^{0}=(\Theta _{1,it}^{0},\cdots ,\Theta _{p,it}^{0})^{\prime }$ is a $p\times 1$ vector of slope coefficients, $ \Theta _{0,it}^{0}=\lambda _{i}^{0\prime }f_{t}^{0}$ is an intercept term that exhibits a factor structure with $r_{0}$ factors, and $e_{it}$ is the error term. Here, we assume $r_{0}$ is a fixed integer that does not change as $\left( N,T\right) \rightarrow \infty .$ Let $\Lambda ^{0}=\left( \lambda _{1}^{0},\cdots ,\lambda _{N}^{0}\right) ^{\prime }$ and $F^{0}=\left( f_{1}^{0},\cdots ,f_{T}^{0}\right) ^{\prime }$. Let $Y=\left\{ Y_{it}\right\} ,$ $X_{j}=\left\{ X_{j,it}\right\} ,$ $\Theta _{j}^{0}=\{\Theta _{j,it}^{0}\}$ and $E=\left\{ e_{it}\right\} ,$ all of which are $N\times T$ matrices. Then we can rewrite ((ref)) in matrix form as

equation[equation omitted — 92 chars of source]

We assume that the slope coefficients follow time-varying latent group structures, viz.,

equation*[equation* omitted — 98 chars of source]

where $\left\{ G_{kt}\right\} _{k\in K_{t}}$ forms a partition of $[N]$ for each specific time $t$ with $K_{t}$ being the number of groups at time $t$. Moreover, we assume that the group-specific slope coefficients $\alpha _{kt}$ or the memberships change at an unknown time point $T_{1}$, i.e.,

eqnarray*[eqnarray* omitted — 406 chars of source]

with $K^{(1)}$ and $K^{(2)}$ being the number of latent groups before and after the break point, respectively. Let $g_{i}^{(1)}$ and $g_{i}^{(2)}$ respectively denote the individual group indices before and after the break:

equation*[equation* omitted — 165 chars of source]

Let $r_{j}$ be the rank of $\Theta _{j}^{0}$ for $j\in \lbrack p]\cup \{0\}$ . It is easy to see that $\Theta _{j}^{0}$ exhibits a low-rank structure for all $j.$ By the SVD, we have

equation*[equation* omitted — 163 chars of source]

where $\mathcal{U}_{j}^{0}\in \mathbb{R}^{N\times r_{j}}$, $\mathcal{V} _{j}^{0}\in \mathbb{R}^{T\times r_{j}}$, $\Sigma _{j}^{0}=\text{diag}(\sigma _{1,j},\cdots ,\sigma _{r_{j},j})$, $U_{j}^{0}=\sqrt{N}\mathcal{U} _{j}^{0}\Sigma _{j}^{0}$ with each row being $u_{i,j}^{0\prime }$, and $ V_{j}^{0}=\sqrt{T}\mathcal{V}_{j}^{0}$ with each row being $v_{t,j}^{0\prime }$.

Note that we allow $\left\{ \Theta _{it}^{0}\right\} _{i=1}^{N}$ to exhibit latent group structures before and after the break. For a particular $j\in \lbrack p]$, the $N\times T$ matrix $\Theta _{j}^{0}$ may have no group structure before or after the break, or no break, or more or fewer groups after the break. Let $K_{j}^{(1)}$ and $K_{j}^{(2)}$ denote the number of groups before and after the break, respectively, for $\{\Theta _{j,it}^{0}\}_{i=1}^{N}$. Let $\mathcal{G}_{j}^{(\ell )}=\{G_{1,j}^{(\ell )},\cdots ,G_{K_{j}^{(\ell )},j}^{(\ell )}\},$ $\ell =1,2,$ denote the associated latent group structures. Define $N_{k,j}^{(\ell )}=|G_{k,j}^{(\ell )}|$ and $\pi _{k,j}^{(\ell )}=\frac{N_{k,j}^{(\ell )}}{N} $ for $\ell =1,2,$ where $\left\vert A\right\vert $ denotes the cardinality of set $A.$ Further define $\tau _{T}:=\frac{T_{1}}{T}$. We show that $ \Theta _{j}^{0}$ has a low-rank structure in all of the following cases:

description$\Theta _{j}^{0}$ exhibits neither structural break nor group structure.\newline In this case, $K_{j}^{(1)}=K_{j}^{(2)}=1$, and $\Theta _{j,it}^{0}=\alpha _{j}$ $\forall \left( i,t\right) \in \lbrack N]\times \lbrack T]$. Without loss of generality, assume that $\alpha _{j}>0.$ Then by the SVD, we have \begin{align*} & \mathcal{U}_{j}=\frac{1}{\sqrt{N}}\iota _{N}\in \mathbb{R}^{N\times 1},\quad \Sigma _{j}=\alpha _{j},\quad \mathcal{V}_{j}=\frac{1}{\sqrt{T}} \iota _{T}\in \mathbb{R}^{T\times 1}, \\ & U_{j}=\alpha _{j}\iota _{N}\in \mathbb{R}^{N\times 1},\quad V_{j}=\iota _{T}\in \mathbb{R}^{T\times 1}, \end{align*} where $\iota _{d}=\left( 1,\cdots ,1\right) ^{\prime }\in \mathbb{R} ^{d\times 1}$ for any natural number $d$. Obviously, $r_{j}=1$ in Case 1. • $\Theta _{j}^{0}$ exhibits no structural break but a group structure.\newline In this case, $K_{j}^{(1)}=K_{j}^{(2)}=K_{j}$, $ G_{k,j}^{(1)}=G_{k,j}^{(2)}=G_{k,j}$, $N_{k,j}^{(1)}=N_{k,j}^{(2)}=N_{k,j}$, $\pi _{k,j}^{(1)}=\pi _{k,j}^{(2)}=\pi _{k,j}~\forall k\in \left[ K_{j} \right] $, and $\Theta _{j,it}^{0}=\operatornamewithlimits{\sum}\limits _{k\in \lbrack K_{j}]}\alpha _{k,j}\mathbf{1}\left\{ i\in G_{k,j}\right\} $ for $t\in \lbrack T]$. Therefore, we have \begin{align*} & \mathcal{U}_{j,i}=\frac{\sum_{k\in \lbrack K_{j}]}\alpha _{k,j}\mathbf{1} \left\{ i\in G_{k,j}\right\} }{\sqrt{\sum_{k\in \lbrack K_{j}]}N_{k,j}\left( \alpha _{k,j}\right) ^{2}}},\quad \Sigma _{j}=\sqrt{\sum_{k\in \lbrack K_{j}]}\pi _{k,j}\left( \alpha _{k,j}\right) ^{2}},\quad \mathcal{V}_{j}= \frac{1}{\sqrt{T}}\iota _{T}, \\ & u_{i,j}=\sum_{k\in \lbrack K_{j}]}\alpha _{k,j}\mathbf{1}\left\{ i\in G_{k,j}\right\} ,\quad V_{j}=\iota _{T}, \end{align*} where $\mathcal{U}_{j,i}$ is the $i$-th element in $\mathcal{U}_{j}$. Obviously, $r_{j}=1$ in this case. • $\Theta _{j}^{0}$ exhibits both a structural break and a group structure. \
itemize$K_{j}^{(1)}\neq K_{j}^{(2)}$, where we have different numbers of groups before and after the break; • $K_{j}^{(1)}=K_{j}^{(2)}=K_{j}$ and $G_{k,j}^{(1)}\neq G_{k,j}^{(2)}$, where we have the same number of groups before and after the break, but the group memberships change after the break point; • $K_{j}^{(1)}=K_{j}^{(2)}=K_{j}$, $ G_{k,j}^{(1)}=G_{k,j}^{(2)}=G_{k,j}$ for $\forall k\in \lbrack K_{j}]$, and $ \alpha _{k,j}^{(1)}\neq \alpha _{k,j}^{(2)}$ for at least one $k\in \lbrack K_{j}]$, where even though neither the number of groups nor group membership changes after the break, there exists at least one group whose slope coefficients change.

For any positive integer $d$, we use $\mathbf{0}_{d}$ to denote a $d\times 1$ vector of zeros. The following lemma lays down the foundation for break point detection in our model.

lemmaFor any $j\in [p]$ such that $\Theta_{j}^{0}$ lies in Case 3 above, we have $rank(\Theta_{j}^{0})\leq 2$. When $ rank(\Theta_{j}^{0})=2$, we have \begin{itemize} • $\Theta_{j}^{0}=\mathcal{U}_{j}\Sigma_{j}\mathcal{V} _{j}^{\prime}=U_{j}V_{j}^{\prime }$ where $U_{j}=\mathcal{U}_{j}\Sigma_{j}/ \sqrt{T},$ $V_{j}=\sqrt{T}\mathcal{V}_{j}=D_{j}R_{j}$$,$ $D_{j}= \begin{bmatrix} \frac{1}{\sqrt{\tau_{T}}}\iota_{T_{1}} & \mathbf{0}_{T_{1}} \\ \mathbf{0}_{T-T_{1}} & \frac{1}{\sqrt{1-\tau_{T}}}\iota_{T-T_{1}} \end{bmatrix} $ and $R_{j}^{\prime }R_{j}=I_{2};$$\left\Vert \frac{v_{t,j}^{0}}{\left\Vert v_{t,j}^{0}\right\Vert }-\frac{v_{t^{\ast },j}^{0}}{\left\Vert v_{t^{\ast },j}^{0}\right\Vert } \right\Vert =\sqrt{2}$ for any $t\leq T_{1}$ and $t^{\ast }>T_{1}$. \end{itemize}

By Lemma (ref) for Case 3 and the above analyses for Cases 1 and 2, we conclude that $\Theta _{j}^{0}$ is a low-rank matrix with rank equal to or less than $2$. In view of the low-rank structure of the slope matrices, we propose to adopt the NNR to obtain the preliminary estimates below. Moreover, under Case 3, Lemma (ref)(ii) indicates that singular vectors of the slope matrix with rank 2 contain the structural break information.

Estimation

In this section we provide the estimation algorithm. We first assume that the ranks $r_{j}$ for $j\in [p]\cup \{0\}$ are known, and then propose a singular value thresholding (SVT) procedure to estimate them. After we recover the break point and the latent group structures, we propose consistent estimates of the group-specific parameters.

Estimation Algorithm

Given $r_{j},~\forall j\in \lbrack p]\cup \{0\}$, we propose the following four-step procedure to estimate the break point and to recover the latent group structures before and after the break.

description• Nuclear Norm Regularization (NNR). We run the nuclear norm regularized regression and obtain the preliminary estimates as follows: \begin{equation} \{\tilde{\Theta}_{j}\}_{j\in \lbrack p]\cup \{0\}}=\operatorname*{arg\,min}\limits_{\left\{ \Theta _{j}\right\} _{j=0}^{p}}\frac{1}{NT}\left\Vert Y-\sum_{j=1}^{p}X_{j}\odot \Theta _{j}-\Theta _{0}\right\Vert ^{2}+\sum_{j=0}^{p}\nu _{j}\left\Vert \Theta _{j}\right\Vert _{\ast }, \end{equation} where $\nu _{j}$ is the tuning parameter for $j\in \lbrack p]\cup \{0\}$. For each $j$, conduct the SVD: $\frac{1}{\sqrt{NT}}\tilde{\Theta}_{j}=\hat{ \tilde{\mathcal{U}}}_{j}\hat{\tilde{\Sigma}}_{j}\hat{\tilde{\mathcal{V}}} _{j}^{\prime }$, where $\hat{\tilde{\Sigma}}_{j}$ is a diagonal matrix with the diagonal elements being the descending singular values of $\tilde{\Theta} _{j}$. Let $\tilde{\mathcal{V}}_{j}$ consist of the first $r_{j}$ columns of $\hat{\tilde{\mathcal{V}}}_{j},$ and $\tilde{V}_{j}=\sqrt{T}\tilde{\mathcal{V }}_{j}.$ Let $\tilde{v}_{t,j}^{\prime }$ denote the $t$-th row of $\tilde{V} _{j}$ for $t\in \lbrack T]$. • Row- and Column-Wise Regressions. First run the row-wise regressions of $Y_{it}$ on $\left( \tilde{v}_{t,0},\{\tilde{v} _{t,j}X_{j,it}\}_{j\in \lbrack p]}\right) $ to obtain $\{\dot{u} _{i,j}\}_{j\in \lbrack p]\cup \{0\}}$ for $i\in \lbrack N]$. Then run the column-wise regressions of $Y_{it}$ on $\left( \dot{u}_{i,0},\{\dot{u} _{i,j}X_{j,it}\}_{j\in \lbrack p]}\right) $ to obtain $\{\dot{v} _{t,j}\}_{j\in \lbrack p]\cup \{0\}}$ for $t\in \lbrack T].$ Let $\dot{\Theta }_{j,it}=\dot{u}_{i,j}^{\prime }\dot{v}_{t,j}$ for $\left( i,t\right) \in \lbrack N]\times \lbrack T]$ and $j\in \lbrack p]\cup \{0\}$. Specifically, the row- and column-wise regressions are given by \begin{align} \left\{ \dot{u}_{i,j}\right\} _{j\in \lbrack p]\cup \{0\}}& =\operatorname*{arg\,min} _{\{u_{i,j}\}_{j\in \lbrack p]\cup \{0\}}}\frac{1}{T}\sum_{t\in \lbrack T]}\left( Y_{it}-u_{i,0}^{\prime }\tilde{v}_{t,0}-\sum_{j=1}^{p}u_{i,j}^{ \prime }\tilde{v}_{t,j}X_{j,it}\right) ^{2},\quad i\in \lbrack N], \\ \left\{ \dot{v}_{t,j}\right\} _{j\in \lbrack p]\cup \{0\}}& =\operatorname*{arg\,min} _{\{v_{t,j}\}_{j\in \lbrack p]\cup \{0\}}}\frac{1}{N}\sum_{i\in \lbrack N]}\left( Y_{it}-v_{t,0}^{\prime }\dot{u}_{i,0}-\sum_{j=1}^{p}v_{t,j}^{ \prime }\dot{u}_{i,j}X_{j,it}\right) ^{2},\quad t\in \lbrack T]. \end{align} • Break Point Estimation. We estimate the break point as follows: \begin{equation} \hat{T}_{1}=\operatorname*{arg\,min}_{s\in \{2,\cdots ,T-1\}}\frac{1}{pNT}\sum_{j\in \lbrack p]}\sum_{i\in \lbrack N]}\left\{ \sum_{t=1}^{s}\left( \dot{\Theta}_{j,it}- \bar{\dot{\Theta}}_{j,i}^{(1s)}\right) ^{2}+\sum_{t=s+1}^{T}\left( \dot{ \Theta}_{j,it}-\bar{\dot{\Theta}}_{j,i}^{(2s)}\right) ^{2}\right\} , \end{equation} where $\bar{\dot{\Theta}}_{j,i}^{(1s)}=\frac{1}{s}\sum_{t=1}^{s}{\dot{\Theta} _{j,it}}$ and $\bar{\dot{\Theta}}_{j,i}^{(2s)}=\frac{1}{T-s}\sum_{t=s+1}^{T}{ \dot{\Theta}_{j,it}}$. • \textbf{Sequential Testing K-means (STK).} In this step, we estimate the number of groups and the group membership before and after the break by using the STK algorithm. For each $j\in \lbrack p]$, define $\dot{\Theta}_{j,i}^{(1)}=(\dot{\Theta}_{j,i1},\cdots ,\dot{\Theta} _{j,i\hat{T}_{1}})^{\prime }$, $\dot{\Theta}_{j,i}^{(2)}=(\dot{\Theta}_{j,i, \hat{T}_{1}+1},\cdots ,\dot{\Theta}_{j,iT})^{\prime }$, $\dot{\beta} _{i}^{(1)}=\frac{1}{\sqrt{\hat{T}_{1}}}(\dot{\Theta}_{1,i}^{(1)\prime },\cdots ,\dot{\Theta}_{p,i}^{(1)\prime })^{\prime },$ and $\dot{\beta} _{i}^{(2)}=\frac{1}{\sqrt{\hat{T}_{2}}}(\dot{\Theta}_{1,i}^{(2)\prime },\cdots ,\dot{\Theta}_{p,i}^{(2)\prime })^{\prime }$. Let $z_{\varsigma }$ be some predetermined value which will be specified in the next subsection. Given the subsample before and after the estimated break point, initialize $m=1$ and classify each subsample into $m$ groups by the K-means algorithm with group membership obtained as $\hat{\mathcal{G} }_{m}^{(\ell )}:=\{\hat{G}_{k,m}^{(\ell )}\}_{k\in \lbrack m]}$. Next, we construct a suitable test statistic $\hat{\Gamma}_{m}^{(\ell )}$, defined by (ref) in the next subsection, and compare it to its critical value $z_{\varsigma }$ at significance level $\varsigma$ under the null hypothesis of $m$ subgroups, setting $m=m+1$ and moving to the next iteration if $\hat{\Gamma}_{m}^{(\ell )}>z_{\varsigma }$ and stopping the STK algorithm otherwise. Lastly, define $\hat{K}^{(\ell )}=m$ and $\hat{\mathcal{G}}^{(\ell )}=\hat{ \mathcal{G}}_{m}^{(\ell )}$. The next subsection provides a detailed breakdown of these steps in the STK algorithm. \begin{figure}[h] \caption{The flow chart of STK algorithm} \end{figure}

Several remarks are in order. First, the ranks of the intercept and slope matrices are assumed known in Step 1 but consistent estimates for them are proposed in the SVT below. Second, we obtain preliminary estimates by NNR based on the low-rank structure of the intercept and slope matrices in the model. These estimates are consistent in terms of the Frobenius norm but pointwise or uniform convergence for their elements is not established. Nonetheless, SVD can be employed to obtain preliminary estimates of the factors and factor loadings to be used subsequently. Third, row- and column-wise linear regressions are conducted to obtain updated estimates of the factors and factor loadings for which we can establish pointwise and uniform convergence rates. Fourth, using the consistent estimates obtained in the second step, we can estimate the break point in Step 3 consistently by using a binary segmentation process. Fifth, the STK algorithm in Step 4 then yields the estimated number of groups and the group memberships together.

In the latent group literature, it is standard and popular to assume the number of groups in the K-means algorithm is known and then to estimate the number of groups by using certain information criteria. In this case, one needs to consider not only under- and just-fitting cases, but also over-fitting cases. It is well known that the major difficulty with this approach is showing that the over-fitting case occurs with probability approaching zero. The STK algorithm ensures a focus on the under-and just-fitting cases, which helps to avoid the difficulty caused by K-means classification with a larger than true number of groups. In addition, although this sequential algorithm approach is adopted, the error from the previous iteration does not accumulate in the following iterations owing to fact that the classification in each iteration is new and is not based on the K-means outcomes in previous iterations.

The STK algorithm

This subsection describes the K-means algorithm and the construction of the test statistics $\hat{\Gamma}_{m}^{(\ell )}$ that are used in the STK algorithm for $ \ell \in \{1,2\}$.

First, we define the objective function for the K-means algorithm with $m$ clusters at each iteration. Let $a_{k,m}^{(\ell )}$ be a $p\hat{T}_{1}\times 1$ and $p(T-\hat{T}_{1})\times 1$ vector for $\ell =1,2$, respectively. We obtain the group membership with $m$ groups by solving the following minimization problem

equation[equation omitted — 300 chars of source]

which yields the membership estimates for each individual at the $m$-th iteration as

equation[equation omitted — 210 chars of source]

Let $\hat{G}_{k,m}^{(\ell )}:=\{i\in \lbrack N]:\hat{g}_{i,m}^{(\ell )}=k\}$.

Second, we discuss the construction of the test statistic based on the idea of homogeneity test for several subsamples. At iteration $m$, we have $m$ potential subgroups $(\hat{G}_{1,m}^{(\ell )},\cdots ,\hat{G}_{m,m}^{(\ell )})$ after the K-means classification for $\ell =1$ and 2. Let $\mathcal{ \hat{T}}_{1}=[\hat{T}_{1}]$, $\mathcal{\hat{T}}_{2}=[T]\backslash \lbrack \hat{T}_{1}]$, $\mathcal{\hat{T}}_{1,-1}=\mathcal{\hat{T}}_{1}\backslash \{ \hat{T}_{1}\}$, $\mathcal{\hat{T}}_{2,-1}=\mathcal{\hat{T}}_{2}\backslash \{T\}$, $\mathcal{\hat{T}}_{1,j}=\{1+j,\cdots ,\hat{T}_{1}\}$, and $\mathcal{ \hat{T}}_{2,j}=\{\hat{T}_{1}+1+j,\cdots ,T\}$ for some specific $j\in \hat{ \mathcal{T}}_{\ell ,-1}$. Based on these estimated subgroups, we can obtain the estimates of the coefficients, factors and factor loadings for each subgroup in regime $\ell $ as follows:

equation*[equation* omitted — 457 chars of source]

where $\hat{\Lambda}_{k,m}^{(\ell )}=\{\hat{\lambda}_{i,k,m}^{\left( \ell \right) }\}_{i\in \hat{G}_{k,m}^{(\ell )}}$ and $\hat{F}_{k,m}^{(\ell )}=\{ \hat{f}_{t,k,m}^{\left( \ell \right) }\}_{t\in \hat{\mathcal{T}}_{\ell }}.$ For all $i\in \lbrack N]$ and $t\in \lbrack T]$, define the residuals

equation*[equation* omitted — 226 chars of source]

Let $\hat{X}_{i}^{(1)}=(X_{i1},\cdots ,X_{i\hat{T}_{1}})^{\prime }$, $\hat{X} _{i}^{(2)}=(X_{i,\hat{T}_{1}+1},\cdots ,X_{iT})^{\prime }$, and $\hat{T} _{2}=T-\hat{T}_{1}.$ Define

align*[align* omitted — 642 chars of source]

Let $\hat{z}_{it}^{(\ell )\prime }$ be the $t$-th row of $M_{\hat{F} _{k,m}^{(\ell )}}\hat{X}_{i}^{(\ell )}.$ For each subgroup $\hat{G} _{k,m}^{(\ell )}$ with $k\in \lbrack m]$, we follow the lead of pesaran2008testing and ando2015simple and define the following test statistic components

equation[equation omitted — 230 chars of source]

where

align*[align* omitted — 832 chars of source]

$k(\cdot )$ is a kernel function, $S_{T}$ is a bandwidth/truncation parameter, and $\hat{\Omega}_{i,k,m}^{(\ell )}$ is a traditional HAC estimator. Using the components (ref) we now define the test statistic

align[align omitted — 117 chars of source]

We will show that $\hat{\Gamma}_{m}^{(\ell )}$ is asymptotically distributed as the maximum of $m$ independent $\chi ^{2}(1)$ random variables under the null hypothesis that the slope coefficients in each of the $m$ subsamples are homogeneous, whereas it diverges to infinity under the alternative. Let $ z_{\varsigma }$ denote the critical value at significance level $\varsigma ,$ which is calculated from the maximum of $m$ independent $\chi ^{2}(1)$ random variables. We reject the null of $m$ subgroups in favor of more groups at level $\varsigma $ if $\hat{\Gamma}_{m}^{(\ell )}>z_{\varsigma }.$

Rank Estimation

To obtain the rank estimator, we use the low-rank estimators from ((ref) ) and estimate $r_{j}$ by the singular value thresholding (SVT) criterion

equation*[equation* omitted — 247 chars of source]

where $\sigma _{i}\left( A\right) $ denotes the $i$-th largest singular value of $A$ and $N\wedge T=\min (N,T).$ By arguments as used in the proof of Proposition D.1 in chernozhukov2019inference and that of Theorem 3.2 in hong2023profile, we can show that $\mathbb{P}(\hat{r} _{j}=r_{j})\rightarrow 1$ for each $j$ as $\left( N,T\right) \rightarrow \infty .$

Parameter Estimation

Once we obtain the estimated break point, the number of groups and the group membership before and after the estimated break point, we can estimate the group-specific slope coefficients $\{\alpha _{k}^{(\ell )}\}_{k\in \lbrack \hat{K}^{(\ell )}]}$ along with the factors and factor loadings as follows

equation[equation omitted — 327 chars of source]

where $\mathbb{L}\left( \Lambda ,F,\left\{ a_{k}^{(\ell )}\right\} _{k\in \lbrack \hat{K}^{(\ell )}]}\right) =\frac{1}{N\hat{T}_{\ell }}\sum_{k=1}^{ \hat{K}^{(\ell )}}\sum_{i\in \hat{G}_{k}^{(\ell )}}\sum_{t\in \mathcal{\hat{T }}_{\ell }}\left( Y_{it}-\lambda _{i}^{\prime }f_{t}-X_{it}^{\prime }a_{k}^{(\ell )}\right) ^{2}.$ Here, we ignore the fact that the prior- and post-break regimes share the same set of factor loadings and estimate the group-specific parameters separately for the two regimes at the cost of sacrificing some efficiency for the factor loading estimates. Alternatively, we can pool the observations before and after the break to estimate the parameters as

equation*[equation* omitted — 370 chars of source]

where

equation[equation omitted — 363 chars of source]

In either case, as should be clear because of the presence of the group structures, establishment of the asymptotic properties of the post-classification estimators of the group-specific slope coefficients becomes much more involved than in bai2009panel and moon2017dynamic. For this reason, we will focus on the estimates defined in ((ref)).

Asymptotic Theory

This section develops the asymptotic properties of the estimators introduced above.

Basic Assumptions

Define $e_{i}=\left( e_{i1},\cdots ,e_{iT}\right) ^{\prime }$ and $ X_{j,i}=\left( X_{j,i1},\cdots ,X_{j,iT}\right) ^{\prime }.$ Let $V_{j}^{0}$ be a $T\times r_{j}$ matrix with its $t$-th row being $v_{t,j}^{0\prime }$, and $U_{j}^{0}$ be the $N\times r_{j}$ matrix with its $i$-th row being $ u_{i,j}^{0\prime }$. Throughout the paper, we treat the factors $ \{V_{j}^{0}\}_{j\in \lbrack p]\cup \{0\}}$ as random and their loadings $ \{U_{j}^{0}\}_{j\in \lbrack p]\cup \{0\}}$ as deterministic. Let $\mathscr{D} :=\sigma (\{V_{j}^{0}\}_{j\in \lbrack p]\cup \{0\}})$ denote the minimum $ \sigma $-field generated by $\{V_{j}^{0}\}_{j\in \lbrack p]\cup \{0\}}.$ Similarly, let $\mathscr{G}_{t}:=\sigma (\mathscr{D},\left\{ X_{is}\right\} _{i\in \lbrack N],s\leq t+1},\left\{ e_{is}\right\} _{i\in \lbrack N],s\leq t})$. Let $\max_{i}=\max_{i\in \lbrack N]},$ $\max_{t}=\max_{t\in \lbrack T]}\ $and $\max_{i,t}=\max_{i\in \lbrack N],t\in \lbrack T]}.$ Let $M$ and $ C $ be generic bounded positive constants which may vary across lines. \

ass\begin{enumerate} • $\left\{ e_{it},X_{it}\right\}_{t\in [T]}$ are conditionally independent across $i$ given $\mathscr{D}$. • $\mathbb{E}(e_{it}|X_{it},\mathscr{D})=0$. • For each $i$, $\left\{ \left( e_{it},X_{it}\right) ,t\geq 1\right\} $ is strong mixing conditional on $\mathscr{D}$ with the mixing coefficient $\alpha _{i}(\cdot )$ satisfying $\max_{i}\alpha _{i}(z)\leq M\vartheta ^{z}$ for some constant $\vartheta \in \left( 0,1\right) $. • There exists a constant $C>0$ such that $\max_{i}\frac{1}{T} \sum_{t\in \lbrack T]}\left\Vert \xi _{it}\right\Vert ^{2}\leq C~a.s.\quad $ and$\quad \max_{t}\frac{1}{N}\sum_{i\in \lbrack N]}\left\Vert \xi _{it}\right\Vert ^{2}$ $\leq C~a.s.$ for $\xi _{it}=e_{it},$ $X_{it}$ and $ X_{it}e_{it}$. • $\max_{i,t}\mathbb{E}[\left\Vert \xi _{it}\right\Vert ^{q}\big| \mathscr{D}]\leq M~a.s.$ and $\max_{i,i^{\ast },t}\mathbb{E}[\left\Vert X_{it}e_{i^{\ast }t}\right\Vert ^{q}\big|\mathscr{D}]\leq M~a.s.$ for some $ q>8$ and $\xi _{it}=e_{it},$ $X_{it}$ and $X_{it}e_{it}.$ • As $(N,T)\rightarrow \infty $, $\sqrt{N}(\log N)^{2}T^{-1}\rightarrow 0$ and $T(\log N)^{2}N^{-3/2}\rightarrow 0$. \end{enumerate}
assp{\ref*{ass:1}{$^*$} } (i), (iv) and (v) are same as Assumption (ref)(i), (iv) and (v). In addition: \begin{itemize} • $\mathbb{E}(e_{it}|\mathscr{G}_{t-1})=0$ $\forall (i,t)\in \lbrack N]\times \lbrack T]$, and $\max_{i,t}\mathbb{E(}e_{it}^{2}\big| \mathscr{G}_{t-1})\leq M~a.s.$. • $\left\{ e_{it}\right\} _{i\in \lbrack N]}$ is conditionally independent across $t$ given $\mathscr{D}$. \end{itemize}

Assumption (ref)(i) imposes conditional independence on $\left\{ e_{it},X_{it}\right\} _{t\in \lbrack T]}$ across the cross sectional units. Assumption (ref)(ii) is the conditional moment condition. Assumption (ref)(iii) imposes conditional strong mixing conditions along the time dimension. See Prakasa_Rao2009 for the definition of conditional strong mixing and su2013testing for an application in the panel setup. Assumptions (ref)(iv)-(v) impose conditions that restrict the tail behavior of $\xi _{it}$. Note that neither the regressors nor the errors are constrained to be bounded. Assumption (ref) (vi) imposes restrictions on $N$ and $T$ but does not require $N$ and $ T$ to diverge at the the same rate. It is possible to allow $N$ to diverge to infinity faster but not too much faster than $T$, and vice versa.

Assumption (ref) is used for the study of dynamic panel data models. To be specific, Assumption (ref)(ii) requires that the error sequence $ \left\{ e_{it},t\geq 1\right\} $ be a martingale difference sequence (m.d.s.) with respect to the filtration $\mathscr{G}_{t},$ which allows for lagged dependent variables in $X_{it}$. Assumption (ref)(iii) imposes conditional independence of the errors over $t$. The presence of serially correlated errors in dynamic panels typically induces endogeneity, which invalidates least-squares-based PCA estimation.

assrank$(\Theta _{j}^{0})=r_{j}\leq \bar{r}$ for $j\in \lbrack p]\cup \{0\}$ and some fixed $\bar{r},$ and $\max_{j\in \lbrack p]\cup \{0\}}||\Theta _{j}^{0}||_{\max }\leq M$.

Assumption (ref) imposes low-rank conditions on the coefficient matrices, which facilitate the use of NNR in obtaining preliminary estimates in the first step. As discussed in the previous section, we see that the low-rank assumption for the slope matrices is satisfied for the model in Section (ref). Moreover, we follow ma2020detecting and assume the elements of the coefficient matrices are uniformly bounded to simplify the proofs. The boundedness of the slope coefficients is reasonable given that their cardinality does not grow with the sample size. The boundedness assumption for the intercept coefficient can be relaxed at the cost of more lengthy arguments.

assLet $\sigma _{l,j}$ denote the $l$-th largest singular values of $\Theta _{j}^{0}$ for $j\in \lbrack p]\cup \{0\}.$ There exist some constants $C_{\sigma }$ and $c_{\sigma }$ such that \begin{equation*} \infty >C_{\sigma }\geq \lim \sup_{(N,T)\rightarrow \infty }\max_{j\in \lbrack p]}\sigma _{1,j}\geq \lim \inf_{(N,T)\rightarrow \infty }\min_{j\in \lbrack p]}\sigma _{r_{j},j}\geq c_{\sigma }>0. \end{equation*}

Assumption (ref) imposes some conditions on the singular values of the coefficient matrices. These ensure that only pervasive factors are allowed when the matrices are written as a factor structure. The assumption can be readily verified given the latent group structures of the slope coefficients.

Consider the SVD for $\Theta _{j}^{0}$: $\Theta _{j}^{0}=\mathcal{U} _{j}\Sigma _{j}\mathcal{V}_{j}^{\prime }$ $\forall j\in \lbrack p]\cup \{0\}$ . Decompose $\mathcal{U}_{j}=\left( \mathcal{U}_{j,r},\mathcal{U} _{j,0}\right) $ and $\mathcal{V}_{j}=\left( \mathcal{V}_{j,r},\mathcal{V} _{j,0}\right) $ with $\left( \mathcal{U}_{j,r},\mathcal{V}_{j,r}\right) $ being the singular vectors corresponding to nonzero singular values and $ \left( \mathcal{U}_{j,0},\mathcal{V}_{j,0}\right) $ being the singular vectors corresponding to zero singular values. Hence, for any matrix $W\in \mathbb{R}^{N\times T}$, we define

equation*[equation* omitted — 229 chars of source]

where $\mathcal{P}_{j}\left( W\right) $ can be seen as the linear projection of matrix $W$ into the low-rank space with $\mathcal{P}_{j}^{\bot }\left( W\right) $ being its orthogonal space. Let $\Delta _{\Theta _{j}}=\Theta _{j}-\Theta _{j}^{0}$ for any $\Theta _{j}$. Based on the spaces constructed above, with some positive constants $C_{1}$ and $C_{2}$, we define the restricted set for full-sample parameters as follows:

align[align omitted — 449 chars of source]

Lemma (ref) in the online appendix shows that our nuclear norm estimators are in a restricted set larger than ((ref)), which derives from the restriction on the Frobenius norm in the definition of $\mathcal{R} \left( C_{1},C_{2}\right) $. Intuitively, the first restriction in ((ref) )\ means the projection onto the orthogonal low-rank space of the estimator error can be controlled by its projection onto the low-rank space. Theorem (ref) below largely hinges on this property.

assFor any $C_{2}>0$, there are constants $C_{3}$ and $C_{4}$ such that for any $(\{\Delta _{\Theta _{j}}\}_{j\in \lbrack p]\cup \{0\}})$ $ \in \mathcal{R}(3,C_{2})$, we have \begin{equation*} \left\Vert \Delta _{\Theta _{0}}+\sum_{j=1}^{p}\Delta _{\Theta _{j}}\odot X_{j}\right\Vert ^{2}\geq C_{3}\sum_{j\in \lbrack p]\cup \{0\}}\left\Vert \Delta _{\Theta _{j}}\right\Vert ^{2}-C_{4}(N+T)\quad w.p.a.1. \end{equation*}

Assumption (ref) imposes the restricted strong convexity (RSC) condition, which is similar to Assumption 3.1 in chernozhukov2019inference. The latter authors also provide some sufficient conditions to verify such an assumption.

Let $r=\sum_{j\in \lbrack p]\cup \{0\}}r_{j}.$ Define the following $r\times r$ matrices:

equation*[equation* omitted — 228 chars of source]

where $\phi _{it}^{0}=(v_{t,0}^{0\prime },v_{t,1}^{0\prime }X_{1,it},\cdots ,v_{t,p}^{0\prime }X_{p,it})^{\prime },$ and $\psi _{it}^{0}=(u_{i,0}^{0\prime },u_{i,1}^{0\prime }X_{1,it},\cdots ,u_{i,p}^{0\prime }X_{p,it})^{\prime }.$

assThere exist constants $C_{\phi}$ and $c_{\phi}$ such that \begin{align*} &\infty>C_{\phi}\geq \limsup_{T}\max_{t\in[T]}\lambda_{\max}(\Psi_{t})\geq \liminf_{T}\min_{t\in[T]}\lambda_{\min}(\Psi_{t})\geq c_{\phi}>0, \\ &\infty>C_{\phi}\geq \limsup_{N}\max_{i\in [N]}\lambda_{\max}(\Phi_{i})\geq\liminf_{N}\min_{i\in [N]}\lambda_{\min}(\Phi_{i})\geq c_{\phi}>0. \end{align*}

Assumption (ref) is similar to Assumption 8 in ma2020detecting and it imposes some rank conditions.

Asymptotics of NNR Estimators and Singular Vector Estimators

Let $\eta _{N,1}=\frac{\sqrt{\log T}}{\sqrt{N\wedge T}}$ and $\eta _{N,2}= \frac{\sqrt{\log (N\vee T)}}{\sqrt{N\wedge T}}(NT)^{1/q}.$ Let $\tilde{\sigma }_{k,j}$ denote the $k$-th largest singular value of $\tilde{\Theta}_{j}$ for $j\in \lbrack p]\cup \{0\}.$ Our first main result is about the consistency of the first-stage NNR estimators and the second-stage singular vector estimators.

theoremSuppose that Assumptions (ref)--(ref) hold. Then $ \forall j\in \lbrack p]\cup \{0\}$, we have \begin{itemize} • $\frac{1}{\sqrt{NT}}||\tilde{\Theta}_{j}-\Theta _{j}^{0}||=O_{p}(\eta _{N,1})$, $\max_{k\in \lbrack r_{j}]}|\tilde{\sigma} _{k,j}-\sigma _{k,j}|=O_{p}(\eta _{N,1})$, and $||V_{j}^{0}-\tilde{V} _{j}O_{j}||=O_{p}(\sqrt{T}\eta _{N,1})$ where $O_{j}$ is an orthogonal matrix defined in the proof. If in addition Assumption (ref) is also satisfied, then we have • $\operatornamewithlimits{\max}\limits_{i\in \lbrack N]}||\dot{u} _{i,j}-O_{j}u_{i,j}^{0}||=O_{p}(\eta _{N,2}),$ $\operatornamewithlimits{\max} \limits_{t\in \lbrack T]}\left\Vert \dot{v}_{t,j}-O_{j}v_{t,j}^{0}\right \Vert _{2}=O_{p}(\eta _{N,2})$, • $\operatornamewithlimits{\max}\limits_{i\in \lbrack N],t\in \lbrack T]}|\dot{\Theta}_{j,it}-\Theta _{j,it}^{0}|=O_{p}(\eta _{N,2})$. \end{itemize}

Theorem (ref)(i) reports the error bounds for $\tilde{\Theta}_{j},$ $ \tilde{\sigma}_{k,j},$ and $\tilde{V}_{j}$. The $\log T$ term in the numerator of $\eta _{N,1}$ is due to the use of some exponential inequality for (conditional)\ strong mixing processes. Theorem (ref)(ii)--(iii) report the uniform convergence rate of the factor and factor loading estimators. The extra $(NT)^{1/q}$ term in the definition of $\eta _{N,2}$ is due to the nonboundedness of $X_{j,it}$ in Assumption (ref)(v), and it disappears when $X_{j,it}$ is assumed to be uniformly bounded.

Consistency of the Break Point Estimate

Recall that $g_{i}^{(1)}$ and $g_{i}^{(2)}$ denote the true group individual $i$ belongs to before and after the break, respectively. To estimate the break point consistently, we add the following condition.

ass\begin{enumerate} • $\sqrt{\frac{1}{N}\sum_{i\in \lbrack N]}||\alpha _{g_{i}^{(1)}}-\alpha _{g_{i}^{(2)}}||^{2}}=C_{5}\zeta _{NT}$, where $C_{5}$ is a positive constant and $\zeta _{NT}\gg \eta _{N,2}$. • $\tau_{T}:=\frac{T_{1}}{T}\rightarrow \tau \in (0,1)$ as $ T\rightarrow \infty.$ \end{enumerate}

Assumption (ref)(i) imposes conditions on the break size in order to identify the break point. Note that we allow the average break size to shrink to zero at a rate slower than $\sqrt{\frac{\log (N\vee T)}{N\wedge T }}(NT)^{1/q}$. This rate is of much bigger magnitude than the optimal $ \left( NT\right) ^{-1/2}$-rate that can be detected in the panel threshold regressions (PTRs) for several reasons. First, in PTRs, the slope coefficients are usually assumed to be homogeneous so that each individual is subject to the same change in the slope coefficients and one can use the cross-sectional information effectively. In contrast, we allow for heterogeneous slope coefficients here and the change can occur only for a subset of cross section units but not all. In addition, in the presence of latent group structures, we not only allow the slope coefficients of some specific groups to change with group membership fixed, but also allow the slope coefficient to remain the same for some groups while the group memberships change after the break. Second, our break point estimation relies on the binary segmentation idea borrowed from the time series literature where one can allow break sizes of bigger magnitude than $ T^{-1/2} $ in order to identify the break ratio consistently but not the break point consistently. As is apparent, even though we require bigger break sizes, we can estimate the break date consistently by using information from both the cross section and time dimensions. Third, as mentioned above, the additional term $\log (N\vee T)$ in the above rate is mainly due to the use of an exponential inequality and the factor $(NT)^{1/q}$ is due to the fact that we only assume the existence of $q$-th order moments for some random variables.

The following theorem indicates that we can estimate the break date $T_{1}$ consistently.

theoremSuppose Assumptions (ref)--(ref) hold, with the true break point being $T_{1}$ and the estimator defined in ((ref)). Then $\mathbb{P}(\hat{T}_{1}=T_{1})\rightarrow 1$ as $ \left( N,T\right) \rightarrow \infty $.

Theorem (ref) shows that we can estimate the true break date consistently w.p.a.1 despite the fact that we allow the break size to shrink to zero at a certain rate.

Consistency of the Estimates of the Number of Groups and the Latent Group Structures

To study the asymptotic properties of the estimates of the number of groups and the recovery of the latent group structures, we first add the following definition.

definitionFix $K^{(\ell)}>1$ and $m\leq K^{(\ell)}$. The estimated group structure $\hat{\mathcal{G}}_{m}^{(\ell)}$ satisfies the non-splitting property (NSP) if for any pair of individuals in the same true group, the estimated group labels are the same.

Definition (ref) describes the non-splitting property introduced by Jin2022Optimal. The latter authors show that the STK algorithm yields the estimated group structures enjoying the NSP.

To proceed, we add the following assumptions.

ass\begin{enumerate} • Let $k_{s}$ and $k_{s^{\ast }}$ be different group indices. Assume that $\min_{1\leq k_{s}<k_{s^{\ast }}\leq K^{(\ell )}}\left\Vert \alpha _{k_{s}}^{(\ell )}-\alpha _{k_{s^{\ast }}}^{(\ell )}\right\Vert _{2}$ $\geq C_{5}$ for$~\ell \in \{1,2\}.$ • Let $N_{k}^{(\ell )}$ be the number of individuals in group $k\ $ for $k\in \lbrack K^{(\ell )}]$. Define $\pi _{k}^{(\ell )}=\frac{ N_{k}^{(\ell )}}{N}$ for $\ell =1,2.$ Assume $K^{(\ell )}$ is fixed and \begin{equation*} \infty >\bar{C}\geq \lim \sup_{N}\sup_{k\in \lbrack K^{(\ell )}]}\pi _{k}^{(\ell )}\geq \lim \inf_{N}\inf_{k\in \lbrack K^{(\ell )}]}\pi _{k}^{(\ell )}\geq c>0, \ell =1,2. \end{equation*} • For any combination of the collection of $n$ true groups with $ n\geq 2$, we have \begin{equation*} \frac{T_{\ell }}{\sqrt{N}}\sum_{s=1}^{n}N_{k_{s}}^{(\ell )}\left\Vert \sum_{s^{\ast }\in \left[ n\right] ,s^{\ast }\neq s}(\alpha _{k_{s^{\ast }}}^{(\ell )}-\alpha _{k_{s}}^{(\ell )})\right\Vert ^{2}/(\log N)^{1/2}\rightarrow \infty , \ell =1,2. \end{equation*} \end{enumerate}

Remark 3. Assumption (ref)(i)--(ii) is the standard assumption for K-means algorithm, which is comparable to Assumption 4 in su2020strong and greatly facilitates the subsequent analyses. Assumption (ref)(i) assumes that the minimum distance of two distinct groups is bounded away from 0, and Assumption (ref)(ii) imposes that each group has asymptotically non-negligible number of units. For Assumption (ref)(iii), it can be shown to hold under mild conditions. Below we explain this assumption in detail. When $n=2$, it is clear that

align*[align* omitted — 553 chars of source]

by Assumptions (ref)(ii) and (ref)(i)--(ii). When $n>2$, we consider a special case such that $S_{s}=:||\sum_{s^{\ast }\in \left[ n \right] ,s^{\ast }\neq s}(\alpha _{k_{s^{\ast }}}^{(\ell )}-\alpha _{k_{s}}^{(\ell )})||=0$ for some specific $s=s_{0}\in \lbrack n]$. Then it is easy to see $S_{s}$ is non-zero for all $s\in \lbrack n]\backslash \{s_{0}\}$ under Assumption (ref)(i). Hence, if we assume $S_{s}$ is lower bounded by a constant $c$ for any $s\in \lbrack n]\backslash \{s_{0}\}$ , Assumption (ref)(iii) will holds naturally. Similar arguments hold for the other general cases.

assLet $\mathcal{T}_{1}=\left[ T_{1}\right] \ $and $\mathcal{T} _{2}=[T]\backslash \left[ T_{1}\right] $. \begin{itemize} • $\frac{1}{T_{\ell }}\sum_{t\in \mathcal{T}_{\ell }}f_{t}^{0}f_{t}^{0\prime }\overset{p}{\longrightarrow }\Sigma _{F}^{(\ell )}>0$ as $T\rightarrow \infty $. $\frac{1}{N_{k}^{(\ell )}}\Lambda _{k}^{0,\left( \ell \right) \prime }\Lambda _{k}^{0,\left( \ell \right) } \overset{p}{\longrightarrow }\Sigma _{\Lambda ,k}^{\left( \ell \right) }>0$ as $N\rightarrow \infty $, where $\Lambda _{k}^{0,\left( \ell \right) }$ is a stack of $\lambda _{i}^{0}$ for all individuals in group $k\ $and $k\in \lbrack K^{(\ell )}]$. • There exists a constant $C>0$ such that $\max_{i}\frac{1}{ T_{\ell }}\sum_{t\in \mathcal{T}_{\ell }}\left\Vert \xi _{it}\right\Vert ^{2}\leq C~a.s.$ for $\xi _{it}=e_{it},$ $X_{it}$ and $X_{it}e_{it}.$ \end{itemize}

Assumption (ref)(i) imposes some standard assumptions on the factors and factor loadings. Assumption (ref)(ii) is similar as Assumption (ref)(iv), which strengthens Assumption (ref)(iv) to hold for two time regimes.

The next theorem reports the asymptotic properties of the STK estimators.

theoremLet $\varsigma =\varsigma _{N}\rightarrow 0$ at rate $N^{-c}$ for some $c>0$ as $N\rightarrow \infty .$ Suppose that Assumption (ref) and Assumptions (ref)--(ref) hold. Then for $\ell \in \{1,2\}$, we have \begin{itemize} • if $m=K^{(\ell )}$, \begin{itemize} • $\operatornamewithlimits{\max}\limits_{i\in \lbrack N]}\mathbf{1} \{\hat{g}_{i,K^{(\ell )}}^{(\ell )}\neq g_{i}^{(\ell )}\}=0$ w.p.a.1, • $\hat{\Gamma}_{K^{(\ell )}}^{(\ell )}$ is asymptotically distributed as the maximum of $K^{(\ell )}$ independent $\chi ^{2}(1)$ random variables, • $\mathbb{P}(\hat{K}^{(\ell )}\leq K^{(\ell )})\geq 1-\alpha +o(1)$ , \end{itemize} • if $m<K^{(\ell )}$, $\hat{\Gamma}_{m}^{(\ell )}/\log N\rightarrow \infty $ w.p.a.1. Thus $\mathbb{P}(\hat{K}^{(\ell )}\neq K^{(\ell )})\leq \varsigma +o(1)$. \end{itemize}

Theorem (ref) studies the asymptotic properties of the STK algorithm. Since we allow $\varsigma =\varsigma _{N}$ to shrink to zero at rate $N^{-c}, $ the critical value $z_{\varsigma }$ diverges to infinity at rate $\log N$ as $N\rightarrow \infty $ by virtue of the tail properties of $\chi ^{2}(1)$ random variables. At iteration $m$ such that $m<K^{(\ell )}$, w.p.a.1, the test statistics $\hat{\Gamma}_{m}^{(\ell )}$ diverges to infinity at a rate faster than $\log N$, which means the iteration will continue at the $(m+1)$ -th iteration. At iteration $m$ with $m=K^{(\ell )}$, however, we can easily find that $z_{\varsigma }\rightarrow \infty $ while the test statistic $ \hat{\Gamma}_{m}^{(\ell )}$ is stochastically bounded. As a result, the iteration stops w.p.a.1 and we have $\mathbb{P}(\hat{K}^{(\ell )}=K^{(\ell )})\rightarrow 1$. As aforementioned, Theorem (ref) ensures the application of K-means algorithm only for the under-fitting and just-fitting cases and it avoids the theoretical challenge of handling the over-fitting case in the classification.

For dynamic panels, we can focus on Assumption (ref), where the error term is an m.d.s. Under this assumption, the HAC estimator $\hat{\Omega} _{i,k,m}^{(\ell )}$ degenerates to $\frac{1}{\hat{T}_{\ell }}\sum_{t\in \hat{ \mathcal{T}}_{\ell }}\hat{z}_{it}^{(\ell )}\hat{z}_{it}^{(\ell )\prime }\hat{ e}_{it}^{2}$. For static panels, we typically allow for serially correlated errors and employ the HAC estimator, and the results in Theorem (ref) continue to hold with Assumption (ref) replaced by Assumption (ref). For the kernel function and bandwidth, we can follow andrews1991heteroskedasticity and let $k\left( \cdot \right) $ belong to the following class of kernels

align*[align* omitted — 251 chars of source]

See, e.g., andrews1991heteroskedasticity and white2014asymptotic for details.

Distribution Theory for the Group-specific Slope Estimators

For $\ell \in \{1,2\}$, let $\{\hat{\alpha}_{k}^{\ast (\ell )}\}_{k\in K^{(\ell )}}$ be the oracle estimators of the group-specific slope coefficients before and after the break point by using the true break and membership information for all individuals. To study the asymptotic distribution theory for $\{\hat{\alpha}_{k}^{(\ell )}\}_{k\in K^{(\ell )}},~\ell \in \{1,2\}$, we only need to show that for the oracle estimators $ \{\hat{\alpha}_{k}^{\ast (\ell )}\}_{k\in K^{(\ell )}}$ based on Theorems (ref) and (ref) by extending the result of bai2009panel and moon2017dynamic.

To proceed, we add some notation. For $\ell \in \{1,2\}$, we first define the matrix notation for each subgroup. For $j\in \lbrack p]$, let $ X_{j,i}^{(1)}=\left( X_{j,i1},\cdots ,X_{j,iT_{1}}\right) ^{\prime }$, $ X_{j,i}^{(2)}=\left( X_{j,i(T_{1}+1)},\cdots ,X_{j,iT}\right) ^{\prime }$, $ e_{i}^{(1)}=\left( e_{i1},\cdots ,e_{iT_{1}}\right) ^{\prime }$ and $ e_{i}^{(2)}=\left( e_{i(T_{1}+1)},\cdots ,e_{iT}\right) ^{\prime }$. Then we use $\mathbb{X}_{j,k}^{(\ell )}\in \mathbb{R}^{N_{k}^{(\ell )}\times T_{\ell }}$ and $E_{k}^{(\ell )}\in \mathbb{R}^{N_{k}^{(\ell )}\times T_{\ell }}$ to denote the regressor and error matrix for subgroup $k\in \lbrack K^{(\ell )}] $ with each row being $X_{j,i}^{(\ell )}$ and $e_{i}^{(\ell )}$, respectively. Let $\mathcal{X}_{j,k}^{(\ell )}=M_{\Lambda _{k}^{0,(\ell )}} \mathbb{X}_{j,k}^{(\ell )}M_{F^{0,(\ell )}}\in \mathbb{R}^{N_{k}^{(\ell )}\times T_{\ell }}$ with the $\left( i,t\right) $-th entry given by $ \mathcal{X}_{j,k,it}^{(\ell )}.$ Let $\mathcal{X}_{k,it}^{(\ell )}=(\mathcal{ X}_{1,k,it}^{(\ell )},\cdots ,\mathcal{X}_{p,k,it}^{(\ell )})^{\prime }$. Further define

align*[align* omitted — 1,311 chars of source]

Let $\mathbb{W}_{NT,k}^{(\ell )}$ be a $p\times p$ matrix with $(j_{1},j_{2}) $-th entry $\frac{1}{N_{k}^{(\ell )}T_{\ell }}tr\left( M_{F^{0,(\ell )}}\mathbb{X}_{j_{1},k}^{(\ell )\prime }M_{\Lambda _{k}^{0,(\ell )}}\mathbb{X }_{j_{2},k}^{(\ell )}\right) $. Then we define the overall bias term for each subgroup as

equation*[equation* omitted — 200 chars of source]

where $\rho _{k}^{(\ell )}=\sqrt{\frac{N_{k}^{(\ell )}}{T_{\ell }}}$. To state the main result in this subsection, we add the following assumption.

ass\begin{itemize} • As $(N,T)\to\infty$, $T(\log T)N^{-4/3}\to 0$. • $\operatorname*{plim}_{(N,T)\rightarrow \infty }\frac{1}{N_{k}^{(\ell )}T_{\ell }}\sum_{i\in G_{k}^{(\ell )}}\sum_{t\in \mathcal{T}_{\ell }}X_{it}X_{it}^{\prime }>0$\ for $\ell \in \{1,2\}$ and $k\in \lbrack K^{(\ell )}]$. • For $\ell\in\{1,2\}$ and $k\in[K^{(\ell)}]$, separate the $p$ regressors of each subgroups into $p_{1}$ “low-rank regressors" $\mathbb{X} _{j,k}^{(\ell)}$ such that $rank(\mathbb{X}_{j,k}^{(\ell)})=1,~\forall j\in\{1,\cdots,p_{1}\},$ and “high-rank regressors" $\mathbb{X} _{j,k}^{(\ell)}$ such that $rank(\mathbb{X}_{j,k}^{(\ell)})>1,~\forall j\in\{p_{1}+1,\cdots,p\}.$ Let $p_{2}:=p-p_{1}$. These two types of regressors satisfy: \begin{itemize} • Consider the linear combinations $b\cdot \mathbb{X} _{high,k}^{(\ell )}:=\sum_{j=p_{1}+1}^{p}b_{j}\mathbb{X}_{j,k}^{(\ell )}$ for high-rank regressors with $p_{2}$-vectors $b$ such that $\left\Vert b\right\Vert _{2}=1$ and $b=\left( b_{p_{1}+1},\cdots ,b_{p}\right) ^{\prime }$. There exists a positive constant $C_{b}$ such that \begin{equation*} \operatornamewithlimits{\min}_{\left\{ \left\Vert b\right\Vert _{2}=1\right\} }\sum_{n=2r_{0}+p_{1}+1}^{N}\lambda _{n}\left[ \frac{1}{ NT_{\ell }}\left( b\cdot \mathbb{X}_{high,k}^{(\ell )}\right) \left( b\cdot \mathbb{X}_{high,k}^{(\ell )}\right) ^{\prime }\right] \geq C_{b}\quad w.p.a.1. \end{equation*} • For $j\in \lbrack p_{1}]$, write $\mathbb{X}_{j,k}^{(\ell )}=w_{j,k}^{(\ell )}v_{j,k}^{(\ell )\prime }$ with $N_{k}^{(\ell )}$-vectors $w_{j}^{(\ell )}$ and $T_{\ell }$-vectors $v_{j}^{(\ell )}$. Let $ w_{k}^{(\ell )}=(w_{1,k}^{(\ell )},\cdots ,w_{p_{1},k}^{(\ell )})\in \mathbb{ R}^{N\times p_{1}}$, $v^{(\ell )}=(v_{1}^{(\ell )},\cdots ,v_{p_{1}}^{(\ell )})\in \mathbb{R}^{T_{\ell }\times p_{1}}$, $M_{w_{k}^{(\ell )}}=I_{N_{k}^{(\ell )}}-w_{k}^{(\ell )}(w_{k}^{(\ell )\prime }w_{k}^{(\ell )})^{-1}w_{k}^{(\ell )\prime }$ and $M_{v^{(\ell )}}=I_{T_{\ell }}-v^{(\ell )}\left( v^{(\ell )\prime }v^{(\ell )}\right) ^{-1}v^{(\ell )\prime }$. There exists a positive constant $C_{B}$ such that $(N_{k}^{(\ell )})^{-1}\Lambda _{k}^{0,(\ell )\prime }M_{w_{k}^{(\ell )}}\Lambda _{k}^{0,(\ell )}>C_{B}I_{r_{0}}$ and $T_{\ell }^{-1}F^{0,(\ell )\prime }M_{w^{(\ell )}}F^{0,(\ell )}>C_{B}I_{r_{0}}$ w.p.a.1. \end{itemize} • For $\forall j\in \lbrack p]$, $\ell \in \{1,2\}$, and $k\in K^{(\ell )}$, \begin{equation*} \frac{1}{N_{k}^{(\ell )}\left( T_{\ell }\right) ^{2}}\sum_{i\in G_{k}^{(\ell )}}\sum_{t_{1}\in \mathcal{T}_{\ell }}\sum_{t_{2}\in \mathcal{T}_{\ell }}\sum_{s_{1}\in \mathcal{T}_{\ell }}\sum_{s_{2}\in \mathcal{T}_{\ell }}\left\vert Cov\left( e_{it_{1}}\tilde{X}_{j,it_{2}},e_{is_{1}}\tilde{X} _{j,is_{2}}\right) \right\vert =O_{p}(1), \end{equation*} where $\tilde{X}_{j,it}=X_{j,it}-\mathbb{E}\left( X_{j,it}|\mathscr{D} \right) $. \end{itemize}

Assumption (ref) imposes some conditions to help derive the asymptotic distribution theory for the panel model with IFEs which allows for dynamics. Assumption (ref)(i) slightly strengthens Assumption (ref)(vi). Assumption (ref)(ii) is the standard non-collinearity condition for regressors, which is analogous to Assumption 4(i) in moon2017dynamic. Assumption (ref)(iii) is the identification assumption which is comparable to Assumption 4 in moon2017dynamic. With the conditional strong mixing condition in Assumption (ref)(iii), we can verify Assumption (ref)(iv).

The following theorem establishes the asymptotic distribution of $\{\hat{ \alpha}_{k}^{(\ell )}\}_{k\in K^{(\ell )}}.$

theoremSuppose that Assumption (ref) or (ref) and Assumptions (ref)--(ref) hold. For $\ell \in \{1,2\}$, the estimators $\{\hat{\alpha}_{k}^{(\ell )}\}_{k\in K^{(\ell )}}$ are asymptotically equivalent to the oracle estimators $\{\hat{\alpha}_{k}^{\ast (\ell )}\}_{k\in K^{(\ell )}}$, and we have \begin{equation*} \mathbb{W}_{NT}^{(\ell )}\mathbb{D}_{NT}^{(\ell )} \begin{pmatrix} \hat{\alpha}_{1}^{(\ell )}-\alpha _{1}^{(\ell )} \\ \vdots \\ \hat{\alpha}_{K^{(\ell )}}^{(\ell )}-\alpha _{K^{(\ell )}}^{(\ell )} \end{pmatrix} -\mathbb{B}_{NT}^{(\ell )}\rightsquigarrow \mathbb{N}\left( 0,\Omega ^{(\ell )}\right) , \end{equation*} such that $\mathbb{D}_{NT}^{(\ell )}=\text{diag}\left( \sqrt{N_{1}^{(\ell )}T_{\ell }},\cdots ,\sqrt{N_{K^{(\ell )}}^{(\ell )}T_{\ell }}\right) $, $ \mathbb{W}_{NT}^{(\ell )}=\text{diag}\left( \mathbb{W}_{NT,1}^{(\ell )},\cdots ,\mathbb{W}_{NT,K^{(\ell )}}^{(\ell )}\right) $, $\mathbb{B} _{NT}^{(\ell )}=\text{diag}\left( \mathbb{B}_{NT,1}^{(\ell )},\cdots , \mathbb{B}_{NT,K^{(\ell )}}^{(\ell )}\right) $ and $\Omega ^{(\ell )}=\text{ diag}\left( \Omega _{1}^{(\ell )},\cdots ,\Omega _{K^{(\ell )}}^{(\ell )}\right) $.

Theorem (ref) establishes the asymptotic distribution for the estimators of the group-specific slope coefficients before and after the break. It shows that the parameter estimators from our algorithm enjoy the oracle property given the results in Theorems (ref) and (ref). The proof of the above theorem can be done by following moon2017dynamic and lu2016shrinkage.

Alternatives and Extensions

This section first considers an alternative method to estimate the break point and then discusses several possible extensions.

Alternative for Break Point Detection

The algorithm proposed in Section (ref) uses low-rank estimates of $\Theta _{j}^{0}$ to find the break point estimates. However, by Lemma (ref)(ii), we observe that the right singular vector matrix of $\Theta _{j}^{0}$, i.e., $V_{j}^{0}$, contains the structural break information when $r_{j}=2$. For this reason, we can propose an alternative way to estimate the break point under the case where the maximum rank of the slope matrix in the model is 2. Let $\dot{v}_{t,j}^{\ast }:= \frac{\dot{v}_{t,j}}{\left\Vert \dot{v}_{t,j}\right\Vert }$ and $\dot{v} _{t}^{\ast }:=\left( \dot{v}_{t,1}^{\ast \prime },\cdots ,\dot{v} _{t,p}^{\ast \prime }\right) ^{\prime },$ with the true values being $ v_{t,j}^{\ast }:=\frac{O_{j}v_{t,j}^{0}}{\left\Vert O_{j}v_{t,j}^{0}\right\Vert }$ and $v_{t}^{\ast }:=\left( v_{t,1}^{\ast \prime },\cdots ,v_{t,p}^{\ast \prime }\right) ^{\prime }$, respectively. Then Step 3 can be replaced by Step 3* below:

description• Break Point Estimation by Singular Vectors. We estimate the break point as follows: \begin{equation} \tilde{T}_{1}=\operatorname*{arg\,min}_{s\in \{2,\cdots ,T-1\}}\frac{1}{T}\left\{ \sum_{t=1}^{s}\left\Vert \dot{v}_{t}^{\ast }-\bar{\dot{v}}^{\ast (1)s}\right\Vert ^{2}+\sum_{t=s+1}^{T}\left\Vert \dot{v}_{t}^{\ast }-\bar{ \dot{v}}^{\ast (2)s}\right\Vert ^{2}\right\} , \end{equation} where $\bar{\dot{v}}^{\ast (1)s}=\frac{1}{s}\sum_{t=1}^{s}{\dot{v}_{t}} ^{\ast }$ and $\bar{\dot{v}}^{\ast (2)s}=\frac{1}{T-s}\sum_{t=s+1}^{T}{\dot{v }}_{t}^{\ast }$.

The following two theorems establish the consistency of $\dot{v}_{t}^{\ast }$ and $\tilde{T}_{1},$ respectively.

theoremSuppose that Assumptions (ref)--(ref) hold. Then $ \max_{t}\left\Vert \dot{v}_{t}^{\ast }-v_{t}^{\ast }\right\Vert =O_{p}(\eta _{N,2}).$
theoremSuppose that Assumptions (ref)--(ref) hold. Then $ \mathbb{P(}\tilde{T}_{1}=T_{1})\rightarrow 1$ as $\left( N,T\right) \rightarrow \infty .$

Since the singular vectors of the slope matrices contain the structural change information, Theorem (ref) indicates that we can consistently estimate the break point by using the factor estimates instead of the slope matrix estimates in ((ref)). Given Theorem (ref) and Lemma (ref)(iii), we can prove Theorem (ref) with arguments analogous to those used in the proof of Theorem (ref).

Test for the Presence of a Structural Break

Section (ref) considers time-varying latent group structures with one break point. In this subsection, we propose a test for the null that the slope coefficients are time-invariant against the alternative that there's one structural break as assumed in Section (ref).

Since various scenarios can occur once we allow for the presence of a structural break in the latent group structures, the number of groups may or may not change under the alternative and so may some of the group-specific coefficients. First, information on the latent group structures may be ignored and the slope coefficients can be tested for possible time-variation. In this case, we can rewrite $\Theta _{it}^{0}$ as

equation*[equation* omitted — 57 chars of source]

where $\Theta _{i}^{0}:=\frac{1}{T}\sum_{t\in \lbrack T]}\Theta _{it}^{0}$. The null and alternative hypotheses can then be specified as

eqnarray[eqnarray omitted — 165 chars of source]

To construct the test statistics we follow bai1998estimating and consider a sup-$F$ test. Let $\mathcal{T}_{\epsilon }:=\left\{ T_{1}:\epsilon T\leq T_{1}\leq (1-\epsilon )T\right\} ,$ where $ \epsilon >0$ is a tuning parameter that avoids breaks at the end of the sample. Define

equation*[equation* omitted — 169 chars of source]

where

equation*[equation* omitted — 329 chars of source]

$\tilde{\beta}_{i}^{\left( 1\right) }\left( T_{1}\right) $ and $\tilde{\beta} _{i}^{\left( 2\right) }(T_{1})$ are the PCA slope estimators of $\Theta _{i}^{0}$ in the linear panels with IFEs for each individual $i$ with the prior-break observations $\left\{ (i,t):i\in \lbrack N],t\in \lbrack T_{+}]\right\} $ and post-break observations $\left\{ (i,t):i\in \lbrack N],t\in \lbrack T]\backslash \lbrack T_{+}]\right\} $, respectively, \footnote{ See Section (ref) in the appendix for the detail of the PCA estimation in linear panels with IFEs.} and $\hat{\Sigma}_{i}(T_{+})$ is the consistent estimator for the asymptotic variance of $\tilde{\beta} _{i}^{\left( 1\right) }(T_{1})-\tilde{\beta}_{i}^{\left( 2\right) }\left( T_{1}\right) $. Following bai1998estimating, we conjecture that the asymptotic distribution of $\operatornamewithlimits{\sup}\limits_{T_{1}\in \mathcal{T}_{\epsilon }}F_{i}(T_{1})$ depends on a $p$-vector of Wiener processes on $\left[ 0,1\right] ,$ based on which the corresponding distribution of $F_{NT}(1|0)$ can be obtained.

Alternatively, we can estimate the model with latent group structures by assuming the presence of a break point at $T_{1}.$ Then we obtain the estimates of the group-specific parameters $\{\alpha _{j}^{\left( 1\right) }\left( T_{1}\right) \}_{j\in K^{\left( 1\right) }}$ prior to the potential break point $T_{1}$ and those of the group-specific parameters $\{\alpha _{j}^{\left( 2\right) }\left( T_{1}\right) \}_{j\in K^{\left( 2\right) }}$ after the potential break point $T_{1}.$ It is possible to construct a test statistic based on the contrast of these two sets of estimates or the corresponding residual sum of squares (RSS) and then take the supremum over $ T_{1}\in \mathcal{T}_{\epsilon }.$ Evidently, this approach is also quite involved as it is necessary to determine the number of groups before and after the break, $K^{\left( 1\right) }$ and $K^{\left( 1\right) },$ at each $ T_{1}. $ It is not clear how the estimation errors from these estimates and those of the factors and factor loadings with slow convergence rates affect the asymptotic properties of the estimators of the group-specific parameters.

Last, it is also possible to estimate the model with latent group structures under the case of no structural changes to obtain the restricted residuals. If there exists a structural change in the latent group structure, it should be reflected in properties of the restricted residuals obtained under the null. We can then consider the regression of the restricted residuals on the regressors and construct an LM-type test statistic to check the goodness of fit of such an auxiliary regression model as in su2013testing and su2020testing. This analysis is left for future research.

The Case of Multiple Breaks

Section (ref) considers only a one-time structural break in the latent group structures. In practice it is possible to have multiple breaks especially if $T$ is large. Here we generalize the model in Section (ref) to allow for multiple breaks. In this case, we have

equation*[equation* omitted — 271 chars of source]

where $b\geq 1$ denotes the number of breaks.

To estimate the number of breaks and the break points $T_{1},\cdots ,T_{b}$, in principle we can follow the sequential method proposed by bai1998estimating. First, using the full-sample data, we can construct $ F_{NT}(1|0)$ defined in the previous subsection and estimate the break point as in (ref). Second, for each regime before and after the estimated break point, we test the hypothesis in (ref) and estimate the break point for each regime separately. This sequential method is repeated until the null can not be rejected for all subsamples, leading to the break point estimates $\{\hat{T}_{a}\}_{a\in\lbrack \hat{b}]}$ where $\hat{b}$ is the estimated number of breaks. We conjecture that we can establish the consistency of $\hat{b}$ and $\{\hat{T} _{a}\}.$

After the estimated number of breaks and break points are obtained for each subsample

equation*[equation* omitted — 93 chars of source]

where $a\in [\hat{b}+1]$ with $\hat{T}_{0}:=0$ and $\hat{T}_{\hat{b}+1}:=T$, we can continue Step 4 in the estimation algorithm in Section (ref) to obtain the estimated group structure for each subsample.

Monte Carlo Simulations

In this section we report simulation results for the low-rank estimates, break point estimates, group membership estimates and the group number estimates based on 1,000 replications, and tuning parameter $\nu _{j}$ chosen by a procedure similar to that described in chernozhukov2019inference. We focus on the linear panel model $ Y_{it}=\lambda _{i}^{\prime }f_{t}+X_{it}^{\prime }\Theta _{it}+e_{it}$, where $X_{it}=(X_{1,it},X_{2,it})^{\prime }$ and $\Theta _{it}=(\Theta _{1,it},\Theta _{2,it})^{\prime }$.

Data Generating Processes (DGPs)

The following four main DGPs are employed.

description• [Static panel with homoskedasticity] $ X_{1,it}\sim $ $i.i.d.$ $U(-2,2)$, $X_{2,it}\sim i.i.d.$ $U(-2,2)$, and errors $e_{it}\sim i.i.d.$ $\mathbb{N}(0,1)$. For $\Theta _{1}$, we randomly choose the break point $T_{1}$ from $0.4T$ to $0.6T$. • [Static panel with heteroscedasticity] Compared to DGP 1, the errors $e_{it}\sim i.i.d.$ $\mathbb{N} (0,\sigma _{it}^{2})$ with $\sigma _{it}^{2}\sim i.i.d.$ $U(0.5,1)$. The settings for the regressors and break point are the same as those in DGP 1. • [Serially correlated error] Compared to DGP 2, the errors $e_{it}=0.2e_{i,t-1}+\eta _{it},$ where $\eta _{it}\sim i.i.d.$ $\mathbb{N}(0,1)$ and all other settings are the same as in DGP 2. • [\textbf{Dynamic panel}] In this case, $ X_{1,it}=Y_{i,t-1}$ with $Y_{i,0}\sim i.i.d.~\mathbb{N}(0,1)$. $X_{2,it}\sim i.i.d.~U(-2,2)$, and $e_{it}\sim i.i.d.~\mathbb{N}(0,0.5)$.

For each DGP above, we set $r_{0}=0$ and draw $\lambda _{i}$ and $f_{t}$ from $\mathbb{N}(0,1)$ independently. We define the slope coefficient based on three subcases below.

description• In this case, the group membership and the number of groups do not change after the break point and only the value of the slope coefficient changes. We set the number of groups to be 2, the ratio of individuals among the two groups is $N_{1}:N_{2}=0.5:0.5$, and the group membership $G_{1}$ is obtained by a random draw from $[N]$ without replacement. For DGPs 1.1, 2.1, and 3.1, \begin{equation*} \Theta _{1,it}=\Theta _{2,it}=\left\{ \begin{aligned} & 0.1, \quad i\in G_{1}, t\in\{1,\cdots,T_{1} \},\\ & 0.9, \quad i\in G_{2}, t\in\{1,\cdots,T_{1} \},\\ & 0.05, \quad i\in G_{1}, t\in\{T_{1}+1,\cdots,T \},\\ & 0.45, \quad i\in G_{2}, t\in\{T_{1}+1,\cdots,T \}. \end{aligned}\right. \end{equation*} For DGP 4.1, $\Theta _{2,it}$ is same as other DGPs X.1 for X$\in \{1,2,3\}$ , and \begin{equation*} \Theta _{1,it}=\left\{ \begin{aligned} & 0.1, \quad i\in G_{1}, t\in\{1,\cdots,T_{1} \},\\ & 0.7, \quad i\in G_{2}, t\in\{1,\cdots,T_{1} \},\\ & 0.05, \quad i\in G_{1}, t\in\{T_{1}+1,\cdots,T \},\\ & 0.35, \quad i\in G_{2}, t\in\{T_{1}+1,\cdots,T \}. \end{aligned}\right. \end{equation*} • Compared to DGP X.1, the values of the slope coefficients for different groups do not change after the break point, but the group membership changes. The number of groups is 2, the ratio of individuals among the groups is still $N_{1}:N_{2}=0.5:0.5$. Nevertheless, $\{G_{1}^{(1)},G_{2}^{(1)}\}$ is different from $ \{G_{1}^{(2)},G_{2}^{(2)}\}$ so that elements in both $G_{1}^{(1)}$ and $ G_{1}^{(2)}$ are independent draws from $[N]$ without replacement. In addition, for DGPs 1.2, 2.2, and 3.2, \begin{equation*} \Theta _{1,it}=\Theta _{2,it}=\left\{ \begin{aligned} & 0.1, \quad i\in G_{1}^{(1)}, t\in\{1,\cdots,T_{1} \},\\ & 0.9, \quad i\in G_{2}^{(1)}, t\in\{1,\cdots,T_{1} \},\\ & 0.1, \quad i\in G_{1}^{(2)}, t\in\{T_{1}+1,\cdots,T \},\\ & 0.9, \quad i\in G_{2}^{(2)}, t\in\{T_{1}+1,\cdots,T \}. \end{aligned}\right. \end{equation*} For DGP 4.2, $\Theta _{2,it}$ is defined same as other DGPs X.2 for $X\in \{1,2,3\}$, and \begin{equation*} \Theta _{1,it}=\left\{ \begin{aligned} & 0.1, \quad i\in G_{1}^{(1)}, t\in\{1,\cdots,T_{1} \},\\ & 0.7, \quad i\in G_{2}^{(1)}, t\in\{1,\cdots,T_{1} \},\\ & 0.1, \quad i\in G_{1}^{(2)}, t\in\{T_{1}+1,\cdots,T \},\\ & 0.7, \quad i\in G_{2}^{(2)}, t\in\{T_{1}+1,\cdots,T \}. \end{aligned}\right. \end{equation*} • Under this scenario, the number of groups changes after the break. We set $N_{1}^{(1)}:N_{2}^{(1)}=0.5:0.5$ and $ N_{1}^{(2)}:N_{2}^{(2)}:N_{3}^{(2)}=0.4:0.3:0.3$ before and after the break, respectively. Specifically, for DGPs 1.3, 2.3, and 3.3, we have \begin{equation*} \Theta _{1,it}=\Theta _{2,it}=\left\{ \begin{aligned} & 0.1, \quad i\in G_{1}^{(1)}, t\in\{1,\cdots,T_{1} \},\\ & 0.9, \quad i\in G_{2}^{(1)}, t\in\{1,\cdots,T_{1} \},\\ & 0.1, \quad i\in G_{1}^{(2)}, t\in\{T_{1}+1,\cdots,T \},\\ & 0.5, \quad i\in G_{2}^{(2)}, t\in\{T_{1}+1,\cdots,T \},\\ & 0.9, \quad i\in G_{3}^{(2)}, t\in\{T_{1}+1,\cdots,T \}. \end{aligned}\right. \end{equation*} For DGP 4.3, $\Theta _{2,it}$ is defined as in DGP X.3 for X$\in \{1,2,3\}$, and \begin{equation*} \Theta _{1,it}=\left\{ \begin{aligned} & 0.1, \quad i\in G_{1}^{(1)}, t\in\{1,\cdots,T_{1} \},\\ & 0.7, \quad i\in G_{2}^{(1)}, t\in\{1,\cdots,T_{1} \},\\ & 0.1, \quad i\in G_{1}^{(2)}, t\in\{T_{1}+1,\cdots,T \},\\ & 0.4, \quad i\in G_{2}^{(2)}, t\in\{T_{1}+1,\cdots,T \},\\ & 0.7, \quad i\in G_{3}^{(2)}, t\in\{T_{1}+1,\cdots,T \}. \end{aligned}\right. \end{equation*}

Results

Table (ref) reports the proportion of correct rank estimation for the intercept (IFE) and slope matrices based on the SVT in Section 3.3. Note that $r_{0}$ denotes the true rank of the intercept matrix and $r_{1}$ and $ r_{2}$ denote those of the two slope matrices. From Table (ref), we notice that the true ranks of both the intercept and slope matrices can be almost perfectly estimated for the sample sizes under investigation.

table[table omitted — 2,389 chars of source]

Table (ref) reports the results for the break point estimation in Step 3 based on different $(N,T)$ combinations. We summarize some important findings from Table (ref). First, when the group membership and the number of groups do not change as in DGP X.1 for X$\in \left[ 3\right] ,$ the frequency of correct break point estimation may not be 1 especially if $N$ is not large. This suggests that the binary segmentation does not work perfectly in such a scenario. Second, the change of group membership or the number of groups help to identify the break point as reflected in the simulation results for DGP X.2 and X.3 for X$\in \left[ 4 \right] .$ In general, the binary segmentation works well in our setting.

table[table omitted — 880 chars of source]

Table (ref) reports the results for the group membership estimation when the number of groups are either known (infeasible in practice) or estimated from the data (feasible). With known number of groups, the STK algorithm degenerates to the traditional K-means algorithm. The \textquotedblleft Infeasible\textquotedblright\ part of Table (ref) reports the frequency of correct group membership estimation before and after the estimated break point, $G_{B}$ and $G_{A}$, based on the known true number of groups and K-means algorithm. Evidently, the K-means classification exhibits excellent performance in this case. Nevertheless, without prior information on the true number of groups, the STK algorithm is able to estimate the group membership and the number of groups simultaneously. In this case, the frequency of correct estimation of the group membership and that of the number of groups are shown in the \textquotedblleft Feasible\textquotedblright\ part in Table (ref) and in Table (ref), respectively. To implement the STK algorithm with unknown number of groups, we set $\varsigma _{N}=N^{-2}$ to ensure the consistency of the group number estimators. As expected, the performance of the STK algorithm is slightly worse than that of the K-means algorithm with knowledge of the true number of groups. But the performance improves when both $N$ and $T$ increase. Table (ref) suggests that the number of groups can be nearly perfectly estimated in DGPs 1.1, 1.2, 1.3 and 2.1. For the more complicated DGPs (e.g., the dynamic case in DGPs 4.1, 4.2, and 4.3 or the static panel with serially correlated errors in DGPs 3.1, 3.2, and 3.3), the performance is not as good as that in the simple DGPs.

table[table omitted — 3,540 chars of source]
table[table omitted — 1,972 chars of source]

Table (ref) presents more detailed results for the estimation of the number of groups. For DGPs 1.X and DGP 2.X where we have static panels with independent errors, the results show that the group membership and the number of groups can be well estimated with nearly 100% accuracy under different $(N,T)$ combinations. For DGPs 3.X and 4.X where we have static panels with serially correlated errors and dynamic panels, respectively, the frequency of correct estimation of both the group membership and the number of groups is not great when $T$ is small, but gradually approaches unity as $T$ increases. One reason for this is that we need to use HAC estimates of certain long-run variance objects in the STK algorithm and it is well known that a relatively large value of $T$ is required in order for the HAC estimates to be reasonably well behaved in finite samples.

table[table omitted — 5,633 chars of source]

Table (ref) shows results for the post-classification estimator for the first slope coefficient. We follow su2016identifying to define the evaluation criteria as bias and coverage. Specifically, we define the bias to be the weighted versions of bias for slope estimator from all estimated groups, i.e. $\text{Bias}=\sum_{k=1}^{K^{(1)}}\text{Bias}(\alpha _{k,1}^{(\ell )})$ for $\ell \in \{1,2\}$. Similarly, we define the weighted version of the coverage ratio of the 95% confidence interval estimators. The \textquotedblleft Infeasible" panel shows the result assuming the number of groups information is known, and the \textquotedblleft Feasible" panel shows the result without knowing the number of groups information by the STK algorithm. From Table (ref), we notice that the coverage ratio for DGP 1 and 2 is close to 95% under different combinations of $N$ and $T$ for both the \textquotedblleft Infeasible" and \textquotedblleft Feasible" panels, which is due to the higher correct classification ratio. For DGPs 3 and 4, by using the STK algorithm, the coverage ratio is a bit lower for $ T=100$, which is due to the inaccuracy of the group number and membership estimators; but coverage approaches 95% quickly when $T$ doubles.

table[table omitted — 6,365 chars of source]

Empirical Study

The estimation methods were applied to analyze the time-varying latent group structure of real house price changes in Metropolitan Statistical Areas (MSAs) in the United States. Studies of U.S. house price changes are plentiful in the literature. malpezzi1999simple, capozza2002determinants, gallin2006long, and ortalo2006housing all show that the house price changes are closely correlated with real income in the long run. su2023identifying consider a heterogenous spatial panel and show that real income growth affects the U.S. house prices in different ways for different MSAs. In this application, we examine whether there exist latent group structures for the real income growth elasticity of house price changes and whether these structures change over the time dimension.

Model

We consider the following panel data model with IFEs and two-way slope heterogeneity

equation[equation omitted — 127 chars of source]

where the dependent variable $\pi _{it}$ measures the percentage of real house price growth for MSA $i$ at time period $t$. The $\lambda _{i}$ and $f_{t}$ are the individual fixed effects and time fixed effects, the covariate $ginc_{it}$ denotes the percentage of income growth for MSA $i$ at time period $t$, and $ginc_{i,t-1}$ is the lagged value of $ginc_{it}$. Unlike aquaro2021estimation and su2023identifying who consider individual fixed effects and additive two-way fixed effects, respectively, we allow the model to have IFEs. In the above model, we allow the slope parameters $\left( \Theta _{1,it},\Theta _{2,it}\right) $ to exhibit latent group structures along the cross-sectional dimension and an unknown break along the time dimension.

Data

The dataset we use is obtained from aquaro2021estimation, which is the quarterly data for 377 MSAs over 1975 to 2014. To construct the growth rate and the lagged term, we lose two periods of observations, which yields $ T=158 $. Similar to su2023identifying, we deseasonalize the growth rate of real house price and real income. We do not de-factor the variables since our model contains IFEs to control the common shocks.

Empirical Results

We first apply the singular value thresholding to estimate the ranks of $ \Theta _{0}=\{\lambda _{i}^{\prime }f_{t}\},$ $\Theta _{1}=\{\Theta _{1,it}\} $ and $\Theta _{2}=\{\Theta _{2,it}\}.$ The estimates are: $\hat{r}_{0}=1,$ $ \hat{r}_{1}=2,$ and $\hat{r}_{2}=2$. Before applying the proposed estimation algorithm in Section (ref), we first test the presence of a structural break as in Section (ref). As given in bai2003critical, for each individual $i$, the critical value of test statistic $\sup_{T_{1}\in \mathcal{T}_{\epsilon }}F_{i}(T_{1})$ is 15.37. We then construct the sup-F test statistic for each MSA. Results show that $ \operatornamewithlimits{\text{min}}\limits_{i\in \lbrack N]}\sup_{T_{1}\in \mathcal{T}_{\epsilon }}F_{i}(T_{1})=0.0195$ and the final test statistic is $F_{NT}(1|0)=\operatornamewithlimits{\text{max}}\limits_{i\in \lbrack N]}\sup_{T_{1}\in \mathcal{T}_{\epsilon }}F_{i}(T_{1})=2161.65$. Based on this outcome, we reject the null that there is no structural break for slope coefficients $\Theta _{1,it}$ and $\Theta _{2,it}$ at the 1% significance level.

With the presence of a structural break, we apply the proposed multi-stage estimation result in Section (ref) to estimate the break date and numbers of groups before and after the break. The estimated break date is given by $\hat{T}_{1}=51,$ which suggests that the structural break happens at the first quarter in 1988. We conjecture that this break may be related to the catastrophic stock market crash that occurred on October 1987, which is considered to be the first contemporary global financial crisis event.

figure[figure omitted — 276 chars of source]

By setting $\varsigma _{N}=N^{-2}$ for the STK algorithm as in the simulations, we obtain the estimated prior- and post-break numbers of groups given by $\hat{K}^{(1)}=6$ and $\hat{K}^{(2)}=2,$ respectively. As for the group structure, Figures (ref) and (ref) use six and two colors to show the classification results for the 377 MSAs during 1975Q3 to 1987Q4 and 1988Q1 to 2014Q4, respectively. Table (ref) reports the pooled regression results for the full sample in column (1), the subsample before the estimated break point in column (2), and the subsample after the estimated break point in column (3). All the slope estimators are bias-corrected. The pooled regression results in Table (ref) show that real income growth has positive and significant effect on house prices. Comparing the two subsamples before and after the estimated break, we observe that, with a 1 percentage increase in real income growth, the real house price growth rate will increase 0.09 percentage before 1988, which is 0.02 percentage points higher than that after 1988. The slope estimates for the lagged term are similar for the two subsamples.

table[table omitted — 889 chars of source]

To examine the difference for each of the 6 estimated groups before the break, Table (ref) reports the post-classification regression results for each estimated group before the estimated break. Even though the effects of real income growth are positive for all estimated groups, they differ vastly across groups. The effect of real income growth for Group 2 is the highest, followed by Groups 5 and 3, and the effects of real income growth on real house prices in all of these three groups are higher than 0.15 percent. In contrast, the effects of real income growth for the remaining three groups, viz., Groups 1, 4, and 6, are less than 0.07 percent. Similarly, Table (ref) reports the post-classification regression results for each estimated group after the estimated break. The estimated slope coefficients for both groups are statistically significant. Especially, during 1988Q1-2014Q1, the slope estimator for the lagged term in the first group is much higher than that for the second group.

We also applied the C-Lasso algorithm of su2016identifying to estimate the group structure before and after the estimated break point. The C-Lasso approach in conjuction with IC detects 2 groups both before and after the break. In view of the six groups detected by our present algorithm, we conjecture that the difference may due to the smaller time periods before the break.

table[table omitted — 1,034 chars of source]
table[table omitted — 677 chars of source]

Conclusion

This paper considers a linear panel data model with interactive fixed effects and two-way heterogeneity such that the heterogeneity across individuals is captured by latent group structures and the heterogeneity across time is captured by an unknown structural break. We allow the model to have different group numbers, or different group memberships, or just changes in the slope coefficients for some specific groups before and after the break. To estimate the unknown structural break, the number of groups and group memberships before and after the break point, we propose an estimation algorithm with initial nuclear-norm-regularized estimates, followed by row- and column-wise linear regressions. Then, the break point estimator is obtained by binary segmentation and the group structure together with the number of groups are estimated simultaneously using a sequential testing K-means algorithm. We show that the structural break estimator, the group number estimators, and the group membership estimators before and after the break point are all consistent, and the final post-classification slope coefficient estimators enjoy the oracle property.

There are several interesting topics for further research. First, even though we discuss a possible test for the existence of a break in the panel data models with latent group structures, we have not fully worked out the asymptotic theory, a challenge that deserves separate treatment. Second, we assume the presence of a single break in the data and it is interesting to extend our theory to allow for multiple breaks. Third, the present treatment rules out both unit-root-type nonstationarity and nonstochastic trending nonstationarity and it is interesting to extend our theory to allow for nonstationarity. We will pursue these topics in future research.

{ {\ }}

{