EconBase
← All papers

K-Means Panel Data Clustering in the Presence of Small Groups

Mikihito Nishi

arXiv 21 Aug 2025 · Econometrics

arXiv:2508.15408 · PDF · Extracted main text

Abstract

We consider panel data models with group structure. We study the asymptotic behavior of least-squares estimators and information criterion for the number of groups, allowing for the presence of small groups that have an asymptotically negligible relative size. Our contributions are threefold. First, we derive sufficient conditions under which the least-squares estimators are consistent and asymptotically normal. One of the conditions implies that a longer sample period is required as there are smaller groups. Second, we show that information criteria for the number of groups proposed in earlier works can be inconsistent or perform poorly in the presence of small groups. Third, we propose modified information criteria (MIC) designed to perform well in the presence of small groups. A Monte Carlo simulation confirms their good performance in finite samples. An empirical application illustrates that K-means clustering paired with the proposed MIC allows one to discover small groups without producing too many groups. This enables characterizing small groups and differentiating them from the other large groups in a parsimonious group structure.

Citation extraction

24
references
98
in-text mentions
27
distinct cited
0
self-citations
11,149
main-text words

appendix boundary found by appendix_command · 48% of the source is main text. Read the extracted text to check this.

Most heavily cited references

The works this paper leans on most, across its whole bibliography — not restricted to papers in our corpus. Ranked by composite intensity, which combines how often a work is mentioned, how many sections mention it, and how much of that falls in the main text rather than the appendix.

ReferenceIntensityMentionsSectionsMain text
1Lumsdaine, Okui and Wang (2023) Estimation of Panel Group Structure Models with Structural Breaks in Group Memberships and Coefficients1.00094100%
2Bai and Ng (2002) Determining the Number of Factors in Approximate Factor Models1.00083100%
3Bonhomme and Manresa (2015) Grouped Patterns of Heterogeneity in Panel Data0.97325892%
4Sun, Tan, Zhang and Zhu (2025) Homogeneity Pursuit in Clustered Data Analysis When Cluster Sizes Are Small*0.92843100%
5Mehrabani (2023) Estimation and Identification of Latent Group Structures in Panel Data0.87462100%
6Su, Shi and Phillips (2016) Identifying Latent Structures in Panel Data0.87452100%
7Mugnier (2023) A Simple and Computationally Trivial Estimator for Grouped Fixed Effects Models0.81142100%
8Wang, Phillips and Su (2018) Homogeneity Pursuit in Panel Data Models: Theory and Application0.81142100%
9Wang and Su (2021) Identifying Latent Group Structures in Nonlinear Panels0.81142100%
10Chetverikov and Manresa (2022) Spectral and Post-Spectral Estimators for Grouped Panel Data Models0.73732100%

Showing the top 10 of 27 scored citations.