Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
75,100 characters · 19 sections · 75 citation commands
K-Means Panel Data Clustering in the Presence of Small Groups
\onehalfspacing
Modeling cross-sectional heterogeneity is an important task when dealing with panel data, in which researchers face a trade-off between flexibility of models with complete heterogeneity and efficiency obtained from models with homogeneity across units. Recently, assuming a group structure among individuals has been viewed as a parsimonious yet sufficiently flexible way to model heterogeneity. In the group structure model, units within a group share the same parameter, while units belonging to different groups have different characteristics. Many researchers have developed estimation theory for grouped heterogeneity and unknown group memberships. linEstimationPanelData2012 and bonhommeGroupedPatternsHeterogeneity2015 base their estimation strategies on the least-squares (LS) principle and the K-means algorithm. In particular, bonhommeGroupedPatternsHeterogeneity2015 establish consistency and asymptotic normality of the LS estimator in models with group-specific time-varying fixed-effects, or grouped fixed-effects (GFE). In a similar spirit, andoClusteringHugeNumber2017 consider models with group-specific interactive fixed-effects and develop asymptotic theory based on a penalized objective function. suIdentifyingLatentStructures2016, mehrabaniEstimationIdentificationLatent2023, and wangHomogeneitySparsityAnalysis2024 propose using penalized objective functions to estimate group-specific slope coefficients. okuiHeterogeneousStructuralBreaks2021 and lumsdaineEstimationPanelGroup2023 consider models in which group-specific slope parameters and group structure possibly experience structural breaks. The vast literature on the topic also includes hahnPanelDataModels2010, luDeterminingNumberGroups2017, liuIdentificationEstimationPanel2020, wangIdentifyingLatentGroup2021, chetverikovSpectralPostspectralEstimators2022, mugnierSimpleComputationallyTrivial2023, and yuSpectralClusteringVariance2024.
A commonly employed assumption in the above literature is that all latent groups have a size proportional to the number of cross-sections $N$, ruling out cases where one or more groups have a disproportionately small size bonhommeGroupedPatternsHeterogeneity2015,suIdentifyingLatentStructures2016,andoClusteringHugeNumber2017,luDeterminingNumberGroups2017,wangIdentifyingLatentGroup2021,okuiHeterogeneousStructuralBreaks2021,chetverikovSpectralPostspectralEstimators2022,lumsdaineEstimationPanelGroup2023. This assumption, however, may be violated in real data analysis, as evidenced in some studies. For example, andoClusteringHugeNumber2017 apply a clustering algorithm to stock returns data and obtain a group of 2616 cross-sections, and, at the same time, also obtain a group consisting of only 19 units. Another example is aparicio-perezDisentanglingHeterogeneousEffect2025, who investigate the nexus between a country's natural resources endowment and economic growth. In their analysis, they obtain 6 groups, one of which consists of 47 countries, while there is a group of only 2 countries. grunewaldTradeoffIncomeInequality2017 and martinez-zarzosoSearchingGroupedPatterns2020 also document disproportionality in group size. Although some recent studies establish inferential theory allowing for the presence of small groups mehrabaniEstimationIdentificationLatent2023,mugnierSimpleComputationallyTrivial2023,wangHomogeneitySparsityAnalysis2024, the properties of clustering methods are still underexplored in such a situation.
One might think that the assumption of all groups having a size proportional to $N$ is imposed just to simplify the theoretical derivation and is innocuous. But we demonstrate that this is not true, at least as far as the LS estimator and the K-means algorithm are concerned. Specifically, we consider models with group-specific slope coefficients and GFE as considered in bonhommeGroupedPatternsHeterogeneity2015, and study the asymptotic behavior of the LS estimators for slope coefficients, GFE, and group memberships and information criteria for the unknown number of groups.
In terms of the model specification and estimators employed, bonhommeGroupedPatternsHeterogeneity2015 is obviously a precursor to our study, but we extend their analysis in three nontrivial ways. First, we derive sufficient conditions under which the LS estimators are consistent and asymptotically normal, in an environment where at least one group has a size of smaller order than $N$. One of the sufficient conditions implies that a longer sample period $T$ is required as there are smaller groups. This points out a potential limitation of the LS clustering that may arise in short panels with small groups. Second, we study the asymptotic behavior of information criteria for the number of groups. Because bonhommeGroupedPatternsHeterogeneity2015 propose an information criterion without providing a theoretical analysis, the results drawn in the present work are entirely new. We derive a sufficient condition under which the information criterion is inconsistent and underestimates the number of groups. We show that an information criterion proposed in baiDeterminingNumberFactors2002 can satisfy the condition and hence underestimate the number of groups, particularly when there are small groups or the model includes GFE. Although we find that the BIC proposed in bonhommeGroupedPatternsHeterogeneity2015 does not satisfy the condition for underestimation, our simulation shows that it overestimates the number of groups in finite samples when the model does not include GFE. Finally, we propose modified information criteria (MIC) designed not to satisfy the conditions for inconsistency and to improve the information criteria proposed in baiDeterminingNumberFactors2002 and bonhommeGroupedPatternsHeterogeneity2015. A simulation study shows their good performance in the presence of small groups.
The scope of the present work is wider than that of earlier studies which consider the panel model with small groups.\footnote{Note that the “small group" problem we discuss has nothing to do with the “small cluster" problem considered in sunHomogeneityPursuitClustered2025. They consider partitioning $m$ clusters into groups, where $m$ corresponds to $N$ in our notation, and assume that each group has a size proportional to $m$. What they mean by small clusters is that each cluster has a small number of observations, which is analogous to a small $T$ (sample period) in our case, and thus does not refer to small groups in our sense. See Remark (ref) below for details.} wangHomogeneityPursuitPanel2018, mehrabaniEstimationIdentificationLatent2023 and wangHomogeneitySparsityAnalysis2024 propose estimation frameworks based on penalized objective functions and show that their estimators are consistent. However, they study models where heterogeneity manifests itself only through slope coefficients and do not consider models with GFE. Furthermore, we show that a condition assumed in wangHomogeneityPursuitPanel2018 and mehrabaniEstimationIdentificationLatent2023 which ensures the consistency of their estimators for the number of groups does not hold in the presence of small groups. mugnierSimpleComputationallyTrivial2023, who proposes a clustering method for the GFE model, shows that his estimators are consistent as long as each group contains at least two units. However, he mainly assumes homogeneous slope coefficients across cross-sections and only briefly discusses how his results may be extended to the cases where slope parameters are group-specific. Unlike these works, we study the LS estimator and the estimator for the number of groups, allowing for small groups, group-specific slope parameters, and GFE.
In our empirical application, we apply the K-means algorithm coupled with the proposed MIC to examine the determinants of sales growth of Japanese firms. Using the MIC, we discover a small group in our dataset, which is overlooked when we use baiDeterminingNumberFactors2002's (baiDeterminingNumberFactors2002) information criterion. This empirical exercise also illustrates that the MIC enables detecting small groups without producing too many groups, unlike the BIC considered in bonhommeGroupedPatternsHeterogeneity2015. This allows us to characterize the small group and differentiate it from the other large groups while keeping a reasonably parsimonious structure.
The remainder of this paper is organized as follows. In Section (ref), we specify the model and spell out conditions under which we work. Section (ref) gives asymptotic results for the LS estimator and the information criterion for the number of groups in models without GFE. In Section (ref), we extend the analysis given in Section (ref) to models with both group-specific slope parameters and GFE. Section (ref) conducts a simulation study. In Section (ref), an empirical application is presented. Section (ref) concludes the paper. All mathematical proofs are relegated to Appendix.
Notation: For any $m\times n$ matrix $A=(a_{ij})$, $\|A\|\coloneqq(\sum_{i=1}^m\sum_{j=1}^na_{ij}^2)^{1/2}$ denotes the Frobenius norm of $A$. For any square matrix $A$, $\lambda_{\min}(A)$ and $\lambda_{\max}(A)$ denote the minimum and maximum eigenvalues of $A$, respectively. For any positive integer $n$, $[n]\coloneqq\{1,2,\ldots,n\}$ is the set of positive integers up to $n$. $\stackrel{p}{\to}$ and $\stackrel{d}{\to}$ signify convergence in probability and convergence in distribution as $N,T\to\infty$, respectively, where $N$ is the number of cross-sections and $T$ is the length of the sample period.
Consider the following panel data model:
where $x_{it}$ is a $p\times 1$ vector of covariates, $k_i\in[K]$ denotes the nonrandom membership of unit $i$, $\theta_k\in\Theta\subset \mathbb{R}^p$ for $k\in[K]$, and $K$ is the number of groups.\footnote{We can also include individual fixed-effects in model (ref). Including individual fixed-effects does not change the conclusion given later as long as assumptions given below are modified suitably and regressor $x_{it}$ does not include a lagged dependent variable. We work with model (ref) for notational simplicity.} Let $\boldsymbol{\gamma}_N \coloneqq (k_1,\ldots,k_N)\in [K]^N$ be the set of memberships of $N$ cross-sections and $\boldsymbol{\theta}$ denote the set of all $\theta_k$'s. We use $K^0$, $\theta_k^0$ and $k_i^0$ to denote the true values of $K$, $\theta_k$, and $k_i$, respectively. Similarly, $\boldsymbol{\theta}^0$ and $\boldsymbol{\gamma}_N^0$ are the sets of $\theta_k^0$'s and $k_i^0$'s, respectively. Also let $G_k^0\coloneqq\{i\in[N]:k_i^0=k\}$ be the set of cross-sections belonging to the true $k$-th group and $N_k^0\coloneqq|G_k^0|=\sum_{i=1}^N1\{k_i^0=k\}$ be the size (cardinality) of $G_k^0$.
We use the least-squares approach, or K-means, to estimate parameters $\boldsymbol{\theta}^0$ and $\boldsymbol{\gamma}_N^0$. Given a prespecified value $K$ for the number of groups, the least-squares estimator is the solution to the following minimization problem:
where $\widehat{\boldsymbol{\gamma}}_N^{(K)}=(\widehat{k}_1^{(K)},\ldots,\widehat{k}_N^{(K)})$ is the set of estimated memberships. Let $\widehat{\sigma}^2(K, \widehat{\boldsymbol{\gamma}}_N^{(K)})\coloneqq (NT)^{-1}\sum_{i=1}^N\sum_{t=1}^T\bigl(y_{it}-x_{it}'\widehat{\theta}_{\widehat{k}_i^{(K)}}^{(K)}\bigr)^2$ be the averaged least-squares SSR obtained under $K$ groups.
Suppose that the true group sizes, $N_k^0$, can be expressed as $N_k^0=\tau_kN^{\alpha_k}$ with $\tau_k>0$ and $\alpha_k\in[0,1]$ such that $N_k^0\geq1$ for all $k\in[K^0]$.\footnote{$N=\sum_{k=1}^{K^0}N_k^0$ is implicitly assumed. This condition may require $\tau_k$ to depend on $N$ (i.e., $\tau_{N,k}$), but we omit this dependence on $N$ for notational simplicity. The results and conclusion below do not change as long as there exist some constants $c$ and $C$ such that $c,C\in(0,\infty)$ and $c<\tau_{N,k}<C$ for all $N$ and $k$.} Without loss of generality, we assume $\alpha_1\geq\alpha_2\geq\cdots\geq\alpha_{K^0}$.
$m$ is the number of groups with asymptotically nonnegligible relative size (proportional to $N$). We require that at least one group have a size proportional to $N$ ($m\geq1$). The conventional situation where all groups have a nonnegligible size ($m=K^0$) is excluded.
Assumption (ref)(a) requires that the parameter space be compact, a standard assumption in the literature. Assumptions (ref)(b) and (c) are also standard moment conditions. Assumption (ref)(d) essentially requires that $x_{it}\varepsilon_{it}$ be weakly correlated cross-sectionally and temporarily. Assumption (ref)(e) is similar to the usual rank condition in OLS estimation. In particular, this condition implies that $(N_k^0T)^{-1}\sum_{i\in G_k^0}\sum_{t=1}^Tx_{it}x_{it}'$ is uniformly positive definite for sufficiently large $N$ and $T$. Assumption (ref)(f) is similar.\footnote{In fact, (f) implies (e) under Assumption (ref)(a), but we keep (e) as an assumption for convenience.} Assumption (f) is essentially the same as Assumption 1(iv) of lumsdaineEstimationPanelGroup2023.
The first part of Assumption (ref)(g) with $\alpha_{K^0}=1$ is assumed in Corollary 1 of bonhommeGroupedPatternsHeterogeneity2015. In the second part, $T$ must diverge faster as $\alpha_{K^0}$ is closer to 0, which implies that a smaller $\alpha_{K^0}$ needs to be compensated by a larger $T$. In particular, $T$ must diverge faster than $N$ if $\alpha_{K^0}=0$. Assumption (ref)(h) guarantees the identification of groups and is called the group separation condition in the literature. Assumption (ref)(i) requires that $T^{-1}\sum_{t=1}^T\|x_{it}\|^2$ be bounded by some (large) $M$ with arbitrary probability for sufficiently large $T$. In Assumption (ref)(j), $T^{-1}\sum_{t=1}^Tx_{it}\varepsilon_{it}$ has an arbitrarily thin tail for sufficiently large $T$. Assumptions (ref)(i) and (j) hold if $x_{it}$ and $x_{it}\varepsilon_{it}$ satisfy certain tail and mixing conditions; see bonhommeGroupedPatternsHeterogeneity2015.
The following assumption is exploited to derive the asymptotic distribution of $\widehat{\boldsymbol{\theta}}^{(K^0)}$.
The following theorem, which is understood up to group relabeling, is obtained assuming that $K=K^0$ is known.
bonhommeGroupedPatternsHeterogeneity2015 and lumsdaineEstimationPanelGroup2023 derive the same results as Theorem (ref)(i) and (ii), but they assume that all groups have a size proportional to $N$ (i.e., $\alpha_k=1$ for all $k$). Theorem 1 extends their results to models with small groups. A key condition is Assumption (ref)(g), which requires that $T$ be sufficiently large adaptively to the size of the smallest $K^0$-th group. We cannot ensure consistency and asymptotic normality without this assumption. This warns us about a potential limitation of K-means that may arise in panel data with small $T$ and small groups. Such a situation does not seem unusual, as researchers often apply clustering methods to panels of short to moderate length bonhommeGroupedPatternsHeterogeneity2015,suIdentifyingLatentStructures2016,martinez-zarzosoSearchingGroupedPatterns2020.
Now, we turn to the problem of estimating $K^0$ by using the information criterion. Referring to baiDeterminingNumberFactors2002 and baiPanelDataModels2009, bonhommeGroupedPatternsHeterogeneity2015 propose the following criterion:
where $n(K)$ is the number of parameters under $K$ groups, $\widetilde{\sigma}^2$ is a consistent estimate of $\sigma^2=\operatorname*{\mathrm{p}\!\lim} (NT)^{-1}\sum_{i=1}^N\sum_{t=1}^T\varepsilon_{it}^2$, and $h_{NT}>0$ is a penalty term satisfying $h_{NT}\to0$ as $N,T\to\infty$.\footnote{$\widetilde{\sigma}^2$ is typically constructed as $\widetilde{\sigma}^2=\{NT-n(K_{\max})\}^{-1}\times NT\widehat{\sigma}^2(K_{\max},\widehat{\boldsymbol{\gamma}}_N^{(K_{\max})})$ baiDeterminingNumberFactors2002,bonhommeGroupedPatternsHeterogeneity2015.} $n(K)=N+pK$ in model (ref). The estimated number of groups is
for some $K_{\max}\geq K^0$.
Proposition (ref) has two parts. Part (i) gives a sufficient condition on $h_{NT}$ for $\widehat{K}(h_{NT})$ not to overestimate. The second part is an inconsistency result. Condition (ref) tells us what kind of $h_{NT}$ we should not use. In the next subsection, we will check whether popular penalty choices satisfy (ref) and (ref).
We consider two examples. The first one is the penalty proposed in baiDeterminingNumberFactors2002, of the form
Clearly, $h_{NT}^{\mathrm{BN}}$ satisfies (ref), and hence $\widehat{K}(h_{NT}^{\mathrm{BN}})$ asymptotically does not overestimate. If $N/T\to c\in[0,\infty)$, condition (ref) is not satisfied under $h_{NT}=h_{NT}^{\mathrm{BN}}$, because $N^{1-\alpha_{m+1}}h_{NT}^{\mathrm{BN}}=N^{1-\alpha_{m+1}}(\ln N)/N\to0$ (unless $\alpha_{m+1}=0$). If $N/T\to\infty$, then $N^{1-\alpha_{m+1}}h_{NT}^{\mathrm{BN}}=N^{1-\alpha_{m+1}}(\ln T)/T$. Recalling Assumption (ref)(g), we can easily verify that, when $\alpha_{K^0}\geq1/2$, condition (ref) holds if $\nu>(1-\alpha_{m+1})^{-1}$ and $N/T^{\nu'}\to c \in (0,\infty]$ for some $\nu'\in[(1-\alpha_{m+1})^{-1},\nu)$, where $\nu$ is the number appearing in Assumption (ref)(g). When $\alpha_{K^0}<1/2$ and $N/T\to\infty$, and under Assumption (ref)(g), condition (ref) holds if, for example, $T=cN^{\delta}$ for some $c>0$ and some $\delta\in(1-2\alpha_{K^0},1-\alpha_{m+1}]$. Under these conditions, $\widehat{K}(h_{NT}^{\mathrm{BN}})$ asymptotically underestimates. Intuitively speaking, $\widehat{K}(h_{NT}^{\mathrm{BN}})$ cannot capture a small amount of information coming from cross-sections that belong to small groups, so that the small groups can be overlooked.
In sum, $\widehat{K}(h_{NT}^{\mathrm{BN}})$ can underestimate the number of groups if $N/T\to\infty$, as long as the regularity conditions hold. This is another result that warns us about a pitfall we may face when dealing with short panels. Furthermore, the above analysis reveals that $\widehat{K}(h_{NT}^{\mathrm{BN}})$ can be inconsistent if $N$ is disproportionately large relative to $T$, even if $T$ itself is “large" in the absolute sense. We emphasize that not only the value of $(N,T)$ but also the ratio $N/T$ matters.
Another example is the BIC-type penalty proposed in the supplementary materials to baiPanelDataModels2009 and bonhommeGroupedPatternsHeterogeneity2015 (in models with interactive fixed effects):
It can be easily verified that (ref) holds under $h_{NT}=h_{NT}^{\mathrm{BIC}}$. Moreover, $h_{NT}=h_{NT}^{\mathrm{BIC}}$ does not satisfy (ref) unless $\alpha_{m+1}=0$ and $N=\exp(T^{\beta})$ for some $\beta>1$, which is a rare situation in typical applications. However, as pointed out by bonhommeGroupedPatternsHeterogeneity2015, $\widehat{K}(h_{NT}^{\mathrm{BIC}})$ may overestimate the number of groups in finite samples. Indeed, our simulation (reported in Section (ref)) documents this tendency, indicating that $h_{NT}=h_{NT}^{\mathrm{BIC}}$ is not suitable unless a large sample is available.\footnote{This is not the case when we consider models with GFE; see Section (ref).}
The above analysis provides us with an insight into what kind of penalty may be useful. On the one hand, $\widehat{K}(h_{NT}^{\mathrm{BN}})$ can be inconsistent when $N/T\to \infty$, because $h_{NT}^{\mathrm{BN}}$ shrinks toward zero so slowly that (ref) holds. On the other hand, the shrinking rate of $h_{NT}^{\mathrm{BIC}}$ is fast enough to prevent (ref) from being satisfied, but this rate is so fast that $\widehat{K}(h_{NT}^{\mathrm{BIC}})$ overestimates the number of groups in finite samples. Therefore, a penalty that performs well needs to: (i) satisfy (ref), (ii) not satisfy (ref), (iii) shrink toward zero faster than $h_{NT}^{\mathrm{BN}}$ when $N/T\to\infty$, and (iv) shrink more slowly than $h_{NT}^{\mathrm{BIC}}$. Given this observation, we propose the following penalty:
It is straightforward to check that $h_{NT}^{\mathrm{MIC}1}$ satisfies (ref) but does not satisfy (ref) if $\alpha_{m+1}>0$.\footnote{If $\alpha_{m+1}=0$, it suffices to replace the denominator of $h_{NT}^{\mathrm{MIC}1}$ with $N^{1+\epsilon}$ for some small $\epsilon>0$.} $h_{NT}^{\mathrm{MIC}1}$ takes the same form as $h_{NT}^{\mathrm{BN}}$ when $N\leq T$, and is asymptotically equivalent to $h_{NT}^{\mathrm{BN}}$ if $N/T\to c\in [0,\infty)$. This is because our simulation shows that $h_{NT}^{\mathrm{BN}}$ performs well in this case. When $N/T\to\infty$, $h_{NT}^{\mathrm{MIC}1}=0.5\ln(NT)/N \leq (\ln N)/N \ll (\ln T)/T=h_{NT}^{\mathrm{BN}}$. Finally, $h_{NT}^{\mathrm{MIC}1} \gg h_{NT}^{\mathrm{BIC}}$. $h_{NT}^{\mathrm{MIC}1}$ satisfies all the required conditions.\footnote{A scaling constant used in the definition of $h_{NT}^{\mathrm{MIC}1}$ is not important asymptotically. We adjust a scaling constant just to preserve continuity of $h_{NT}^{\mathrm{MIC}1}$ in $(N,T)$. Although the choice of a scaling constant may affect the finite-sample performance, we do not explore in this direction.} In our simulation, $\widehat{K}(h_{NT}^{\mathrm{MIC}1})$ shows a good finite-sample performance; see Section (ref).
Since the introduction by bonhommeGroupedPatternsHeterogeneity2015, panel data models with GFE have drawn researchers' attention chetverikovSpectralPostspectralEstimators2022,mugnierSimpleComputationallyTrivial2023. In this section, we consider the following model:
where $\mu_{kt}\in\mathcal{M}\subset\mathbb{R}$ for $k\in[K]$ and $t\in[T]$. Here, $\mu_{k_it}$ is introduced as GFE into model (ref). To simplify the discussion, we assume $\mu_{kt}$'s are nonrandom. Define $\mu_k\coloneqq(\mu_{k1},\ldots,\mu_{kT})'$ for all $k$, and let $\boldsymbol{\mu}$ denote the set of all $\mu_k$'s. As earlier, we use the superscripts $0$ to refer to the true parameter values.
The least-squares estimator under $K$ groups is
With some abuse of notation, we let $\widehat{\sigma}^2(K, \widehat{\boldsymbol{\gamma}}_N^{(K)})= (NT)^{-1}\sum_{i=1}^N\sum_{t=1}^T\bigl(y_{it}-x_{it}'\widehat{\theta}_{\widehat{k}_i^{(K)}}^{(K)} - \widehat{\mu}_{\widehat{k}_i^{(K)}t}^{(K)}\bigr)^2$ denote the averaged least-squares SSR obtained from (ref) with $K$ groups.
Introducing GFE requires modifying Assumptions (ref)-(ref).
\setcounter{asm}{0}
Assumption (ref) requires that all groups be of divergent size as $N\to\infty$. This is because, for each $k\in[K]$, we can only use cross-sectional information to estimate $\mu_{kt}$ for each $t=1,\ldots,T$; that is, a consistent estimation of $\mu_{kt}$ is achieved only when $N_k^0\to \infty$.
Assumption (ref)(e') is essentially the same as Assumption S.2(a) of bonhommeGroupedPatternsHeterogeneity2015.
The following theorem establishes the asymptotic normality of the LS estimator from (ref) with $K=K^0$.
Theorem (ref) is a generalization of Corollary S.2 of bonhommeGroupedPatternsHeterogeneity2015 to the case of panels with small groups.
As before, we use $\mathrm{IC}(K,h_{NT})$ defined in (ref) as the criterion to determine the number of groups. Note that the definition of $\mathrm{IC}(K,h_{NT})$ in this section differs from that considered in Section (ref) in two respects. First, the first term is $\widehat{\sigma}^2(K, \widehat{\boldsymbol{\gamma}}_N^{(K)})= (NT)^{-1}\sum_{i=1}^N\sum_{t=1}^T\bigl(y_{it}-x_{it}'\widehat{\theta}_{\widehat{k}_i^{(K)}}^{(K)} - \widehat{\mu}_{\widehat{k}_i^{(K)}t}^{(K)}\bigr)^2$. Second, the number of parameters is $n(K)=N+(p+T)K$.
A key condition is (ref), similar to (ref). Condition (ref), however, differs from (ref) by a multiplicative factor of $T$. This factor emerges because the number of parameters is now $n(K)=N+(p+T)K$. The difference between the second terms of $\mathrm{IC}(K_1,h_{NT})$ and $\mathrm{IC}(K_2,h_{NT})$ ($K_1> K_2$), say, is $\widetilde{\sigma}^2(n(K_1) - n(K_2))h_{NT} = \widetilde{\sigma}^2(p+T)(K_1-K_2)h_{NT} \asymp Th_{NT}$, so that $Th_{NT}$ plays the role of a penalty in effect.
Let us investigate whether penalties considered in Section (ref) satisfy (ref). Recall that $h_{NT}^{\mathrm{BN}}$ is defined by (ref). When $N/T\to\infty$, we have $TN^{1-\alpha_{m+1}}h_{NT}^{\mathrm{BN}} = TN^{1-\alpha_{m+1}}(\ln T)/T \to \infty$. When $N/T\to c\in[0,\infty)$, then $TN^{1-\alpha_{m+1}}h_{NT}^{\mathrm{BN}} \asymp TN^{1-\alpha_{m+1}}(\ln N)/N \to \infty$. Therefore, $h_{NT}^{\mathrm{BN}}$ necessarily satisfies (ref), which implies that $\widehat{K}(h_{NT}^{\mathrm{BN}})$ asymptotically underestimates in model (ref). Furthermore, $h_{NT}^{\mathrm{BN}}$ satisfies (ref) even when $\alpha_{m+1}=\alpha_{K^0}=1$ (i.e., there is no small group), although we exclude this case by Assumption (ref). Indeed, our simulation shows that $\widehat{K}(h_{NT}^{\mathrm{BN}})$ underestimates the number of groups in GFE model (ref) even when all groups have a size proportional to $N$, indicating that underestimation by $\widehat{K}(h_{NT}^{\mathrm{BN}})$ is not a matter confined to the small-group situation.
Next, consider $h_{NT}^{\mathrm{BIC}}$. When $N/T\to c\in(0,\infty]$, we have $TN^{1-\alpha_{m+1}}h_{NT}^{\mathrm{BIC}}=TN^{1-\alpha_{m+1}}\ln(NT)/NT \leq \ln(N^2)/N^{\alpha_{m+1}}\to0$, and hence (ref) does not hold in this case. On the other hand, $h_{NT}^{\mathrm{BIC}}$ can satisfy (ref) and hence underestimate the number of groups when $N/T\to0$. For example, it is straightforward to check that $TN^{1-\alpha_{m+1}}h_{NT}^{\mathrm{BIC}} \to \infty$ if $T=\exp(N^{\alpha_{m+1}+\epsilon})$ for some $\epsilon>0$. In our simulation reported in Section (ref), $\widehat{K}(h_{NT}^{\mathrm{BIC}})$ performs well in a setting where $T$ is not too small or too large relative to $N$.
Finally, we consider constructing a penalty that fits model (ref). Since $Th_{NT}$ essentially behaves as a penalty, we suggest the following penalty, which is obtained by simply dividing $h_{NT}^{\mathrm{MIC}1}$ by $T$:
When $N>T$, $h_{NT}^{\mathrm{MIC}2}=h_{NT}^{\mathrm{BIC}}$. When $N\leq T$, $h_{NT}^{\mathrm{MIC}2}$ is still asymptotically equivalent to $h_{NT}^{\mathrm{BIC}}$ if $\ln T = O(\ln N)$. However, if $T$ satisfies $\ln T/\ln N \to\infty$, then $h_{NT}^{\mathrm{MIC}2} \ll h_{NT}^{\mathrm{BIC}}$ and thus $\widehat{K}(h_{NT}^{\mathrm{MIC}2})$ tends to be larger than $\widehat{K}(h_{NT}^{\mathrm{BIC}})$. In a typical application where $T$ is not too large relative to $N$, one will often have $\widehat{K}(h_{NT}^{\mathrm{MIC}2})=\widehat{K}(h_{NT}^{\mathrm{BIC}})$.
In this section, we evaluate the finite-sample performance of the information criteria for the number of groups and the LS estimator calculated under $K=K^0$. We consider the following three data generating processes (DGPs) in which the number of groups is $K^0=3$:
DGP 1 (Static panel). $y_{it}=\theta_{k_i,1}x_{it,1} + \theta_{k_i,2}x_{it,2} + \varepsilon_{it}$, where $x_{it,1},x_{it,2},\varepsilon_{it}\sim \ \mathrm{i.i.d.} \ N(0,1)$, $(\theta_{1,1},\theta_{1,2})=(3,-3)$, $(\theta_{2,1},\theta_{2,2})=(1,-2)$, and $(\theta_{3,1},\theta_{3,2})=(4,-1)$.
DGP 2 (Dynamic panel). $y_{it}=\theta_{k_i,1}x_{it,1} + \theta_{k_i,2}y_{i,t-1} + \varepsilon_{it}$, where $(\theta_{1,1},\theta_{1,2})=(3,0.2)$, $(\theta_{2,1},\theta_{2,2})=(1,0.5)$, and $(\theta_{3,1},\theta_{3,2})=(4,0.8)$.
DGP 3 (GFE). $y_{it}=\theta_{k_i,1}x_{it,1} + \theta_{k_i,2}x_{it,2} + \mu_{k_it} + \varepsilon_{it}$, where $(\theta_{1,1},\theta_{1,2})=(3,-3)$, $(\theta_{2,1},\theta_{2,2})=(1,-2)$, $(\theta_{3,1},\theta_{3,2})=(4,-1)$, $\mu_{1t}=4t/T$, $\mu_{2t}=2t/T$, and $\mu_{3t}=4$ for all $t$.
For each DGP, we consider $N\in\{60,90,120\}$ and set $N_1^0=N/K^0$, $N_3=\lfloor c_\alpha N^{\alpha}\rfloor$, where $\alpha\in\{0.2,0.3,\ldots,1\}$, and
and $N_2=N - (N_1+N_3)$. Here, scaling constant $c_\alpha$ is chosen so that $N_2$ and $N_3$ are not too small when $\alpha=1$ and $N_3$ is strictly increasing in $\alpha$. Group 3 is the smallest group unless $\alpha=1$. For DGPs 1 and 3, we apply within-transformation to the data before conducting K-means clustering to demonstrate that the theoretical results derived in the present work are robust to the presence of individual fixed-effects and the routinely applied transformation, as long as the regressors does not include a lagged dependent variable.\footnote{If the regressors include a lagged dependent variable as in DGP 2, and if the model includes individual fixed-effects, then one will need to apply bias correction or IV estimation; see, e.g., hahnAsymptoticallyUnbiasedInference2002, phillipsBiasDynamicPanel2007, and the supplementary material to bonhommeGroupedPatternsHeterogeneity2015, Section S.4.1. Models that include both individual fixed-effects and a lagged dependent variable on the right-hand side are beyond our scope, so we do not apply within-transformation for DGP 2.} The number of replications is 100.
In the computational aspect, we follow Algorithm 1 of bonhommeGroupedPatternsHeterogeneity2015 to numerically solve (ref) and (ref). Specifically, we employ the following iterative algorithm:
Since it is well-known that the above iterative procedure depends on the initial value and can converge to a local minimum, we repeat Algorithm (ref) over 1000 initializations.
For DGPs 1-2, we analyze the performance of $\widehat{K}(h_{NT}^{\mathrm{BN}})$, $\widehat{K}(h_{NT}^{\mathrm{BIC}})$, and $\widehat{K}(h_{NT}^{\mathrm{MIC}1})$. For DGP 3, we consider $\widehat{K}(h_{NT}^{\mathrm{BN}})$, $\widehat{K}(h_{NT}^{\mathrm{BIC}})$, and $\widehat{K}(h_{NT}^{\mathrm{MIC}2})$. Information criteria are calculated over $2\leq K \leq K_{\max} = 5$.
First, we study the performance of information criteria when $T$ is small relative to $N$. Specifically, we consider $T\in\{10,20\}$. Figure (ref) displays the mean of estimated number of groups for each penalty as a function of $\alpha$, calculated for DGP 1. When $h_{NT}^{\mathrm{BN}}$ is used, $\widehat{K}(h_{NT}^{\mathrm{BN}})$ correctly estimates the number of groups to be 3 on average for $\alpha\geq 0.7$ and $T=10$, but it tends to underestimate when $\alpha\leq 0.6$ and $T=10$. This tendency to underestimate is somewhat mitigated when $T=20$, but $\widehat{K}(h_{NT}^{\mathrm{BN}})$ still underestimates the number of groups for $\alpha\leq0.5$. This implies that $\widehat{K}(h_{NT}^{\mathrm{BN}})$ tends to overlook small groups. It is noteworthy that, if $T$ is fixed, the problem of underestimation is more severe for larger $N$, implying that only increasing the number of cross-sections does not improve (in fact, aggravates) the estimation of the number of groups. These observations corroborate the theoretical analysis given in Section (ref). In contrast, $\widehat{K}(h_{NT}^{\mathrm{BIC}})$ overestimates the number of groups on average for all $(N,T,\alpha)$. In particular, it estimates the number of groups to be 5 in all replications for $N=90$ and $N=120$ (irrespective of the value of $T$). For the MIC, $\widehat{K}(h_{NT}^{\mathrm{MIC}1})$ shows reasonably good performance for all $(N,T,\alpha)$. Although it can overestimate or underestimate when $\alpha$ is small or $T$ is small relative to $N$, its mean stays within a close neighborhood of the true number 3, unlike $\widehat{K}(h_{NT}^{\mathrm{BN}})$ and $\widehat{K}(h_{NT}^{\mathrm{BIC}})$.
Figure (ref) shows the simulation results for DGP 2. The general tendency is the same as in the case of DGP 1. $\widehat{K}(h_{NT}^{\mathrm{BN}})$ performs well only when $\alpha$ is close to 1 and $T$ is not too small compared to $N$, and tends to underestimate the number of groups otherwise. $\widehat{K}(h_{NT}^{\mathrm{BIC}})$ overestimates for all $(N,T,\alpha)$. $\widehat{K}(h_{NT}^{\mathrm{MIC}1})$ can correctly estimate the number of groups on average.
In Figure (ref), we show the performance of information criteria for DGP 3. Note that $h_{NT}^{\mathrm{BIC}}=h_{NT}^{\mathrm{MIC}2}$ because we consider the model with GFE and cases with $N>T$. As predicted in Section (ref), $\widehat{K}(h_{NT}^{\mathrm{BN}})$ underestimates the number of groups for all $(N,T,\alpha)$. In fact, it estimates the number of groups to be 2 in all replications. For $h_{NT}^{\mathrm{BIC}}$, $\widehat{K}(h_{NT}^{\mathrm{BIC}})$ performs well for $(N,T)=(60,10)$ but tends to overestimate for $(N,T)=(90,10)$ and $(N,T)=(120,10)$, with this tendency stronger for larger $N$. Therefore, $\widehat{K}(h_{NT}^{\mathrm{BIC}})$ can overestimate the number of groups when $T$ is small relative to $N$, although this tendency is weaker than in models without GFE.\footnote{Overestimation by the BIC estimator in the GFE model with $N$ large relative to $T$ is also reported in janysMentalHealthAbortions2024.} When $T$ increases to 20, $\widehat{K}(h_{NT}^{\mathrm{BIC}})$ shows good performance for all $(N,T,\alpha)$, indicating that the BIC-type penalty is a reasonable choice in models with GFE when $T$ is not too small compared to $N$.
In this subsection, we analyze the behavior of $\widehat{K}(h_{NT})$ when $T$ is larger than $N$. To do this, we consider cases with $T/N\in\{1.5,3\}$ for each DGP and $N$, and fix the size of the third group at $\alpha=0.3$. Because $h_{NT}^{\mathrm{MIC}1}=h_{NT}^{\mathrm{BN}}$ for DGPs 1-2, and $h_{NT}^{\mathrm{MIC}2}$ is asymptotically equivalent to $h_{NT}^{\mathrm{BIC}}$ for DGP 3, we only consider $\widehat{K}(h_{NT}^{\mathrm{BN}})$ and $\widehat{K}(h_{NT}^{\mathrm{BIC}})$ in this subsection. The results are summarized in Table (ref).
For DGP 1, we observe three facts. First, $\widehat{K}(h_{NT}^{\mathrm{BN}})=\widehat{K}(h_{NT}^{\mathrm{MIC}1})$ is close or equal to the true value 3 on average thanks to the simulation design where $T$ is larger than $N$. Second, increasing $N$ and $T$ simultaneously and increasing $T$ only (with $N$ fixed) improve the performance of $\widehat{K}(h_{NT}^{\mathrm{BN}})=\widehat{K}(h_{NT}^{\mathrm{MIC}1})$. Third, the mean of $\widehat{K}(h_{NT}^{\mathrm{BIC}})$ approaches the true value $K^0=3$ as $T$ increases with $N$ fixed but approaches the upper bound, $K_{\max}=5$, as $N$ increases. This indicates that the overestimation by $\widehat{K}(h_{NT}^{\mathrm{BIC}})$ cannot be solved even in moderately large samples if GFE is not incorporated into the model. The same observations hold for DGP 2.
In DGP 3, which includes GFE, $\widehat{K}(h_{NT}^{\mathrm{BN}})$ underestimates the number of groups for all $(N,T)$ considered. For the BIC, $\widehat{K}(h_{NT}^{\mathrm{BIC}})$ precisely estimates the number of groups to be 3 on average. This, in conjunction with the results given in Section (ref), leads us to conclude that $h_{NT}=h_{NT}^{\mathrm{BIC}}$ is a reasonable choice in the GFE model (ref) if $T$ is not too small compared to $N$. However, recall that $\widehat{K}(h_{NT}^{\mathrm{BIC}})$ can underestimate the number of groups if $T$ is too large relative to $N$, in which case $h_{NT}=h_{NT}^{\mathrm{MIC}2}$ is preferred; see Section (ref).
Next, we examine how the group size affects the performance of the K-means estimator calculated given the correct number of groups. Since (true) Groups 1 and 2 have sizes proportional to $N$ for all $\alpha\in\{0.2,\ldots,1\}$, we present simulation results for Group 3 only. For each DGP, we calculate the root mean squared error (RMSE). Specifically, for DGPs 1-2, the RMSE is defined by
and for DGP 3,
Figure (ref) reports the mean RMSE over 100 replications. For DGP 1 (Figure (ref)), there are three points to be noted. First, increasing $N$ with $T$ fixed improves the estimation accuracy when $\alpha\geq0.7$, but this improvement is only marginal. Second, increasing $N$ while fixing $T$ can lead to a larger estimation error for small $\alpha$. In particular, $N=120$ gives by far the largest RMSE on average when $T=10$ and $\alpha=0.2$. Third, this large estimation error can be mitigated by increasing $T$. These observations are consistent with our theoretical results. The same comment applies to the cases of DGPs 2-3.
We also evaluate the effect of the presence of a small group on classification performance. We calculate the proportion of perfect classification (PPC) defined as
See Figure (ref). We observe three facts. First, increasing $N$ while fixing $T=10$ decreases PPC irrespective of the value of $\alpha$. Second, for models without GFE, a smaller $\alpha$ can lead to a smaller PPC when $T=10$, but the effect of $\alpha$ is not remarkable. In contrast, PPC is highly sensitive to the group size when $T=10$ in the model with GFE. Third, PPC takes values close to 1 for all $\alpha$ when $T=20$, indicating that the group size has a minor effect on classification performance if $T$ is not too small relative to $N$.
In this section, we apply K-means clustering to investigate the determinants of sales growth of Japanese firms. Our empirical application is inspired by lumsdaineEstimationPanelGroup2023, who explain that sales growth is a key measure of corporate performance and hence it is important for economists, investors, and decision makers to understand how firms' variables affect sales growth. In particular, the authors are interested in the relationship between leverage and sales growth because there are competing views on the significance of leverage in relation to corporate performance myersDeterminantsCorporateBorrowing1977,millerLeverage1991.
As lumsdaineEstimationPanelGroup2023 argue, it is documented in the literature that firms' responses to market-wide shocks exhibit a group pattern of heterogeneity. Motivated by this fact, we study the determinants of sales growth of Japanese firms using the following model:
where $y_{it}$ is sales growth, $x_{it}$ is the set of regressors, and $\eta_i$ is the individual fixed effects. Following lumsdaineEstimationPanelGroup2023, we let $x_{it}$ include leverage (LEV), the logarithm of total assets (TA), Tobin's q (TQ), cash flow (CF), property, plant and equipment (PPE), and return on assets (ROA). Lagged regressors $x_{i,t-1}$ are used in the regression to alleviate possible endogeneity lumsdaineEstimationPanelGroup2023.
The variables used in this section are obtained from the eol database. After excluding firms that contain missing variables and outlying sales growth,\footnote{Outliers are defined as sales growth that fall outside $3$ standard deviations neighborhood of its mean.} 417 firms remain in our dataset. The sample period is 2005-2020.
We first determine the number of groups using the information criterion, $\mathrm{IC}(K,h_{NT})$. $\mathrm{IC}(K,h_{NT})$ is calculated over $2\leq K \leq K_{\max} = 8$ using different choices for $h_{NT}$. Table (ref) shows the selected number of groups and the sizes of the estimated groups obtained from each $\widehat{K}(h_{NT})$-group clustering. When $h_{NT}=h_{NT}^{\mathrm{BN}}$, the estimated number of groups is 2, the minimum possible value, while $h_{NT}=h_{NT}^{\mathrm{BIC}}$ yields the upper bound value $\widehat{K}(h_{NT}^{\mathrm{BIC}})=8$, as expected from our numerical experiment. In contrast, our proposed $h_{NT}=h_{NT}^{\mathrm{MIC}1}$ leads to a 3-group model. Furthermore, the smallest group (labeled Group 3) contains only 3 firms. The fact that this small group is overlooked when we use $h_{NT}=h_{NT}^{\mathrm{BN}}$ is consistent with the analyses given in the preceding sections. Although the 8-group model suggested by the BIC leads to unbiased estimation of group heterogeneity (as long as $8\geq K^0$), the estimates may be unnecessarily inefficient (unless $K^0=8$) and interpretation of the group structure is difficult. Because $h_{NT}=h_{NT}^{\mathrm{MIC}1}$ leads to a reasonably parsimonious model and allows us to detect a small group, we will now proceed with the 3-group model.
Shown in Table (ref) are the slope coefficient estimates. For each regressor, the sign of the estimated coefficient is generally identical across groups, while the magnitude of the estimated slope differs remarkably across groups. For Group 1, all regressors but PPE have a significant effect on sales growth. While leverage, total assets and ROA have a negative effect, Tobin's q affects sales growth positively. The facts that total assets and ROA negatively affect sales growth and that Tobin's q exhibits a positive effect are largely consistent with the results of lumsdaineEstimationPanelGroup2023, who investigate the determinants of sales growth using US firms data. For Group 2, all estimated slopes are significant and of larger magnitude than those of Group 1. For Group 3, only the coefficients on the log of total assets and ROA are significant, probably due to the small sample size for Group 3. Group 3 is notably differentiated from the other groups by the magnitudes of the coefficients on the total assets and ROA. For example, compared with Group 1, the coefficient on the total assets is more than 40 times larger, and that on ROA is more than 30 times larger, which reveals that there is remarkable heterogeneity across groups.
An interesting question arising from the above 3-group model is how we can characterize the small group (Group 3). Group 3 consists of (ferrous and nonferrous) metal producers, and so our clustering analysis reveals that there is a small group in the metal industry. Let us also characterize Group 3 based on observables. Figure (ref) presents the time path of the averages of sales growth and regressors, where the averages are taken within each group. Up to the sampling error (mainly due to the extremely small sample size of Group 3 with $N_3=3$), there are notable differences between the group averages of total assets and PPE; namely, Group 3 can be characterized by high total assets and low PPE, or a low PPE/total assets ratio.
In this article, we analyzed the behavior of the K-means/LS estimator relaxing the conventional assumption of all groups having a size proportional to $N$. Our framework allows for smaller groups whose sizes are of order $N^{\alpha}$ with $\alpha<1$, and we derived conditions under which the LS estimators for the slope parameter and GFE are consistent and asymptotically normal, assuming that the true number of groups is known. In the case of an unknown number of groups, we derived conditions under which the information criterion for the number of groups is inconsistent. Our theoretical and numerical analyses showed that an information criterion proposed by baiDeterminingNumberFactors2002 can underestimate the number of groups, particularly when there are small groups or the model includes GFE. The BIC proposed by bonhommeGroupedPatternsHeterogeneity2015 was found to overestimate the number of groups in models without GFE, although it is a reasonable choice if the model includes GFE and $T$ is not too small or too large relative to $N$. We proposed information criteria that are designed not to satisfy the derived conditions for inconsistency and to improve the above precursors. A simulation study confirmed their good performance in finite samples. In the simulation study, we also evaluated the finite-sample performance of the LS estimator given the correct number of groups. We found that increasing $N$ with $T$ fixed does not necessarily improve, or can worsen, classification performance and the accuracy of parameter estimation for small groups, a problem that becomes more severe as the sizes of those groups are smaller. Our empirical application illustrates that our modified information criterion enables detecting and characterizing small groups in a reasonably parsimonious group structure.