EconBase
← Back to paper

Heterogeneous Grouping Structures in Panel Data

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

85,454 characters · 14 sections · 49 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Heterogeneous Grouping Structures in Panel Data

\affil[1]{ King's Business School , King's College London}

abstractIn this paper we examine the existence of heterogeneity within a group, in panels with latent grouping structure. The assumption of within group homogeneity is prevalent in this literature, implying that the formation of groups alleviates cross-sectional heterogeneity, regardless of the prior knowledge of groups. While the latter hypothesis makes inference powerful, it can be often restrictive. We allow for models with richer heterogeneity that can be found both in the cross-section and within a group, without imposing the simple assumption that all groups must be heterogeneous. We further contribute to the method proposed by su2016identifying, by showing that the model parameters can be consistently estimated and the groups, while unknown, can be identifiable in the presence of different types of heterogeneity. Within the same framework we consider the validity of assuming both cross-sectional and within group homogeneity, using testing procedures. Simulations demonstrate good finite-sample performance of the approach in both classification and estimation, while empirical applications across several datasets provide evidence of multiple clusters, as well as reject the hypothesis of within group homogeneity. \newline{{Keywords}}: Panel Data Models, Latent Grouping Structures, Heterogeneity tests. \newline {JEL Classifications}: C18, C23, C38, C51, C55.

\affil{King's Business School , King's College London}

Introduction

The use and analysis of panel datasets and associated models is a vital research topic in both econometric theory and applications; these models are frequently applied in many disciplines across economics and other social sciences. This type of dataset allows the consideration of both cross-sectional and time dimensions, and provides a rich source of empirical information. In its majority, panel data analysis relies heavily on the assumption of homogeneity across units making inference more powerful.

Nevertheless, such an assumption is often restrictive and frequently rejected; see, among others, chudik2015common, phillips2007transition, pesaran2006estimation, and browning2014dynamic. Ignoring heterogeneity, across units, can result in biased estimation and misleading inference; see, among others, Chapter 1 in baltagi2008econometric, moulton1986random, and moulton1987diagnostics. On the other hand, allowing complete heterogeneity leads to potentially inefficient panel data estimation and inference. To address this issue, researchers often consider intermediate ways of exploring panel data such as panel structure models.

Panel structure models rely on the hypothesis that all, or a subset of, parameters are heterogeneous across groups but homogeneous within a group where neither the number of groups nor the individuals' group-membership is known. Examples of these models can be found in phillips2007transition who accommodate the hypothesis of convergence of clubs; in this case different countries belong to different groups and behave differently with the group structures being latent rather than assumed to be observable.

The identification of latent structures is not an easy task. It is a computationally expensive procedure and it can be even infeasible to try all the possible permutations of units across groups. A part of the literature considers that units in the dataset are grouped using an external observable classification; see, e.g., BESTER2016197. This approach may suffer, as suitable external variables can be difficult to find in many empirical settings. Furthermore, a wrong choice for the external classifier may again lead to misleading inference. For this reason, it can be argued that it is more useful for the researcher to use data-driven classification procedures such as k-means, see, e.g., sarafidis2015partially, bonhomme2015grouped, lin2012estimation, and ando2016panel ; these studies consider k-means in linear panel data models, with an unknown grouping structure. Although unsupervised classification is an appealing approach and has been theoretically proven to be asymptotically consistent, see, e.g., pollard1981strong, alternative approaches based on the Classifier Lasso (c-lasso), proposed by su2016identifying (SSP) and extended by su2018identifying and su2019sieve, have been recently introduced in the literature. In particular, SSP propose a novel estimation technique which involves penalisation and simultaneously classifies the data and estimates the underlying parameters consistently. meh2022 provides a framework for joint estimation and identification of latent grouping structures in panel data models using a pairwise fusion penalized approach. mammen2022estimation presents a new approach to grouping fixed effects in a linear panel model to reduce their dimensionality and ensure identifiability, by using unsupervised non-parametric density based clustering, cluster patterns including their location and number are not restricted.

A different line of work is that of ren2022matrix who propose a group matrix network autoregression (GMNAR) model, which assumes that the subjects in the same group share the same set of model parameters. In a similar vein, zhu2022simultaneous study dynamic behaviours of heterogeneous individuals observed in a network, where the dynamic patterns are characterised by a network vector autoregression model with a latent grouping structure, where group-wise network effects and time-invariant fixed-effects can be incorporated. wu2022 develop methods to recover sparsity patterns and grouping structures in large panels with both individual and time effects.

In a low dimensional setting, where $T>>N,p$, wanga2022panel generalise their framework and consider panel model with interactive fixed effects such that individual heterogeneity is captured by latent grouping structure and time heterogeneity is captured by an unknown structural break. liu2022panel propose a methodology for identifying and estimating explosive bubbles in mixed-root panel autoregressions with a latent grouping structure.

Full homogeneity placed within groups, is a very common assumption in the literature, but this can be a potentially restrictive hypothesis. It is conceivable and, as we will argue, often the case, in practice, that each group contains units with heterogeneous parameters whose mean is group specific and differs across groups.

To fix ideas let the parameter of interest across $N$ panel units be $\boldsymbol{\beta}_{i} = (\boldsymbol{\beta}_{1,i}, \boldsymbol{\beta}_{2,i}, \ldots, \boldsymbol{\beta}_{p, i})^{\prime}$, where $p$ is the number of covariates within each unit $i=1,\ldots,N$. Traditionally, under full homogeneity, $\boldsymbol{\beta }_{i}=\boldsymbol{\beta}$. This is full homogeneity. A heterogeneous extension is $\boldsymbol{\beta}_{i}=\boldsymbol{\beta}+\boldsymbol{\eta}_{i}$, where $\boldsymbol{\eta}_{i}$ is an i.i.d. process and focus is placed on estimation of and inference for $\boldsymbol{\beta}$. Grouping structures specify that $\boldsymbol{\beta}_{i}=\boldsymbol{\alpha}_{k}$, whereby unit $i$ belongs to group $k$, where $k=1,\ldots,K$ and $K$ is usually assumed finite. Focus here is placed on $\boldsymbol{\alpha}_{k}$. Our setting extends this by specifying that $\boldsymbol{\beta}_{i}=\boldsymbol{\alpha}_{k}+\boldsymbol{\eta}_{i}.$ Clearly $\boldsymbol{\eta}_{i}=\boldsymbol{0}_{p\times1}$ and $\boldsymbol{\alpha}_{k}=\boldsymbol{\alpha}$, are both assumptions that are empirically verifiable and potentially invalid. The case where neither holds has not been explored in the literature and is the main focus, and contribution, of this paper.

Following the work of SSP, we extend the Classifier-Lasso framework to the above case and further propose an estimation technique based on the k-means clustering method allowing for within-group heterogeneity. Our focus is mainly to showcase consistency of the estimation of $\boldsymbol{\beta}_{i}$ and $\boldsymbol{\alpha}_{k}$ providing both theoretical and small sample evidence, rather than the identification of an exact grouping structure, which under our form of heterogeneity is obviously not consistently possible. We suggest the use of a simple unsupervised classifier, like the k-means, to determine the grouping structure and, once determined, we estimate the unit-specific parameters with penalised methods.

Our work extends SSP and contributes to the literature as follows: first, we allow for a less restrictive model, second we allow for richer cross-sectional heterogeneity, and finally, we provide a data-driven approach to classify each cross-section. While the latter is a gain of our modelling approach, heterogeneity is still restrictive and, in certain cases, contradicts the basic principles of k-means clustering. In particular, when the level of heterogeneity is high, causing the data to be more noisy cross-sectionally, the clustering process tends to be misleading since it violates one of the three basic assumptions of k-means; that is, the variance between individuals must be constant and small. Because k-means minimises the within cluster deviation (in terms of squared Euclidean distances), the optimisation method converges to local minima when the variance between individuals (cross-sections) is large.

The latter is inconvenient when the empirical data is noisy or even correlated, but there are other unsupervised classifiers which deal with these types of cases; see, e.g., leisch1999bagged who re-estimates the data in different partitions using bootstrapping. In this paper, we do not focus on such a task since our approach does not cover cases where data exhibit autocorrelation or endogeneity. Obviously, an alternative to our approach would be the construction of factors instead of groups (clusters). This would lower the dimension of the data by choosing $K$ factors which explain considerably the level of heterogeneity in the cross-sections; see, e.g., su2018identifying for more details.

A further and important contribution of our work is a set of testing procedures, exploring the validity of $\boldsymbol\eta_{i}=\boldsymbol0$ and $\boldsymbol\alpha_{k}=\boldsymbol\alpha$, using testing procedures in the former case. We use our methods in a variety of empirical settings to explore the validity of the restrictions and find that in most cases both are not valid, further showcasing the utility of our framework.

The rest of the paper is organised as follows. Sections (ref) and (ref) outline our theoretical considerations. In Section (ref) we present extensive Monte Carlo simulations and discussion, Section (ref) presents several empirical examples in support of our method. We conclude in Section (ref). Technical proofs, of the main (auxiliary) results, and descriptions of the empirical datasets are relegated to the Appendix.

Notation

For any vector $\boldsymbol{x}\in\mathbb{R}^{n}, $ we denote the $\ell_{p}$-, and $\ell_{\infty}$ as $\left\| \boldsymbol{x} \right\| _{p}=\left( \sum_{i=1}^{n} | x_{i} |^{p}\right) ^{1/p}$, $\lVert \boldsymbol{x} \rVert_{\infty} = \max_{i=1,\ldots,n}|x_{i}|$. Further, $\boldsymbol{1}\{\cdot\}$ denotes an indicator function. We use "$\to_{P}$" to denote convergence in probability. For two deterministic sequences $a_{n}$ and $b_{n}$ we define asymptotic proportionality, "$\asymp$", by writing $a_{n} \asymp b_{n}$ if there exist constants $0 < a_{1} \leq a_{2}$ such that $a_{1}b_{n} \leq a_{n} \leq a_{2}b_{n}$ for all $n \geq1$. For any set $A$, $|A|$ denotes its cardinality, while $A^{c}$ denotes its complement. Define $\boldsymbol{A}$, a $n\times m $ matrix, we denote $\Lambda_{\min}(\boldsymbol{A})$ as the smallest eigenvalue of $\boldsymbol{ {A}}$, and $\Lambda_{\max}(\boldsymbol{A})$ as the largest eigenvalue of $\boldsymbol{ {A}}$.

Theoretical considerations

We consider the following balanced panel model:

align[align omitted — 329 chars of source]

where $\mu_{i}^{0}$ is an individual fixed effect, $\boldsymbol{x}_{it}^{\ast }=\left( {x}_{1,it}^{\ast},\ldots,{x}_{p,it}^{\ast}\right) ^{\prime}$is a $(p\times1)$ vector of covariates, with $E(x_{l,it}^{\ast}x_{j,it}^{\ast})=0$, $\varepsilon_{it}^{\ast}=(\varepsilon_{1t},\ldots,\varepsilon_{Nt}^{\ast})$ is the idiosyncratic error, $\boldsymbol{\alpha}^{0}=(\boldsymbol{\alpha}_{1} ^{0},\ldots,\boldsymbol{\alpha}_{K}^{0})^{\prime}$ a $K\times p$ matrix of group specific centres for each covariate $j$, where $K$ denotes the number of groups considered. $\boldsymbol{G}_{k}^{0}$, $k=1,\ldots,K$, are sets of indices, denoting groups of units. Note that in our analysis, both $N$ and $T$ can be large and $N$ can possibly grow faster than $T$. Further, $\boldsymbol{\alpha }_{j}^{0}\neq\boldsymbol{\alpha}_{k}^{0}$ for any $j\neq k,$ $\boldsymbol{G} _{k}^{0}\subset\{1,2,\ldots,N\}$, $\cup_{k=1}^{K}\boldsymbol{G}_{k} ^{0}=\{1,2,\ldots,N\}$, and $\boldsymbol{G}_{k}^{0}\cap\boldsymbol{G}_{j} ^{0}=\varnothing$ for any $j\neq k$, $j=1,\ldots,K$. We denote the cardinality of the set $\boldsymbol{G}_{k}$ as $N_{k}=|\boldsymbol{G}_{k}^{0}|$. The slope heterogeneity, $\boldsymbol{\eta}_{i}=(\eta_{i,1},\ldots,\eta_{i,p})^{\prime}$ is considered to be unknown and allowed to be group specific, along with the cluster membership. Intuitively, $K$ corresponds to the number of groups and units (indexed by $i$) within the same group share a common slope $\boldsymbol{\alpha}_{k}$ but are allowed to differ by $\boldsymbol{\eta} _{i},i\in\boldsymbol{G}_{k},k=1,\ldots,K$. Further the grouping membership is common across different covariates.

Often in the panel data literature, a common characteristic can be proxied by factors (or interactive effects), e.g., as in bai2009panel, because of their ability to effectively summarise information in large data sets. While existing work suggests that the grouping of different units in a panel model can be effective under cross-sectional dependence and heteroscedasticity where $N<T$, see, e.g., bai2020standard, and can even accommodate time variation, see e.g. su2019sieve. In this paper, we allow for a degree of within-group heterogeneity, implying that grouping does not account for the entirety of cross-sectional heterogeneity. While we are agnostic about the cross-sectional and within-group idiosyncratic effect, we do assume an a-priori knowledge of the existence of grouping.

For the linear model in (ref) with ${E}\left( \varepsilon _{it}^{\ast}\mid\boldsymbol{x}_{it}^{\ast},\mu_{i}^{0}\right) =0$, we have $\widehat{\mu}_{i}=\bar{y}_{i.}^{\ast}-\boldsymbol{\beta}_{i}^{\prime} \bar{\boldsymbol{x}}_{i}^{\ast}$, where $\bar{y}_{i.}^{\ast}=\frac{1}{T} \sum_{t=1}^{T}y_{it}^{\ast},{y}_{it}=y_{it}^{\ast}-\bar{y}_{i}^{\ast}$, and ${\bar{\boldsymbol{x}}}_{i}$, $\boldsymbol{x}_{it}$, $\bar{\varepsilon}_{it}$, and $\varepsilon_{it}$ are analogously defined. Then (ref) becomes

align[align omitted — 134 chars of source]

Following su2016identifying and motivated by the literature on fused Lasso e.g. tibshirani2005sparsity, we minimise the following penalised profile likelihood (PPL) to estimate the parameters of interest.

equation[equation omitted — 364 chars of source]

where $\lambda>0$ is the regularisation parameter and $\boldsymbol{{\beta =(\boldsymbol{\beta}}}_{1},\ldots,\boldsymbol{{\beta}}_{N}\boldsymbol{{)} }^{\prime}$. Then, (ref) is minimised for some value of $(\lambda,K)$ and produces $(\boldsymbol{\widehat{\beta}} ,\boldsymbol{\widehat{\alpha}})^{\prime}$. We generalise su2016identifying (c-lasso hereafter) by allowing for heterogeneity within groups.

We show that both the group and the unit-specific parameter can be consistently estimated even when the underlying generating process is (ref). We obtain estimates of the (unit)group-specific parameters by using either the c-lasso iterative procedure or k-means estimation, see e.g. macqueen1967, pollard1981strong.

The penalty term in (ref) takes a mixed additive-multiplicative form, introduced in su2016identifying. In the traditional Lasso literature, e.g. tibshirani1996regression, the penalty term is additive, and imposes a degree of sparsity in favour of a parsimonious model at the cost of a degree of bias, whereas in this setting, parsimony results from grouping. The multiplicative nature of the penalty term in (ref) is essential as it produces $N$ additive terms for each of the $K$ separate penalties, such that $\{i=1\ldots,N,\;k=1,\ldots ,K:\;i\in\boldsymbol{G}_{k},\boldsymbol{\beta}_{i}^{0}\in \boldsymbol{\mathcal{B}}_{k}\subseteq\boldsymbol{\mathcal{B}}\}$, where, for each group $k$, $\boldsymbol{\mathcal{B}}_{k}\subseteq\boldsymbol{\mathcal{B} }$ and $\boldsymbol{{\beta}}^{0}\in\boldsymbol{\mathcal{B}}\subseteq \mathbb{R}^{N\times p}$. The parameters are assumed to exhibit a certain grouped pattern, the number of unknown slope parameters in $\{\boldsymbol{{\beta}}_{i}\}$ is of order $O(K)$, instead of $O(N)$, where, $K(<N)$ is typically a fixed constant in empirical applications.

Contrary to c-lasso, we do not impose homogeneity within any of the groups $\boldsymbol{\mathcal{B}}_{k}$, hence the unit-specific $\boldsymbol{{\beta}}_{i}^{0}$ are not penalised down to a specific group-centre $\boldsymbol{{\alpha}}^{0}_{k}$, simply because $\boldsymbol{{\beta}}_{i}^{0}\neq\boldsymbol{{\alpha}}^{0}_{k}$, which is more indicative of a {fused-Lasso} penalty. The component of summation is essential in order to extract information from the $N$ cross-sectional units, resulting in the identification of both $\{ \boldsymbol{{\beta}}_{i}^{0} \}$ and $\{ \boldsymbol{{\alpha}}^{0}_{k} \}$. The grouping, while fixed, is not known a-priori.

We present the necessary assumptions that ensure consistency of the proposed estimator:

assumption\begin{enumerate}[nolistsep] • $\left\{ \boldsymbol{x}_{it}\right\} $ is a $p$-dimensional ergodic sequence of r.v.'s such that $E\left( \boldsymbol{x}_{it}|\mathcal{F} _{t-1}\right) =0$ a.s., $\forall\;i=1,\ldots,N$, $\sup_{i,t}E\left( {x}_{j,it}^{2}|\mathcal{F}_{t-1}\right) =\sigma^{2}>0$ $\,$ $\sup_{i}\sup_{t}E\left( |{x}_{j,it}|^{2+\nu}\right) <\infty$, for some $\nu>0$ and $j=1,\ldots, p$, where $\mathcal{F}_{t-1}=\left\{ {x}_{j, it-1}, \ldots, {x}_{j, i1}\right\}$ is the information set at time $t-1$. • $\boldsymbol{\Sigma}_{\boldsymbol{x}}=\mathbb{E}\left[\boldsymbol{x}_{it} \boldsymbol{x}_{it}^{\prime} \mid \mathcal{F}_{t-1}\right] $, we assume that $ \boldsymbol{\Sigma}_{\boldsymbol{x}}=\operatorname{plim}_{T,N \rightarrow \infty} \frac{1}{NT} \sum_{i=1}^N \sum_{t=1}^T\boldsymbol{x}_{it}\boldsymbol{x}_{it}'$ exists and is positive definite, such that $\Lambda_{\min}(\boldsymbol{\Sigma}_{\boldsymbol{x}})>m, \; m>0 $, where $\widehat{\boldsymbol{\Sigma}}_{\boldsymbol{x}}=\frac{1}{NT} \sum_{i=1}^N \sum_{t=1}^T\boldsymbol{x}_{it}\boldsymbol{x}_{it}'$ is the sample covariance matrix, and $\Lambda_{\min}(\boldsymbol{\Sigma}_{\boldsymbol{x}})$ the smallest eigenvalue of $\boldsymbol{\Sigma}_{\boldsymbol{x}}$. • $\left\{ \varepsilon_{it}\right\} $ is an ergodic sequence of r.v.'s such that $E\left(\varepsilon_{i t}\right)=0, E\left(\varepsilon_{i t}^2\right)=\sigma_{\varepsilon_i}^2$ and $E\left(\varepsilon_{i t}^{2+\nu}\right)<\infty$ for $i=1, \ldots, N ; t=1, \ldots, T$ and for some $\nu>0$. Let $\mathcal{F}_{-\varepsilon_{i t}}$ be the $\sigma$-field of all stochastic elements in the panel data model, apart from $\varepsilon_{i t}$. Then, $E\left(\varepsilon_{i t} \mid \mathcal{F}_{-\varepsilon_{i t}}\right)=0$. • $\{\varepsilon_{it}\}$ is uncorrelated with past, as well as future realisations of $\{\boldsymbol{\boldsymbol{x}_{it}}\}$, such that $E({\varepsilon}_{it}|\boldsymbol{x}_{i1},\ldots,\boldsymbol{x}_{iT})=0$, $\forall\;i=1,\ldots,N,\;t=1,\ldots,T$ \end{enumerate}
remarkIn Assumption (ref).(ref) and (ref).(ref) we assume that $\{\boldsymbol{x}_{it} \}$ and $\{ \varepsilon_{it}\}$ have finite $2+\nu$ moments for some $\nu>0.$ In Assumption (ref).(ref) we ensure positive definiteness of the covariance matrix, guaranteeing $\left\|\frac{1}{NT} \sum_{i=1}^N\sum_{t=1}^T\left[\boldsymbol{x}_{it} \boldsymbol{x}_{it} ^{\prime}-\boldsymbol\Sigma_{\boldsymbol{x}}\right]\right\|_2=o_P(1) .$ Note that in the current framework we do not accommodate panels with cross-sectional dependence. However, the latter can be relaxed at the cost of imposing stronger boundedness conditions on the largest eigenvalue of the variance covariance matrix of $\{\varepsilon_{it}\}$.
assumption\begin{enumerate}[nolistsep] • $E\left(\boldsymbol{\eta}_i\right)=0 \text { and } E\left(\boldsymbol{\eta}_i \boldsymbol{\eta}_i^{\prime}\right)=\boldsymbol{\Sigma}_{\eta \eta, i}$$E(\boldsymbol{\eta}_{i}\boldsymbol{\eta}_{j}')=0$ for ${i\neq j= 1, \ldots, N}$. • $E\left( \boldsymbol{x}_{it} \boldsymbol{\eta }_{i}\right)=0 $. \end{enumerate}
remarkSubstituting $\boldsymbol\beta_{i}^0$ in (ref), we have that $y_{it}={\boldsymbol{\alpha}_k^0}'\boldsymbol{ x}_{it }+\boldsymbol{ x}_{it }\boldsymbol{ \eta}'_{i }+\varepsilon_{it}$. Under Assumption (ref) we obtain a consistent estimation of ${\boldsymbol{\alpha}_k^0}$.
remarkIn Assumption (ref).(ref), weak exogeneity can be considered in our framework by adding lagged dependent variables, for example considering dynamic panel models. Note that in the case of dynamic panel specification, we may maintain the Assumption (ref).(ref), however we can no longer assume that $E\left(\boldsymbol{\eta}_i y_{i, t-1}\right)=\boldsymbol{0}$, where $y_{i, t-1}$ can be defined through continuous substitutions, such that $ y_{i, t-1}=\sum_{j=0}^{\infty}\left(\eta_{i 1}\right)^j \boldsymbol{x}_{i, t-j-1}^{\prime}\left(\boldsymbol{\alpha}_k^0+\boldsymbol\eta_{i 2}\right)+\sum_{j=0}^{\infty}\left(\eta_{i 1}\right)^j u_{i, t-j-1}, $ and $\boldsymbol{\eta}_i=\left(\eta_{i 1}, \boldsymbol{\eta}_{i 2}^{\prime}\right)^{\prime}$. It follows that $E\left(\boldsymbol{\eta}_i y_{i, t-1}\right) \neq \boldsymbol{0}$. The violation of the independence between the covariates and the individual effects, $\boldsymbol{\eta}_i$, implies that $y_{i t}|y_{i, t-1}$, and $\boldsymbol{x}_{i t}$ will result to inconsistent estimates of $\boldsymbol{\alpha}^0_k$, even for sufficiently large $T$ and $N$.
assumption\begin{enumerate}[nolistsep] • The number of clusters, $K$, is fixed and $N_{k}/N\rightarrow\tau_{k} \in(0,1)$ for each $k=1,\ldots,K$ as $N\rightarrow\infty$, where $N_{k}= |\boldsymbol{G}^0_k|$ in cluster $k$. • $\lambda\asymp\log{T}^{-\alpha}$, for some $\alpha\in(0,1/2]$ as $(N, T) \rightarrow\infty$. • $T \lambda^{2} /(\ln T)^{6+2 \nu} \rightarrow\infty$ and $\lambda(\ln T)^{\nu} \rightarrow0$ for some $\nu>0$ as $(N, T) \rightarrow\infty $. • $N^{- 1 / 2} T^{-1}(\ln T)^{9} \rightarrow0$ and $N^{2} T^{1-q / 2} \rightarrow c \in[0, \infty)$ as $(N, T) \rightarrow\infty$. \end{enumerate}
remarkAssumption (ref).(ref) imposes conditions on $\lambda$, all of which hold if $\lambda\propto T^{-\delta} \quad$ for any $\delta\in(0,1 / 2)$. Assumption (ref).(ref) is needed to ensure some higher-order terms vanish asymptotically.

We further introduce an additional assumption on the second moment of $\boldsymbol{\eta}_{i}$, which serves as a group-uniqueness condition towards cluster identifiability.c-lasso

assumption\begin{enumerate}[nolistsep] • $\int\|\boldsymbol{\eta}_{i} \|_{2}^{2} P(d \boldsymbol{\eta} _{i})<\infty$ and the probability measure $P$ has a continuous density $f$ on $\mathbb{R}^{p}$. • The group centres are pairwise different, such that $\boldsymbol{\alpha }^{0}_{k}\neq\boldsymbol{\alpha}^{0}_{j}, $ and groups are pairwise disjoint such that $\boldsymbol{G}_{j}^{0} \cap\boldsymbol{G}_{k}^{0} = \varnothing$, for $j\neq k , j,k=1,\ldots, K$. \end{enumerate}
remarkAssumption (ref).(ref) is an identifiability assumption and imposes finiteness conditions on $\boldsymbol{ {\eta}}_{i}$ for each $i=1,\ldots, N$.

With the following theorems, Theorem (ref) and (ref), we provide asymptotic guarantees for the proposed estimator, the c-lasso estimator and the feasible k-means estimator. The following theorems establish consistency of the adjusted penalised profile likelihood estimates of the slope parameter ${\boldsymbol{\beta}}$ and the and group-specific parameter ${\boldsymbol{\alpha}}$.

We consider two types of penalised estimators, the first is c-lasso. We obtain the set of estimated parameters $(\boldsymbol{\ddot{ \beta} }, \boldsymbol{\ddot{ \alpha}})$, by minimising the objective in (ref), $\arg\min_{\boldsymbol{\beta}\in\mathbb{R}^{N\times p}, \boldsymbol{\alpha}\in\mathbb{R}^{p\times K}} Q_{iNT,\lambda}\left( \boldsymbol{{\beta}},\boldsymbol{{\alpha}}\right) $ for a choice of tuning parameter, $\lambda$. In the following theorem we establish asymptotic consistency of c-lasso under heterogeneous groups.

theoremConsider Assumptions (ref) and (ref), the linear model in (ref) and the minimisation of (ref). Then, for a tuning parameter $\lambda=o(1)$, we write \begin{align} \left\| {\boldsymbol{\ddot{\beta}}}_{i}- \boldsymbol{\beta}^{0}_{i}\right\| _{2} & =O_{P}(T^{-1/2}+\lambda),\\ \Vert\boldsymbol{\ddot{\alpha}} _{k}-\boldsymbol{\alpha} _{k}^{0} \Vert_{2} & =O_{P}\left( T^{-1/2} \right) , \\ (\boldsymbol{\ddot{\alpha}}_{1}, \ldots, \boldsymbol{\ddot{\alpha} }_{K}) - ( \boldsymbol{\alpha}_{1}^{0}, \ldots, \boldsymbol{\alpha}_{K}^{0}) & = O_{P}(T^{-1/2}) \end{align} for $l\neq k, \; l=1,\ldots, K$, where $({\boldsymbol{\ddot{\beta}}}, \boldsymbol{\ddot{\alpha}})$ minimise the c-lasso objective, in (ref).

Similarly to c-lasso, one can use alternative estimators for $\boldsymbol{{\alpha}}$. An estimator to be considered is the k-means. Let $Q_{1,NT}={(NT)}^{-1}\sum_{i=1}^{N}\sum_{i=1}^{T}\frac{1}{2}\left( y_{it}-\boldsymbol{\beta}_{i}^{\prime}\boldsymbol{x}_{it}\right) ^{2}$. The second methodology involves the minimisation of:

align[align omitted — 221 chars of source]

where $f^{KM}(\boldsymbol{\beta})=\arg\min_{\boldsymbol{\alpha}}\sum _{i}\left\Vert \boldsymbol{\beta}_{i}-\boldsymbol{\alpha}_{k}\right\Vert _{2}^{2}$ is the k-means solution. Then ${\boldsymbol{\widehat{\beta}} }=\arg\min_{\boldsymbol{\beta}\in\mathbb{R}^{N\times p}}Q_{2,iNT,\lambda }\left( \boldsymbol{{\beta}}\right) $ and $\boldsymbol{\widehat{\alpha} }=f^{KM}(\boldsymbol{\widehat{\beta}})$. Theorem (ref) has similar asymptotic properties to the c-lasso, to Theorem (ref), but with slightly slower convergence rates.

theoremConsider Assumptions (ref)--(ref), the linear model in (ref) and the minimisation in (ref). Then, for a tuning parameter $\lambda=o(1)$, we write \begin{align} \left\| {\boldsymbol{\widehat{\beta}}}_{i} - \boldsymbol{\beta}^{0} _{i}\right\| _{2} & =O_{P}(T^{-1/2}+\lambda),\\ \Vert\boldsymbol{\widehat{\alpha}} _{k}-\boldsymbol{\alpha} _{k}^{0}\Vert_{2} & =O_{P}\left( T^{-1/2} +N^{-1/2}\right) ,\\ (\boldsymbol{\widehat{\alpha}}_{1}, \ldots, \boldsymbol{\widehat{\alpha}}_{K}) - ( \boldsymbol{\alpha}_{1}^{0}, \ldots, \boldsymbol{\alpha}_{K}^{0}) & =O_{P}\left( T^{-1/2}+N^{-1/2}\right) , \end{align} where by Assumption (ref) (ref), $ \ln{T}^{9}/T = o(N^{1/2} ) $ and $(\boldsymbol{\widehat{\beta}}, \widehat{\boldsymbol{\alpha}})$ are centres obtained through the k-means Lasso iterative procedure, outlined in (ref).
remarkTheorem (ref) establishes point-wise convergence of the estimated slope parameter $\boldsymbol{\widehat{\beta}} _{i}$. The asymptotic rate which guarantees consistency for $\boldsymbol{\widehat{\alpha}}$ can be $O_{P}(T^{-1/2})$ which is the rate using c-lasso in the case where $N$ is of larger order of magnitude than $T$, i.e. $T=o(N)$. The latter, while sub-optimal, is of interest as clustering is a simple method in terms of asymptotic results and can deal with empirical datasets with a large cross-section

An alternative estimator for the grouping centres

In this section we consider a feasible way to identify the cluster structure. We devise a simple method, where we obtain the estimates of the unit-specific parameters with ols \[ \boldsymbol{\widetilde{\beta}}_{i}=\left( \boldsymbol{x}_{i}^{\prime }\boldsymbol{x}_{i}\right) ^{-1}\boldsymbol{x}_{i}^{\prime}\boldsymbol{y} _{i}. \] and minimise ${Q}_{N}(\boldsymbol{G}_{k},\boldsymbol{\alpha} )=\frac{1}{N}\sum_{i=1}^{N}\left( \boldsymbol{\beta}_{i}-\boldsymbol{\alpha }_{k}\right) ^{2} $ with respect to $\boldsymbol{\alpha}$, for $i=1,\ldots,N,\;k=1,\ldots,K$.

It becomes clear that $\arg\min_{\boldsymbol{\alpha}} {Q}_{N}(G_{k}, \boldsymbol{\alpha}) $ imposes an infeasible problem as there has to be knowledge of $\boldsymbol{\beta}_{i}$. Therefore, we consider a unit-by-unit estimation to obtain $\boldsymbol{\widehat{\beta}}_{i}$, using (ref). Substituting (ref) in (ref), we have that

align[align omitted — 329 chars of source]

where $\boldsymbol{x}_{i}=(x_{i1},\ldots,x_{iT})^{\prime}$, $\boldsymbol{\varepsilon}_{i}=(\varepsilon_{i1},\ldots,\varepsilon _{iT})^{\prime}$ and $\boldsymbol{\psi}_{i}=\left( \boldsymbol{x}_{i} ^{\prime}\boldsymbol{x}_{i}\right) ^{-1}\boldsymbol{x}_{i}^{\prime }\boldsymbol{\varepsilon}_{i}$. Then, a feasible k-means loss can be constructed, $\widehat{Q}_{N}(G_{k}, \boldsymbol{\alpha} )$, where the unit-by-unit estimates can be clustered instead of $\boldsymbol{\beta}_{i}$,

align[align omitted — 206 chars of source]

Note that, minimising (ref), the grouping structure relies on the estimation of ${{\boldsymbol{\beta}}}_{i}^{\ast}$. We showcase, with the following example, that ${Q}_{N}(\cdot)$ can be minimised with respect to some grouping centre $\boldsymbol{\alpha}= \left(\boldsymbol{\alpha}_{j, 1}, \ldots, \boldsymbol{\alpha}_{j, K} \right)'$, $j=1, \ldots, p$.

In the following theorem we show consistency of the estimated centres obtained by clustering $\boldsymbol{ \widehat{\beta}}$, in (ref), through k-means, we call this method feasible k-means

theoremConsider Assumptions (ref)--(ref), and (ref) then the following holds \begin{align} \Vert\boldsymbol{\widetilde{\alpha}} _{k} -\boldsymbol{\alpha} _{k}^{0}\Vert_{2} & =O_{p}(N^{-1/2}), \end{align} where $\boldsymbol{\widetilde{\alpha}} _{k} $ is the minimiser of (ref).

Naturally, Theorem (ref) relies heavily on Assumption (ref), where $( \boldsymbol{ {\alpha}}^{0}_{1}, \ldots, \boldsymbol{ {\alpha}}^{0}_{K})^{\prime}$ is a unique minimiser of the population $Q(\cdot)_{N}$, therefore, $(\boldsymbol{\widetilde{\alpha}} _{1}, \ldots, \boldsymbol{\widetilde{\alpha}} _{K})^{\prime}$ minimises $\widehat{Q}_{N} (G_{k}, \boldsymbol{\alpha} )$.

Classification under group heterogeneity

In this section we discuss the classification consistency of the different estimating techniques.

definitionLet estimators for the true regression coefficient of unit $i$ be given by $\boldsymbol{\widehat{\beta}}_{i}$, $i=1,\ldots,N$, and for the set of cluster centres by $\boldsymbol{\widehat{\alpha}}_{k}$, $k=1,\ldots,K.$ Consider $\boldsymbol{G}_{k}^{0}\subset\{1,2,\ldots,N\}$, $\cup_{k=1}^{K} \boldsymbol{G}_{k}^{0}=\{1,2,\ldots,N\}$, and $\boldsymbol{G}_{k}^{0} \cap\boldsymbol{G}_{j}^{0}=\varnothing$ for any $j\neq k$, $j=1,\ldots,K$. Then we define $\boldsymbol{\widehat{G}}_{k}\subseteq\{1,2,\ldots,N\}$ as the set of units whose associated estimated centre, $\boldsymbol{\widehat{\alpha }}_{k}$ satisfies $p\lim\boldsymbol{\widehat{\alpha}}_{k}=\boldsymbol{\alpha }_{k}^{0}$, where a unit $i$ is associated with centre $\boldsymbol{\widehat{\alpha}}_{k}$ if $\arg\min_{j}\left\Vert \boldsymbol{\widehat{\beta}}_{i}-\boldsymbol{\widehat{\alpha}}_{j}\right\Vert =\boldsymbol{\widehat{\alpha}}_{k}$.

Due to the presence of $\boldsymbol{\eta}_{i}$ in the generating process in (ref), the probabilities of events of the form $\mathcal{I} _{1,i}=\{i\in\boldsymbol{\widehat{G}}_{k}|i\notin\boldsymbol{G}_{k}^{0}\}$, and $\mathcal{I}_{2i}=\{i\notin\boldsymbol{\widehat{G}}_{k}|i\in \boldsymbol{G}_{k}^{0}\}$, respectively, become difficult to control. The literature usually defines classification consistency as the property where $\Pr(\cup_{i}\mathcal{I}_{1,i})\rightarrow0$ and $\Pr(\cup_{i}\mathcal{I} _{2,i})\rightarrow0.$ In fact in our case it is possible to show the converse. In particular, we have the following Lemma

lemmaConsider Assumptions (ref)--(ref), the model in (ref) and the minimisation in (ref). Then, \begin{align} \lim_{N,T\rightarrow\infty}\Pr(\cup_{i}\mathcal{I}_{1,i}) & >0,\;and\\ \lim_{N,T\rightarrow\infty}\Pr(\cup_{i}\mathcal{I}_{2,i}) & >0. \end{align}
remarkIn Lemma (ref) we show that classification consistency is not guaranteed due to the underlying generating process of $\boldsymbol{ {\beta} }_{i}$. Further we show that the probability of Type I and II errors is larger than 0 as opposed to approaching 0 as $(N,T)\to\infty$. It is further important to note that similarly to Lemma (ref), this result can be shown using alternative estimators such as the c-lasso and the feasible k-means.

Heterogeneity Inference

The possibility of the presence of both a group structure as well as within group heterogeneity necessitates an exploration of data dependent methods for heterogeneity inference. Two issues arise. First determining whether there is heterogeneity within clusters and the second is, under heterogeneity, how to determine the number of clusters. It is important to have the right sequencing in this investigation. We propose to start with a test of the following null hypothesis

align[align omitted — 190 chars of source]

Under $H_{0}$, methods such as c-lasso can be used to determine the number of clusters and obtain consistent estimates of their centres. These methods can then be used as discussed in Section (ref) below. If the null is not rejected one reverts to existing methods and find ways to determine the number of clusters under heterogeneity. Existing methods associated to k-means cluster analysis can be used and are discussed in Section (ref). In the following discussion we keep matters simple by not allowing for more refined features such as weak heterogeneity (defined as heterogeneity that is asymptotically negligible) or heterogeneity in a proportion of clusters (although we discuss cluster specific tests below). All such extensions are possible but beyond the scope of the current paper.

Testing for the existence of heterogeneity

The main interest in this section is to verify the existence of different types of heterogeneity, i.e. cross-sectional, as well as within group. We follow the two-step method described in Section (ref) to extract feasibly both the group and unit-specific parameter and show that using the unit-by-unit estimate as a preliminary estimator of $\boldsymbol{\beta}_{i}$ can lead to identifiable groups via k-means, while the dimensionality of the cross-sectional parameters is allowed to diverge.

We consider the model in (ref) and simplify our analysis so that $\boldsymbol{\beta} = ({\beta}_{1}, \ldots, {\beta}_{N})^{\prime}$, and $\boldsymbol{\alpha} = ({\alpha}_{1}, \ldots,{\alpha} _{k})$ are $N\times 1$ and $K\times 1$ vectors respectively. To investigate the existence of heterogeneity in the cross section, and, consequently, within a group ${G}_k$ we construct test statistics under the null hypothesis in (ref). Our methodology proceeds in distinct steps; We first obtain the estimates of $\widetilde{\beta}_{i}$ through ols on $\boldsymbol{{y}}_{i}|\boldsymbol{x}_{i}$ and uncover the grouping structure using k-means on ${\widetilde{\beta}}_{i}$ to obtain ${\widetilde{\alpha}}_k,$ the minimiser of (ref). We then obtain $\boldsymbol{\widehat{\varepsilon}}_{i}=\boldsymbol{{y}}_{i} -\boldsymbol{x}_{i}{\widetilde{\alpha}}_k $ where $i\in \widehat G_{k}$. Note that here it is sensible to assume that $\widehat G_{k}\to_P G_{k}^0$. The latter is clear under $H_0$, contrary, under $H_1$, where the results of Lemma (ref) verify an alternative statement. Then under the Assumption that $E({\varepsilon}_{it}|{x}_{it})=0$, and $E({\eta}_{it}|{x}_{it})=0$, $T^{-1/2} \sum_{i} T^{-1}\boldsymbol{x}_{i}^{\prime}{\varepsilon}_{it} \sim N(0,\boldsymbol{\sigma}_{x\varepsilon})$, where $\boldsymbol{\sigma }_{x\varepsilon} = N^{-1} \sum_{i} T^{-1}\boldsymbol{x}_{i}^{\prime }{\varepsilon}_{it} {\varepsilon}_{it} ^{\prime}\boldsymbol{x}_{i}$. Recall from (ref) that

align[align omitted — 192 chars of source]

where $\boldsymbol{{\varepsilon}}_{i}= ( \boldsymbol{{\varepsilon}}_{i1}, \ldots, \boldsymbol{{\varepsilon}}_{iT} )^{\prime}$. A feasible solution to (ref) would be to replace $\boldsymbol{{\varepsilon}}_{i}$ with $\boldsymbol{\widehat{\varepsilon}}_{i} .$ One can then construct the cross-sectional heterogeneity test statistic,

align[align omitted — 370 chars of source]

where $\widehat{\sigma}_{i} =\frac{1}{T-1}\sum_{i} \boldsymbol{\widehat{\varepsilon}}_{i} \boldsymbol{\widehat{\varepsilon}}_{i} ^{\prime}$ a consistent estimate of $\sigma_i=E\frac{1}{T}\boldsymbol{{\varepsilon}}_{i} \boldsymbol{{\varepsilon}}_{i} ^{\prime}$.

Further, a poolability-type test would rely on the $R^{2}$ statistic, where one uses the residuals resulting from unit-by-unit ols, see e.g. (ref), to asses the fit against a pooled ols regression, e.g. $ \widehat{\beta}_{\textsc{ols}} = ( \sum_{i=1}^{N} \sum_{t=1}^{T} \boldsymbol{x} _{it} ^{\prime}\boldsymbol{x}_{it} ) ^{-1} ( \sum_{i=1}^{N} \sum_{t=1}^{T} \boldsymbol{x}_{it} ^{\prime}y_{it} ), $ Further, under standard assumptions of independence, see e.g Assumption (ref), the null hypothesis in (ref) will be rejected once the unit-by-unit ols reports larger $R^{2}$ than the corresponding $R^{2}$ of the pooled ols, which indicates that treating the data as heterogeneous increases the overall panel fit.

To that end we construct the $R^{2}$ statistic, resulting from the auxiliary regressions of $\boldsymbol{\widehat{\varepsilon}}_{{i}}|\boldsymbol{x} _{i}$, such that

align[align omitted — 373 chars of source]

where $ \boldsymbol{\widehat{v}}_{i}\boldsymbol{(}\beta\boldsymbol{)=\widehat{\varepsilon}}_{i}-\widehat{\gamma}\left( \beta\right) \boldsymbol{x}_{i} $ and $ \widehat{\gamma}\left( \beta\right) ={\boldsymbol{\widehat{\varepsilon}}_{i} ^{\prime}\boldsymbol{x}_{i}}({\boldsymbol{x}_{i}^{\prime }\boldsymbol{x}_{i}})^{-1}, $ lastly, ${\sigma_{i, R}}$ can be defined similarly to ${\sigma_{i}}$. Under $H_0$, both $\widehat{R}^{2}_{N} , \; \widehat{S}^{2}_{N}$ are asymptotically $\chi_1^2$ distributed, while consistent under the alternative. The latter is the key finding of the following theorem.

Then, we have the following theorem giving the asymptotic distribution of $S_{N}$ and $R_{N}$, defined in (ref) and (ref).

theoremConsider Assumptions (ref), (ref) and Assumption (ref). Then, under the null \begin{align} \widehat{S}^{2}_{N} \overset{d}{\to} \chi_{1}^{2},\quad\widehat{R}^{2}_{N} \overset{d}{\to} \chi_{1}^{2}, \end{align} where $\widehat{S}^{2}_{N}$, and $\widehat{R}^{2}_{N} $ are defined in (ref) and (ref) respectively.
remarkWhile testing for cross-sectional heterogeneity is an important first step towards uncovering a grouping structure, it is also of interest to investigate the existence of heterogeneity within a group $k$. More specifically, in (ref)--(ref) we can substitute $\boldsymbol{\widehat{\varepsilon}}_{{i,{k}}}$ instead of $\boldsymbol{\widehat{\varepsilon} }_{{i}}$ such that \begin{align} \widehat{S}_{\widehat{K}}^{2} & = \left( \frac{1}{\sqrt{\widehat N_{k}}} \sum_{i, \; i\in \widehat G_k} \frac{ \widehat{t}_{k}^2({\widetilde{\alpha}}_k) -1 }{\widehat{\sigma}_{k} }\right) ^{2} ,\quad \widehat{R}^{2}_{\widehat{K}} =\frac{1}{\sqrt{\widehat N}_{k}} \sum_{i, \; i\in \widehat G_k} \frac{ (T\widehat{r}_{k}^{2}-1 ) }{\widehat{\sigma}^{2} _{\widehat{r}_{k}, }} , \; $ i=1, \ldots, N$, \end{align} where $\widehat{t}_{k}({\widetilde{\alpha}}_k) = {\boldsymbol{\widehat{\varepsilon}}_{i} ^{\prime}\boldsymbol{x}_{i}}{\widehat{\sigma}_{k}({\widetilde{\alpha}})\left( \boldsymbol{x} _{i}^{\prime}\boldsymbol{x}_{i}\right) ^{-1/2}},$ where ${\widetilde{\alpha}}_{k}$ is the minimiser of (ref). Further, $\widehat{\sigma}_{k}=\frac{1}{T-1}\sum_{i\in G_{k}} \boldsymbol{\widehat{\varepsilon}}_{{i} }\boldsymbol{\widehat{\varepsilon}}_{i}^{\prime}, $ and $\widehat{r}^{2}_{k} = \sum_{ i\in G_{k}}({\widetilde{\alpha}}_{k}\boldsymbol{x}_{i} ^{\prime}\boldsymbol{x}_{i}{\widetilde{\alpha}}_{k})^{\prime}({\boldsymbol{\widehat{\varepsilon} }_{{i}} \boldsymbol{\widehat{\varepsilon}}_{{i}}^{\prime} })^{-1} $ and $\widehat{\sigma}_{\widehat{r},k} $ can be defined similarly to $\widehat{\sigma}_{k} $ for $1\leq i\leq N.$ Note that using Lemma (ref) under $H_0 $, $\widehat{G}_k\to_P G_k^{0}$ and $\widehat{N}_k\to_P N_k$ and the test-statitsics follow $\chi_1^2$. The statements in (ref) can be shown following a similar analysis to show Theorem (ref).
remarkUnder $H_1$ the test statistics (ref), (ref), as well as their group-specific counterparts, (ref) are consistent. The statement in (ref) can accommodate a reasonably larger dimensionality in terms of the number of covariates, $p$, as long as the clustering remains the same across covariates.

Determination of the number of groups under heterogeneity

Allowing for different types of heterogeneity, the identifiability of groups becomes a difficult task. First, because in real datasets there is no a-priory knowledge of the necessity of grouping, the number of groups formed or the grouping structure. Second, in their majority grouping methods require spherical groups, this way allowing only for a degree of heterogeneity, or are density based techniques, which in turn require the knowledge of the empirical distribution for each group. Further, versatile clustering techniques, such as k-means, appear effective, but in cases of no-grouping at all appear inappropriate as it forces grouping. We address that issue by estimating the number of groups using techniques known in the statistical literature. The most commonly used is the Gap Statistic, see e.g. gap, and can be constructed as follows,

align[align omitted — 249 chars of source]

for ${i\neq j }= 1, \ldots, N$ and $k=1, \ldots, K$, where $\boldsymbol{d}$ is the square euclidean distance from a centre in group $k$, $W_{k}$ is the pooled within cluster sum of squares around the cluster means, $E\left[ \log(W_{k}) \right] $ denotes the expectation under a sample of size $N$ is drawn from a reference distribution and $N_{k}= |G_{k}|.$ Under Assumption (ref) and well separated groups the number of clusters ${K}^{\ast }$ maximises $\text{Gap}_{N}(k), $ defined in (ref), such that

align[align omitted — 115 chars of source]

We consider a feasible alternative to maximise (ref) with respect to $\hat{K}$, such that for ${i\neq j },$

align[align omitted — 187 chars of source]

where $\boldsymbol{\widehat{\beta}}_i$ an estimator of $\boldsymbol{\beta}_i$. Note that $\boldsymbol{\widehat{\beta}}_i$ can be a mazimiser of either $Q_{iNT,\lambda}\left( \boldsymbol{{\beta}}\right)$, $Q_{2,iNT,\lambda}\left( \boldsymbol{{\beta}}\right) $, or it can maximise the unit-wise least squares objective, such that $\arg\min_{\boldsymbol{\beta}_i\in \mathbb{R}^p}Q_{1,NT}, \; Q_{1,NT}={(NT)}^{-1}\sum_{i=1}^{N}\sum_{i=1}^{T}\frac{1}{2}\left( y_{it}-\boldsymbol{\beta}_{i}^{\prime}\boldsymbol{x}_{it}\right) ^{2}$, denoted (ref), (ref) and (ref) respectively. Consider Assumptions (ref), (ref), and (ref), then $\log\widehat{W}_{k} \to _P \log W_{k} $ where ${W}_k$ is defined in (ref) and $\widehat{W}_k$ in (ref).

The gap statistic is an empirical tool to identify feasibly the number of clusters taking into account the within-cluster dispersion, given various clustering algorithms. In Section (ref) we explore the small sample properties of the method using simulations, and report the frequency of selecting the correct numbers of groups, where it is known. Further, as the findings of gap suggest, the gap statistic outperforms other methods in the literature, even in the case of no grouping, i.e. $K=1$.

Simulation study

In this section we explore the finite sample properties of the different estimators explored in Section (ref) and (ref) and make comparisons. A detailed description of the methods considered can be found in (ref). Further, this section is divided in two sub-sections: In Section (ref) we explore the small sample properties of the different methods in terms of mean squared error and miss-classification rate, and in Section (ref) we focus on the performance of the tests proposed in Section (ref).

Parameter estimation (misclassification) error

To evaluate the finite-sample performance of the classification and estimation procedure, we consider 1000 simulations of model (ref), where the number of groups assumed are $K=2$, $\boldsymbol{G}_{k}$ is a membership vector for group $k=1,\ldots, K$, the simulated true centres are $\boldsymbol{\alpha} =(0.5, 2)^{\prime}$ and $\eta_{i}\sim N(0, 0.2)$. Further, $x_{it}\sim_{\text{i.i.d}} N(0,1)$, $\varepsilon_{it}\sim_{\text{i.i.d}} N(0,1)$ and $(\beta_{1},\ldots, \beta_{N/2})^{\prime}\in\boldsymbol{G}_{1}, \; (\beta_{N/2+1}, \ldots, \beta_{N})^{\prime}\in\boldsymbol{G}_{2},$ where $N=[20,50,100,200], \; T=[50,100,200,500]$.

We consider various different methods and evaluate the consistency of the grouped parameters $\beta_{i}$, and individual group centres, as well as the quality of grouping under homogeneous and heterogeneous parameters. Namely we consider the c-lasso of su2016identifying reported as SSP, the c-lasso estimator using k-means to obtain the group centres, reported as Km and the c-lasso without imposing within class homogeneity, reported as H-SSP. The penalty parameter for the penalisation throughout methods was carried out via cross validation scheme, from a parameter grid, $\lambda \in[0.125, 0.25, 0.5, 1, 2]T^{-1/3} $. We carried out a k-fold cross validation scheme without re-shuffling, where $k=10$.

To asses the consistency of the individual parameters we report the average mean squared error (MSE) across units under both the null and the alternative, e.g. (ref). We report these findings in Table (ref). In Table (ref) we report the average MSE of the centres of the first group to save space, e.g. $1000^{-1} \sum_{i=1}^{1000} (\alpha_{1,i}- \widehat{\alpha}_{1,i})$, where the estimation of the individual $K$ centres is carried out using four different methods, descried below.

To quantify the quality of grouping we use a dissimilarity measure, Rand Index, e.g. rand1971objective. The Rand index, ranges between 0 and 1, with 0 indicating that the two data groupings do not agree on any pair of points and 1 indicating that the data groupings are exactly identical to some permutation. More specifically, given the following partitions $\boldsymbol{G}= \{\boldsymbol{G}_{1}, \ldots, \boldsymbol{G}_{r} \}$, of $\boldsymbol{\beta} $ into $r$ subsets and $\boldsymbol{\widehat{G}}=\{ \boldsymbol{\widehat{G}}_{1}, \ldots, \boldsymbol{\widehat{G}}_{s} \}$, a partition of $\boldsymbol{\beta} $ into $s$ subsets, where $1<r, s\leq N$, we write the following measure \[ RI = \frac{a+b}{{\binom{N }{2 }}}\equiv\frac{TP+TN}{TP+TN+FP+FN}, \] where $N$ is the number of units to be clustered, TP is the number of true positives, TN is the number of true negatives, FP is the number of false positives, and FN is the number of false negatives. Further, $a= |\boldsymbol{B}^{*}|, \; \boldsymbol{B}^{*} = \{ (\beta_{i},\beta_{j}) | \beta_{i},\beta_{j} \in\boldsymbol{G}, \; \beta_{i},\beta_{j} \in \boldsymbol{\widehat{G}}\}$ and $b = |\boldsymbol{B}^{*}|, \; \boldsymbol{B} ^{*} = \{ (\beta_{i}, \beta_{j} ) | \beta_{i}\in\boldsymbol{I_{1}},\beta _{j}\in\boldsymbol{G}_{2}, \;\beta_{i}\in\boldsymbol{\widehat{G}}_{1} ,\beta_{j}\in\boldsymbol{\widehat{G}}_{2} \}$, where the former declares the number of pairs of elements in $\{\beta_{1}, \ldots, \beta_{N}\}$ that are in the same subset in $\boldsymbol{G}$ and in the same subset in $\boldsymbol{\widehat{G}}$, and the latter declares the number of pairs of elements in $\{\beta_{1}, \ldots, \beta_{N}\}$ that are in different subsets in $\boldsymbol{G}$ and in different subsets in $\boldsymbol{\widehat{G}}$. We report these findings in Table (ref) both under homogenous and heterogeneous groups.

Further, since heterogeneity is allowed within a group, we consider additional metrics to ensure that the grouping is not forced but indeed latent structures can be found in the cross section. We then estimate the number of groups that can be found using the unit-by-unit (UE) estimate, ${\beta}^{\ast}_i$ obtained in (ref), reporting the frequency of selecting the true number of groups, $K_{0}=2$ throughout 1000 replications, through the gap statistic, e.g. gap. We report these findings in Table (ref) both under homogenous and heterogeneous coefficients.

table[table omitted — 1,939 chars of source]
table[table omitted — 2,166 chars of source]

As it is expected, in Table (ref), the MSE decreases as the sample size increases, while SSP reports a higher MSE compared to all other methods as both $N,T$ increase in the cases of heterogeneity (see right panel of Table (ref)), while it also decreases with an increase of $T$. The latter is due to the fact that SSP does not allow within group heterogeneity , forcing the cross-sectional parameters to be penalised down to their corresponding group centre, while heterogeneity persists within a group, as the finding of Table (ref) indicate. $Km$ reports the smallest MSE across methods in cases of heterogeneity, In the cases of homogeneity (left panel of Table (ref)), SSP is the best performing method, as it is expected, while all methods report smaller MSEs as $T$ increases.

In Table (ref) all methods reported appear equivalent, while as $T$ increases the error approaches zero. In Table (ref) we report the average RI (see (ref)) throughout 1000 replications. As it is expected, under the null (see left panel of Table (ref)), all methods perform are equivalently, i.e. most methods report average accuracy around or equal 100%. In the case both when generated under the alternative hypothesis, i.e. heterogeneity across units, SSP reports the smallest accuracy throughout samples, but as the sample size $T$ increase, accuracy increases as well. The latter is a common finding across methods.

Table (ref) reports the average frequency of selecting the correct number of groups via the gap statistic. Table (ref) compliments the results of Table (ref), verifying that (1) clustering the cross sectional units is necessary, and (2) while clustering is necessary, through Table (ref) we verify that accuracy of grouping is on average high throughout samples. Finding (1) can be observed both generating under the null and the alternative as the average frequency of selecting two groups is close to 1 and increases with the sample size $T$.

table[table omitted — 1,709 chars of source]
table[table omitted — 779 chars of source]

Testing for the existence of heterogeneity

We simulate the model in (ref) and take the one-way within transformation, see (ref), with $x_{it}\sim_{\text{i.i.d}} N(0,1)$, $\varepsilon_{it}\sim_{\text{i.i.d}} N(0,1)$, $\boldsymbol{\alpha} =(0.5, 2)^{\prime}$, $\eta_{i}\sim N(0,0.2). $ We generate under the null hypothesis of (ref), to report the size of the test in (ref) on the right panel of Table (ref) and further power of the test can be reported by generating under the alternative, see the left panel of Table (ref). The reported quantities of Table (ref) can be similarly described for the test in (ref) where we assess the goodness of fit under clustered effects. \newline In Table (ref) we report the average frequency of selecting $K=2$ groups when the true number of groups is $K_{0}=2$ throughout 1000 replications of model (ref), while in Table (ref) we report the average frequency of selecting $K=1$ groups when the true number of groups is $K_{0}=1$. Further we report four different measures of identifying the number of latent groups in a vector of estimates (ref), the maximum number of clusters considered is $K_{\max}=5$. Namely, we consider the gap statistic, e.g. gap, reported as 'gap', the Silhouette statistic e.g. sil, reported as 'Sil', the Calinski-Harabasz statistic, e.g. ch reported as 'CH' and the Davies-Bouldin, e.g. db reported as 'DB'. \newline We are interested in investigating the existence of heterogeneity and for that matter the existence of groups in two separate cases; First, we investigate the existence of heterogeneity in the cross-section, using the unit-by-unit ols estimator in (ref) (reported as CS in Tables (ref)--(ref)) and second we investigate the existence of heterogeneity after grouping using the feasible k-means estimator (reported as WG in Tables (ref)--(ref)) minimising (ref). All the tests are carried out at $5\%$ level of significance. In Table (ref) there is clear indication that asymptotically, as $T$ increases, the test in (ref) rejects the null for both types of heterogeneity frequently, showing that indeed heterogeneity still persists both before and more importantly after grouping. Similarly to Table (ref), in Table (ref) the goodness of fit both before and after grouping is appears better as $T$ increases, showing that grouping is necessary even when heterogeneity persists within each group. The latter is an interesting finding coupled with the findings of Table (ref) --(ref), because there is an implication that in the cases where the number of groups $K=2$, asymptotically the gap measure identifies the correct number of groups 100% o the times, while when there is no grouping at all, the rate at which the gap statistic selects the correct number of groups is 100% throughout $N,\; T$. Therefore, there is clear indication of grouping structure in the cross-section, and the findings of Tables (ref)--(ref) show that heterogeneity clearly exists within the different groups as well.

table[table omitted — 1,954 chars of source]
table[table omitted — 1,634 chars of source]
table[table omitted — 2,293 chars of source]
table[table omitted — 2,117 chars of source]

Empirical study

In this section we discuss the results of the empirical analysis, and provide a detailed discussion of the datasets we used, in the Appendix.

In Table (ref) we report the statistics in (ref) and (ref) respectively, with p-values and report statistical significance at different levels using $(*)$. Note that ${***}$ signifies 1% significant, ${**} $ signifies 5% significant, and ${*} $ signifies 10% significant. Note that the datasets used in the empirical analysis have different number of variables, hence the critical value for the $\chi_{d}^{2}$ with which the statistics (ref) and (ref) are compared, changes, where $d$ are the degrees of freedom. Namely, for $d=1: c=3.841$, $d=2: c=5.991$, $d=3: c=7.815$, $d=4: c=9.488$.

Across datasets we use the feasible k-means estimator based on the unit-wise estimation in (ref) and estimate the number of clusters using the gap statistic, see e.g. gap. The optimal number of clusters is selected from the range ${K} \in\{ 1, \ldots, 20 \}$ to maximise the criterion in (ref). Further we obtain the solution of (ref) based on the optimal number of clusters $\widehat{K}$ and use (ref) to evaluate whether there is remainder heterogeneity collectively across groups and asses the fit of the proposed clustering method using (ref).

In our empirical application we use eleven datasets, in which for the means of clustering we have standardised both $\boldsymbol{X}$ and $\boldsymbol{y}$, the number of variables for each dataset is indicated with $p$. Throughout the datasets, we observe substantial heterogeneity across countries was observed in all the macroeconomic variables, it is important then to construct groups that allow for within group heterogeneity and at the same time, keeping the groups identifiable. It is observable that heterogeneity within groups exists at most datasets, within groups. The only dataset that indicates near-homogeneous groups is the dataset of the Italian regions, where the F-statistic reported is statistically significant at 10% level of significance. Further, the "US MSA's"," Gravity" and "Gasoline Demand" datasets report one group. This finding for the "US MSAs", at least, can appear odd because the Metropolitan housing areas are 381, and there are grounds for grouping, as the housing market corresponds to different incomes, while for the rest of the two a similar hypothesis can be adopted since income is different across different groups of countries. In these datasets strong heterogeneity (to noisy to find a signal, i.e. groups) in the cross-section is observed, indicative of no distinct grouping. Therefore the gap statistic fails to produce optimal number of groups since it is based on the hypothesis that the groups while heterogeneous, are clearly distinct.

table[table omitted — 1,573 chars of source]

Within-group analysis

We test for heterogeneity within each group, using (ref), across the eleven different datasets considered in Table (ref). First, we note that across datasets heterogeneity has been observed within groups. It is interesting to point out that our test reports different levels of significance across different variables and groups, implying borderline homogeneity for some groups, if the test reports p-values $\approx10\%$. While our findings do not suggest borderline homogeneity, the findings of the "Savings" dataset are worth discussing. Specifically, we find that in the first group of ${S}_{t}$ (the ratio of savings to GDP), the p-value is $0.0003<<1\%$, while in the second group a similar result applies. Regarding the second variable ${ I} _{t}$ (the CPI-based inflation rate) the p-value of the first group is $0.00026$, and for the second group it is $0.00014<<1\%$, while for ${GDP}_{t}$ (the per capita GDP growth rate) similar results are found. These findings are indicative of heterogeneity within each group for the different variables, as the test results appear significant at 1%, with p-values significantly lower than the nominal rate. For the third variable, ${R}_{t}$ (the real interest rate), p-values for the first and second group are $0.04$, $0.04<5\%$, which indicate that our tests report statistically significant results at the $5\%$ level of significance. The latter indicates that interest rates across countries (economic units) can be indicative of two groups and further those two groups appear less heterogeneous than the rest.

We use the same analysis in the remainder of the datasets and find, across these datasets, that all groups rejected the null hypothesis at $1\%$ level of significance.

Discussion

We explore the idea of different types of heterogeneity in large panel data models where the number of units are allowed to diverge faster than the sample size. We follow the work of su2016identifying, and use several approaches to identify and estimate latent grouping structures in panel data, developing panel penalized profile likelihood methods for classification and estimation. More importantly, we argue that, often, cross-sectional heterogeneity is not alleviated through grouping, creating a degree of bias and further identification issues in the model. In this paper, we explore the asymptotic properties of such a model, and emphasize on the identification of heterogeneous components within a group. We devise a testing procedure that identifies heterogeneity, for which the statistic is asymptotically chi-squared under the hypothesis of homogeneity, while it is consistent otherwise.

These techniques combined, provide a general approach to grouping and estimating panel models with unknown groups, heterogeneity within and across groups, and an unknown number of groups. We use known in the literature statistics, e.g. gap, to identify the number of groups, while the method is forcing neither grouping, nor within group heterogeneity.

Simulations show that the different approaches considered have good finite-sample performance and can be implemented in empirical datasets. Further, the two tests considered have good size and power under the null hypothesis. We use eleven empirical applications which reveal the advantages of data-determined identification of latent grouping structures in empirical panel modeling, while most importantly our tests show that the different groups identified across datasets are heterogeneous.

This work can be extended in a number of different directions, which we leave for future research. First, it may be appealing to consider a more general framework that allows the number $(K)$ of groups to grow with the sample size. The theoretical framework used in this paper can allow $K$ to grow with $N$ at a slower rate. Second, an extension to panel models with endogeneity will raise new statistical and computational challenges. Lastly, and in our view, a more appealing issue is considering mixed panel models, allowing for more heterogeneity in the sense of noise, while grouping structures can be considered as signals. It is then natural to devise different testing procedures to identify both the heterogeneous components and the signals, such as LM-type of tests for the former and simple t-tests for the latter. This framework though welcomes issues such as different degrees of bias to control, while joint estimation procedures can come with an additional mathematical and computational challenge, empirically though this framework is more realistic and general.