Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
85,454 characters · 14 sections · 49 citation commands
Heterogeneous Grouping Structures in Panel Data
\affil[1]{ King's Business School , King's College London}
\affil{King's Business School , King's College London}
The use and analysis of panel datasets and associated models is a vital research topic in both econometric theory and applications; these models are frequently applied in many disciplines across economics and other social sciences. This type of dataset allows the consideration of both cross-sectional and time dimensions, and provides a rich source of empirical information. In its majority, panel data analysis relies heavily on the assumption of homogeneity across units making inference more powerful.
Nevertheless, such an assumption is often restrictive and frequently rejected; see, among others, chudik2015common, phillips2007transition, pesaran2006estimation, and browning2014dynamic. Ignoring heterogeneity, across units, can result in biased estimation and misleading inference; see, among others, Chapter 1 in baltagi2008econometric, moulton1986random, and moulton1987diagnostics. On the other hand, allowing complete heterogeneity leads to potentially inefficient panel data estimation and inference. To address this issue, researchers often consider intermediate ways of exploring panel data such as panel structure models.
Panel structure models rely on the hypothesis that all, or a subset of, parameters are heterogeneous across groups but homogeneous within a group where neither the number of groups nor the individuals' group-membership is known. Examples of these models can be found in phillips2007transition who accommodate the hypothesis of convergence of clubs; in this case different countries belong to different groups and behave differently with the group structures being latent rather than assumed to be observable.
The identification of latent structures is not an easy task. It is a computationally expensive procedure and it can be even infeasible to try all the possible permutations of units across groups. A part of the literature considers that units in the dataset are grouped using an external observable classification; see, e.g., BESTER2016197. This approach may suffer, as suitable external variables can be difficult to find in many empirical settings. Furthermore, a wrong choice for the external classifier may again lead to misleading inference. For this reason, it can be argued that it is more useful for the researcher to use data-driven classification procedures such as k-means, see, e.g., sarafidis2015partially, bonhomme2015grouped, lin2012estimation, and ando2016panel ; these studies consider k-means in linear panel data models, with an unknown grouping structure. Although unsupervised classification is an appealing approach and has been theoretically proven to be asymptotically consistent, see, e.g., pollard1981strong, alternative approaches based on the Classifier Lasso (c-lasso), proposed by su2016identifying (SSP) and extended by su2018identifying and su2019sieve, have been recently introduced in the literature. In particular, SSP propose a novel estimation technique which involves penalisation and simultaneously classifies the data and estimates the underlying parameters consistently. meh2022 provides a framework for joint estimation and identification of latent grouping structures in panel data models using a pairwise fusion penalized approach. mammen2022estimation presents a new approach to grouping fixed effects in a linear panel model to reduce their dimensionality and ensure identifiability, by using unsupervised non-parametric density based clustering, cluster patterns including their location and number are not restricted.
A different line of work is that of ren2022matrix who propose a group matrix network autoregression (GMNAR) model, which assumes that the subjects in the same group share the same set of model parameters. In a similar vein, zhu2022simultaneous study dynamic behaviours of heterogeneous individuals observed in a network, where the dynamic patterns are characterised by a network vector autoregression model with a latent grouping structure, where group-wise network effects and time-invariant fixed-effects can be incorporated. wu2022 develop methods to recover sparsity patterns and grouping structures in large panels with both individual and time effects.
In a low dimensional setting, where $T>>N,p$, wanga2022panel generalise their framework and consider panel model with interactive fixed effects such that individual heterogeneity is captured by latent grouping structure and time heterogeneity is captured by an unknown structural break. liu2022panel propose a methodology for identifying and estimating explosive bubbles in mixed-root panel autoregressions with a latent grouping structure.
Full homogeneity placed within groups, is a very common assumption in the literature, but this can be a potentially restrictive hypothesis. It is conceivable and, as we will argue, often the case, in practice, that each group contains units with heterogeneous parameters whose mean is group specific and differs across groups.
To fix ideas let the parameter of interest across $N$ panel units be $\boldsymbol{\beta}_{i} = (\boldsymbol{\beta}_{1,i}, \boldsymbol{\beta}_{2,i}, \ldots, \boldsymbol{\beta}_{p, i})^{\prime}$, where $p$ is the number of covariates within each unit $i=1,\ldots,N$. Traditionally, under full homogeneity, $\boldsymbol{\beta }_{i}=\boldsymbol{\beta}$. This is full homogeneity. A heterogeneous extension is $\boldsymbol{\beta}_{i}=\boldsymbol{\beta}+\boldsymbol{\eta}_{i}$, where $\boldsymbol{\eta}_{i}$ is an i.i.d. process and focus is placed on estimation of and inference for $\boldsymbol{\beta}$. Grouping structures specify that $\boldsymbol{\beta}_{i}=\boldsymbol{\alpha}_{k}$, whereby unit $i$ belongs to group $k$, where $k=1,\ldots,K$ and $K$ is usually assumed finite. Focus here is placed on $\boldsymbol{\alpha}_{k}$. Our setting extends this by specifying that $\boldsymbol{\beta}_{i}=\boldsymbol{\alpha}_{k}+\boldsymbol{\eta}_{i}.$ Clearly $\boldsymbol{\eta}_{i}=\boldsymbol{0}_{p\times1}$ and $\boldsymbol{\alpha}_{k}=\boldsymbol{\alpha}$, are both assumptions that are empirically verifiable and potentially invalid. The case where neither holds has not been explored in the literature and is the main focus, and contribution, of this paper.
Following the work of SSP, we extend the Classifier-Lasso framework to the above case and further propose an estimation technique based on the k-means clustering method allowing for within-group heterogeneity. Our focus is mainly to showcase consistency of the estimation of $\boldsymbol{\beta}_{i}$ and $\boldsymbol{\alpha}_{k}$ providing both theoretical and small sample evidence, rather than the identification of an exact grouping structure, which under our form of heterogeneity is obviously not consistently possible. We suggest the use of a simple unsupervised classifier, like the k-means, to determine the grouping structure and, once determined, we estimate the unit-specific parameters with penalised methods.
Our work extends SSP and contributes to the literature as follows: first, we allow for a less restrictive model, second we allow for richer cross-sectional heterogeneity, and finally, we provide a data-driven approach to classify each cross-section. While the latter is a gain of our modelling approach, heterogeneity is still restrictive and, in certain cases, contradicts the basic principles of k-means clustering. In particular, when the level of heterogeneity is high, causing the data to be more noisy cross-sectionally, the clustering process tends to be misleading since it violates one of the three basic assumptions of k-means; that is, the variance between individuals must be constant and small. Because k-means minimises the within cluster deviation (in terms of squared Euclidean distances), the optimisation method converges to local minima when the variance between individuals (cross-sections) is large.
The latter is inconvenient when the empirical data is noisy or even correlated, but there are other unsupervised classifiers which deal with these types of cases; see, e.g., leisch1999bagged who re-estimates the data in different partitions using bootstrapping. In this paper, we do not focus on such a task since our approach does not cover cases where data exhibit autocorrelation or endogeneity. Obviously, an alternative to our approach would be the construction of factors instead of groups (clusters). This would lower the dimension of the data by choosing $K$ factors which explain considerably the level of heterogeneity in the cross-sections; see, e.g., su2018identifying for more details.
A further and important contribution of our work is a set of testing procedures, exploring the validity of $\boldsymbol\eta_{i}=\boldsymbol0$ and $\boldsymbol\alpha_{k}=\boldsymbol\alpha$, using testing procedures in the former case. We use our methods in a variety of empirical settings to explore the validity of the restrictions and find that in most cases both are not valid, further showcasing the utility of our framework.
The rest of the paper is organised as follows. Sections (ref) and (ref) outline our theoretical considerations. In Section (ref) we present extensive Monte Carlo simulations and discussion, Section (ref) presents several empirical examples in support of our method. We conclude in Section (ref). Technical proofs, of the main (auxiliary) results, and descriptions of the empirical datasets are relegated to the Appendix.
For any vector $\boldsymbol{x}\in\mathbb{R}^{n}, $ we denote the $\ell_{p}$-, and $\ell_{\infty}$ as $\left\| \boldsymbol{x} \right\| _{p}=\left( \sum_{i=1}^{n} | x_{i} |^{p}\right) ^{1/p}$, $\lVert \boldsymbol{x} \rVert_{\infty} = \max_{i=1,\ldots,n}|x_{i}|$. Further, $\boldsymbol{1}\{\cdot\}$ denotes an indicator function. We use "$\to_{P}$" to denote convergence in probability. For two deterministic sequences $a_{n}$ and $b_{n}$ we define asymptotic proportionality, "$\asymp$", by writing $a_{n} \asymp b_{n}$ if there exist constants $0 < a_{1} \leq a_{2}$ such that $a_{1}b_{n} \leq a_{n} \leq a_{2}b_{n}$ for all $n \geq1$. For any set $A$, $|A|$ denotes its cardinality, while $A^{c}$ denotes its complement. Define $\boldsymbol{A}$, a $n\times m $ matrix, we denote $\Lambda_{\min}(\boldsymbol{A})$ as the smallest eigenvalue of $\boldsymbol{ {A}}$, and $\Lambda_{\max}(\boldsymbol{A})$ as the largest eigenvalue of $\boldsymbol{ {A}}$.
We consider the following balanced panel model:
where $\mu_{i}^{0}$ is an individual fixed effect, $\boldsymbol{x}_{it}^{\ast }=\left( {x}_{1,it}^{\ast},\ldots,{x}_{p,it}^{\ast}\right) ^{\prime}$is a $(p\times1)$ vector of covariates, with $E(x_{l,it}^{\ast}x_{j,it}^{\ast})=0$, $\varepsilon_{it}^{\ast}=(\varepsilon_{1t},\ldots,\varepsilon_{Nt}^{\ast})$ is the idiosyncratic error, $\boldsymbol{\alpha}^{0}=(\boldsymbol{\alpha}_{1} ^{0},\ldots,\boldsymbol{\alpha}_{K}^{0})^{\prime}$ a $K\times p$ matrix of group specific centres for each covariate $j$, where $K$ denotes the number of groups considered. $\boldsymbol{G}_{k}^{0}$, $k=1,\ldots,K$, are sets of indices, denoting groups of units. Note that in our analysis, both $N$ and $T$ can be large and $N$ can possibly grow faster than $T$. Further, $\boldsymbol{\alpha }_{j}^{0}\neq\boldsymbol{\alpha}_{k}^{0}$ for any $j\neq k,$ $\boldsymbol{G} _{k}^{0}\subset\{1,2,\ldots,N\}$, $\cup_{k=1}^{K}\boldsymbol{G}_{k} ^{0}=\{1,2,\ldots,N\}$, and $\boldsymbol{G}_{k}^{0}\cap\boldsymbol{G}_{j} ^{0}=\varnothing$ for any $j\neq k$, $j=1,\ldots,K$. We denote the cardinality of the set $\boldsymbol{G}_{k}$ as $N_{k}=|\boldsymbol{G}_{k}^{0}|$. The slope heterogeneity, $\boldsymbol{\eta}_{i}=(\eta_{i,1},\ldots,\eta_{i,p})^{\prime}$ is considered to be unknown and allowed to be group specific, along with the cluster membership. Intuitively, $K$ corresponds to the number of groups and units (indexed by $i$) within the same group share a common slope $\boldsymbol{\alpha}_{k}$ but are allowed to differ by $\boldsymbol{\eta} _{i},i\in\boldsymbol{G}_{k},k=1,\ldots,K$. Further the grouping membership is common across different covariates.
Often in the panel data literature, a common characteristic can be proxied by factors (or interactive effects), e.g., as in bai2009panel, because of their ability to effectively summarise information in large data sets. While existing work suggests that the grouping of different units in a panel model can be effective under cross-sectional dependence and heteroscedasticity where $N<T$, see, e.g., bai2020standard, and can even accommodate time variation, see e.g. su2019sieve. In this paper, we allow for a degree of within-group heterogeneity, implying that grouping does not account for the entirety of cross-sectional heterogeneity. While we are agnostic about the cross-sectional and within-group idiosyncratic effect, we do assume an a-priori knowledge of the existence of grouping.
For the linear model in (ref) with ${E}\left( \varepsilon _{it}^{\ast}\mid\boldsymbol{x}_{it}^{\ast},\mu_{i}^{0}\right) =0$, we have $\widehat{\mu}_{i}=\bar{y}_{i.}^{\ast}-\boldsymbol{\beta}_{i}^{\prime} \bar{\boldsymbol{x}}_{i}^{\ast}$, where $\bar{y}_{i.}^{\ast}=\frac{1}{T} \sum_{t=1}^{T}y_{it}^{\ast},{y}_{it}=y_{it}^{\ast}-\bar{y}_{i}^{\ast}$, and ${\bar{\boldsymbol{x}}}_{i}$, $\boldsymbol{x}_{it}$, $\bar{\varepsilon}_{it}$, and $\varepsilon_{it}$ are analogously defined. Then (ref) becomes
Following su2016identifying and motivated by the literature on fused Lasso e.g. tibshirani2005sparsity, we minimise the following penalised profile likelihood (PPL) to estimate the parameters of interest.
where $\lambda>0$ is the regularisation parameter and $\boldsymbol{{\beta =(\boldsymbol{\beta}}}_{1},\ldots,\boldsymbol{{\beta}}_{N}\boldsymbol{{)} }^{\prime}$. Then, (ref) is minimised for some value of $(\lambda,K)$ and produces $(\boldsymbol{\widehat{\beta}} ,\boldsymbol{\widehat{\alpha}})^{\prime}$. We generalise su2016identifying (c-lasso hereafter) by allowing for heterogeneity within groups.
We show that both the group and the unit-specific parameter can be consistently estimated even when the underlying generating process is (ref). We obtain estimates of the (unit)group-specific parameters by using either the c-lasso iterative procedure or k-means estimation, see e.g. macqueen1967, pollard1981strong.
The penalty term in (ref) takes a mixed additive-multiplicative form, introduced in su2016identifying. In the traditional Lasso literature, e.g. tibshirani1996regression, the penalty term is additive, and imposes a degree of sparsity in favour of a parsimonious model at the cost of a degree of bias, whereas in this setting, parsimony results from grouping. The multiplicative nature of the penalty term in (ref) is essential as it produces $N$ additive terms for each of the $K$ separate penalties, such that $\{i=1\ldots,N,\;k=1,\ldots ,K:\;i\in\boldsymbol{G}_{k},\boldsymbol{\beta}_{i}^{0}\in \boldsymbol{\mathcal{B}}_{k}\subseteq\boldsymbol{\mathcal{B}}\}$, where, for each group $k$, $\boldsymbol{\mathcal{B}}_{k}\subseteq\boldsymbol{\mathcal{B} }$ and $\boldsymbol{{\beta}}^{0}\in\boldsymbol{\mathcal{B}}\subseteq \mathbb{R}^{N\times p}$. The parameters are assumed to exhibit a certain grouped pattern, the number of unknown slope parameters in $\{\boldsymbol{{\beta}}_{i}\}$ is of order $O(K)$, instead of $O(N)$, where, $K(<N)$ is typically a fixed constant in empirical applications.
Contrary to c-lasso, we do not impose homogeneity within any of the groups $\boldsymbol{\mathcal{B}}_{k}$, hence the unit-specific $\boldsymbol{{\beta}}_{i}^{0}$ are not penalised down to a specific group-centre $\boldsymbol{{\alpha}}^{0}_{k}$, simply because $\boldsymbol{{\beta}}_{i}^{0}\neq\boldsymbol{{\alpha}}^{0}_{k}$, which is more indicative of a {fused-Lasso} penalty. The component of summation is essential in order to extract information from the $N$ cross-sectional units, resulting in the identification of both $\{ \boldsymbol{{\beta}}_{i}^{0} \}$ and $\{ \boldsymbol{{\alpha}}^{0}_{k} \}$. The grouping, while fixed, is not known a-priori.
We present the necessary assumptions that ensure consistency of the proposed estimator:
We further introduce an additional assumption on the second moment of $\boldsymbol{\eta}_{i}$, which serves as a group-uniqueness condition towards cluster identifiability.c-lasso
With the following theorems, Theorem (ref) and (ref), we provide asymptotic guarantees for the proposed estimator, the c-lasso estimator and the feasible k-means estimator. The following theorems establish consistency of the adjusted penalised profile likelihood estimates of the slope parameter ${\boldsymbol{\beta}}$ and the and group-specific parameter ${\boldsymbol{\alpha}}$.
We consider two types of penalised estimators, the first is c-lasso. We obtain the set of estimated parameters $(\boldsymbol{\ddot{ \beta} }, \boldsymbol{\ddot{ \alpha}})$, by minimising the objective in (ref), $\arg\min_{\boldsymbol{\beta}\in\mathbb{R}^{N\times p}, \boldsymbol{\alpha}\in\mathbb{R}^{p\times K}} Q_{iNT,\lambda}\left( \boldsymbol{{\beta}},\boldsymbol{{\alpha}}\right) $ for a choice of tuning parameter, $\lambda$. In the following theorem we establish asymptotic consistency of c-lasso under heterogeneous groups.
Similarly to c-lasso, one can use alternative estimators for $\boldsymbol{{\alpha}}$. An estimator to be considered is the k-means. Let $Q_{1,NT}={(NT)}^{-1}\sum_{i=1}^{N}\sum_{i=1}^{T}\frac{1}{2}\left( y_{it}-\boldsymbol{\beta}_{i}^{\prime}\boldsymbol{x}_{it}\right) ^{2}$. The second methodology involves the minimisation of:
where $f^{KM}(\boldsymbol{\beta})=\arg\min_{\boldsymbol{\alpha}}\sum _{i}\left\Vert \boldsymbol{\beta}_{i}-\boldsymbol{\alpha}_{k}\right\Vert _{2}^{2}$ is the k-means solution. Then ${\boldsymbol{\widehat{\beta}} }=\arg\min_{\boldsymbol{\beta}\in\mathbb{R}^{N\times p}}Q_{2,iNT,\lambda }\left( \boldsymbol{{\beta}}\right) $ and $\boldsymbol{\widehat{\alpha} }=f^{KM}(\boldsymbol{\widehat{\beta}})$. Theorem (ref) has similar asymptotic properties to the c-lasso, to Theorem (ref), but with slightly slower convergence rates.
In this section we consider a feasible way to identify the cluster structure. We devise a simple method, where we obtain the estimates of the unit-specific parameters with ols \[ \boldsymbol{\widetilde{\beta}}_{i}=\left( \boldsymbol{x}_{i}^{\prime }\boldsymbol{x}_{i}\right) ^{-1}\boldsymbol{x}_{i}^{\prime}\boldsymbol{y} _{i}. \] and minimise ${Q}_{N}(\boldsymbol{G}_{k},\boldsymbol{\alpha} )=\frac{1}{N}\sum_{i=1}^{N}\left( \boldsymbol{\beta}_{i}-\boldsymbol{\alpha }_{k}\right) ^{2} $ with respect to $\boldsymbol{\alpha}$, for $i=1,\ldots,N,\;k=1,\ldots,K$.
It becomes clear that $\arg\min_{\boldsymbol{\alpha}} {Q}_{N}(G_{k}, \boldsymbol{\alpha}) $ imposes an infeasible problem as there has to be knowledge of $\boldsymbol{\beta}_{i}$. Therefore, we consider a unit-by-unit estimation to obtain $\boldsymbol{\widehat{\beta}}_{i}$, using (ref). Substituting (ref) in (ref), we have that
where $\boldsymbol{x}_{i}=(x_{i1},\ldots,x_{iT})^{\prime}$, $\boldsymbol{\varepsilon}_{i}=(\varepsilon_{i1},\ldots,\varepsilon _{iT})^{\prime}$ and $\boldsymbol{\psi}_{i}=\left( \boldsymbol{x}_{i} ^{\prime}\boldsymbol{x}_{i}\right) ^{-1}\boldsymbol{x}_{i}^{\prime }\boldsymbol{\varepsilon}_{i}$. Then, a feasible k-means loss can be constructed, $\widehat{Q}_{N}(G_{k}, \boldsymbol{\alpha} )$, where the unit-by-unit estimates can be clustered instead of $\boldsymbol{\beta}_{i}$,
Note that, minimising (ref), the grouping structure relies on the estimation of ${{\boldsymbol{\beta}}}_{i}^{\ast}$. We showcase, with the following example, that ${Q}_{N}(\cdot)$ can be minimised with respect to some grouping centre $\boldsymbol{\alpha}= \left(\boldsymbol{\alpha}_{j, 1}, \ldots, \boldsymbol{\alpha}_{j, K} \right)'$, $j=1, \ldots, p$.
In the following theorem we show consistency of the estimated centres obtained by clustering $\boldsymbol{ \widehat{\beta}}$, in (ref), through k-means, we call this method feasible k-means
Naturally, Theorem (ref) relies heavily on Assumption (ref), where $( \boldsymbol{ {\alpha}}^{0}_{1}, \ldots, \boldsymbol{ {\alpha}}^{0}_{K})^{\prime}$ is a unique minimiser of the population $Q(\cdot)_{N}$, therefore, $(\boldsymbol{\widetilde{\alpha}} _{1}, \ldots, \boldsymbol{\widetilde{\alpha}} _{K})^{\prime}$ minimises $\widehat{Q}_{N} (G_{k}, \boldsymbol{\alpha} )$.
In this section we discuss the classification consistency of the different estimating techniques.
Due to the presence of $\boldsymbol{\eta}_{i}$ in the generating process in (ref), the probabilities of events of the form $\mathcal{I} _{1,i}=\{i\in\boldsymbol{\widehat{G}}_{k}|i\notin\boldsymbol{G}_{k}^{0}\}$, and $\mathcal{I}_{2i}=\{i\notin\boldsymbol{\widehat{G}}_{k}|i\in \boldsymbol{G}_{k}^{0}\}$, respectively, become difficult to control. The literature usually defines classification consistency as the property where $\Pr(\cup_{i}\mathcal{I}_{1,i})\rightarrow0$ and $\Pr(\cup_{i}\mathcal{I} _{2,i})\rightarrow0.$ In fact in our case it is possible to show the converse. In particular, we have the following Lemma
The possibility of the presence of both a group structure as well as within group heterogeneity necessitates an exploration of data dependent methods for heterogeneity inference. Two issues arise. First determining whether there is heterogeneity within clusters and the second is, under heterogeneity, how to determine the number of clusters. It is important to have the right sequencing in this investigation. We propose to start with a test of the following null hypothesis
Under $H_{0}$, methods such as c-lasso can be used to determine the number of clusters and obtain consistent estimates of their centres. These methods can then be used as discussed in Section (ref) below. If the null is not rejected one reverts to existing methods and find ways to determine the number of clusters under heterogeneity. Existing methods associated to k-means cluster analysis can be used and are discussed in Section (ref). In the following discussion we keep matters simple by not allowing for more refined features such as weak heterogeneity (defined as heterogeneity that is asymptotically negligible) or heterogeneity in a proportion of clusters (although we discuss cluster specific tests below). All such extensions are possible but beyond the scope of the current paper.
The main interest in this section is to verify the existence of different types of heterogeneity, i.e. cross-sectional, as well as within group. We follow the two-step method described in Section (ref) to extract feasibly both the group and unit-specific parameter and show that using the unit-by-unit estimate as a preliminary estimator of $\boldsymbol{\beta}_{i}$ can lead to identifiable groups via k-means, while the dimensionality of the cross-sectional parameters is allowed to diverge.
We consider the model in (ref) and simplify our analysis so that $\boldsymbol{\beta} = ({\beta}_{1}, \ldots, {\beta}_{N})^{\prime}$, and $\boldsymbol{\alpha} = ({\alpha}_{1}, \ldots,{\alpha} _{k})$ are $N\times 1$ and $K\times 1$ vectors respectively. To investigate the existence of heterogeneity in the cross section, and, consequently, within a group ${G}_k$ we construct test statistics under the null hypothesis in (ref). Our methodology proceeds in distinct steps; We first obtain the estimates of $\widetilde{\beta}_{i}$ through ols on $\boldsymbol{{y}}_{i}|\boldsymbol{x}_{i}$ and uncover the grouping structure using k-means on ${\widetilde{\beta}}_{i}$ to obtain ${\widetilde{\alpha}}_k,$ the minimiser of (ref). We then obtain $\boldsymbol{\widehat{\varepsilon}}_{i}=\boldsymbol{{y}}_{i} -\boldsymbol{x}_{i}{\widetilde{\alpha}}_k $ where $i\in \widehat G_{k}$. Note that here it is sensible to assume that $\widehat G_{k}\to_P G_{k}^0$. The latter is clear under $H_0$, contrary, under $H_1$, where the results of Lemma (ref) verify an alternative statement. Then under the Assumption that $E({\varepsilon}_{it}|{x}_{it})=0$, and $E({\eta}_{it}|{x}_{it})=0$, $T^{-1/2} \sum_{i} T^{-1}\boldsymbol{x}_{i}^{\prime}{\varepsilon}_{it} \sim N(0,\boldsymbol{\sigma}_{x\varepsilon})$, where $\boldsymbol{\sigma }_{x\varepsilon} = N^{-1} \sum_{i} T^{-1}\boldsymbol{x}_{i}^{\prime }{\varepsilon}_{it} {\varepsilon}_{it} ^{\prime}\boldsymbol{x}_{i}$. Recall from (ref) that
where $\boldsymbol{{\varepsilon}}_{i}= ( \boldsymbol{{\varepsilon}}_{i1}, \ldots, \boldsymbol{{\varepsilon}}_{iT} )^{\prime}$. A feasible solution to (ref) would be to replace $\boldsymbol{{\varepsilon}}_{i}$ with $\boldsymbol{\widehat{\varepsilon}}_{i} .$ One can then construct the cross-sectional heterogeneity test statistic,
where $\widehat{\sigma}_{i} =\frac{1}{T-1}\sum_{i} \boldsymbol{\widehat{\varepsilon}}_{i} \boldsymbol{\widehat{\varepsilon}}_{i} ^{\prime}$ a consistent estimate of $\sigma_i=E\frac{1}{T}\boldsymbol{{\varepsilon}}_{i} \boldsymbol{{\varepsilon}}_{i} ^{\prime}$.
Further, a poolability-type test would rely on the $R^{2}$ statistic, where one uses the residuals resulting from unit-by-unit ols, see e.g. (ref), to asses the fit against a pooled ols regression, e.g. $ \widehat{\beta}_{\textsc{ols}} = ( \sum_{i=1}^{N} \sum_{t=1}^{T} \boldsymbol{x} _{it} ^{\prime}\boldsymbol{x}_{it} ) ^{-1} ( \sum_{i=1}^{N} \sum_{t=1}^{T} \boldsymbol{x}_{it} ^{\prime}y_{it} ), $ Further, under standard assumptions of independence, see e.g Assumption (ref), the null hypothesis in (ref) will be rejected once the unit-by-unit ols reports larger $R^{2}$ than the corresponding $R^{2}$ of the pooled ols, which indicates that treating the data as heterogeneous increases the overall panel fit.
To that end we construct the $R^{2}$ statistic, resulting from the auxiliary regressions of $\boldsymbol{\widehat{\varepsilon}}_{{i}}|\boldsymbol{x} _{i}$, such that
where $ \boldsymbol{\widehat{v}}_{i}\boldsymbol{(}\beta\boldsymbol{)=\widehat{\varepsilon}}_{i}-\widehat{\gamma}\left( \beta\right) \boldsymbol{x}_{i} $ and $ \widehat{\gamma}\left( \beta\right) ={\boldsymbol{\widehat{\varepsilon}}_{i} ^{\prime}\boldsymbol{x}_{i}}({\boldsymbol{x}_{i}^{\prime }\boldsymbol{x}_{i}})^{-1}, $ lastly, ${\sigma_{i, R}}$ can be defined similarly to ${\sigma_{i}}$. Under $H_0$, both $\widehat{R}^{2}_{N} , \; \widehat{S}^{2}_{N}$ are asymptotically $\chi_1^2$ distributed, while consistent under the alternative. The latter is the key finding of the following theorem.
Then, we have the following theorem giving the asymptotic distribution of $S_{N}$ and $R_{N}$, defined in (ref) and (ref).
Allowing for different types of heterogeneity, the identifiability of groups becomes a difficult task. First, because in real datasets there is no a-priory knowledge of the necessity of grouping, the number of groups formed or the grouping structure. Second, in their majority grouping methods require spherical groups, this way allowing only for a degree of heterogeneity, or are density based techniques, which in turn require the knowledge of the empirical distribution for each group. Further, versatile clustering techniques, such as k-means, appear effective, but in cases of no-grouping at all appear inappropriate as it forces grouping. We address that issue by estimating the number of groups using techniques known in the statistical literature. The most commonly used is the Gap Statistic, see e.g. gap, and can be constructed as follows,
for ${i\neq j }= 1, \ldots, N$ and $k=1, \ldots, K$, where $\boldsymbol{d}$ is the square euclidean distance from a centre in group $k$, $W_{k}$ is the pooled within cluster sum of squares around the cluster means, $E\left[ \log(W_{k}) \right] $ denotes the expectation under a sample of size $N$ is drawn from a reference distribution and $N_{k}= |G_{k}|.$ Under Assumption (ref) and well separated groups the number of clusters ${K}^{\ast }$ maximises $\text{Gap}_{N}(k), $ defined in (ref), such that
We consider a feasible alternative to maximise (ref) with respect to $\hat{K}$, such that for ${i\neq j },$
where $\boldsymbol{\widehat{\beta}}_i$ an estimator of $\boldsymbol{\beta}_i$. Note that $\boldsymbol{\widehat{\beta}}_i$ can be a mazimiser of either $Q_{iNT,\lambda}\left( \boldsymbol{{\beta}}\right)$, $Q_{2,iNT,\lambda}\left( \boldsymbol{{\beta}}\right) $, or it can maximise the unit-wise least squares objective, such that $\arg\min_{\boldsymbol{\beta}_i\in \mathbb{R}^p}Q_{1,NT}, \; Q_{1,NT}={(NT)}^{-1}\sum_{i=1}^{N}\sum_{i=1}^{T}\frac{1}{2}\left( y_{it}-\boldsymbol{\beta}_{i}^{\prime}\boldsymbol{x}_{it}\right) ^{2}$, denoted (ref), (ref) and (ref) respectively. Consider Assumptions (ref), (ref), and (ref), then $\log\widehat{W}_{k} \to _P \log W_{k} $ where ${W}_k$ is defined in (ref) and $\widehat{W}_k$ in (ref).
The gap statistic is an empirical tool to identify feasibly the number of clusters taking into account the within-cluster dispersion, given various clustering algorithms. In Section (ref) we explore the small sample properties of the method using simulations, and report the frequency of selecting the correct numbers of groups, where it is known. Further, as the findings of gap suggest, the gap statistic outperforms other methods in the literature, even in the case of no grouping, i.e. $K=1$.
In this section we explore the finite sample properties of the different estimators explored in Section (ref) and (ref) and make comparisons. A detailed description of the methods considered can be found in (ref). Further, this section is divided in two sub-sections: In Section (ref) we explore the small sample properties of the different methods in terms of mean squared error and miss-classification rate, and in Section (ref) we focus on the performance of the tests proposed in Section (ref).
To evaluate the finite-sample performance of the classification and estimation procedure, we consider 1000 simulations of model (ref), where the number of groups assumed are $K=2$, $\boldsymbol{G}_{k}$ is a membership vector for group $k=1,\ldots, K$, the simulated true centres are $\boldsymbol{\alpha} =(0.5, 2)^{\prime}$ and $\eta_{i}\sim N(0, 0.2)$. Further, $x_{it}\sim_{\text{i.i.d}} N(0,1)$, $\varepsilon_{it}\sim_{\text{i.i.d}} N(0,1)$ and $(\beta_{1},\ldots, \beta_{N/2})^{\prime}\in\boldsymbol{G}_{1}, \; (\beta_{N/2+1}, \ldots, \beta_{N})^{\prime}\in\boldsymbol{G}_{2},$ where $N=[20,50,100,200], \; T=[50,100,200,500]$.
We consider various different methods and evaluate the consistency of the grouped parameters $\beta_{i}$, and individual group centres, as well as the quality of grouping under homogeneous and heterogeneous parameters. Namely we consider the c-lasso of su2016identifying reported as SSP, the c-lasso estimator using k-means to obtain the group centres, reported as Km and the c-lasso without imposing within class homogeneity, reported as H-SSP. The penalty parameter for the penalisation throughout methods was carried out via cross validation scheme, from a parameter grid, $\lambda \in[0.125, 0.25, 0.5, 1, 2]T^{-1/3} $. We carried out a k-fold cross validation scheme without re-shuffling, where $k=10$.
To asses the consistency of the individual parameters we report the average mean squared error (MSE) across units under both the null and the alternative, e.g. (ref). We report these findings in Table (ref). In Table (ref) we report the average MSE of the centres of the first group to save space, e.g. $1000^{-1} \sum_{i=1}^{1000} (\alpha_{1,i}- \widehat{\alpha}_{1,i})$, where the estimation of the individual $K$ centres is carried out using four different methods, descried below.
To quantify the quality of grouping we use a dissimilarity measure, Rand Index, e.g. rand1971objective. The Rand index, ranges between 0 and 1, with 0 indicating that the two data groupings do not agree on any pair of points and 1 indicating that the data groupings are exactly identical to some permutation. More specifically, given the following partitions $\boldsymbol{G}= \{\boldsymbol{G}_{1}, \ldots, \boldsymbol{G}_{r} \}$, of $\boldsymbol{\beta} $ into $r$ subsets and $\boldsymbol{\widehat{G}}=\{ \boldsymbol{\widehat{G}}_{1}, \ldots, \boldsymbol{\widehat{G}}_{s} \}$, a partition of $\boldsymbol{\beta} $ into $s$ subsets, where $1<r, s\leq N$, we write the following measure \[ RI = \frac{a+b}{{\binom{N }{2 }}}\equiv\frac{TP+TN}{TP+TN+FP+FN}, \] where $N$ is the number of units to be clustered, TP is the number of true positives, TN is the number of true negatives, FP is the number of false positives, and FN is the number of false negatives. Further, $a= |\boldsymbol{B}^{*}|, \; \boldsymbol{B}^{*} = \{ (\beta_{i},\beta_{j}) | \beta_{i},\beta_{j} \in\boldsymbol{G}, \; \beta_{i},\beta_{j} \in \boldsymbol{\widehat{G}}\}$ and $b = |\boldsymbol{B}^{*}|, \; \boldsymbol{B} ^{*} = \{ (\beta_{i}, \beta_{j} ) | \beta_{i}\in\boldsymbol{I_{1}},\beta _{j}\in\boldsymbol{G}_{2}, \;\beta_{i}\in\boldsymbol{\widehat{G}}_{1} ,\beta_{j}\in\boldsymbol{\widehat{G}}_{2} \}$, where the former declares the number of pairs of elements in $\{\beta_{1}, \ldots, \beta_{N}\}$ that are in the same subset in $\boldsymbol{G}$ and in the same subset in $\boldsymbol{\widehat{G}}$, and the latter declares the number of pairs of elements in $\{\beta_{1}, \ldots, \beta_{N}\}$ that are in different subsets in $\boldsymbol{G}$ and in different subsets in $\boldsymbol{\widehat{G}}$. We report these findings in Table (ref) both under homogenous and heterogeneous groups.
Further, since heterogeneity is allowed within a group, we consider additional metrics to ensure that the grouping is not forced but indeed latent structures can be found in the cross section. We then estimate the number of groups that can be found using the unit-by-unit (UE) estimate, ${\beta}^{\ast}_i$ obtained in (ref), reporting the frequency of selecting the true number of groups, $K_{0}=2$ throughout 1000 replications, through the gap statistic, e.g. gap. We report these findings in Table (ref) both under homogenous and heterogeneous coefficients.
As it is expected, in Table (ref), the MSE decreases as the sample size increases, while SSP reports a higher MSE compared to all other methods as both $N,T$ increase in the cases of heterogeneity (see right panel of Table (ref)), while it also decreases with an increase of $T$. The latter is due to the fact that SSP does not allow within group heterogeneity , forcing the cross-sectional parameters to be penalised down to their corresponding group centre, while heterogeneity persists within a group, as the finding of Table (ref) indicate. $Km$ reports the smallest MSE across methods in cases of heterogeneity, In the cases of homogeneity (left panel of Table (ref)), SSP is the best performing method, as it is expected, while all methods report smaller MSEs as $T$ increases.
In Table (ref) all methods reported appear equivalent, while as $T$ increases the error approaches zero. In Table (ref) we report the average RI (see (ref)) throughout 1000 replications. As it is expected, under the null (see left panel of Table (ref)), all methods perform are equivalently, i.e. most methods report average accuracy around or equal 100%. In the case both when generated under the alternative hypothesis, i.e. heterogeneity across units, SSP reports the smallest accuracy throughout samples, but as the sample size $T$ increase, accuracy increases as well. The latter is a common finding across methods.
Table (ref) reports the average frequency of selecting the correct number of groups via the gap statistic. Table (ref) compliments the results of Table (ref), verifying that (1) clustering the cross sectional units is necessary, and (2) while clustering is necessary, through Table (ref) we verify that accuracy of grouping is on average high throughout samples. Finding (1) can be observed both generating under the null and the alternative as the average frequency of selecting two groups is close to 1 and increases with the sample size $T$.
We simulate the model in (ref) and take the one-way within transformation, see (ref), with $x_{it}\sim_{\text{i.i.d}} N(0,1)$, $\varepsilon_{it}\sim_{\text{i.i.d}} N(0,1)$, $\boldsymbol{\alpha} =(0.5, 2)^{\prime}$, $\eta_{i}\sim N(0,0.2). $ We generate under the null hypothesis of (ref), to report the size of the test in (ref) on the right panel of Table (ref) and further power of the test can be reported by generating under the alternative, see the left panel of Table (ref). The reported quantities of Table (ref) can be similarly described for the test in (ref) where we assess the goodness of fit under clustered effects. \newline In Table (ref) we report the average frequency of selecting $K=2$ groups when the true number of groups is $K_{0}=2$ throughout 1000 replications of model (ref), while in Table (ref) we report the average frequency of selecting $K=1$ groups when the true number of groups is $K_{0}=1$. Further we report four different measures of identifying the number of latent groups in a vector of estimates (ref), the maximum number of clusters considered is $K_{\max}=5$. Namely, we consider the gap statistic, e.g. gap, reported as 'gap', the Silhouette statistic e.g. sil, reported as 'Sil', the Calinski-Harabasz statistic, e.g. ch reported as 'CH' and the Davies-Bouldin, e.g. db reported as 'DB'. \newline We are interested in investigating the existence of heterogeneity and for that matter the existence of groups in two separate cases; First, we investigate the existence of heterogeneity in the cross-section, using the unit-by-unit ols estimator in (ref) (reported as CS in Tables (ref)--(ref)) and second we investigate the existence of heterogeneity after grouping using the feasible k-means estimator (reported as WG in Tables (ref)--(ref)) minimising (ref). All the tests are carried out at $5\%$ level of significance. In Table (ref) there is clear indication that asymptotically, as $T$ increases, the test in (ref) rejects the null for both types of heterogeneity frequently, showing that indeed heterogeneity still persists both before and more importantly after grouping. Similarly to Table (ref), in Table (ref) the goodness of fit both before and after grouping is appears better as $T$ increases, showing that grouping is necessary even when heterogeneity persists within each group. The latter is an interesting finding coupled with the findings of Table (ref) --(ref), because there is an implication that in the cases where the number of groups $K=2$, asymptotically the gap measure identifies the correct number of groups 100% o the times, while when there is no grouping at all, the rate at which the gap statistic selects the correct number of groups is 100% throughout $N,\; T$. Therefore, there is clear indication of grouping structure in the cross-section, and the findings of Tables (ref)--(ref) show that heterogeneity clearly exists within the different groups as well.
In this section we discuss the results of the empirical analysis, and provide a detailed discussion of the datasets we used, in the Appendix.
In Table (ref) we report the statistics in (ref) and (ref) respectively, with p-values and report statistical significance at different levels using $(*)$. Note that ${***}$ signifies 1% significant, ${**} $ signifies 5% significant, and ${*} $ signifies 10% significant. Note that the datasets used in the empirical analysis have different number of variables, hence the critical value for the $\chi_{d}^{2}$ with which the statistics (ref) and (ref) are compared, changes, where $d$ are the degrees of freedom. Namely, for $d=1: c=3.841$, $d=2: c=5.991$, $d=3: c=7.815$, $d=4: c=9.488$.
Across datasets we use the feasible k-means estimator based on the unit-wise estimation in (ref) and estimate the number of clusters using the gap statistic, see e.g. gap. The optimal number of clusters is selected from the range ${K} \in\{ 1, \ldots, 20 \}$ to maximise the criterion in (ref). Further we obtain the solution of (ref) based on the optimal number of clusters $\widehat{K}$ and use (ref) to evaluate whether there is remainder heterogeneity collectively across groups and asses the fit of the proposed clustering method using (ref).
In our empirical application we use eleven datasets, in which for the means of clustering we have standardised both $\boldsymbol{X}$ and $\boldsymbol{y}$, the number of variables for each dataset is indicated with $p$. Throughout the datasets, we observe substantial heterogeneity across countries was observed in all the macroeconomic variables, it is important then to construct groups that allow for within group heterogeneity and at the same time, keeping the groups identifiable. It is observable that heterogeneity within groups exists at most datasets, within groups. The only dataset that indicates near-homogeneous groups is the dataset of the Italian regions, where the F-statistic reported is statistically significant at 10% level of significance. Further, the "US MSA's"," Gravity" and "Gasoline Demand" datasets report one group. This finding for the "US MSAs", at least, can appear odd because the Metropolitan housing areas are 381, and there are grounds for grouping, as the housing market corresponds to different incomes, while for the rest of the two a similar hypothesis can be adopted since income is different across different groups of countries. In these datasets strong heterogeneity (to noisy to find a signal, i.e. groups) in the cross-section is observed, indicative of no distinct grouping. Therefore the gap statistic fails to produce optimal number of groups since it is based on the hypothesis that the groups while heterogeneous, are clearly distinct.
We test for heterogeneity within each group, using (ref), across the eleven different datasets considered in Table (ref). First, we note that across datasets heterogeneity has been observed within groups. It is interesting to point out that our test reports different levels of significance across different variables and groups, implying borderline homogeneity for some groups, if the test reports p-values $\approx10\%$. While our findings do not suggest borderline homogeneity, the findings of the "Savings" dataset are worth discussing. Specifically, we find that in the first group of ${S}_{t}$ (the ratio of savings to GDP), the p-value is $0.0003<<1\%$, while in the second group a similar result applies. Regarding the second variable ${ I} _{t}$ (the CPI-based inflation rate) the p-value of the first group is $0.00026$, and for the second group it is $0.00014<<1\%$, while for ${GDP}_{t}$ (the per capita GDP growth rate) similar results are found. These findings are indicative of heterogeneity within each group for the different variables, as the test results appear significant at 1%, with p-values significantly lower than the nominal rate. For the third variable, ${R}_{t}$ (the real interest rate), p-values for the first and second group are $0.04$, $0.04<5\%$, which indicate that our tests report statistically significant results at the $5\%$ level of significance. The latter indicates that interest rates across countries (economic units) can be indicative of two groups and further those two groups appear less heterogeneous than the rest.
We use the same analysis in the remainder of the datasets and find, across these datasets, that all groups rejected the null hypothesis at $1\%$ level of significance.
We explore the idea of different types of heterogeneity in large panel data models where the number of units are allowed to diverge faster than the sample size. We follow the work of su2016identifying, and use several approaches to identify and estimate latent grouping structures in panel data, developing panel penalized profile likelihood methods for classification and estimation. More importantly, we argue that, often, cross-sectional heterogeneity is not alleviated through grouping, creating a degree of bias and further identification issues in the model. In this paper, we explore the asymptotic properties of such a model, and emphasize on the identification of heterogeneous components within a group. We devise a testing procedure that identifies heterogeneity, for which the statistic is asymptotically chi-squared under the hypothesis of homogeneity, while it is consistent otherwise.
These techniques combined, provide a general approach to grouping and estimating panel models with unknown groups, heterogeneity within and across groups, and an unknown number of groups. We use known in the literature statistics, e.g. gap, to identify the number of groups, while the method is forcing neither grouping, nor within group heterogeneity.
Simulations show that the different approaches considered have good finite-sample performance and can be implemented in empirical datasets. Further, the two tests considered have good size and power under the null hypothesis. We use eleven empirical applications which reveal the advantages of data-determined identification of latent grouping structures in empirical panel modeling, while most importantly our tests show that the different groups identified across datasets are heterogeneous.
This work can be extended in a number of different directions, which we leave for future research. First, it may be appealing to consider a more general framework that allows the number $(K)$ of groups to grow with the sample size. The theoretical framework used in this paper can allow $K$ to grow with $N$ at a slower rate. Second, an extension to panel models with endogeneity will raise new statistical and computational challenges. Lastly, and in our view, a more appealing issue is considering mixed panel models, allowing for more heterogeneity in the sense of noise, while grouping structures can be considered as signals. It is then natural to devise different testing procedures to identify both the heterogeneous components and the signals, such as LM-type of tests for the former and simple t-tests for the latter. This framework though welcomes issues such as different degrees of bias to control, while joint estimation procedures can come with an additional mathematical and computational challenge, empirically though this framework is more realistic and general.