Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
127,819 characters · 21 sections · 91 citation commands
Heterogeneous structural breaks in panel data models
Panel data sets have become increasingly popular in economics and finance, because they allow us to make use of variations over both time and individual dimensions in a flexible way. However, when we analyze panel data, it is important to take into account structural changes, such as financial crises, technological progress, economic transitions, etc. occurring during the time periods covered in the data. This is because these events may influence the relationships between economic variables, causing breaks in the parameters of panel data models. An important issue is that the time points and/or the impact of the changes are likely to vary significantly across individuals. For example, despite the wide-ranging effects of the 1997 Asian financial crisis, which resulted in the restructuring of most Asian economies, both China and India were largely unaffected park&yang&shi&jiang2010, radelet&sachs2000.
Similarly, for the recent European debt crisis after 2009, not all European countries have been affected. Even within the Eurozone, some countries fell into crisis much earlier than the others, and the impact of the crisis on the economic structure also differed across countries, with Central European countries appearing to be much less affected than the main victims in Southern Europe, such as Italy and Spain claeves&vasicek&2014. Together, these observations suggest that structural breaks are not necessarily common across all observational units and breaks might occur at different time points across units and/or be of heterogeneous size. Importantly, existing break detection techniques for panel data primarily focus on only common structural breaks.
This paper provides a new model and estimation procedure that allows us to detect heterogeneous structural breaks. We consider a linear panel data model in which the coefficients are heterogeneous and time varying. We model individual heterogeneity via a grouped pattern, and allow the group membership structure (i.e., which individuals belong to which group) to be unknown and estimated from the data. For each group, we then allow common structural breaks in the coefficients, while the number of breaks, breakpoints, and/or break sizes can differ across groups.
To estimate this model, we employ a hybrid procedure of the grouped fixed effects (GFE) method proposed by bonhomme&manresa2015 and the adaptive group fused Lasso (AGFL) in qian&su2016.\footnote{Lasso stands for “least absolute shrinkage and selection operator”, as introduced by tibshirani1996.} We title our procedure Grouped AGFL (GAGFL). The idea of our estimation approach is to use GFE to estimate the group memberships, and simultaneously for each group to employ AGFL to detect the breaks and estimate the coefficients. The GFE approach estimates the group memberships by minimizing the sum of squared residuals. We choose the GFE approach for classification because it facilitates theory and guarantees that all units are categorized to one of the groups. For its part, AGFL estimates the break dates by minimizing a penalized least squares objective function, where the penalty term is the norm of the difference between the values of the time-varying parameters in adjacent time periods. It is particularly useful in our context because it estimates break dates jointly with the coefficients, and it allows breaks to occur in every period.
Of course, break detection, which is interesting in its own right, also improves the estimation of coefficients and group memberships by automatically pooling the time periods between two breaks. This hybrid estimation approach allows us to: 1) consistently estimate the latent group membership structure, 2) automatically determine the number of breaks and consistently estimate the breakpoints for each group in one joint step, and furthermore 3) consistently estimate the regression coefficients with group-specific structural breaks. Computationally, we develop an iterative method to compute the estimates by combining the K-means bonhomme&manresa2015 algorithm with Lasso. This algorithm works particularly well when the number of groups is not large, which is typically the case in many applications bonhomme&manresa2015,vogt2017classification,su&shi&phillips2016,lu&su2017.
Imposing a grouped pattern is a convenient way to model individual heterogeneity. Nevertheless, while in principle we may be able to analyze each individual separately by allowing complete individual heterogeneity, such an approach is not practical for two reasons. First, individual estimation can be inefficient as it does not make use of cross-sectional information at all, especially given that the length of time series in a typical panel data set is not very large wang&zhang&paap. Second, analyzing individuals separately fails to capture any common pattern, which prevents learning from the experience of other individuals. A grouped pattern of heterogeneity is then especially useful because it allows us to estimate the coefficient parameters (and the breaks) more efficiently and thereby provides insights into the similarity of individual units. Moreover, it does not require a priori knowledge about the determinants of the group structure, and we can obtain an intuition about the underlying mechanism that causes heterogeneity by investigating the estimated group structure. Imposing a grouped heterogeneity pattern also enables us to estimate coefficients for every period, which further allows us to detect consecutive structural breaks.
We may conveniently characterize individual time-varying heterogeneity using a group pattern, such that individual units within a given group share the same time-varying paths, in many applications. For example, bonhomme&manresa2015 showed in an analysis of the democracy--income relationship that the time-varying unobserved heterogeneity demonstrates a group pattern and countries assemble depending on whether and when they experienced democratic transitions. Elsewhere, ando&bai2016 found that the styles of US mutual funds and asset return performance in mainland China both featured a group pattern of heterogeneity. In addition, hahn&moon2010 provided sound foundations for group structure in game theoretic or macroeconomic models where a multiplicity of equilibria is expected and bonhomme&lamadon&manresa2017 argued that group structure can be a good discrete approximation, even if individual heterogeneity is continuous.
It is important to consider heterogeneity and structural breaks jointly. Ignoring heterogeneous structural breaks can lead to the incorrect detection of breakpoints and inconsistent slope coefficient estimates. The potential harm arising may be demonstrated in the following two scenarios. First, we consider pooling individuals with different breakpoints.\footnote{This includes the special case of pooling individuals with and without structural breaks.} In this case, smaller breaks characterize the pooled data. Hence, it is more difficult to detect breaks in finite samples, and the failure of break detection further leads to incorrect coefficient estimates. This issue is also highlighted by baltagi&qu&kao2016, who showed via simulation that their break detection method (assuming common breaks) suffers from a significant loss of accuracy when some series contain no break. Second, even if all individuals have breaks at the same time, ignoring heterogeneity in break size and pooling all individuals may average out the break, and thus act against the correct detection of breaks, again leading to inconsistent coefficient estimates. Our proposed GAGFL method can simultaneously address both these problems because it allows us to identify which individuals are (un)affected by the breaks and to detect the group-specific breaks in one step.
We examine the finite sample performance of our method via Monte Carlo simulation and compare it with other break detection techniques assuming common breaks. GAGFL performs well in heterogeneous panels with structural breaks occurring at possibly different points in finite samples. First, we are able to estimate group membership precisely, and the clustering accuracy improves as the length of the time series increases. Second, we are able to estimate the correct number of breaks and the true breakpoints for each group. We also find that ignoring heterogeneity in breaks (even if we account for the heterogeneity in coefficients) leads to an inconsistent estimation of the number of breaks, along with less accurate breakpoint estimates. This also results in inaccurate coefficient estimates. Importantly, because of the accurate estimation of both groups and breaks, our coefficient estimates are much more precise than those produced by common break detection methods.
We illustrate our method by revisiting the relationship between income and democracy analyzed by acemoglu&johnson&robinson&yared2008. We mainly compare our results with the GFE estimates produced by bonhomme&manresa2015 and find that our method provides a compatible but different grouped pattern. On the one hand, we can distinguish between countries with stable, early and late transition democracies, and this result is consistent with that of bonhomme&manresa2015. On the other hand, we identify countries that have a fluctuating income--democracy relation toward the middle and/or end of the period not captured by GFE. The different grouped pattern arises because the groups formed by GAGFL draw on the whole coefficient vector that allows for structural changes. Thus, the grouped pattern produced by GAGFL not only reflects the magnitudes of the differences in the coefficient estimates, but also the heterogeneity in structural change.
The remainder of this paper is organized as follows. Section (ref) reviews the related literature. Section (ref) describes the model. Section (ref) explains the estimation method and establishes its asymptotic properties. Section (ref) discusses the models with individual-specific fixed effects. In Section (ref), we examine various extensions of our basic model and our estimation method and Section (ref) considers the finite sample properties of our estimator via simulation. To demonstrate our method using actual data, we analyze the association between income and democracy in Section (ref). Section (ref) provides some concluding remarks. All technical details are in the Appendix, and a supplementary file contains additional theoretical results, further simulation studies, and an added application on the determinants of the savings rate.
Individual heterogeneity in panel data has been widely documented in empirical studies, especially heterogeneity in slope coefficients (see, for example, su&chen2013 and durlauf&kourtellos&minkin2001 for cross-country evidence and browning2007 for ample microeconomic evidence). Various estimation techniques have been developed to account for cross-sectional slope heterogeneity, including random coefficient models swamy1970, mean group estimation pesaran&smith95, GFE bonhomme&manresa2015, classifier-Lasso su&shi&phillips2016, and Panel-CARDS wang&phillips&su2016, among others.
Our work most closely relates to recent research concerning latent group structure in panel data.\footnote{Note that bester&hansen2013 and ando&bai2015 also considered panel models with group structure, but they assume that the group structure is known.} sun2005 considered a finite-mixture model. lin&ng2012 and su&shi&phillips2016 examined cases in which the slope coefficients of individuals belong to an unknown latent group. bonhomme&manresa2015 introduced a GFE estimator that allows the fixed effect parameters to have a time-varying grouped pattern. Note that hahn&moon2010 discussed a similar estimator in the context of the multiple equilibria problem in empirical industrial organization, ando&bai2016 studied unobserved group factor structures, and vogt2017classification considered nonparametric regression models with latent group structure. However, these authors all assume that the panel is stable without structural change. We contribute to this literature by studying latent group estimation in the presence of unknown heterogeneous structural breaks. su&wang&jin2015 also considered grouped patterns with structural instability. They modeled instability using continuous time-varying slope coefficients, whereas we consider discontinuous structural breaks.
While testing and dating structural breaks in panel data has become a popular research topic, many of the existing techniques employ the assumption of homogeneous slope coefficients. For example, bai2010 studied the estimation of a common break in means and variances for panel data, deWachter&tzavalis2012 developed a break detection test for panel autoregressive models, and kim2011 proposed an estimation procedure for a common deterministic time-trend break in large panels. In other work, baltagi&kao&liu2017 introduced a procedure for breakpoint estimation in panel models for (non)stationary regressors and error terms, and qian&su2016 proposed estimation procedures for panels with common breaks in slope coefficients using AGFL. li&qian&su later extended this idea to interactive fixed effects models. We differ from these studies by considering heterogeneous structural breaks occurring for some individuals and/or with different sizes.
Despite ample evidence of cross-sectional heterogeneity in panel data, only a few studies have considered structural breaks in heterogeneous panels, which is the topic of this study. A closely related work is baltagi&qu&kao2016, which considered breakpoint estimation in heterogeneous panels with and without cross-sectional dependence. Our work differs in three main aspects. First, baltagi&qu&kao2016 considered structural breaks common to all individuals. They documented in their simulation study that their break detection method loses significant accuracy if a break only affects a part of the individuals. We explicitly model heterogeneous breaks that only impact a proportion of unknown units, and estimate the range of impacts from the data. Incorporating heterogeneity in structural breaks also allows us to estimate the breakpoints more accurately than baltagi&qu&kao2016. Second, given the number of common breaks, baltagi&qu&kao2016 estimated the breakpoints and then the slope coefficients in separate steps. In contrast, our method simultaneously estimates the number of breaks, breakpoints, and slope coefficients in one joint step.\footnote{To improve finite sample performance, we may re-estimate the slope coefficients after obtaining the breakpoints from the Lasso procedure.} Note that pooling the units may make it more difficult to determine the number of breaks in the presence of heterogeneity because it also dilutes the break sizes. Finally, we model individual heterogeneity via the grouped pattern, while baltagi&qu&kao2016 allowed for individual-specific coefficients. We believe the parsimony of group specification in our approach may improve efficiency and achieve a faster convergence rate. We compare our method with baltagi&qu&kao2016 in the simulation to confirm this.
In other work, pauwels&chan&griffoli2012 proposed a testing procedure for the presence of structural breaks that allows breaks to occur to only a proportion of individuals. Our approach complements pauwels&chan&griffoli2012 by estimating the breakpoints and slope coefficients, and providing insights into which particular individuals are (un)affected by the breaks. Compared with individual time series estimation, we make use of cross-section variation and thus are able to estimate consecutive structural breaks. For their part, schnatter&kaufmann considered clustering multiple time series using the finite-mixture method in a Bayesian framework. They assume that the group memberships, characterized by the belonging probabilities, either depend on some exogenous observables through a specific (e.g., logistic) function or are equal to the group size. In contrast, we model group memberships via discrete indicators and allow them to be fully unrestricted. Finally, our work is also closely related to studies utilizing shrinkage estimators to estimate structural breaks (e.g. lee&seo&shin2016, qian&su2016, cheng&liao&schorfheide2016, li&qian&su). However, none of these studies considered heterogeneous panels with possibly group-specific structural breaks.
In this section, we describe the model setting. We consider a linear panel data model with heterogeneous and time-varying coefficients. The heterogeneity is restricted to have a grouped pattern, and structural breaks characterize the time-varying nature of the coefficients. The discussion of models with individual fixed effects is deferred to Section (ref).
Suppose that we have panel data $\{ \{ y_{it}, x_{it} \}_{t=1}^T\}_{i=1}^N$, where $y_{it}$ is a scalar dependent variable and $x_{it}$ is a $k\times 1$ vector of regressors, typically including the first element being 1. As usual, $t$ and $i$ denote time period and observational unit, respectively. The number of cross-sectional observations is $N$ and the length of the time series is $T$. We consider the following linear panel data model:
where $\epsilon_{it}$ is the error term with zero mean. The unknown coefficient $\beta_{i,t}$ is heterogeneous across individuals and changes over time. We put a structure on $\beta_{i,t}$, namely a grouped pattern of heterogeneity and structural breaks, to make it estimable and this also facilitates interpretation of the model.
We assume that $\beta_{i,t}$ is group specific. Suppose that observational units can be divided into $G$ groups. Let $\mathbb{G} = \{1, \dots ,G\}$ be the set of groups where $g_i \in \mathbb{G}$ indicates the group membership of unit $i$. Units in the same group share the same time-varying coefficient $\beta_{g,t}$, where $g \in \mathbb{G}$. Our model can be rewritten as
We use this representation of our model hereafter. Imposing a grouped pattern offers a sensible and convenient way to model individual heterogeneity because it allows us to capture heterogeneity in a flexible way while keeping the model parsimonious, permitting us to take advantage of cross-sectional variation in break and coefficient estimation bai2010. The group membership structure $\{g_i\}_{i=1}^N$ is unknown and to be estimated.
We also assume that for each group $g$, the time-varying pattern of coefficients $\{\beta_{g,1},\ldots,\beta_{g,T}\}$ can be characterized by structural breaks. For each group, there are $m_g$ breaks and $\mathcal{T}_{m_g, g} = \{T_{g,1}, \dots, T_{g,m_g} \}$ denotes a set of break dates, where both the number of breaks $m_g$ and break dates $\mathcal{T}_{m_g,g}$ are group specific and unknown, and we estimate them from the data. The value of coefficient $\beta_{g,t}$ changes only at a break date and remains the same in the period between any two break dates. Let $\alpha_{g,j}$, $j= 1, \dots, m_g$ be the value of coefficients until the $j$-th break date and $\alpha_{g,m_g+1}$ be the value of coefficients in the last period:
where we define $T_{g,0} =1$ and $T_{g,m_g+1} = T+1$. Note that we allow consecutive structural breaks to occur in two adjacent periods because we make use of cross-sectional observations via grouping, and this clearly includes the end-of-sample breaks andrews2003 as a special case. The true number of breaks for each group, $m_g^0$, is permitted to increase as $T$ increases, and the minimum break size ($\min_{g \in \mathbb{G}, 1 \le j \le m_g^0+1} \| \alpha_{g, j+1} - \alpha_{g, j}\|$) is allowed to contract to zero as $(N,T)\to\infty$.
Our model of heterogeneous breaks permits identification of the number of groups and breaks under regularity conditions in the sense that 1) an ignored break cannot be incorporated by increasing the number of groups and; 2) distinct groups cannot be modeled as a homogeneous pool with structural breaks. The identification of groups separately from breaks is relatively straightforward because we cannot incorporate structural breaks by increasing the number of groups. To illustrate the identification of structural breaks, we consider a simple case in which the model includes only an intercept $\beta_t$ with one break, and there are two groups which differ only in break dates, $T_{1,1} \neq T_{2,1}$: $$ y_{it}=\beta_{1} I(t<T_{j,1})+ \beta_{2} I(t\geq T_{j,1}) + \epsilon_{it},\quad \textrm{for $i$ such that $g_i = j$ and}\ j = 1,2. $$ These two groups cannot be treated as one group with two breaks because observations between $T_{1,1}$ and $T_{2,1}$ correspond to distinct intercepts for the two groups. Next, we consider two groups with a break at a common date but of different sizes. This clearly leads to the separation of two groups we cannot merge because the coefficients are heterogeneous in at least one of the regimes. We generalize these arguments to cases with heterogeneous breaks at different points and of distinct sizes and to cases with multiple breaks. Therefore, we cannot accommodate the underestimation of the number of groups by increasing the number of breaks as long as groups are separated.
This model is very general. If the parameter $\beta_{g,t}$ is constant over time, the model reduces to that considered in su&shi&phillips2016 and lin&ng2012. If $\beta_{g_i,t}$ is homogeneous over individuals, it then reduces to panel models with common structural breaks as in qian&su2016 and baltagi&kao&liu2017. It also includes the grouped fixed effect model by bonhomme&manresa2015 as a special case where only the intercept is allowed to have a time-varying grouped pattern.
In this section, we explain our estimation method and consider its asymptotic properties. Our estimation method is a hybrid of GFE by bonhomme&manresa2015 and AGFL by qian&su2016.
We first introduce some notation. Let $\beta$ be the vector stacking all $\beta_{g,t}$ such that $\beta = (\beta_{1,1}' , \dots, \beta_{1,T}', \beta_{2,1}', \dots, \beta_{G,T}')$. Let $\mathcal{B} \subset \mathbb{R}^k$ be the parameter space for each $\beta_{g,t}$. The parameter space for $\beta$ is $\mathcal{B}^{GT}$. Let $\gamma$ be the vector of $g_i$s such that $\gamma = \{g_1, \dots, g_N\}$. Note that $\mathbb{G}^N$ is the parameter space for $\gamma$.
We estimate $(\beta, \gamma)$ by minimizing a penalized least squares objective function. We employ an iterative procedure in which we iterate the estimation of $\gamma$ by minimizing the sum of squared errors for each unit and the estimation of $\beta$ by applying the AGFL to each group. We state our estimation method assuming that $G$ is known in this section and discuss how to select the number of groups $G$ in Section (ref).
We propose to estimate model (ref) by minimizing the following penalized least squares objective function:
The second term on the right-hand side of (ref) is the penalty term, where $\lambda$ is a tuning parameter whose choice is discussed in Section (ref), and $\dot w_{g,t}$ is a data-driven weight defined by
with $\kappa$ being a user specific constant and $ \dot \beta $ being a preliminary estimate of $ \beta$. We can use the GFE-type estimate for $\dot \beta$ as follows:
The resulting coefficient and group membership estimates are both consistent (see Theorem (ref) and Appendix (ref)), although the coefficient estimates vary for each period because the least squares objective function in (ref) does not include the penalty term.\footnote{Note that, strictly speaking, these estimators are not covered by bonhomme&manresa2015 because the coefficients are heterogeneous and change over time.}
Despite the fact that the estimated group memberships and slope coefficients are jointly obtained by solving the optimization problem (ref), we can better understand this objective function by conditioning. Given the coefficients, the estimate of $g_i$ ($i$'s group membership) is a group that yields the smallest sum of squared residuals for $i$. Note that the second term in (ref) does not depend on $g_i$ and does not have a direct effect on the estimation of $g_i$ (of course, it indirectly affects the estimation through the coefficients; see the more detailed discussion following Algorithm (ref)).
Next, given the group membership structure, our estimation problem is the same as applying the AGFL by qian&su2016 to each group. The penalty term in (ref) is proportional to the norm of the difference between the coefficients in adjacent periods. As is well known in the Lasso literature, this type of penalty term (called $L_1$ penalty) has a sparsity property and provides us estimates with the properties that $\hat \beta_{g,t} = \hat \beta_{g,t-1}$ for some $g$ and $t$. Estimated break dates are periods at which $\hat \beta_{g,t} - \hat \beta_{g,t-1} \neq 0$. Let $\hat{\mathcal{T}}_{\hat{m}_g,g}$ denote the set of estimated break dates for group $g$ such that $\hat{\mathcal{T}}_{\hat{m}_g,g} = \{ t \in \{ 2, \dots, T \} \mid \hat \beta_{g,t} - \hat \beta_{g,t-1} \neq 0\}$, then we can estimate the number of breaks for group $g$ by the cardinality of $\hat{\mathcal{T}}_{\hat{m}_g,g}$.
The consistency of preliminary estimates $\dot{\beta}$ plays an important role in break detection as it yields appropriates weights ($\dot w_{g,t}$'s). Note that if $\beta_{g,t} - \beta_{g,t-1} =0$, then $\dot \beta_{g,t} - \dot \beta_{g,t-1} $ is likely to be close to zero and $\dot w_{g,t}$ is likely to be large, resulting in a heavy penalty. This in turn allows us to achieve the consistent estimation of break dates. The use of consistent estimators as initial values in our iterative algorithm given below facilitates the convergence. However, the preliminary coefficient estimate fails to capture the structural breaks, and as a result, the group and coefficient estimates produced by the GFE-type objective function (ref) are also less accurate than those produced by (ref).
To obtain the minimizer in (ref), we propose to use the following iterative algorithm.
Step 1 applies AGFL by qian&su2016 for each estimated group. Step 2 updates the group membership based on the least squares objective function. We assign each unit to one of the groups that gives the smallest sum of squared residuals. The penalty term does not contribute to the estimation of group membership directly because it does not depend on $i$. However, it does indirectly improve the group estimation of (ref) in finite samples, because it forces us to pool the time points between two breaks. Such pooling obviously makes use of more observations in estimating the slope coefficients, leading to more precise coefficient estimates and to a better group structure estimate.
We choose the preliminary GFE-type estimate of the grouped pattern as the initial grouping $\gamma^{(0)}$, namely $\gamma^{(0)}=\dot{\gamma}$. The estimation of $(\dot \beta, \dot \gamma)$ may be implemented by an algorithm similar to Algorithm (ref), in which the objective function in (ref) is replaced by $\sum_{i=1}^N \sum_{t=1}^T (y_{it} - x_{it} ' \beta_{g^{(s)}_i ,t} )^2 /(N T)$. Given $\dot{\gamma}$ is consistent (see Appendix (ref)), this makes the convergence fast. Note that although individual-invariant regressors are in principle allowed in model (ref) as long as there are periods without a structural break, this preliminary GFE-type estimation prevents including individual-invariant regressors (e.g., some global variables that are common to all countries) when $x_{it}$ includes a constant term, because these regressors are multicollinear with the constant. A possible solution is to replace the adaptive Lasso penalty by the SCAD penalty fan&li2001, which does not require preliminary estimates, yet has similar properties to the adaptive Lasso.
As a preliminary consistent estimate is used as an initialization, this algorithm converges quickly, and the value of the objective function does not increase over the iterations. To implement GFE estimation of (ref), we draw a large number of initial values at random and select the estimate that minimizes the objective function. The main computational effort involves trialing a large number of starting values for the preliminary GFE estimates, and the computation time increases linearly with this number.\footnote{In our simulation experiment with $(N,T)$=(100,40), one estimation based on 100 starting values takes roughly 1.4 seconds, and that based on 1,000 starting values takes roughly 12.5 seconds of CPU time.} This algorithm works well when the number of groups is not large and there are no outliers. When the number of groups is large, a more robust difference-of-convex functions programming chu2017 can be considered.
We derive the asymptotic properties of GAGFL. We first show that the difference between GAGFL and AGFL applied to each group under known group memberships is asymptotically small. As a result, the break dates can be estimated consistently by our method. The asymptotic distribution of the coefficient estimator is the same as the least squares estimator under known group memberships and known break dates. In this section, we use the following notation. Let $M$ denote a generic universal constant. Parameters with superscript 0 are the true values.
We make the following assumptions to derive the asymptotic results. The first set of assumptions (Assumptions (ref)--(ref)) is for GFE-type estimation. These assumptions are similar to those used in bonhomme&manresa2015, but the difference arises because we consider more general models.
Assumption (ref).(ref) requires that the parameter space of slope coefficients is compact, as in most econometric literature. Assumption (ref).(ref) states that the regressors are exogenous. Note that it does not exclude the cases with $E (\epsilon_{it} x_{is}) \neq 0$ for $t \neq s$, and thus predetermined regressors are allowed. Assumptions (ref).(ref) and (ref).(ref) restrict the magnitude of the variability and the degree of dependence in the data. For example, when the data are i.i.d. over time, $\epsilon_{it}$ and $x_{it}$ are independent, and $\epsilon_{it} x_{it} $ has fourth-order moments, then these two assumptions are satisfied. Assumption (ref).(ref) imposes a condition on the fourth-order moment of $x_{it}$.
Assumption (ref) is an identification condition. Assumption (ref).(ref) provides an identification condition for the coefficients. Roughly speaking, this assumption states that there is no multicollinearity problem in any group structure. Assumption (ref).(ref) provides an identification condition for group membership and may be called the “group separation condition.” $ D_{g\tilde g i}$ is a measure of the distance between predicted values of $y_{it}$ under group $g$ and $\tilde g$ for unit $i$ and the conditions state that it is bounded away from zero.
Assumption (ref) is used to bound the maximum clustering error by placing restrictions on the tail behavior of the variables and the dependence structure. For example, Assumption (ref).(ref) may be violated when the tail of the distribution of $x_{i t}$ is so heavy that it does not possess a finite fourth moment. Another form of violation occurs when $x_{i t}$ exhibits a unit root. Assumption (ref).(ref) requires that the mixing coefficient $\alpha [t]$ decays exponentially fast. This assumption enables us to apply exponential inequalities in bonhomme&manresa2015 (which is based on rio2017asymptotic).
To examine the impact of the estimation error in the group structure on the coefficient estimates, we define
Note that $\mathring \beta$ is the estimator of $\beta$ when the group memberships (i.e., $\gamma^0$) are known. Denote $N_g$ as the number of units in group $g$, i.e. $N_g = \sum_{i=1}^N \mathbf{1} \{ g_i^0 = g\}$ for $g \in \mathbb{G}$. With the assumptions stated above, Lemma (ref) states that the impact of the estimation error in the group structure is limited, and thus the difference between $\hat \beta$ and $\mathring \beta$ is small.
Note that $\delta$ in the theorem can be arbitrarily large, and we obtain this “super-consistency” result because we model heterogeneity via discrete grouped pattern with the number of groups being fixed and finite. This also enables us to detect the breaks even when the true group memberships are unknown.
Next, we show that our method detects breaks correctly. Recall that $m_g$ denotes the number of breaks for group $g$ and $m_g^0$ is the true number of breaks for this group. $\mathcal{T}_{m_g,g} = \{ T_{g,1}, \dots, T_{g ,m} \}$ denotes a set of break dates and $\mathcal{T}_{m_g^0 ,g}^0 = \{T_{g,1}^0, \dots, T_{g, m_g^0}^0\}$ is the set of true break dates. $\mathcal{T}_{m_g^0, g}^{0c}$ denotes the complement of $\mathcal{T}_{m_g^0 ,g}^0$, representing the set of time with no breaks. Note that $\beta_{g,t}^0 - \beta_{g, t-1}^0 =0$ if $t \in \mathcal{T}_{m_g^0, g}^{0c}$. Let $\alpha_{g,1} = \beta_{g,1}$ and $\alpha_{g,j} = \beta_{g, T_{g,j-1}}$ for $j = 2, \dots, m_g^0+1$. Let $ J_{\min} = \min_{g \in \mathbb{G}, 1 \le j \le m_g^0} \| \alpha_{g, j+1} - \alpha_{g, j}\| $ be the minimum break size. We make the following assumption that is similar to Assumption A.2 in qian&su2016.
Assumption (ref).(ref) allows the total number of breaks of all groups $\sum_{g\in\mathbb{G}}m_g^0$ to diverge to infinity at a slow rate. Assumptions (ref).(ref) and (ref).(ref) are used to show the consistent detection of breaks by AGFL. Assumption (ref).(ref) states that break sizes are not very small so that all breaks can be identified. Note that it still allows the break size to tend to zero as long as the rate is slower than $\sqrt{N}$.
In the following, we show that our method can consistently select the number of breaks and estimate the breakpoints even when the group membership is unknown. Consistent break detection in AGFL requires that the weights be adaptive and correctly estimated, which further requires the preliminary estimates used to construct the weights to be consistent. Hence, we first show that the preliminary estimator is $\sqrt{N}$-consistent.
This theorem guarantees that the weight in the penalty, $\dot w_{g,t}$, becomes large when there is no break (i.e., $\beta_{g,t}-\beta_{g,t-1}$), which further enables us to detect periods without breaks. Let $\hat \theta_{g,t} = \hat \beta_{g,t} - \hat \beta_{g,t-1}$.
This theorem states that our method consistently spots dates on which there is no break. This theorem requires that $N / T^{\delta} \to 0$ for some $\delta$. As $\delta$ is arbitrary, this condition is satisfied as long as $N$ is of geometric order $T$ but it does not hold if $N$ is of exponential order $T$. Note that the probability in this theorem is unconditional in the sense that it does not depend on knowing the true group membership. We also establish selection consistency as below.
This theorem states that the number of breaks and break dates are estimated consistently as $N$ and $T$ go to infinity at an appropriate rate. As above, the probabilities are both unconditional, implying that the selection consistency is established even when we do not know the group structure.
Fundamentally, Theorems (ref) and (ref) both result from the super-consistency of group membership estimation and consistency of AGFL applied to each group with known group memberships. The error introduced by estimating group memberships is negligible (Lemma (ref)) and GAGFL is asymptotically equivalent to applying AGFL to each group with known group structure. The AGFL estimator applied to each true group is pointwise consistent with the convergence rate of $1/\sqrt{N}$ (see Lemma (ref) in Appendix (ref) or Theorem 3.2(ii) of qian&su2016). We can thus obtain the pointwise consistency of our GAGFL estimator under unknown group structure.
Lastly, we present the asymptotic distribution of $\hat \beta$. Put simply, this is the same as that of the least squares estimator for each group within each no-break regime with known group memberships and break dates. To state the theorem, we introduce the following notation. Denote $I_{g,j}$ as the number of time periods between $T_{g,j}^0$ and $T_{g,j+1}^0$ for group $g$, and denote $I_{\min}$ as the minimum number of periods for which there is no break, namely $ I_{\min} = \min_{g \in \mathbb{G}, 1 \le j \le m_g^0+1} \| T_{g,j}^0 - T_{g,j-1}^0 \| $. Let $ \Sigma_{x,g,j} = \operatorname*{plim}_{T, N_g \to \infty} 1/(N_gI_{g, j})\sum_{g_i^0 = g} \sum_{t=T_{g,j}^0}^{T_{g,j+1}^0 -1} x_{it} x_{it}', $ and
Let $\Omega_{g,g}$ be a $(m_g^0+1) k \times (m_g^0+1) k $ matrix whose $(j,j')$-th $k\times k$ block is $\Omega_{g, g, j', j'}$, and $\Sigma_{x,g}$ be a $ (m_g^0 +1) k \times (m_g^0+1) k$ block diagonal matrix whose $t$-th diagonal block is $\Sigma_{x,g ,j}$. Further, let $\Omega$ be a $\sum_{g=1}^G (m_g^0 +1) k \times \sum_{g=1}^G (m_g^0 +1) k $ matrix whose $(g,h)$-th $ (m_g^0+1)k\times (m_h^0+1)k$ block is $\Omega_{g, h}$, and $\Sigma_{x}$ be a $ \sum_{g=1}^G (m_g^0 +1) k \times \sum_{g=1}^G (m_g^0 +1) k $ block diagonal matrix whose $g$-th diagonal block is $\Sigma_{x,g}$. We require the following extra assumptions for the asymptotic distribution.
Assumption (ref) simply states that the standard assumptions for least squares are satisfied for each group and each span of periods between two breaks. The purpose of introducing $D$ is to analyze a finite dimensional vector of linear combinations of elements of $\hat \alpha$. Note that the dimension of $\hat \alpha$ is potentially increasing in $T$ and its asymptotic distribution is hard to discuss. We instead examine finite dimensional objects. Assumption (ref) is a technical assumption on the break sizes.
The following shows that the GAGFL slope coefficient estimator of group $g$ in regime $j$ is consistent and asymptotically normal with the convergence rate of $\sqrt{N_gI_{g,j}}$.
We point out that these theoretical results also hold if we replace $\hat \beta$ with $\beta^{(0)}$ from Algorithm (ref). This is because the initial group assignment in the algorithm, $\dot \gamma$, is consistent. Note that for $\hat \beta$, the consistency of $\dot \gamma$ is not fundamental (although it does provide rapid numerical convergence). What is crucial is the consistency of $\dot \beta $, because it guarantees that $\dot w_{g,t}$'s have appropriate orders, which makes it possible to detect breaks consistently. In this paper, we focus on $\hat \beta$ because we find that the iterative estimator $\hat \beta_{g,t}$ typically outperforms $\beta^{(0)}$ in finite samples due to possible refinement.
To implement GAGFL, we need to choose the tuning parameter in AGFL for each group, and at the same time specify the number of groups. We discuss these two issues in turn.
First, to select the tuning parameter $\lambda$ of the Lasso penalty in AFGL, we follow qian&su2016 to minimize the following information criterion (IC): $$ IC(\lambda)=\frac{1}{NT}\sum_{j=1}^{m+1}\sum_{t=T_j-1+1}^{T_j}\sum_{i=1}^N(y_{it}-x'_{it}\hat{\alpha}_{g_i,j})^2+\rho_{NT}k(m_{\lambda}+1), $$ where $\hat{\alpha}_{g_i,j}$ is the post-Lasso estimate of the coefficients for each of $g$ and $j$, $m_\lambda$ is the number of breaks associated with the tuning parameter $\lambda$, $\rho_{NT}$ determines the amount of penalty on the number of breaks, and we choose $\rho_{NT}=c\ln(NT)/\sqrt{NT}$ with $c=0.05$ following qian&su2016. We verify via simulation and applications that the performance of our method is robust to the choice of $c$ as long as it lies in a reasonable range.
Next, we discuss how to choose the number of groups. In this paper, we focus on using an information criterion to determine the number of groups, $G$.\footnote{The literature also suggests a Lagrange multiplier test lu&su2017 to determine $G$.} We follow bonhomme&manresa2015 to consider the following Bayesian information criterion (BIC):
where $\hat{\sigma}^2$ is a scaling parameter and can be obtained by an estimate of the variance of $\epsilon_{it}$, and $n_p(G)$ is the total number of estimated coefficients. The information criterion represents a tradeoff between model fitness and the number of parameters. One caution is that this tradeoff is more complicated in our model because increasing $G$ does not always lead to a larger number of parameters and a better fit in our case. It is actually possible that a larger value of $G$ results in fewer breaks in each group, so that the total number of parameters decreases and the model fit worsens. Nevertheless, we find that more parameters correspond to a better fit in our simulation experiments. Furthermore, we can show that with an appropriate choice of the tuning parameter in the penalty terms, the model that coincides with the data generation process produces the lowest information criterion value.
In principle, it is possible to replace $\hat{\alpha}_{g_i,j}$ in (ref) by initial estimates of coefficients ($\dot \beta$) that are fully time varying as defined in (ref). An advantage of using initial estimates to compute the BIC is that both the number of parameters and model fit are monotonically increasing in $G$. However, the disadvantage is that less efficient coefficient estimates may result in less accurate selection of $G$ in finite samples. Simulation results (provided in Section S.1.1 of the supplement) indicate that the BIC based on the final estimates outperforms that based on the initial estimates.
In this section, we consider a model with individual-specific fixed effects. The estimation uses first-differenced data and requires separate theoretical analysis. Here, we provide the assumptions and summarize the theoretical properties, while the proofs and additional details are in the supplement.
In addition to the time-varying grouped fixed effect, we allow an additive time-invariant individual fixed effect $\alpha_i$ arbitrarily correlated with the regressors. We consider the following model:
The individual fixed effects, $\alpha_i$, can be eliminated by first differencing:
where $\Delta y_{it}=y_{it}-y_{it-1}$ for $i=1,\ldots,N$ and $t=2,\ldots, T$. The model is estimated by applying the GAGFL method to the transformed data. {Note that in our setting, it is possible to identify fully time-varying coefficients $\beta_{g,t}$, even though there are only $T-1$ time periods available because of differencing. To see this, suppose that group membership is known. Then, (ref) can be understood as a cross-sectional regression model with dependent variable $\Delta y_{it}$ and regressors $(x_{it}', x_{i,t-1}')'$ for a given $t$. Both $\beta_{g,t}$ and $\beta_{g,t-1}$ can be estimated from the cross-sectional regression. The actual identification argument can be more involved because group memberships are unknown. Nevertheless, as shown by this rough argument, the availability of cross-sectional information is the key to identifying the coefficients even when they are fully time varying.}
The GAGFL estimator for fixed effects models is defined as
The results of Section (ref) are not directly applicable to fixed effects models. For example, both $\beta_{g,t}$ and $\beta_{g,t-1}$ enter the equation and we cannot analyze two time periods separately even when there is a break between these two periods. Nevertheless, we can extend our analysis to fixed effects models. We first state the modified assumptions needed.
This assumption can be implied by Assumption (ref) except for (ref).(ref). However, we continue to make this assumption to facilitate our theory.
Under these new sets of assumptions\footnote{ Assumption (ref) can be implied by Assumptions (ref) and (ref). Nevertheless, we introduce this assumption to directly apply the results of qian&su2016 for AGFL in the presence of individual fixed effects.}, we show that the GAGFL estimator applied to first-differenced models has asymptotic properties similar to those in the case of level models without individual-specific intercepts. As above, we first show that the estimation error in the group structure has a limited impact on the estimation of the coefficients. Let
Then, we can state exactly the same theorems as Theorems (ref) and (ref), but under Assumptions (ref), (ref), (ref), (ref), and (ref), regarding the break detection. Finally, under the correct estimation of break and group memberships, we can similarly show that the coefficient estimates of GAGFL are asymptotically equivalent to the least squares estimates under the true breakpoints and group memberships. This asymptotic equivalence and the proofs of theorems in this section are in the supplement.
There are various directions in which to extend our proposed method. We present those extensions that are relevant in practice and analyzed relatively easily. In particular, we consider models in which a subset of coefficients are fully time varying and models in which a subset of coefficients are homogeneous and/or time invariant. Note that the two extensions considered here are indeed special cases of our model. However, our estimation method needs modification to incorporate their special features.
\paragraph{Models with fully time-varying coefficients and time fixed effects.}
Our model has variants with fully time-varying coefficients in special cases, but the estimation needs modification. Without loss of generality, assume that the first $k_1$ explanatory variables $x_{1,it}$ have fully time-varying coefficients, $\beta_{1, g}$, and the remaining $k-k_1$ covariates $x_{2,it}$ correspond to coefficients, $\beta_{2, g, t}$, which may exhibit structural breaks. Then, we can estimate the model by minimizing the following adjusted objective function
The resulting coefficient estimates $\hat\beta_{g,t}$ satisfy (ref) in Lemma (ref) under the same set of assumptions. Similar theorems such as Theorems (ref) and (ref) hold for the $k-k_1$ subset of the estimated coefficients $\hat{\beta}_{2,g,t}$ with slightly modified assumptions. A complete theoretical analysis of this algorithm is in the supplement, in which we also apply it to an application. An important special case is a model with group-specific time fixed effects, i.e.
where $\lambda_{g,t}$ is a group-specific time effect for group $g$ and period $t$, $z_{it}$ is the explanatory variables, and $\delta_{g_i,t}$ is the associated coefficients characterized by structural breaks. We assume that $\lambda_{g,t}$ changes at every period, and we do not penalize a change in $\lambda_{g,t}$.
\paragraph{Models with partially homogeneous and/or time-invariant coefficients.}
Our model can also incorporate the case of partially time-invariant coefficients. Without loss of generality, assume that the first $k_1$ explanatory variables $x_{1,it}$ have time-invariant coefficients, $\beta_{1, g}$, and the remaining $k-k_1$ covariates $x_{2,it}$ correspond to time-varying coefficients, $\beta_{2, g, t}$. Then, we can estimate the model by minimizing the following adjusted objective function
The iterative algorithm remains the same except that the AGFL only penalizes a part of the coefficient vector $\beta_{2,g,t}$.
If the coefficients of $x_{1,it}$ are not only time invariant but also homogeneous, then the objective function in this case becomes
To find a minimizer for this objective function, we adjust the iterative algorithm by separating the estimation of $\beta_1$ and $\beta_{2,g,t}$ such that a penalized minimization is applied only to $\beta_{2,g,t}$ and an estimate of $\beta_1$ is updated using all the observations rather than observations in a group.
This section evaluates the finite sample performance of the proposed GAGFL estimator. In particular, we investigate if GAGFL can correctly classify units and whether it can effectively detect structural breaks. We also examine the importance of taking heterogeneity of structural breaks into account.
We study four data generation processes differing in the distribution of errors, the specification of fixed effects, and the inclusion of lagged dependent variables.
DGP.1 is our benchmark case. DGP.2 exhibits serial correlation in the errors. DGP.3 includes individual fixed effects, and thus first-differenced data are used as in (ref). DGP.4 follows a dynamic panel model. For each DGP, we consider two cross-sectional sample sizes, $N=(50,100)$, and three lengths of time series, $T=(10, 20, 40)$. In total, we have six combinations of cross-sectional sample size and length of time series.
We estimate the model by GAGFL. We also examine the performance of the penalized least squares (PLS) by qian&su2016 ignoring heterogeneity, and that of common break detection in heterogeneous panels by baltagi&qu&kao2016 (hereafter BFK). Note that PLS is equivalent to GAGFL with $G=1$.
For GAGFL, the number of groups is chosen by minimizing the BIC defined in Section (ref). For both GAGFL and PLS, we follow qian&su2016 for the computational method and choice of tuning parameters. To estimate the breaks for a given grouped pattern, we employ the block-coordinate descent algorithm for solving the PLS. The tuning parameter $\lambda$ is selected by minimizing the information criterion in the interval of $[0.01,100]$, where the upper bound leads to breaks in all time points while the lower bound leads to no breaks. To construct the weights $\{\dot{\omega}_{g,t}\}$, we set $\kappa=2$ following the adaptive Lasso literature. In DGP.3, first differencing is employed to eliminate individual fixed effects for GAGFL and PLS (see Section (ref)).
BFK detects breaks by minimizing the sum of squared residuals over distinct breakpoints, where the residuals are from the individual time series estimation. The method is applied to detect multiple common breaks occurring for all individual units, although the slope coefficients are allowed to be individual specific. To implement this method, we need to specify the number of breaks, and we set the number of breaks equal to three, the true number for the pooled data.\footnote{Groups 1 and 2 share a break at the same time $\lfloor5T/6\rfloor$, and each of them experiences another break at different times $\lfloor T/2\rfloor$ and $\lfloor T/3\rfloor$, respectively. Thus, there are in total three breaks for the pooled data.} We employ the approach discussed by baltagi&qu&kao2016 (see also bai97et,bai2010) that estimates multiple breakpoints sequentially. As the BFK method estimates individual time series separately, no transformation is needed under DGP.3. In this case, breaks are only allowed in the slope coefficient $\beta$, but not in the intercept.
We evaluate the performance of the proposed method from five different perspectives: selecting the right number of groups, clustering, determining the number of breaks, the break date estimates, and the coefficient estimates.
First, we evaluate the performance of the BIC in determining the number of groups by computing the empirical probability of selecting a particular number. Second, we measure the clustering accuracy by the average of the misclassification frequency ($\widehat{g}_i\neq g^0_i$) across replications. Let $I(\cdot)$ be the indicator function. The misclassification frequency (MF) is the ratio of misclassified units to the total number of units, i.e. $ \textrm{MF} = 1/N\sum_{i=1}^N I(\widehat{g}_i\neq g^0_i). $
The third and fourth criteria concern break estimation. The third criterion is the average frequency of correctly estimating the number of breaks. The fourth criterion is the average Hausdorff error of break date estimates. This measure is also used by qian&su2016. The Hausdorff error is the Hausdorff distance (HD) between the estimated break dates and the true set of dates, i.e. $$ \textrm{HD}(\widehat{T}_{g,\widehat{m}}^0,T_{g,m^0}^0)\equiv\max\{\mathcal{D}(\widehat{T}_{g,\widehat{m}}^0,T_{g,m^0}^0), \mathcal{D}(T_{g,m^0}^0,\widehat{T}_{g,\widehat{m}}^0)\}, $$ where $\mathcal{D}(A,B)\equiv\sup_{b\in B}\inf_{a\in A}|a-b|$ for any set $A$ and $B$. We report the Hausdorff error multiplied by 100 and divided by $T$ for each group, i.e. $100\times \textrm{HD}(\widehat{T}_{g,\widehat{m}}^0,T_{g,m^0}^0)/T$, averaged across the replications.
Lastly, we evaluate the accuracy of the coefficient estimates using their root mean squared error (RMSE) and the coverage probability of the two-sided nominal 95% confidence interval. We compute the overall root mean square error for all units at each time as $$ \textrm{RMSE}(\widehat{\beta}_{it}) = \sqrt{\frac{1}{NT}\sum_{i=1}^N\sum_{t=1}^T(\widehat{\beta}_{it}-\beta_{it})^2}. $$ The coverage probability is computed as $$ \textrm{Coverage}(\widehat{\beta}_{it}) = \frac{1}{NT}\sum_{i=1}^N\sum_{t=1}^TI(\widehat{\beta}_{it}-1.96\widehat{\sigma}_{\beta,it}\leq \beta_{it}\leq\widehat{\beta}_{it}+1.96\widehat{\sigma}_{\beta,it}), $$ where $\widehat{\sigma}_{\beta,it}$ is the estimated standard deviation of $\widehat{\beta}_{it}$. We average all evaluation measures across 1,000 replications.
As the implementation of our method requires the specification of the number of groups, we first examine how well the BIC proposed in Section (ref) performs in selecting this number. To implement the BIC, we choose the scaling parameter $\hat{\sigma}^2$ to be the sum of squared residuals obtained by plugging in the estimates from the homogeneous panel, i.e. $G=1$.\footnote{While this may not yield a consistent estimate of the variance of residuals, the main role of $\hat{\sigma}^2$ is to scale the penalty term, such that it is invariant to the variation of the data. In fact, we find that choosing $\hat{\sigma}^2$ as the sum of squared residuals obtained under $G_{max}$ groups can lead to rather unstable results that vary across the specifications of $G_{max}$.}
Table (ref) presents the empirical probability that a particular number of groups, ranging from $G=1$ to 5, is selected according to the BIC. Recall that the true number is three. The results indicate that the BIC can identify the correct group size with a high probability. In DGP.1 and DGP.2, the BIC selected the correct number of groups in more than 97% of the cases, even when $T=10$ and $\sigma_\epsilon=0.75$. In DGP.3, the correct selection rate is also high when $\sigma_\epsilon=0.5$, although it is relatively low when $\sigma_\epsilon=0.75$ and the sample sizes are small, possibly because of first differencing. Introducing lagged dependent variables deteriorates performance slightly, although the correct selection rate still exceeds 86%, even in the worst case. In Section S.1.1 of the supplement, we examine the BIC constructed from initial estimates. This continues to work well but generally performs more poorly than that defined in (ref) using the final estimates.
The previous section suggests that the information criteria are useful in determining the number of groups. Given the true number of groups, we now examine the accuracy of clustering, break detection, and coefficient estimation.
We first investigate the clustering accuracy of the proposed estimator. Table (ref) presents the averaged misclustering frequencies. This shows that the proposed method can correctly classify a large proportion of units. The misclassification frequency generally reduces with $T$ but not with $N$. In DGP.1 and DGP.2, GAGFL can correctly classify more than 95% of individuals even with a relatively short time series ($T=10$) and large errors ($\sigma_\epsilon=0.75$). Allowing individual fixed effects in DGP.3 leads to slightly less accurate clustering. Nevertheless, GAGFL still manages to limit the misclassification frequency to less than 7% for a short $T$, and the rate drops quickly as $T$ increases. Introducing dynamic effects as in DGP.4 scarcely affects the clustering, and more than 96% of individuals can be correctly classified. These results demonstrate that GAGFL can effectively capture grouped patterns of heterogeneity.
We now examine the accuracy of structural break detection. We compare the estimated number of breaks for GAGFL and PLS, and compare the accuracy of the estimated breakpoints for GAGFL, PLS, and BFK. Recall that the latter two approaches both assume common breaks to all units at the same time, although BFK allows individually heterogeneous coefficients.
Table (ref) presents the average frequencies of correctly estimating the number of breaks for each group in our method, and the same frequency for all units in PLS.\footnote{BFK is not compared here because the true number of breaks is assigned to estimate the breakpoints.} Because the panel is heterogeneous, the frequency of PLS is calculated by comparing the estimated number of breaks with the true number of breaks in the pooled data: three in our case. This shows that when errors are of a moderate size ($\sigma_\epsilon=0.5$), our method almost perfectly detects the correct number of breaks (with a correct detection frequency of more than 97%) except in DGP.3. With individual fixed effects in DGP.3, the frequency is about 78% in the small sample with $N=50$ and $T=10$, but this frequency quickly improves to more than 90% when $T=20$ or when $N=100$. In the case with relatively large errors ($\sigma_\epsilon=0.75$), GAGFL still works well, although less accurately than in the case of $\sigma_\epsilon=0.5$. When the sample size is small ($N=50$ and $T=10$), it can correctly estimate the number of breaks in at least 78% of the cases in DGP.1, 92% in DGP.2, 22% in DGP.3, and 83% in DGP.4. Unreported results suggested that when GAGFL fails to detect the correct number of breaks, it typically overestimates them. Again, the correct detection frequency increases quickly with $N$ and $T$. For example, it quickly reaches 90% on average in DGP.3 when $N=100$ and $T=40$. In contrast, the correct detection frequency remains less than 40% for PLS in DGP.1 and DGP.2, and even lower in DGP.3 and DPG.4. Moreover, the frequency does not seem to increase with the sample size.
As another measure of break estimation accuracy in Table (ref) we report the Hausdorff errors between the true break dates and those estimated by GAGFL, PLS, and BFK, conditional on the correction estimation of the number of breaks. We can see that the Hausdorff errors of GAGFL are much smaller than those of PLS and BFK in all cases. Although BFK allows for heterogeneous coefficients, it has difficulty detecting breaks in small samples because heterogeneous breaks become less visible by treating them as common, which is also noted by baltagi&qu&kao2016. These results jointly illustrate the importance of accounting for heterogeneity in breaks, and show that our method can detect the number of breaks correctly and identify the breakpoints precisely, even when the error variance is large.
Finally, we compare the accuracy of the coefficient estimates obtained from GAGFL, PLS, and BFK. Table (ref) presents the average RMSE and coverage probability of the coefficient estimates of the three methods across 1,000 replications.\footnote{To compute the average statistics, we map the estimated coefficients in a regime-group pair to each $(i,t)$ observation based on the estimated group memberships and break dates, and then compute the average RMSE and coverage probability of the entire coefficient vector. These average statistics facilitate comparison with the estimates produced by the two competing methods, PLS and BFK, because they do not produce group-specific estimates. Moreover, the average statistics also reflect the accuracy of membership and break date estimation.} To conserve space, for DGP.4 we report only the results for the coefficient of the lagged dependent variable. In general, we can see that the RMSE of GAGFL is much smaller than that of PLS and BFK. Increasing the sample size reduces the RMSE and improves the coverage probability of GAGFL. In contrast, the RMSE of PLS does not improve as the sample size increases, and its coverage probability remains low. BFK has better coverage probabilities than PLS (although still lower than GAGFL), but at the cost of a much larger RMSE caused by its particularly large standard deviation. This is because BFK uses individual time series estimation, which can be rather inefficient in finite samples.
We consider four extensions of the simulation designs. This section briefly presents the designs and results of the extensions, while additional details are in Sections S.1.2--1.5 of the supplement.
First, we consider the cases where the regressors are group dependent. Allowing for group-dependent regressors does not affect the clustering accuracy, no matter whether the group structure of regressors coincides with the structure of coefficients, except in DGP.3. In DGP.3 with group-dependent regressors, the misclassification frequency tends to be higher than in the case of independent regressors.
Second, we consider the case where breaks are small and the groups are more alike. In this case, clustering and break detection become more difficult, and we detect higher misclassification frequency and less accurate estimates of break dates. However, performance rapidly improves as the sample size increases.
To better understand how cross-sectional variation plays a role in affecting the performance of GAGFL, we consider the case of unequally sized groups, say $N_1:N_2:N_3 = 0.1 : 0.8 : 0.1$, such that some groups contain very few units. In Table (ref), we summarize the average misclassification frequency, the frequency of correct estimation of the number of breaks, and the Hausdorff error of the break date estimates of GAGFL in the presence of small groups.\footnote{To conserve space, Table (ref) provides the results in the leading case of DGP.3, while the more full results are in the supplement.} Although small group size results in less accurate estimation due to a lack of cross-sectional variation, we find that increasing the sample size improves the classification accuracy. In particular, accuracy improves significantly as $N$ increases because cross-sectional variation in the small groups is increased, which improves the coefficient estimates and further indirectly improves classification.
Finally, we consider the case where the break dates are close. We generate breaks in the first group that occur at $\lfloor T/2\rfloor$ and $\lfloor 2T/3\rfloor$, and in the second group at $\lfloor T/3\rfloor $ and $\lfloor T/2\rfloor$, where $\lfloor\cdot\rfloor$ takes the integer part. Now, the difference between the two break dates in both groups is just $\lfloor T/6\rfloor$, i.e. $1$ when $T=10$, $3$ when $T=20$, and $6$ when $T=40$. For the third group, the slope coefficient is stable without a break. The performance of GAGFL is summarized in the bottom panel of Table (ref). This shows that shrinking the interval between the two breaks barely affects the misclassification frequency and the accuracy of break estimation. This indicates that as long as there are sufficient individual units in each group, we can consistently estimate the slope coefficients (and further the groups and breaks), even when the two breaks are consecutive.
We conclude this section by commenting on the iterative feature of the algorithm. On average, the algorithm takes two to three steps to converge in our simulation. The number of iterations increases when $\sigma_\epsilon$ is large but decreases as $T$ increases. It also takes more steps to converge when we generate the data with closer break dates, a small degree of group heterogeneity and breaks, or small groups containing only a few units. Comparing the performance of the iterative and non-iterative estimates, we find the misclassification frequency of the iterative estimates consistently lower than the rate of the non-iterative estimates, suggesting that iteration improves clustering performance, sometimes greatly. Consequently, iterative estimates produce lower RMSEs and higher coverage probabilities than non-iterative estimates.
We apply the proposed GAGFL estimator to revisit the relationship between democracy and income. This analysis was first undertaken by acemoglu&johnson&robinson&yared2008 in a standard panel framework, and revisited by bonhomme&manresa2015 using the GFE approach. As suggested by acemoglu&johnson&robinson&yared2008 and BM, countries experience economic and democratic development at certain “critical junctures”, such as the end of feudalism, the industrialization age, or the process of colonization. These junctures may not only influence the average degree of democracy (captured by the intercept of the democracy--income regression), but also the degree of democracy persistence and the relationship between income and democracy (captured by the slope coefficients). This implies that the effect of income on democracy can shift discontinuously at those junctures. Moreover, the paths of economic and political development also diverge across countries. For example, industrialization may affect a proportion of Western countries, but not Asian countries, at least to a lesser extent or at a later stage. Hence, the democracy--income relationship is likely to shift at different historical junctures across countries, and even for countries affected by the same event, the extent of the effect can be distinct.
We revisit the democracy--income relation by allowing heterogeneous structural breaks in a regression of democracy, measured by the Freedom House index, on its lagged value and the lagged value of the logarithm of GDP per capita, namely
One essential difference from BM's specification is that we allow the slope coefficients to have a grouped pattern and possible structural breaks. Note also that in our specification, $\alpha_{g_i t } $ may exhibit structural breaks and may not change at every time period. We compare our results with those of BM.\footnote{lu&su2017 also studied the same empirical data with group-specific slope coefficients, but they considered individual and time-specific two-way fixed effect panel models.}
We follow BM in using a balanced subsample of the data of acemoglu&johnson&robinson&yared2008 that contains 90 countries for seven periods (five-year frequency over 1970--2000). To ensure similar variation across variables for grouping, we standardize democracy and income by subtracting their overall means and then divide them by the overall standard deviations, respectively.
To implement GAGFL, we choose $\lambda_{\max} = 50$, which results in zero breaks in all groups, and $\lambda_{\min}=0.001$, which results in six breaks in all groups. We then choose the tuning parameter $\lambda$ by searching on the interval $[\lambda_{\min},\lambda_{\max}]$ with 200 evenly-distributed logarithmic grids. We follow qian&su2016 and use the same information criterion as in the simulation with $\rho_{N T} = 0.05\ln(NT)/\sqrt{NT}$ for determining the number of breaks. To determine the number of groups, we employ the BIC as in the simulation. We let the number of groups vary from 1 to 10, and the minimum information criterion corresponds to four groups (i.e. $G=4$), which coincides with the group number specification of BM.
When we estimate (ref) with four groups, we have three groups with structural changes in intercepts and slope coefficients and one group without any break, and name them Groups 4.1--4.4. Figure (ref) displays the estimated grouped pattern when $G=4$, and the estimates of coefficients and structural breaks are reported in Table (ref). We provide the post-Lasso estimates with their corresponding standard errors. The supplement presents the confidence sets of group membership using the method in DzemskiOkui2018. The unit-wise confidence sets suggest that the group membership estimates are reasonably accurate, although the joint sets are wide.
Group 4.1 contains no breaks. A significant feature of this group is strong dynamic persistence in the political system, while the income effect is relatively weak. This group contains a large number of the “high-democracy” group countries in BM such as the US and Switzerland. Nevertheless, the group also includes many countries in BM's “low-democracy” group, such as China, Iran, Cameroon, Guinea and several other African countries. This may appear counterintuitive at first glance, but there are both developed and developing countries in this group whose unit-wise confidence sets (see the supplement) are singleton, suggesting that this result is not an artifact of statistical error. Further examination reveals that this group structure is mainly driven by the persistence of democracy. Although democracy levels within Group 4.1 vary, a common feature is that their political systems are highly persistent, as reflected in the high value of $\theta_{1}$ (= 0.8655). This strong persistency separates countries of this group from other groups containing breaks in their democracy level. When we set $G=6$ corresponding to the second-lowest value of BIC, this group will be segmented (see the supplement). A large discrepancy in democracy and income levels across countries also explains the small intercept and weak income effect, $\theta_{2}$.
Group 4.2 is characterized by one structural break occurring at an early stage of the sample period (the mid-late 1970s). The average democracy level, persistence, and income effect all increase after this break. This group consists mainly of “early transition” countries in BM such as Greece, Nepal, Spain, and Thailand. It also contains a few “low-democracy” countries, including Burundi, Republic of Congo, Togo, etc. A further examination of these low-democracy countries shows that they all have a break in their democracy level at the beginning of the sample period, even though the level remains low in general. In this sense, the break and coefficient estimates produced by GAGFL well capture the political junctures in the mid-late 1970s of these countries.
Group 4.3 also exhibits one structural break, but in the mid- to late 1980s. Again, the average democracy level, persistence, and income effect increase after the break, but to a lesser extent when compared with Group 4.2. A large proportion of this group are “late transition” countries whose democratic reforms occurred at a later stage. It also includes countries that transitioned to highly democratic countries at multiple junctures (e.g., South Korea, Argentina, and Brazil) and those whose democracy level fluctuated greatly in the latter period (i.e., Sierra Leone and Nigeria).
Finally, Group 4.4 exhibits two breaks, one at the beginning of the 1980s and one at the beginning of the 1990s. After each break, dynamic democratic persistence decreases from 0.4796 to 0.3879 and then to 0.0860, and the income effect first strengthens from 0.3019 to 0.8562, and then weakens with an insignificant estimate $-0.0551$. Most countries in this group experienced changes/fluctuations in democracy in the middle and late phase of the period. However, none has a singleton confidence set (see the supplement) and thus it might be difficult to make a strong argument about this group.
We also examine the democracy--income relation under alternative specifications. First, we consider the results with $G=6$ which is the second-lowest value of BIC. In this case, the group with stable coefficients under $G=4$, i.e. Group 4.1, is further divided. Recall that Group 4.1 contains countries with different levels of democracy. When we set $G=6$, the lowly and highly democratic countries are clearly separated (e.g., China and the US), although their slope coefficients are all stable over time. Second, we consider the setup with a fully time-varying intercept, where only slope coefficients are penalized but the $\alpha_{g_i,t}$s are not. The results are generally in line with those of regime-specific estimation (in Table (ref)). These indicate that average democracy does not fluctuate in every time period, but exhibits at most one discontinuous break during the sample period. Moreover, the fully time-varying intercepts are much less significant than those reported in Table (ref). Thus, in this particular application, the fully time-varying estimation may be rather inefficient compared to regime-specific estimation. We provide the complete results of these alternative specifications in the supplement.
In general, our procedure detects a large degree of heterogeneity in the income effect of democracy. Although compatible to some extent, our grouping differs from bonhomme&manresa2015. This is mainly because we form groups based on the whole coefficient vector and allow structural change. The grouped pattern thus not only reflects the magnitude difference of the coefficient estimates, but also the heterogeneity in structural breaks. This allows us to provide additional insights not captured by GFE. In particular, we show that the persistence of the democracy level and the income effect both exhibit significant structural breaks in several of the groups, and that the breakpoints and the size of breaks differ markedly across groups.
We also apply GAGFL to another application to study the determinants of the cross-country differences in savings behavior as in su&shi&phillips2016 and the analysis is in the supplement. We find more heterogeneity than su&shi&phillips2016 in the sense that countries differ in their transition points, the level and stability of their savings rates, and the effects of various determinants. One group of Asian countries features one structural break coinciding with the start of the 1997 Asian financial crisis, while the crisis hardly affects the remaining countries, and they thus have relatively stable coefficients. This application again confirms the importance of incorporating heterogeneous structural breaks.
Structural change in the relationships between variables often characterizes large panels with long time series because of important events, such as financial crises, technological progress, economic transition, etc. The effect of these events often differs across individuals. Some can be greatly influenced, while others in the sample may not be affected at all by events. Even for those affected, the impacts can be heterogeneous. Importantly, failing to incorporate heterogeneity in structural breaks leads to incorrect breakpoint detection and imprecise coefficient estimates.
In this paper, we propose a new model and estimation procedure for panel data with heterogeneous structural breaks. We model individual heterogeneity via a latent grouped pattern such that individuals within the group share the same regression coefficients. For each group, we allow common structural breaks in the coefficients, while the number of breaks, the breakpoints, and/or break sizes can differ across groups. We develop a hybrid procedure of the GFE estimator and AGFL to estimate the model. With the proposed procedure, we can 1) consistently estimate the latent group membership structure, 2) automatically determine the number of breaks and consistently estimate the breakpoints for each group, and 3) consistently estimate the regression coefficients with group-specific structural breaks.
An interesting extension of our work would be to allow the grouped pattern to change at disjoint time intervals. Such a flexible framework is desirable because the impact of significant events may completely change the economic mechanisms governing individuals and their group memberships. For example, the impact of the most recent global financial crisis was so enormous that it reshaped economies and thus changed the group structure of the world. One of the main difficulties is to restrict appropriately the structural breaks, such that there are sufficient observations within each regime to estimate the grouped pattern. This presents a challenge for future research.