Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
88,182 characters · 18 sections · 93 citation commands
Unobserved Grouped Heteroskedasticity and Fixed Effects
Agents that exhibit unobserved differences pose a significant challenge in econometric models by confounding the true relationships between variables of interest. Panel data provides a solution where it is possible to make intuitive, though restrictive, assumptions on the form of heterogeneity such as with two way fixed effects or interactive fixed effects (see wooldridge:2010 and bai:2009, respectively). Another approach is group fixed effects where the unobservable is a discrete variable that labels individuals into groups and allows within-group parameters that vary across time (bm:2015). These methods require estimating the heterogeneity, which could be harmed by ignoring important features of the marginal distribution of confounding latent variables. I focus on the case of a discrete latent variable that induces group heterogeneous fixed effects and variances in the model, which is written as a linear grouped fixed effects model with unobserved groupwise heteroskedasticity. The heterogeneity that fits this model should generally be closer to the true groupings, rather than one estimated solely focused on the group fixed effects.
A linear model with group fixed effects is written as
for $i = 1,\dots, N$ and $t = 1, \dots, T$ where the variable $g_i = 1, \dots, G$ are discrete and take a known or estimable $G$ number of values, are unobserved and may be arbitrarily correlated with the exogenous variables $x_{it}$. The group membership variables $g_i$ and group-specific time effects $\alpha_{g_i t}$ are exogenous and estimated from the data. It is assumed that there are a relatively small number of groups and group membership of individuals does not change over time and that individuals in the same group $g$ follow the same time path of effects $\alpha_{gt}$. This model and a least squares estimation procedure were introduced in bm:2015 and several alternative estimators have since been proposed (chetverikov:2022,mugnierPWD:2022).
The key identification assumption shared by all of these proposals is that the groups are separated, specifically that groups can be distinguished by the group fixed effects $\alpha_{g} = (\alpha_{gt})_t$ for all $g = 1,\dots,G$. From an error component perspective: $v_{it} = y_{it} - x_{it}'\theta = \alpha_{g_i t} + u_{it}$ has the group fixed effect term $\alpha_{g t}$ as the conditional mean given $g_i = g$ so that identification requires the conditional means of the error components to be separated. Therefore, in order to learn group memberships, one must consider partitions or groupings of $v_{i} = y_{i} - x_{i}\theta = (y_{it} - x_{it}'\theta)_t$ into $G$ groups that uniquely satisfy the moment conditions implied by this model. Mathematically, if the group means $\alpha_{g}$ and parameters $\theta$ were known, this amounts to sorting individuals $i=1,\dots,N$ into a group where their $v_{i}$ is closest to the corresponding mean $\alpha_{g}$. However, the error component may also exhibit heteroskedasticity with respect to the unobserved groups, which is not a feature that has received attention in the literature.
Unobserved group heteroskedasticity can be expressed as
where $\sigma_{g} > 0$ for any $g=1,\dots,G$. Using worker's earnings as an example, this expression implies different variability in wages across groups, which could arise when workers belong to different ability tiers. Workers that are part of a highly skilled group may display more years of schooling, experience, and consistently high earnings, with little variation. On the other hand, workers in a lower skilled, but improving ability group may show consistently lower earnings at first, but may start to demand for compensation as they improve with a larger variance of success, perhaps by being promoted or remaining as a subordinate.
This paper extends the group fixed effects model to incorporate group heteroskedasticity. Simulation evidence suggests that ignoring this feature may lead to poor estimation of group memberships and severe finite-sample bias and large standard errors of parameter estimators. Intuitively, assigning according to just sample group means ignores the possibility that outlier individuals from a high variance group appears closer to another group, which may result in an erroneous assignment. I introduce the “weighted grouped fixed-effects” (WGFE) estimator, where the name serves as an analogy to the estimators connection to the GFE estimator and resemblance to the weighted least squares estimator. The WGFE estimator simultaneously estimates the parameters and finds the optimal grouping by organizing units with outcomes net the effect of covariates into groups that they are most similar to, which is determined by nearness to the mean and variance of each group using a modified Mahalanobis distance. As with GFE estimation, without covariates this reduces to a clustering problem, but usual $k$-means algorithms do not apply since they are known to perform poorly when clusters have heterogeneous variances. Using the proposed distance function, I make use of recent advances in latent factor discovery by yang:2022 to develop an algorithm adapted to handle heterogeneous group variance. In fact, the $k$-means algorithm itself is a special case of their algorithm when the clusters are uniformly distributed and have the homogeneous variance. The WGFE estimator is equivalent to the GSR estimator of boot:2022, where the WGFE estimator arises from concentrating out the group variance parameters in their setup. These estimators for the GFE model with group heteroskedasticity were found independently.
I provide an asymptotic theory for the WGFE estimator where I allow the number of individuals $N$ and time periods $T$ to approach infinity simultaneously. This estimator shares the property with the GFE estimator that $N$ can grow much faster than $T$ as opposed to the individual fixed-effects estimator which suffers from bias of order $1/T$ as $N/T$ converges to a constant, also known as the incidental parameter problem (nickell:1981,arellano-hahn:2007). The WGFE and GFE estimators are both consistent and asymptotically normal as $N/T^\nu$ tends to zero for some $\nu >0$ provided the groups are well separated and errors $u_{it}$ satisfy tail and dependence conditions.
For identification WGFE requires a stronger notion of separability between groups in order to ensure consistency of group assignments. The inclusion of second moment information must exclude population groupings where differences in group variances are relatively larger than differences in group means in a way that there is significant overlapping support. In other words, groups must be mean-separated as a function of relative differences in the group variances. If this is achieved, the quality of group assignments improves faster than GFE as $T$ approaches infinity and the WGFE estimator is expected to outperform estimators based solely on first moment classification when $T$ is small, which is a common property across panel data sets. GFE remains consistent for large $N$ and $T$ so the performance gap should close for large and long panels. It is also shown that the WGFE estimator is asymptotically equivalent to GFE when there is no group heteroskedasticity and the groups are equally weighted and that the WGFE objective function is bounded by the GFE objective, suggesting a test of group homoskedasticity. Future work involving an arbitrary number of features may involve these separability type assumptions and connections to GFE so this may be a starting point to provide tools to study discrete heterogeneity in depth.
I showcase the WGFE estimator in two illustrations. The first is an application to study the effect of national income on democracy, adding to the insights found by bm:2015. Following gleditsch-ward:2006 and ahlquist-wibbels:2012, it is empirically observed that there are grouped patterns of democratization that transition in both time and space. I find evidence where groupings determined by the WGFE show contiguous regions on similar time paths using a subset of a balanced panel of 90 countries from 1970-2000. Our findings do not contradict that of bm:2015 or acemoglu:2008 that income has little or no effect on democratization, however it strengthens the evidence and equivalently the correlation between the group variable, income and democracy. This supports the assertion that shared historical factors may have set countries on different time paths such as the formation of religious institutions and economic and defensive pacts; see huntington:1991.
The second empirical example revisits the question of the effects of unionization on wages using data from the biennial Panel Study of Income Dynamics between 2001 and 2019. Unobserved ability is widely accepted as an important omitted variable that would explain employers selecting more able workers while facing higher labor costs stemming from a union contract. Many studies find the effect of unionization is overestimated without accounting for individual-specific unobserved ability (freeman:1984,robinson:1989,card:1996). Using WGFE, workers across various unobserved ability groups experience different rates of expansion and deterioration of worker bargaining power and real earnings across the sample period. Using traditional fixed effects simply shows an aggregate decline, which keeps the different unobserved segments hidden.
Section (ref) presents the issues of identification with different notions of separability, the estimator and computational approach. Section (ref) develops the asymptotic theory of the infeasible WGFE estimator and the WGFE estimator including consistency of parameter estimators and sample group assignments, and asymptotic normality. In Section (ref), I compare the WGFE and GFE estimators in a simulation study and use our approach to study the two empirical applications on unionization-wages and income-democracy.
Modeling with latent group heterogeneity in panel models has received attention as a useful alternative to standard fixed effects approaches. sun:2005 estimate a multinomial logistic regression with unobserved groupings, but known number of groups $G$ via maximum likelihood. lin:2012 in one of their proposed methods develop a $k$-means algorithm to group individuals based on the regression estimate of one of $G$ groups. When (multiple) groupings are known, bester:2016 propose an estimator for a random effects model that assumes individuals share a fixed-effect at some (undetermined) level of grouping. hahn:2010 consider a class of game theoretic models and show that the incidental parameter problem vanishes quickly as $N$ approaches infinity since the number of equilibria are predicted to be finite, hence the support of the fixed-effects is finite. In the case of multidimensional heterogeneity, cheng:2019 consider individuals who may belong to different groups modeled by several latent variables. None of these papers consider time-varying heterogeneity. mugnier:2022 introduce the single index nonlinear GFE model and estimation that does allow for time-varying heterogeneity, but not the possibility of group heteroskedasticity.
When the number of groups are unknown it must be inferred from the data. bm:2015 follow bai:2002 and bai:2009 by using a Bayesian information criteria (BIC) and setting a maximum number of groups. Recently there have been a number of papers that use penalization to classify and estimate parameters. su:2016 propose the C-Lasso related to group Lasso that classifies and shrinks individual coefficients to the unobserbed group coefficients. lu:2017 continue with a testing procedure to determine the number of groups based on C-Lasso. su:2018 extend the C-Lasso to dynamic panel data models with interactive fixed-effects and cross sectional dependence and propose a BIC to estimate the number of factors and groups. ando:2016 consider grouped factor structure and penalization for coefficient estimates and their $C_p$-type criteria to estimate the number of groups and number of factors in each group. mehrabani:2022 take these penalization approaches further by allowing the number of groups to diverge along with the sample dimensions. mugnierPWD:2022 propose a simple nuclear-norm regularized estimator that estimates the number of groups along with parameters.
An equivalent estimator to WGFE was found independently by boot:2022, where their identification conditions are similar in requiring strong separability. zhang:2020 propose a Bayesian approach that treats the number of groups as a parameter to be estimated and use a Dirichlet process prior that allows for infinitely many groups and allows for group heteroskedasticty, but require strictly normal errors while WGFE imposes no such structure. Along these lines, kim:2019 give evidence that in the normal errors case that accounting for group heteroskedasticity in the income and democracy application shows the correlation between the estimate group variable and income is larger than predicted by GFE estimation and that the total number of groups estimated might depend on the specification of group heteroskedasticity, which coincides with the prediction using WGFE.
Finite mixture models are closely related to our approach, see the textbook by mclachlan:2004 for a comprehensive review. In these models the data is not required to label individuals into groups and group membership is instead estimated. It is typical to assume that the group distributions belong to the same family e.g. Gaussian mixture models so group heteroskedasticity is modeled. For identifiability of models with covariates there needs to be restrictions the values covariates can take on and on the interaction between covariates and the latent variable; see kasahara:2009 for a result on identification of discrete choice models with panel data where both the group distribution and choice probabilities are nonparametrically specified. I allow covariates to be arbitrarily correlated with the latent variable as in typical fixed-effects and leave group membership probabilities unrestricted. Mixture models can incorporate fixed-effects, see deb:2013 who provide a solution to the incidental parameter problem for Gaussian and Poisson families. heckman:1984 apply them to duration models where they take the distribution of the unobservable nonparametrically, while imposing a parametric assumption on the group distributions. A related setting is Markov-Switching models for time series, which can be seen as a mixture model, see fruhwirth:2006. Related to switching models, $k$-means clustering is similar to greedy approximations of step function signals, which can be thought of as a estimating a switching model of intercepts with no covariates or parametric assumptions on the error, see rivero:2023. In this setting, the user does not need to specify the number of groups and instead could be determined via penalty methods.
The Expectation Maximization (EM) algorithm commonly used in the maximum likelihood estimation of mixture models can be seen as a clustering algorithm, see redner:1984. In fact, in the case of a Gaussian mixture, the Mahalanobis distance is a key part of the algorithm, which suggests a connection to our clustering algorithm which also incorporates second moment information. Indeed GFE is also the maximizer of the pseudo-likelihood of a Gaussian mixture model, where the mixing probabilities are individual-specific and unrestricted e.g. independent of covariates.
An inspiration for considering unobserved group heteroskedasticity is the the $k$-means problem of assigning a number of individuals into a finite number of groups; see (maccqueen:1967,lloyd:1982,forgy:1965). The $k$-means criterion is a least squares criterion, which implicitly assumes that the groups are separated enough in mean and are of identical variance. When the data violates any of these assumptions, then group assignment at the sample level may be compromised. See the recent survey ahmed:2020 on the performance of the standard $k$-means algorithm in these contexts.
The estimator is interestingly tied to optimal measure transport theory (villani:2009, santambrogio:2015), specifically the Wasserstein (Kantorovich) distance between probability distributions; see kolouri:2017 for more on Wasserstein distance and its applications. The Wasserstein distance function defines a distance function between probability measures and so permits a Fr\'{e}chet mean of distributions that is also itself a distribution known as the Wasserstein barycenter, see ac:2011. The WGFE estimator for a model without covariates can be seen as the minimization problem of optimally forming two parameter (location-scale) groups such that these groups have a barycenter with minimum variance over other possible groupings and barycenters. See galichon:2018 for an introduction to optimal transport for economists.
The slope parameters are contained in the vector $\theta\in\Theta \subset \mathbb{R}^p$ and group-specific time-effects $\alpha_{gt}\in \mathcal{A} \subset \mathbb{R}$ for any $g=1,\dots,G$ and $t = 1,\dots,T$ and $\alpha \in \mathcal{A}^{GT}$. The group assignment parameters are categorical: $g_i\in\{1,\dots,G\}$ for every $i = 1,\dots,N$ and the space of all possible partitions using the $g_i$ is denoted as $\gamma = (g_1, \dots, g_N) \in \{1,\dots,G\}^N =\Gamma_G^N$. The covariates $x_{it}$ can contain lagged outcomes along with strictly exogenous regressors. The covariates are also allowed to be arbitrarily correlated to the time-effects $\alpha_{gt}$. Note the absence of time subscript implies a vector of time profiles or time series for the individual. Denote $ x \mapsto \left\lVertx\right\rVert$ as the Standard Euclidean norm of finite-dimensional vector $x$.
Group fixed effects models assume that group membership is unknown and estimated from data. Identification of group memberships in these class of models require that groups are “separated” according to some criteria. It is sufficient to assume that group fixed effects for any pair of groups are different in the mean squared sense. This means that groups can be distinguished by the comparison of their group-specific time series of effects.
Denote the true parameters with a superscript “0”. Formally for any $(g,\widetilde{g}) \in \{1,\dots,G\}^2$ such that $g\neq \widetilde{g}$:
This is a separability assumption and implies a rule for determining the group of any individual using their time series information. For $i\in\{1,\dots,N\}$ and supposing that $g_i^0 = g$, I will show that the mean squared distance between $v_i = y_i - x_i\theta^0$ and $\alpha_{h}^0$ for any $h \in \{1,\dots,G\}$ is minimized when $h = g$. Recall the true model (ref) and consider the following difference we want to show is bounded away from zero:
where the $o_p(1)$ results from an exogeneity condition applied on $\alpha_g^0 - \alpha_h^0$ that will be made in more detail in Section (ref). The converse is also true-- for any $i\in\{1,\dots,N\}$, if we assume the strict inequality holds for some $g$, it uniquely minimizes the distance between $v_i$ and the group fixed effect $\alpha_{g}$ of group $g$. Therefore, since groups are separated, an individual must belong to a group $g$ if their observed $v_i = y_i - x_i\theta$ is closest to $\alpha_{g}$ compared to the other group fixed effects, thus identifying group assignments for each individual provided parameters are known.
Geometrically this requires that clusters or conditional distributions given groupings have different means. This does not restrict overlap between groups, notably how a particular group variance causes large intersections with other groups. Moreover even when clusters have many features that are heterogeneous this condition delivers identification of group memberships.
The inclusion of group heteroskedasticity in the model brings more information, but when groups must be estimated it necessitates steeper requirements to detect them in the population. A natural criteria can be based on the Mahalanobis distance, which normalizes distances to the cluster mean based on the “dispersion” or noise of the group. Letting $\sigma_g^2$ denote the variance of any group $g$, a candidate rule may be written for all $i \in \{1,\dots,N\}$: $g_i^0 = g$ if and only if, for all $h \neq g$,
For this rule to be valid a stronger separability condition is required that involves both group fixed effects and variances. Following similar calculations on this inequality, we arrive at
where once again we would like this to be asymptotically bounded away from zero. A sufficient condition is the following “strong” separability of groups:
which explicitly requires that the model noise is finite, i.e., $\underset{{T\to\infty}}{\textnormal{plim}}\, \sum_{t=1}^T u_{it}^2 < \infty$ for all $i \in\{1,\dots,N\}$ and that there are no degenerate groups: $\sigma_f >0$ for all $f = 1,\dots,G$. In words, the difference between variances is scaled by the $F$-statistic between the noisiest group and the least. A model that has a large range between group variances therefore requires a large discrepancy between group fixed effects in order to identify groups. As an example with no covariates, with $T=2$, two normally distributed groups, fixing the mean-squared error of the two group fixed effects to 4 and one of the groups variances to 1, (ref) implies a valid range of variances for the other group in the interval $(0.2,2)$ where the groups would be identifiable. Figure (ref) displays a visual on how these groups may appear when the variable standard error is close to 2 and 0.2. According to this example the constraint of strong separability does not seem to place a large restriction on the types of clusters WGFE may handle.
Taking the rule (ref) to data may perform poorly by creating and then favoring groupings with enormous variance. To counteract this, a simple additive penalty may be imposed that vanishes asymptotically with larger $T$ so that strong separability is preserved. Following yang:2022, I propose the following correction:
so that noisy groups are penalized and less noisy groups are assigned primarily by closeness to group mean.
This reasoning is easily generalizable to groups with non diagonal heterogeneous covariances $\mathbf{\Sigma}_{g_i^0}$ (see Appendix (ref)) and potentially to incorporate other moments of the distributions. An extension up to $K$ moments might involve a distance function that takes arguments as the first $K$ moments and then a separability condition that requires increasingly steeper restrictions on how far the group means must be as a function of the other $K-1$ moments.
The weighted grouped fixed-effects (WGFE) estimator for model (ref) is defined as
where
and $P_g^i(\gamma) = \mathds{1}\{g_i = g\}$ and $P_g^\gamma = N^{-1}\sum_{i=1}^N P_{g}^i$.
The objective function is a departure from the pooled least squares criteria of bm:2015 and to $k$-means when covariates are absent. It is instead a minimization of a weighted transformation of within-group least squares that is shown to coincide to choosing $\widehat{g}_i = \widehat{g}_i(\theta,\alpha)$ that satisfies (ref) for every $i \in \{1,\dots,N\}$ in yang:2022. It is a version of the GSR estimator independently developed by boot:2022 which concentrates out the optimal group variance estimator. Because of this results using these estimators will appear similar.
Let $\sigma_{g_i}(\theta,\alpha,\gamma) = \sqrt{\widehat{Q}_{g_i}(\theta,\alpha,\gamma) }$. With a fixed grouping scheme $\gamma^*$ the first-order conditions with respect to $\theta$ and $\alpha$ are
for all $ t=1,\dots,T$, $g = 1,\dots, G$. With these groupings, $\theta^*$ is the solution to the nonlinear equation
in the system of equations with
where $\overline{y}_{{g} t}$ and $\overline{x}_{{g} t}$ are the within-group averages of the respective variables according to grouping $\gamma^*$. The form of (ref) displays similarities to weighted least squares and solvable by the Newton-Raphson method (dennis:1996) or by fixed-point iteration.
The partitioning problem of assigning $N$ units into $G < N$ groups is NP-hard (aloise:2008,dasgupta:2008) and, with additional parameters such as $\theta$, rules out exhaustive search \footnote{clausen:2003, brusco:2006 and aloise:2012 for some global and exact methods that are prohibitive in my application.}. I follow convention and propose a heuristic algorithm based on Lloyd's ($k$-means) algorithm (lloyd:1982) that after initializing will alternate between assignment of individuals into groups and then updating parameters based on those assignments.
This is a variance augmented $k$-means algorithm similar to gradient descent where part of the set of parameters take on discrete values and each iteration improves on the function value until a local minimum is reached (see forgy:1965, maccqueen:1967 and lloyd:1982). The algorithm converges quickly, however a complete application of Algorithm (ref) involves repeating it many times with a randomly drawn starting value and choosing the result with the smallest local minima. I emphasize that this is a heuristic method sensitive to initialization and has no guarantees of finding the global minimizer. Heuristic algorithms perform well enough; see brusco:2007 and in particular the class of variable neighborhood search (VNS) algorithms (hansen:2010). These algorithms combine both stochastic initialization and deterministic search to comb the domain for the smallest local minimum. I adapt a VNS algorithm described in the Supplementary Appendix of bm:2015 and originally from pacheco:2003 to incorporate covariates and utilize it in simulations and the empirical applications since it can be more efficient at finding minima versus repeated applications of Algorithm (ref). See Appendix (ref) for a full description of the VNS Algorithm (ref). Despite using the VNS algorithm in this paper, I often refer to Algorithm (ref) as it provides intuition for identification and the asymptotic theory and still plays a role in the improved VNS algorithm.
The rule (ref) appears in the assignment step and serves as a balanced criteria to assign an individual into a group it is closest in group fixed effect relative to the noise in the group. While very noisy groups make the first term smaller, they are penalized by the second term preventing assignment solely based on the standardized distance. This is not the first clustering criteria incorporating second moment information, see zhao:2015. The Mahalanobis distance can be used, however too much normalization without the penalty term in (ref) may result in an assignment rule that favors the noisiest group and the clustering objective function becomes trivial, see krishnapuram:1999.
The GFE estimator solves the pooled least squares problem
and the assignment step of GFE estimation is given by
The estimators have a closed-form for given grouping $\widetilde{g}$:
and
The GFE criterion function weighs the residual variance the same across all groups while the WGFE criterion separates them by weighing them according to the size of groups. If groups had homogeneous variance and size, then WGFE and GFE estimation would be equivalent.
Furthermore, they are connected in a way economists might describe risky behavior. The GFE criterion is to the risk neutral agent while the WGFE is to the risk averse agent. The following proposition is a consequence of Jensen's inequality.
The square of the WGFE objective function is pointwise dominated by the GFE objective function in both the sample and population level. The inequality is sharp at the population level provided that errors are group homoskedastic and groups are uniform. Proposition (ref) along with asymptotic equivalence of WGFE and GFE estimation under group homoskedasticity suggests constructing a test of group homoskedasticity and uniformness based on the test statistic \[ \tau_{NT} = d_{NT}\left(\widehat{Q}^{GFE}\left(\widehat{\theta},\widehat{\alpha},\widehat{\gamma}\right) - \left[\widehat{Q}^{WGFE}\left(\widehat{\theta},\widehat{\alpha},\widehat{\gamma}\right)\right]^2\right) \geq 0 \] where it is evaluated at the WGFE estimator and $d_{NT} \to \infty$. Under the null that groups are homoskedastic and consistency of criterion functions and estimators, we should have $\tau_{NT} \to_p 0$ as $N,T$ approaches infinity. Indeed, in a sufficiently large sample if $\tau_{NT}$ is exceedingly far from zero we may reject group homoskedasticity and uniformity.
Consider the data generating process \[ y_{it} = x_{it}'\theta^0 + \alpha_{g_i^0 t}^0 + u_{it}, \] where $g_i^0 \in \{1,\dots,G\}$ denotes the true group membership for each individual $i$ and zero superscripts denote true values. Assume that the true number of groups is known so $G= G^0$. I assume throughout that we have a random sample $\{(y_{it}, x_{it})\}$ over $i$ and the unobserved $g_i^0$ are a random sample from a mass function $\mathbb{P}(g_i^0 = g)$ with support on its entire domain $\{1, \dots, G^0\}$. The triple $\{(y_{i}, x_{i},g_i^0)\}_{i =1}^N$ is regarded as a random sample from some joint distribution.
This section provides conditions that enable the WGFE estimators of group membership converge to their true values and use this to show that the WGFE parameter estimators are asymptotically equivalent to their infeasible version where the true group memberships are known. I establish conditions for which this infeasible estimator is consistent and asymptotically normal using standard theory of extremal estimation (whitney:1986) and use asymptotic equivalence of the estimators to connect these properties to the WGFE estimator. In this setting both $N$ and $T$ are allowed to approach infinity, but $T$ may grow much slower than $N$ to the effect of $N/T^\nu \to 0$ for some $\nu \gg 1$ as will be shown.
Let $\left(\widetilde{\theta}, \widetilde{\alpha}\right)$ be the infeasible version of the WGFE estimator where the true group memberships are known:
where
and $P_g^i = \mathds{1}\{g_i^0 = g\}$ and $P_g = N^{-1}\sum_{i=1}^N P_{g}^i$.
Establishing the asymptotic distribution of the GFE estimator follows a similar strategy since the infeasible version is simply a least squares estimator of a pooled regression with group dummy variables. However, ((ref)) is not a pooled least squares estimator so I must show this infeasible version of the estimator has the property of consistency and asymptotic normality under some conditions.
The proposed population criterion function is
where $P_g^0 = \mathbb{P}(g_i^0 = g)$ and
The first two assumptions are standard, while $(c)$ and ($d$) are non degeneracy conditions of each group indexed by $\{1,\dots,G\}$. I require all groups to have many observations in them and thus group variances are bounded away from zero. Assumption (ref)($e$) is a full rank condition analogous to those on design matrices for the basic linear models. A sufficient condition for $(e)$ is that for each $g=1,\dots,G$ and setting $\widetilde{X}_i = X_i - \mathbb{E}\left[X_i\vert g_i^0\right]$ as the stacked $x_{it} - \mathbb{E}\left[x_{it}\vert g_i\right]$ for each $i$, the matrix \[ \mathbb{E}\left[\widetilde{X}_i'\widetilde{X}_i\vert g_i^0 = g\right] \] has full rank. Since these matrices are symmetric, they are positive definite and so is their sum. Hence the minimum eigenvalue of the sum is positive. This is similar to fixed-effects identification where the within-transformation removes the time mean from each individual. Instead, this is at the group level which would require removing the time dependent group mean for which individual $i$ belongs.
Similar to fixed-effects estimation, the following consistency result does not require $T$ to be large, but large $T$ will be needed when groupings are unknown.
The result follows from the consistency of M-estimators due to the definition of $\left(\widetilde{\theta},\widetilde{\alpha}\right)$ as minimizers of $\widetilde{Q}$, uniform weak converge of $\widetilde{Q}$ to $Q$ and the property that the unique minimizer of $Q$ is the true parameter values.
With $\sqrt{N}$-consistency of the infeasible WGFE estimator, I turn to the asymptotic distribution of the estimator. Under a few more standard assumptions, the infeasible WGFE estimator is asymptotically normally distributed as $N,T \to \infty$. This follows from the standard asymptotic normality theorem for M-estimators. First, define the estimator for the group variance when groupings are known as $\widetilde{\sigma}_{g_i^0}^2(\widetilde{\theta},\widetilde{\alpha})$. Since the infeasible estimator is consistent then $\widetilde{\sigma}_{g_i^0}^2(\widetilde{\theta},\widetilde{\alpha}) \to_p \sigma_{g_i^0}^2$ as $N$ approaches infinity using a weak law of large numbers. The following assumptions are sufficient to characterize the asymptotic distribution of $\left(\widetilde{\theta},\widetilde{\alpha}\right)$.\\
Assumption (ref)($a$) allows covariates that are strictly exogenous and also time lagged outcomes. The following establishes the infeasible WGFE estimator is asymptotically normal.
The form of the asymptotic variance of the infeasible estimator is different than the pooled least squares case in that observations are weighted according to the noise in their group. It appears similar to weighted least squares however the weight is the standard deviation instead of the variance.
This section establishes conditions for the consistency of the WGFE estimator $\widehat{\theta}$ of $\theta^0$. Remarkably these assumptions are identical to those made in bm:2015 for consistency of the GFE estimator.\\[1cm]
To summarize from previous work, weak dependence conditions in the form of Assumptions (ref)($b,c,d$) are imposed. These assumptions are related to those made in stock:2002 and bai:2002 on large factor models. Assumption (ref) ($b$) allow for lagged outcomes or any predetermined regressor as covariates e.g. $\mathbb{E}\left[u_{it}\vert x_{it},x_{it-1},u_{it-1}\right] = 0$. Assumptions (ref) ($b$) and $(d)$ are conditions on time series dependence of the errors and covariates while Assumption (ref)($c$) limits cross-sectional dependence. Assumption (ref) ($e$) is akin to Assumption (ref) ($e$), so it is similar to a full rank condition found in other linear regression models. This sample version is stronger since it requires sufficient variation in the covariates across time and individuals within groups generated by any grouping scheme $\widetilde{\gamma}\in \Gamma_G^N$.
The proof largely follows that of bm:2015 with additional algebraic steps taken given the additional complexity of the WGFE criterion.
In this section I discuss the properties of the WGFE estimator for group membership in a large $N$ and $T$ setting. I refer to the assignment rules (ref) and $\eqref{gfe:assignment}$ as the WGFE and GFE assignments, respectively. I will first present the simple case of two groups for intuition and then will move to the general case where I provide conditions for any number of groups $G$ and with covariates.
Consider a simplification of the main model where there are no covariates, $G= 2$ and errors are independent, but drawn from different normal distributions depending on $g_i^0$ with variance $\sigma_{g_i^0}^2$. Formally, \[ y_{it} = \alpha_{g_i^0}^0 + u_{it}, \>\>\> u_{it} \sim N(0,\sigma_{g_i^0}^2), \>\>\> u_{it}\perp\!\!\!\!\perp u_{jt} \text{ for }i\neq j. \] Assume that $\sigma_1 \geq \sigma_2$. The probability of incorrectly assigning an individual into group 2 while they belong to group 1 using the WGFE estimator is
This expression on the left hand side is the distribution function of some generalized chi-squared distribution and has no closed-form; see davies:1980. Note that if $\sigma_1 = \sigma_2$ then this is the example for GFE estimation in bm:2015. In the case that $\alpha_{1}^0 = \alpha_{2 }^0$ are equal: \[ \mathbb{P}\left(\widehat{g}_i(\alpha^0)\right) = 2 \vert g_i^0 =1) = \mathbb{P} \left(\frac{\sum_{t=1}^T u_{it}^2 - T}{\sqrt{2T}} < \frac{\sigma_1\sigma_2 - T}{\sqrt{2T}} \right) \approx \Phi\left(\frac{\sigma_1\sigma_2 - T}{\sqrt{2T}}\right) \longrightarrow 0 \] as $T\to\infty$ where $\Phi$ is the standard normal cdf. However if $\sigma_2 >\sigma_1$ then this is no longer the case since this assignment rule can't distinguish between the smaller density with $\sigma_1$ and the larger density with $\sigma_2$ and always assigns to the larger density. Essentially the leading term is now negative and the inequality will reverse leading to
so despite the additional information in the assignment rule this demonstrates the need for separation of group fixed effects.
There is no closed form for the misclassification probability (ref) so comparisons of asymptotic performance of assignments is given numerically. In Figure (ref), group 1 is assumed to have a variance greater than or equal to group 2 in all the examples. In the top left plot we see again that GFE and WGFE assignment are equivalent when variances are identical. In the top right plot, when there is separation in both mean and variance, the probability of misclassification goes to zero at a faster rate for WGFE assignment. In the bottom plot there is weak separation in means and in variances and the WGFE significantly outperforms GFE assignment.
In the case that the true group variance is strictly less than the missclassified group variance, I run into a similar problem to (ref), however convergence is still possible. In Figure (ref), I fix $\sigma_1 = 1$, $\sigma_2 = 1.5$ and $\alpha_1^0 = 0$ and I take $\alpha_2^0 = 0.5,1,1.5,2$. When there is weak separation in the $\alpha$'s, misclassification is almost surely going to occur for large $T$ since the WGFE assignment can't distinguish between the two groups due to overlap and every point will be classified into the larger variance group. However, as there is more separation, classification improves and there will be convergence. Note that the GFE assignment also benefits from more separation, but the WGFE assignment benefits more in improving misclassification when the variance of the true group is smaller as it takes into account both first and second moment information.
In this example of two groups, there will be a group with a smaller variance and so the misclassification probability is susceptible to the issues in Figure (ref). Therefore to achieve desirable simultaneous convergence properties of assignment to both groups, there needs to be additional restrictions on how separated the groups are. In light of the figures and the mathematical expressions, there is a first order importance on mean separation between groups. While variance separation lends to advantages for the WGFE assignment, if the means aren't far enough then this advantage ceases and becomes an obstacle in convergence for one side of group assignments. This is precisely the requirement of strong separability (ref) for identification of groups in WGFE estimation.
Let $\eta > 0$ and define a neighborhood $\mathcal{N}_\eta$ around the true parameter values $\theta^0$ and $\alpha^0$ as the subset of $(\theta,\alpha)\subset \Theta \times \mathcal{A}^{GT}$ that satisfy $ \left\lVert\theta - \theta^0\right\rVert < \eta$ and $T^{-1}\sum_{t=1}^T (\alpha_{gt}^0 - \alpha_{gt})^2 < \eta$ for all $g = 1,\dots,G$. Define
and denote the discrete probability measure induced by a random partition $\gamma$ as $\lambda(\gamma) \in \mathring{\Delta}^G$ where $\mathring{\Delta}^G$ is the interior of the $G$ dimensional probability simplex. The form of the objective function (ref) prevents empty groups from being optimal hence the optimal grouping must lie in $\mathring{\Delta}^G$.
\phantom{a}
The assumptions made here do not deviate much from those in bm:2015 besides a strong separability assumption in part ($e$) that requires a specific bound depending on other features of the model. Note that is a stronger version of (ref) reflecting the fact that the group standard errors must also be estimated and it plays a role in showing the sample probability of misassigning individuals converges to zero. The denominator of $C_{g,\widetilde{g}}$ is bounded away from zero since the groupings are selected from the interior of the probability simplex.
In Assumptions (ref)($b$) and ($c$) I strengthen restrictions on the dependence and tail properties of the error, respectively. I assume the error is strongly mixing with faster-than-polynomial decay rate and with tails that also decay faster than any polynomial. The mixing property is stronger than ergodicity and says that, eventually, the process will forget its history. Mathematically, for any stochastic process $X_t$ on some probability space $(\Sigma,\mathcal{F},\mathbb{P})$ define the function $\alpha[s]$ as a strongly mixing coefficient: \[ \alpha[s] = \sup\{|\mathbb{P}(A\cap B) - \mathbb{P}(A)\mathbb{P}(B)|: t\geq0, A \in X_0^t, B \in X_{t+s}^\infty\} \] where $X_a^b$ denote the sub $\sigma$-algebra of $\mathcal{F}$ specified between times $a$ and $b$. If $\alpha[s] \to 0$ as $s\to\infty$, then $X_t$ is said to be a strongly mixing process. I also assume that the group time-effects are strongly mixing and contemporaneously uncorrelated with the error $u_{it}$. These assumptions enable us to use exponential inequalities for weakly dependent processes and allows us to bound misclassification probabilities.\footnote{For more on asymptotic theory of weakly dependent process see rio:2017 and for useful exponential inequalities see merlevede:2011. This condition for strongly mixing was established in rosenblatt:1956 to prove a central limit theorem.} Assumption (ref) ($d$) is a condition on the distribution of the covariates. A sufficient condition would be if the covariates have bounded support or, alternatively, satisfy similar dependence and tail conditions as the error term. In their supplementary appendix, bm:2015 discuss in detail the latter condition since the strong mixing property may not hold with lagged outcomes.
The following proposition establishes an asymptotic equivalence of WGFE estimation to infeasible WGFE estimation and consistency of group membership as $N,T$ approach infinity, but $T$ may grow slower than $N$.
With conditions gauranteeing the asymptotic equivalence of WGFE estimator and the infeasible WGFE estimator, the asymptotic distribution of the WGFE estimator for $\theta^0$ and $\alpha^0$ is that of the infeasible version where groups are known.
Simulation follows bm:2015 where they revisit the Acemoglu, Johnson, Robinson, & Yared (2008) study on the association between income and democracy. The data generating distribution is \[ y_{it} = \widehat{\theta}_1 y_{i t-1} + \widehat{\theta}_2 x_{it} + \widehat{\alpha}_{\widehat{g}_i t} + u_{it} \] where $y_{it}$ is an index of democracy and $x_{it}$ is log-income per capita. I split this DGP into two depending on the estimates, groupings and variance of normally distributed errors with zero mean:
I generate data for different specifications of total number of groups $G$ while holding the income vector $x_{it}$ fixed. I collect a sample of 1,000 GFE and WGFE estimates for each specification. In Figure (ref), the values of the heterogeneous variances and sizes are provided for DGP 2. In Figure (ref), I compare the root mean-squared error of the parameter estimates using WGFE, GFE, and two way fixed effects. In Figure (ref), I compare the average rates of misclassification across different $G$ for each DGP. We see that for DGP 1 there is not much difference between WGFE and GFE estimation as expected since there is no group heteroskedasticity. However, for DGP 2 it appears that WGFE is a better performer when variances are heterogeneous across groups. Unlike groupwise heteroskedasticity in simple regression with observed groups, ignoring it in this context not only affects standard errors, but also impacts parameter estimates since the estimators are functions of group assignments. If groups are different in size and variance, then assignment based solely on mean discrepancies is vulnerable to errors. It can also be interpreted as omitted variable bias if some incorrect assignments are determined since the latent variable is not properly controlled for.
In studies on the effects of unionization on earnings it is often argued that controlling for unobserved ability/skill is essential since employers who face a union contract may select more able candidates than those employers who do not (abowd:1982). A solution for this selection problem would be a fixed-effects approach, but measurement error resulting from misclassification of unionized jobs is more pronounced in panel analysis, which has been shown to harm estimates. freeman:1984 show that estimates from panel studies are generally smaller than those from cross sectional studies across different changing union status groups. For the Panel Study of Income Dynamics (PSID), estimates range from 8%--23% and they might be affected by union misclassification. card:1996 estimate a discrete proxy for ability that is used to estimate five separate models for each ability level, controlling for misclassification in the process. He finds at “high” levels of ability there is negative bias when ignoring high skill heterogeneity versus the positive bias found with low ability workers. In this section, I will revisit the effect of unionization using data from the PSID on $N = 1,158$ workers who are the heads of their household between 2001 and 2019 ($T= 10$, odd numbered years). I draw comparisons to both of these studies by showing the WGFE estimate lies between the pooled and fixed effects estimates as postulated by freeman:1984 and the estimated groups follow a similar correlation pattern to skill groups found in card:1996.
It is natural to assume that employers form impressions of ability of workers based on comparisons to other candidates for the position and so it is reasonable that this impression is part of a discrete set. For example, the employer may classify candidates as “avoid hire”, “maybe hire” or “strong hire”. For jobs that are not covered by unions, but compensate well, employers may be more selective and there may be a complicated correlation structure between working at a union job and skill. It is also reasonable that this discrete classification is correlated with wages across time, which may be due to the bargaining power of unions and overall economic conditions. For example, jobs with a strong union may be more robust to poor economic conditions such as a recession or worsening wage inequality. In addition the response from each skill group to earnings shocks might be more variable suggesting unobserved group heteroskedasticity.
I control for this discrete unobserved variable by estimating a linear model of unionization and wages with a group fixed-effects term:
where $logwage_{it}$ is the log real labor earnings of worker $i$ in year $t$ in 2001 USDs, $union_{it}$ is a dummy variable indicating if worker $i$ had a job at time $t$ that was under a union contract and $x_{it}$ contains time-varying and time-invariant characteristics such as years of schooling, if they are not white, if they are female, number of years of experience at current job and if there job is classified as “blue collar”. I also include a time-varying dummy variable indicating if the worker resides in the south, along with their marital status and age.
Figure (ref) shows summary statistics for the data. About 16% of the sample are female and 33% are nonwhite. Most individuals have 12 years of schooling with 26% of the sample only completing grade school, 47% completing some college and 20% with postgraduate work. The sample contains workers who first appear with ages between 19 and 64 years and are followed for 20 years. In this panel, years of schooling is observed to be time-varying as some individuals returned to school or started working during their schooling. Figure (ref) displays median incomes across different groups. Clearly, unions have a positive impact on real earnings despite the erosion of numbers of unionized jobs post 2007 as shown in Figure (ref). Across sex, race, blue collar occupation, this effect persists.
Estimating model (ref) with pooled least squares and two way fixed effects estimation indicates union membership corresponds to 9.32% and 22% increases in real earnings, respectively. This observation is surprising since freeman:1984 pointed out the opposite ordering and worker bargaining power seems to be waning, especially compared to the proportion of unionized jobs in the older studies. One empirical explanation is that of card:1996 where unobserved ability has a complicated correlation structure with unionized job selection. This may be due to a modern labor market with more variable job outcomes due to technological advancements and the need for the associated skills. The worker has options beyond the traditional union job and so I observe in the sample consistent high earners who are compensated well outside of a union contract. Therefore, the more able may not value a union as much since employers make more competitive offers in their ability category. Wage inequality may also play as a secondary factor as the dominating wages this group might earn may induce negative correlation with union job employment, explaining why these estimates are much larger than pooled OLS estimates.
With WGFE a Bayesian Information Criteria\footnote{See the supplementary appendix of bm:2015 and bai:2002 for the BIC method of estimating the number of groups.} is used to estimate $\widehat{G} = 7$ and find that a union contract raises wages by 14.6%, which is between the pooled least squares and two way fixed effects. Figure (ref) plots the group medians of income and proportion of union jobs. Figure (ref) shows nonlinear correlation where union jobs and earnings are positively correlated with unionization until earnings are much larger when it becomes negatively correlated. Ignoring the green and purple curves for now, focus in order of income by red, orange, blue, teal and gold. The lowest union proportions among these are red, orange and gold, where gold is the lowest income group. Teal and blue have higher earnings and the highest proportions of union jobs in these groups. Therefore at the highest stratum of income, union status is likely to be lower and as we fall in the income tiers the union proportions will increase until the levels of income are low enough where union status falls again. The green and pink group are the least stable groups that experience sharp declines in union status which amounts to a 10% fall from over 20% to under 10% of all jobs. At the same time, pink has a significant fall in real earnings while green is more volatile in their earnings time series.
In this section I revisit acemoglu:2008 where they incorporate country fixed-effects to study the link between income and democracy in nations. They find that the positive effects of income on democracy vanish with the inclusion of country fixed effects arguing that this is consistent with unobserved historical factors that initiated divergent economic and political paths across nations. The purpose of this section is to make comparisons and find insights on how to interpret results using these clustering methods since the results are reinforced by GFE or WGFE.
Following this paper and the grouped fixed-effects approach of bm:2015, I consider a regression of the Freedom house index of democracy on time-lagged democracy, lagged log-GDP per capita and a grouped fixed-effect term $\alpha_{g_i t}$, which captures group-specific time varying heterogeneity within group $g_i = g$ for some $g = 1,\dots,G$. The model is written as \[ democracy_{it} = \theta_1 \, democracy_{i t-1} + \theta_2 \, logGDPpc_{i t-1} + \alpha_{g_i t} + u_{it}. \]
Figure (ref) shows the GFE and WGFE estimates and standard errors of the lagged democracy and income coefficients over different total number of groups $G$ using the 1970-2000 balanced subsample of acemoglu:2008. For reference, the results of the pooled regression ($G=1$) is reported under the GFE columns. First note the decrease in both estimates from the pooled estimate, which is consistent with some positive correlation between the unobservable and lagged democracy possibly due to historical/political events. WGFE standard errors are smaller than the GFE standard errors of the estimator of the lag term coefficient with the exception of $G=2$. However, for the income term in larger group numbers the standard errors for GFE are smaller. This can likely be attributed to multicollinearity between the unobservable, that represents additional second moment information and income. In other words the latent factor that is uncovered is more correlated with income, which was also a result found by kim:2019 in their Bayesian approach to group heteroskedasticity. As for the parameter estimates, the WGFE estimates of the lagged democracy term appear to be more robust to the number of groups $G$ by being more consistent, possibly indicating that the latent variable is being better controlled for in the regression.
In Figure (ref) I have estimated the time effects using WGFE and GFE estimation in the case of $G=4$ groups. I follow the main discussion of bm:2015 and also report the alternative specification of $G=5$ in the Appendix (ref) (as in kim:2019). We see heterogeneous time patterns that are clearly distinct from each other. We also see two, well separated time paths in blue and orange, which are known as the low and high democracy groups following the convention of previous work. The high-democracy groups include most European Union countries, the USA, and also Colombia, Japan and India. As for the low-democracy groups both assignment rules result in China, Iran, and many African countries. These two groups run parallel to each other across time.
In contrast, there are two additional groups detected that exhibit transitions from low-democracy to high-democracy in the sample time periods. Keeping with convention, I identify the group that makes the transition sooner as the early-transition group (purple) and the other as the late-transition group (green). The group assignments between these groups are different between the two methods, however both detect Argentina, South Korea, Spain and Greece as early transitioners and Panama, Romania and Taiwan as late transitioners. Differences of the WGFE assignments come from the early and late transitioners poaching members from either the high-democracy or low-democracy groups as determined by GFE assignments. For a complete list of assignments and differences between WGFE and GFE see Appendix (ref), Figure (ref).
The WGFE may provide more consistent groupings for GFE based on historical account. For instance, the Dominican Republic had unstable democratic institutions throughout the 1960's and their first elected president was the proxy president for their last dictator, who remained in power until 1978. His regime was marked by poor human rights and civil liberties where restrictions were placed on opposition parties. Despite these factors and how the Freedom index measures democracy, GFE assigns the Dominican Republic to the high democracy group while WGFE places the country in the early transitioners. This and some other examples (Greece, Cyprus and Turkey) display a potential advantage of WGFE assignment detecting groupings more robust to the instability that might be found in transitioning groups.
Turning attention to the group fixed effects themselves, the WGFE effects of the low and high democracy groups appear smoother compared to GFE. The late transitioners follow the low democracy group close up until 1985 when it sharply rises. The early transitioners follow a more stable upward trend to the high democracy group. As for the income and democracy plots, the WGFE assignments of groups display more positive correlation between the latent variable and income than GFE. This might explain the multicollinearity issue raised for the coefficient estimate of income using the WGFE estimator.
Concerning variances, the late and early transitioners are found to be much different with WGFE estimates where the early transitioners have the most variable democracy outcomes. The stable groups still have the lowest variances, but they have increased from the GFE version. As for group sizes, Huntington's theory suggests democracy was on the rise in this period, so we should observe that most countries are either democracies or early transitioners. From Figure (ref) and (ref) we see that democracy under WGFE is more dominant in the sample period with more early transitioners, as opposed to predictions from GFE which displays much less early transitioners and a large group of late transitioners. WGFE appears more in line with Huntington's theory in this sense.
In Figure (ref) we see even more local clustering of democracies and early transitioners. Indeed, South America seems to be immersed in transition while Africa being unstable in the sample period having a mix of early, late transitioners with low democracy countries. In the Eastern Mediterranean region I see Turkey, Greece and Cyprus all transitioning as opposed to having Turkey as a high democracy country. In Asia, I see transitioners sandwiched by low democracy China and high democracies Australia, India and Japan. I emphasize the absence of any structure imposed on groupings so this apparent geographical clustering is all the result of estimation.
\newgeometry{top=0.5in,bottom=0.5in,right =1in,left =1in}