Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
59,167 characters · 13 sections · 0 citation commands
Discretizing Unobserved Heterogeneity
\global\long \global\long \global\long \global\long \global\long
In both reduced-form and structural work in economics, it is common to model unobserved heterogeneity as a small number of discrete types. Various estimation strategies are available, including discrete-type random-effects (as in Keane and Wolpin, 1997, and many other applications) and grouped fixed-effects (as recently studied by Hahn and Moon, 2010, and Bonhomme and Manresa, 2015). These methods require the researcher to jointly estimate individual heterogeneity and model parameters.\footnote{Also related, nonparametric maximum likelihood methods (e.g., Heckman and Singer, 1984) rely on joint estimation of the distribution of heterogeneity and the parameters.} In addition, little is known about their properties when individual heterogeneity is not discrete in the population.\footnote{In a network context, Gao et al. (2015) provide results for stochastic blockmodels under continuous heterogeneity.} In this paper, we study two-step discrete estimators for panel data, and provide conditions for their validity when heterogeneity is continuous.
We focus on two-step grouped fixed-effects (GFE) estimators. In a first step, we classify individuals based on a set of individual-specific moments, using the kmeans clustering algorithm. The aim of the kmeans classification is to group together individuals whose latent types are most similar.\footnote{Buchinsky et al. (2005) also propose to group individuals in a first step using kmeans.} In a second step, we estimate the model by allowing for group-specific heterogeneity. This second step is similar to fixed-effects (FE) estimation, albeit it involves a smaller number of parameters that are group-specific instead of individual-specific. We analyze the properties of these two-step estimators in panel data models where heterogeneity is continuous. Hence, in contrast with existing theoretical justifications for discrete-type methods, here we use discrete heterogeneity as a dimension reduction device rather than as a substantive assumption about population unobservables.
Our approach is targeted to environments with two key properties. First, unobserved heterogeneity is a function of a low-dimensional latent variable. We do not restrict this latent type to be discrete. In many economic models, agents' heterogeneity in preferences or technology is driven by a low-dimensional type, which enters the model nonlinearly and may affect multiple outcomes. As an example, we study a model of participation in the labor market where the worker's utility is a function of her productivity type, which in turn determines her wage. GFE provides a tool to exploit such nonlinear factor structures.
Second, the first-step moments satisfy an injectivity condition, which requires any two individuals with the same population moments to have the same type. The choice of moments is important to ensure good performance. In examples, we show how suitable moments arise naturally. In models with exogenous covariates, we propose and analyze the use of conditional moments to recover latent types.
Our setup also covers models where heterogeneity varies over time. Unlike additive FE methods and interactive FE methods based on linear factor structures (Bai, 2009), GFE does not require heterogeneity to take an additive or interactive form. As an illustration, we compare GFE and FE estimators in a probit model where heterogeneity is a nonlinear function of a time-invariant factor loading and a time-specific factor.
Our main results are large-$N,T$ asymptotic expansions of two-step GFE estimators under time-invariant and time-varying continuous heterogeneity. In both settings, GFE is consistent as the number of groups grows with the sample size, under conditions that we provide. We find that, when the population heterogeneity is not discrete, estimating group membership induces an incidental parameter bias, similarly to FE methods. Moreover, since discreteness is an approximation in our setting, GFE is affected by approximation error. We propose a simple data-driven rule for the number of groups that controls the approximation error, and discuss how to reduce incidental parameter bias for inference.
The outline of the paper is as follows. We introduce the setup and two-step GFE estimators in Section (ref), study their asymptotic properties in Section (ref), and outline several extensions in Section (ref). The main proofs may be found in the appendix, and the supplemental material contains additional results.
We consider a panel data setup, where we denote outcome variables and exogenous covariates as $Y_i{=}(Y_{i1}',...,Y_{iT}')'$ and $X_{i}{=}\left(X_{i1}',...,X_{iT}'\right)'$, respectively, for $i{=}1,...,N$. In our theory we cover two models. In the first one, unobserved heterogeneity is time-invariant. In this case, the conditional log-density of $Y_i$ given $X_i$ is given by:\footnote{In models with first-order dependence, we assume that $Y_{i0}$ is observed and we condition on it. Higher-order dependence can be accommodated similarly. In dynamic settings, $Y_{it}$ may contain sequentially exogenous covariates in addition to outcome variables.}
and the log-density of exogenous covariates $X_i$ takes the form: $$\ln g_i(\mu_{i0})=\sum_{t=1}^T\ln g(X_{it}\,|\, X_{i,t-1},\mu_{i0}),$$ where $\theta_0$ is a vector of common parameters, and $\alpha_{i0}$ and $\mu_{i0}$ are individual-specific parameters. We leave the form of $g$ unrestricted, and in estimation we will use a conditional likelihood approach based on $f_i$ alone. In other words, in applications the researcher only needs to specify the parametric form of $f_i(\alpha_{i0},\theta_0)$ in ((ref)). However, the heterogeneity $\mu_{i0}$ in covariates plays an important role in our theory.
In the second model, unobserved heterogeneity varies over time. Such variation in unobservables over calendar time (e.g., business cycle), age (e.g., life cycle), counties, or markets, is of interest in many applications. In the time-varying case, log-densities take the form:
where $\alpha_{i0}=(\alpha_{i10}',...,\alpha_{iT0}')'$ and $\mu_{i0}=(\mu_{i10}',...,\mu_{iT0}')'$. In both models we are interested in estimating $\theta_0$, as well as average effects depending on $\alpha_{10},...,\alpha_{N0}$.
GFE relies on two key assumptions that we now present. We defer the presentation of regularity conditions until Section (ref). First, we assume that unobserved heterogeneity is a function of a low-dimensional vector $\xi_{i0}$.
We will refer to $\xi_{i0}$ as an individual type, and to $d$ as the dimension of heterogeneity. The researcher does not need to know $d$, $\alpha$, or $\mu$ in applications. In models with time-varying unobserved heterogeneity, Assumption (ref) requires unobservables to follow a factor structure. The link between $\alpha_{it0}$, $\xi_{i0}$ and $\lambda_{t0}$ may be nonlinear, the linear structure $\alpha_{it0}=\xi_{i0}'\lambda_{t0}$ (Bai, 2009) being covered as a special case. Moreover, the dimension of $\lambda_{t0}$ is unrestricted. Our theory will show that the performance of two-step GFE crucially relies on $\xi_{i0}$ being low-dimensional, a leading case being $d=1$. We provide examples in the next subsection.
Second, we rely on individual-specific moment vectors $h_i$ that are informative about the types $\xi_{i0}$. We state this formally as our second main assumption, where $\|\cdot\|$ denotes an Euclidean norm.
Assumption (ref) requires the individual moment vector $h_i$ to be informative about $\xi_{i0}$, in the sense that, for large $T$, $\xi_{i0}$ can be uniquely recovered from $h_i$. Neither $\varphi$ nor $\psi$ (which may depend on $\theta_0$) need to be known to the econometrician. Intuitively, injectivity guarantees that one can separate the types of two individuals $\xi_{i0}$ and $\xi_{i'0}$ by comparing their moments $h_i$ and $h_{i'}$. For example, an average $h_i=\frac{1}{T}\sum_{t=1}^Th(Y_{it},X_{it})$ will, under Assumption (ref) and suitable regularity conditions, converge as $T$ tends to infinity to a function $\varphi(\xi_{i0})$ of the type $\xi_{i0}$. We require $\varphi$ to be injective.
The convergence rate in Assumption (ref) requires appropriate conditions on the serial dependence of $Y_{it}$ and $X_{it}$. In models with time-varying heterogeneity, $\varphi$ will also depend on the $\lambda_{t0}$ process. In such models, Assumption (ref) requires the moments to be informative about $\xi_{i0}$, and not $\lambda_{t0}$. Injectivity is a key requirement for consistency of two-step GFE estimators. More generally, the choice of moments $h_i$ is important for finite-sample performance.
To illustrate the framework we now describe two examples, for which we will provide illustrative simulations in Subsection (ref). First, consider a dynamic model of wages $W_{it}^*$ and labor force participation $Y_{it}$:
where the wage $W_{it}^*$ is only observed when $i$ works, $U_{it}$ are i.i.d. standard normal, independent of the past $Y_{it}$'s and $\alpha_{i0}$, and $V_{it}$ are i.i.d. independent of all $U_{it}$'s, $Y_{i0}$, and $\alpha_{i0}$. Here the same scalar expected payoff $\alpha_{i0}=\xi_{i0}$, unobserved to the econometrician, drives the wage and the decision to work. Individuals have common preferences denoted by the utility function $u$, the cost function $c$ is state-dependent, and both $u$ and $c$ are unknown to the econometrician.
In this setting, GFE provides a natural approach to exploit the functional link between $\alpha_{i0}$ and $u(\alpha_{i0})$, and to learn about the type $\alpha_{i0}$ using both wages and participation. For instance, when $h_i=(\overline{W}_i,\overline{Y}_i)'$, where $\overline{Z}_{i}=\frac{1}{T}\sum_{t=1}^TZ_{it}$ denotes the individual mean of $Z_{it}$, injectivity is satisfied under mild conditions, provided $ \overline{W}_i=\alpha_{i0}\overline{Y}_i+o_p(1)$ and ${\limfunc{plim}}_{T\rightarrow\infty}\,\overline{Y}_i>0$.
Fixed-effects (FE) is a possible approach to estimate $\theta_0$ in ((ref)). However, a conventional FE estimator would treat $\alpha_{i0}$ and $u_{i0}= u(\alpha_{i0})$ as unrelated parameters, so the FE estimate of $\theta_0$ would be solely based on the binary participation decisions. Another strategy would be to rely on discrete-type random-effects methods, which are typically based on joint estimation. In contrast, we implement GFE in two steps with no need for iterative estimation, and we justify the estimator in environments where heterogeneity is not restricted to be discrete.
As a second example, consider the following probit model with time-varying heterogeneity:
where $U_{it}$ are i.i.d. standard normal, independent of all $V_{it}$'s, $\alpha_{it0}$'s, and $\mu_{it0}$'s, and $V_{it}$ are i.i.d. independent of all $\alpha_{it0}$'s and $\mu_{it0}$'s. Under Assumption (ref), $\alpha_{it0}$ and $\mu_{it0}$ depend on a low-dimensional vector $\xi_{i0}$ of factor loadings, so $\alpha_{it0}={\alpha}(\xi_{i0},\lambda_{t0})$ and $\mu_{it0}={\mu}(\xi_{i0},\lambda_{t0})$. Here $d$ is the dimension of the type $\xi_{i0}$ governing both $\alpha_{it0}$ and $\mu_{it0}$.
To motivate why, in static models with covariates such as ((ref)), $\alpha_{it0}$ and $\mu_{it0}$ may depend on a common low-dimensional type $\xi_{i0}$, suppose that, in every period, agent $i$ chooses $X_{it}$ based on expected utility or profit maximization. She observes $\xi_{i0}$ and $\lambda_{t0}$ --- which enter outcomes through $\alpha_{it0}$ --- and takes her decision before the i.i.d. shock $U_{it}$ is realized. In such a case, $X_{it}$ will be a function of $\xi_{i0}$ and $\lambda_{t0}$, as well as idiosyncratic factors $V_{it}$ in the agent's information set. Here we assume that the agent's information set, and primitives such as preferences or costs, do not include other $i$-specific elements beyond $\xi_{i0}$.\footnote{This example is reminiscent of Mundlak's (1961) classic analysis of farm production functions, where soil quality $\xi_{i0}$ is observed to the farmer but latent to the analyst.}
When $\alpha(\cdot,\cdot)$ is additive or multiplicative in its arguments, model ((ref)) can be estimated using two-way FE (Fern\'andez-Val and Weidner, 2016) or interactive FE (Bai, 2009, Chen et al., 2020), respectively. However, when $\alpha(\cdot,\cdot)$ is unknown, these fixed-effects estimators are inconsistent in general. In contrast, GFE will remain consistent when unobservables are unknown nonlinear functions of factor loadings $\xi_{i0}$ and factors $\lambda_{t0}$, and injectivity holds. Taking ${h}_i=(\overline{Y}_i,\overline{X}_i')'$ as moments in model ((ref)), injectivity is satisfied when types have monotone effects on the heterogeneity components.\footnote{To see this, consider the case where $\alpha_{it0}$ is the only component of heterogeneity (i.e., $\mu_{it0}=0$ in ((ref))), and take $h_i=\overline{Y}_i$. Letting $G$ denote the cdf of $-(V_{it}'\theta_0+U_{it})$, injectivity will hold when $\alpha(\cdot,\cdot)$ is strictly increasing in its first argument and $G$ is strictly increasing, since then $\varphi(\xi)= {\limfunc{plim}}_{T\rightarrow\infty}\,\frac{1}{T}\sum_{t=1}^TG(\alpha(\xi,\lambda_{t0}))$ is strictly increasing.} More generally, in Assumption (ref) we require that the latent type $\xi_{i0}$ can be asymptotically recovered from a moment vector whose dimension is not growing with the sample size.
Two-step GFE consists of a classification step and an estimation step.
\paragraph{First step: classification.} We rely on the individual-specific moments $h_i$ to learn about the individual types $\xi_{i0}$. Specifically, we partition individuals into $K$ groups, corresponding to group indicators $\widehat{k}_i\in\{1,...,K\}$ , by computing:
where $\{k_i\}$ are partitions of $\{1,...,N\}$ into $K$ groups, and $\widetilde{h}(k)$ is a vector. Note that $\widehat{h}(k)$ is simply the mean of $h_i$ in group $\widehat{k}_i=k$.
In the kmeans optimization problem ((ref)), the minimum is taken with respect to all possible partitions $\{k_i\}$. Fast and stable optimization methods such as Lloyd's algorithm are available, although computing a global minimum may be challenging; see Bonhomme and Manresa (2015) for references. Following the literature, we will focus on the asymptotic properties of the global minimum and abstract from optimization error. Lastly, note that the quadratic loss function in ((ref)) can accommodate weights on different components of $h_i$, although here for simplicity we present the unweighted case.
\paragraph{Second step: estimation.} We maximize the log-likelihood function with respect to common parameters $\theta$ and group-specific effects $\alpha$, where the groups are given by the $\widehat{k}_i$ estimated in the first step. We define the two-step GFE estimator as:
Note that, in contrast to fixed-effects (FE) maximum likelihood, this second step involves a maximization with respect to $K$ group-specific parameters instead of $N$ individual-specific ones. In models with time-varying heterogeneity, $\alpha(k)$ will simply be a vector $(\alpha_1(k)',...,\alpha_T(k)')'$.
\paragraph{Choice of $K$.} Two-step GFE estimation requires setting a number of groups $K$. We propose a simple data-driven selection rule based on the first step. The convergence rate of the kmeans estimator (and the rate of the GFE estimator) will be governed by two quantities: the kmeans objective function $\widehat{Q}(K)=\frac{1}{N}\sum_{i=1}^N\|h_i-\widehat{h}(\widehat{k}_i)\|^2$, which decreases as $K$ gets larger and the group approximation becomes more accurate, and the variability $V_h=\mathbb{E}[\|h_i-\varphi(\xi_{i0})\|^2]$ of the moment $h_i$, which does not depend on $K$. We take the smallest $K$ that guarantees that $\widehat{Q}(K)$ is of the same or lower order as $V_h$. That is, letting $\widehat{V}_{h}=V_h+o_p(1/T)$, we suggest setting:
where $\gamma\in(0,1]$ is a user-specified parameter.\footnote{When $h_i=\frac{1}{T}\sum_{t=1}^Th(Y_{it},X_{it})$ and observations are independent over time, one may take $\widehat{V}_{h}=\frac{1}{NT^2}\sum_{i=1}^N \sum_{t=1}^T \|h(Y_{it},X_{it})-h_i\|^2$. With dependent data, one can use trimming or the bootstrap to estimate $V_h$ (Hahn and Kuersteiner, 2011, Arellano and Hahn, 2007).} In the simulations in the next subsection we will set $\gamma=1$, although smaller $\gamma$ values corresponding to larger $K$'s will also be supported by our theory.
To illustrate the performance of GFE in models where heterogeneity follows a nonlinear factor structure, we present the results of a small-scale simulation study based on our two examples ((ref)) and ((ref)). In both cases, we assume that the type $\xi_{i0}$ governing heterogeneity is scalar. We compare the bias of GFE to that of FE and interactive FE estimators. In the supplemental material, we provide details on the simulations and report additional results.
In Figure (ref), we compare GFE and FE in model ((ref)), using a CRRA functional form: $u(\alpha)=\frac{{e^{\alpha}}^{(1-\eta)}-1}{1-\eta}$, with a risk aversion parameter $\eta\in\{1,2\}$. We focus on the difference in costs $c(0;\theta_0)-c(1;\theta_0)$, which measures the degree of state dependence in participation decisions. We take $h_i=(\overline{W}_i,\overline{Y}_i)'$ as moments for GFE, and report average parameter estimates over 1000 simulations. We set $N{=}1000$ and vary T between 5 and 50. We find that FE is more biased than GFE for both values of risk aversion. This is consistent with wages and participation providing informative moments about the latent type in this setting.
In Figure (ref) we compare GFE, FE, and interactive FE in model ((ref)) with $X_{it}$ scalar, using a CES specification: $\alpha_{it0}=\left(a\xi_{i0}^\sigma +(1-a)\lambda_{t0}^{\sigma}\right)^{\frac{1}{\sigma}}$, for $\sigma{\in}\{-10,0,1,10\}$ and $a{=}0.5$, and $\mu_{it0}=\alpha_{it0}$. The factors $\lambda_{t0}$ and the individual loadings $\xi_{i0}$ enter heterogeneity in a nonlinear way. We show estimates of ${\theta}_0$ for various estimators: GFE, FE with additive individual and time effects, and interactive FE with a single multiplicative factor. We use $(\overline{Y}_i,\overline{X}_i)'$ as moments for GFE. Note that both $\overline{Y}_i$ and $\overline{X}_i$ are informative about $\xi_{i0}$ in this data generating process. We report parameter averages over 1000 simulations, for $N{=}1000$. We find that, while GFE, FE, and interactive FE are all biased, the bias of GFE is smaller across all $\sigma$ values.\footnote{Large-$N,T$ theory implies that additive and interactive FE are consistent when $\sigma=1$ and $\sigma=0$, respectively. Figure (ref) shows that, despite being large-$N,T$ consistent in these specifications, in our simulations, additive and interactive FE have larger biases than GFE for the $N$ and $T$ values we consider. }
In this section we provide asymptotic expansions for two-step GFE estimators. Our first result is a rate of convergence for kmeans. Let us define the approximation error one would make if one were to discretize the latent types $\xi_{i0}$ directly, as:
where, similarly to ((ref)), the minimum is taken with respect to all partitions $\{k_i\}$ and vectors $\widetilde{\xi}(k)$. In the asymptotic analysis we let $T=T_N$ and $K=K_N$ tend to infinity jointly with $N$.
The bound in Lemma (ref) has two terms: an $O_p(1/T)$ term that depends on the number of periods used to construct the moments $h_i$, and an $O_p\left(B_{\xi}(K)\right)$ term that reflects the presence of an approximation error. The rate at which $B_{\xi}(K)$ tends to zero depends on the dimension of $\xi_{i0}$. Graf and Luschgy (2002, Theorem 5.3) provide explicit characterizations in the case where $\xi_{i0}$ has compact support.\footnote{See Graf and Luschgy (2002, p. 875) for a discussion of the compact support assumption.} For example, the following lemma implies that $B_{\xi}(K)=O_p(K^{-2})$ when $\xi_{i0}$ is one-dimensional, and $B_{\xi}(K)=O_p(K^{-1})$ when $\xi_{i0}$ is two-dimensional.
We now use these results to study the properties of GFE in models with time-invariant and time-varying heterogeneity, in turn. We use the shorthand notation $\mathbb{E}_{Z}(W)$ and $\mathbb{E}_{Z=z}(W)$ for the conditional expectations of $W$ given $Z$ and $Z=z$, respectively. In the time-varying case, we denote as $\lambda_0$ the process of $\lambda_{t0}$'s, and as $\mathbb{E}_{\lambda_0=\lambda}(W)$ the conditional expectation of $W$ given $\lambda_0=\lambda$. We use a similar notation for variances. Finally, $\|M\|$ denotes the spectral norm of a matrix $M$.
To state our first main theorem, where heterogeneity is time-invariant, we make the following assumptions, where $\ell_{it}(\alpha_{i},\theta)=\ln f(Y_{it}\,|\, Y_{i,t-1},X_{it},\alpha_{i},\theta)$, $\ell_i(\alpha_{i},\theta)=\frac{1}{T}\sum_{t=1}^T\ell_{it}(\alpha_{i},\theta)$, and $\overline{\alpha}(\theta,\xi)={\limfunc{argmax}}_{\alpha}\,\mathbb{E}_{\xi_{i0}=\xi}(\ell_{i}(\alpha,\theta))$ for all $\theta,\xi$.
In part ((ref)) in Assumption (ref) we treat heterogeneity as random in order to use Lemma (ref), which requires $\xi_{i0}$ to be i.i.d. draws from a distribution. However, note we do not restrict how $\alpha_{i0}$ and $\mu_{i0}$ depend on each other. Moreover, while our results require asymptotic stationarity of the time-series processes, the theorem could be extended to allow for nonstationary initial conditions.
In part ((ref)) we require strict concavity of the log-likelihood as a function of $\alpha$. Concavity holds in a number of nonlinear panel data models such as probit and logit models, tobit, Poisson, or multinomial logit; see Fern\'andez-Val and Weidner (2016) and Chen et al. (2020). One can show that Theorem (ref) continues to hold without concavity, under an identification condition and an assumption bounding the derivatives of the empirical GFE objective function. Importantly, note that $H^{-1}$ is the asymptotic variance of the FE estimator. As a result, $H$ being positive definite rules out models that are not identified under FE, such as a linear model with a time-invariant covariate and a heterogeneous intercept.
In part ((ref)) we introduce the target log-likelihood $\frac{1}{NT}\sum_{i=1}^N\sum_{t=1}^T\ell_{it}(\overline{\alpha}(\theta,\xi_{i0}),\theta)$ (Arellano and Hahn, 2007), which we will show approximates the GFE log-likelihood in large samples under our assumptions (note that $\overline{\alpha}(\theta_0,\xi_{i0}){=}\alpha_{i0}$). In part ((ref)) we require some moments to be bounded asymptotically.
We now state our first main result, where we denote, evaluating all quantities at true values $(\theta_0,\alpha_{i0})$ and omitting the dependence from the notation:
The first three terms in ((ref)) also appear in large-$N,T$ expansions of FE estimators (e.g., Hahn and Newey, 2004).\footnote{In the supplemental material we provide a similar expansion for GFE estimators of average effects $M_0=\frac{1}{NT}\sum_{i=1}^N \sum_{t=1}^Tm\left(X_{it},\alpha_{i0},\theta_0\right)$, which are functions of both common parameters and individual heterogeneity.} Similarly to FE, GFE is subject to incidental parameter $O_p(1/T)$ bias. This contrasts with the properties of GFE estimators under discrete heterogeneity (e.g., Hahn and Moon, 2010, Bonhomme and Manresa, 2015). Indeed, when heterogeneity is not restricted to have a small number of points of support, classification noise affects the properties of second-step estimators in general. This motivates using bias reduction techniques for inference analogous to those used in FE, as we will discuss in the next section.
The $O_p(K^{-\frac{2}{d}})$ term in ((ref)) reflects the approximation error, which depends on the number of groups. Setting $K=\widehat{K}$ according to ((ref)) guarantees that the approximation error is $O_p(1/T)$. Formally, we have the following result.
Under Corollary (ref), the biases of FE and GFE have the same order of magnitude. However, the required value of $K$ depends on the dimension $d$ of individual heterogeneity. Specifically, when $\xi_{i0}$ follows a continuous distribution of dimension $d$, setting $K$ proportional to or greater than $\min(T^{\frac{d}{2}},N)$ will ensure that the approximation error is $O_p(1/T)$. For small $d$ (e.g., when $d=1$) this will typically require a small number of groups (of the order of $T^{\frac{1}{2}}$).
GFE can have advantages compared to FE, for two reasons. First, the two-step method can allow researchers to select moments that are particularly informative about the unobserved heterogeneity. To provide intuition, consider a setting where the number of groups is sufficiently large for the approximation error to be of smaller order compared to $1/T$, yet $K/N$ tends to zero. We have the following.
Corollary (ref) shows that the first-order asymptotic bias of GFE is the difference between two terms. The bias is zero when $h_i$ is an injective function of $\xi_{i0}$; i.e., when $\varepsilon_i=h_i-\varphi(\xi_{i0})=0$. More generally, the bias can be expanded in $\varepsilon_i$, and it is small when moments provide accurate estimates of the latent types. Moreover, the first term on the right-hand side of ((ref)) coincides with the bias of FE (e.g., Arellano and Hahn, 2007). The form of ((ref)) implies that the biases of FE and GFE are equal when the moments are the FE estimates $h_i=\widehat{\alpha}_i(\theta_0)$, however other moment choices can lead to smaller biases. From this perspective, GFE provides flexibility to use well-suited proxies of the latent types. As an example, our simulations of the labor force participation model ((ref)) show that, by jointly exploiting wages and participation to construct moments that are informative about the latent type, GFE can have smaller bias than FE (and smaller mean squared error as well, as shown in the supplemental material).
A second advantage of GFE comes from the use of grouping, and from the resulting regularization. Indeed, individual FE estimates can be highly variable whenever the number of parameters per individual is large. In such cases, reducing the number of parameters through grouping can improve performance. For instance, the ability to handle multiple components of heterogeneity is central to the performance of GFE in models with time-varying unobserved heterogeneity. This is the case we focus on next.
To state our second main theorem, where heterogeneity is time-varying, we make the following assumptions, where $\ell_{it}(\alpha_{it},\theta)=\ln f(Y_{it}\,|\, Y_{i,t-1},X_{it},\alpha_{it},\theta)$, $\ell_i(\alpha_{i},\theta)=\frac{1}{T}\sum_{t=1}^T\ell_{it}(\alpha_{it},\theta)$, and $\overline{\alpha}^t(\theta,\xi)={\limfunc{argmax}}_{\alpha}\,\mathbb{E}_{\xi_{i0}=\xi,\lambda_0=\lambda}(\ell_{it}(\alpha,\theta))$.\footnote{Note that $\overline{\alpha}^t(\theta,\xi_{i0})$ depends on the process $\lambda_{0}$ in addition to the type $\xi_{i0}$, although we leave the dependence on $\lambda_0$ implicit in the notation. In a static model, $\overline{\alpha}^t(\theta,\xi_{i0})$ is a function of $\xi_{i0}$ and $\lambda_{t0}$, while in a dynamic model it also depends on the history of the time effects $(\lambda_{t0},\lambda_{t-1,0},...)$.}
In part ((ref)) in Assumption (ref), we impose a stronger concavity condition than in Assumption (ref).\footnote{In particular, we use part ((ref)) in Assumption (ref) to establish consistency. Note that this condition can be restrictive in models with time-varying random coefficients.} The other parts are similar to Assumption (ref), except part ((ref)) where we require regularity of certain conditional expectations and variances.
We next state our second main result, where, differently from Theorem (ref), $s_i$ in ((ref)) and $H$ in ((ref)) are now evaluated at $(\theta_0,\alpha_{it0})$, and expectations are conditional on $(\xi_{i0},\lambda_0)$.
Theorem (ref) shows that GFE is consistent as $N,T,K$ tend to infinity and $K/N$ tends to zero. This requires no parametric assumption about how $\xi_{i0}$ and $\lambda_{t0}$ affect individual and time heterogeneity, unlike additive or interactive FE methods.
To give intuition, consider the probit model ((ref)) with time-varying unobservables. Under Assumption (ref), in the first step, GFE consistently estimates an injective function $\varphi_{i0}=\varphi(\xi_{i0})$ of the type. One can then rewrite the outcome equation in ((ref)) as $Y_{it}=\boldsymbol{1}\left\{X_{it}'\theta_0+\alpha(\psi(\varphi_{i0}),\lambda_{t0})+U_{it}\geq 0\right\}$, where $\psi$ is the function introduced in Assumption (ref), and $\alpha_{it0}=\alpha(\psi(\varphi_{i0}),\lambda_{t0})$ is simply a time-varying function of $\varphi_{i0}$. In the second step, GFE estimates this function by including group-time indicators in the probit regression.
As in Theorem (ref), the expansion in Theorem (ref) features a combination of incidental parameter bias and approximation error. When using the rule ((ref)) for $K$, the approximation error is of the same or lower order compared to 1/T. However, the $O_p(K/N)$ term is a new contribution relative to the time-invariant case, which reflects the estimation of $KT$ group-specific parameters using $NT$ observations. As an example, when $d=1$ and $K$ is chosen of the order of $T^{\frac{1}{2}}$, the $O_p$ terms in ((ref)) are $O_p(1/T+T^{\frac{1}{2}}/N)$.\footnote{When $N/T^{\frac{3}{2}}\rightarrow 0$, one could obtain a faster rate in ((ref)) by choosing another rule for $K$.} Although this rate of convergence can be fast when $N$ is sufficiently large relative to $T$, it is too slow to apply conventional bias-reduction methods for inference. In the next section, under the additional assumption that time heterogeneity $\lambda_{t0}$ is low-dimensional, we describe how to obtain a faster convergence rate by grouping both individuals and time periods.
In models with time-invariant heterogeneity, Corollary (ref) can be used to characterize the asymptotic distribution of GFE estimators. However, as in FE, the presence of the $O_p(1/T)$ term in ((ref)) shifts the distribution of $\widehat{\theta}$ away from $\theta_0$ whenever $T$ is not large relative to $N$. A variety of methods are available to bias-correct FE estimators and construct asymptotically valid confidence intervals; see Arellano and Hahn (2007) for a review. Consider the setup of Corollary (ref), under the additional assumption that the $O_p(1/T)$ term in ((ref)) is equal to $C/T+o_p(1/T)$ for some constant $C$. In this case, one can show that half-panel jackknife (Dhaene and Jochmans, 2015) gives asymptotically valid inference based on GFE as $N$ and $T$ tend to infinity at the same rate.\footnote{In particular, half-panel jackknife is valid under the conditions of Corollary (ref), which requires taking $\gamma=o(1)$ in our rule ((ref)) for $K$ in order for the approximation error to be of small order. Deriving primitive conditions for the validity of half-panel jackknife and other bias-reduction methods for other choices of $K$ is left for future work.} The distribution of the bias-corrected GFE estimator is then asymptotically normal centered at the truth, and the asymptotic variance $H^{-1}$ can be consistently estimated by replacing the expectations in ((ref)) and ((ref)) by group-specific means.
In settings where heterogeneity varies over time, it can be desirable to group not only individuals as in ((ref)), but also time periods (or alternatively counties or markets, depending on the application). We now describe such a method, and discuss its potential for performing inference in models with time-varying heterogeneity. In the two-way GFE approach, we classify time periods based on cross-sectional moments $w_t=\frac{1}{N}\sum_{i=1}^N w(Y_{it},X_{it})$, and compute:
where $\{l_t\}$ are partitions of $\{1,...,T\}$ into $L$ groups. Given the group indicators $\widehat{k}_i$ and $\widehat{l}_t$, we then maximize $\sum_{i=1}^N\sum_{t=1}^T\ln f(Y_{it}\,|\, X_{it},\alpha(\widehat{k}_i,\widehat{l}_t),\theta)$, with respect to $\theta$ and the $KL$ group-specific parameters $\alpha(k,l)$.
Two-way GFE estimators can be expanded similarly to Theorem (ref), under two main additional assumptions: the model is static and observations are independent across $i$ and $t$, and the dimensions $d_{\lambda}$ of time heterogeneity $\lambda_{t0}$ and $d$ of individual heterogeneity $\xi_{i0}$ are both small. Then, for $s_{i}$ and $H$ as in Theorem (ref), we show in the supplemental material that:
Suppose $d=d_{\lambda}=1$, and $K$ is given by ((ref)) with $\gamma$ asymptotically constant, with an analogous choice for $L$. Then the $O_p$ term in this expansion can be shown to be $O_p(1/T+1/N)$. We leave to future work the formal study of the validity of bias reduction methods for inference, such as two-way split panel jackknife (Fern\'andez-Val and Weidner, 2016), as $N$ and $T$ tend to infinity at the same rate.
Our theory shows that the dimension $d$ of heterogeneity plays a key role in the properties of GFE. While models with scalar latent types $\xi_{i0}$, such as model ((ref)) of wages and labor force participation, are not uncommon in economics, many applications involve conditioning covariates. Under Assumptions (ref) and (ref), the moments $h_i$ should, asymptotically, be injective functions of all the heterogeneity coming from both $Y_i$ and $X_i$. However, when $X_i$ depends on multiple components of heterogeneity, this might lead to a large dimension $d$.
We now show that GFE can still perform well under a weaker form of injectivity. Consider the case where Assumption (ref) is replaced by $\alpha_{i0}={\alpha}(\xi_{i0})$ and $\mu_{i0}={\mu}(\xi_{i0},\nu_{i0})$, where $\nu_{i0}$ is another latent component that affects covariates. Moreover, instead of requiring injectivity for both $\xi_{i0}$ and $\nu_{i0}$, let us maintain Assumption (ref), which only requires $h_i$ to be injective for $\xi_{i0}$. In other words, $h_i$ needs to be directly informative about the unobserved heterogeneity component $\xi_{i0}$ that appears in the conditional distribution of $Y_i$ given $X_i$. We show in the supplemental material that, under regularity conditions otherwise similar to those of Corollary (ref), the convergence rate of GFE is unaffected by the dimension of $\nu_{i0}$. Specifically, when $K=\widehat{K}$ is given by ((ref)) with $\gamma=O(1)$ (which adapts to the dimension of $\xi_{i0}$ and not the one of $\nu_{i0}$), we have:
To prove ((ref)) we assume that the rate condition $T^{1+\frac{d}{2}}=O(N)$ holds, where $d$ is the (small) dimension of $\xi_{i0}$.\footnote{In the supplemental material, we provide an asymptotic expansion for GFE in a linear homoskedastic model under a small approximation error, as in Corollary (ref). The argument requires no restriction on the relative rates of $N$ and $T$. Interestingly, in this case the asymptotic variances of GFE and FE differ, since the within-group variation in $\nu_{i0}$ tends to decrease the variance, yet the expansion features an additional score term compared to Theorem (ref).}
In models with time-varying conditioning covariates, a simple way to target moments to $\xi_{i0}$ is to construct $h_i$ using the {conditional distribution} of $Y_{i}$ given $X_{i}$. To see this, consider a static model $f(Y_{it}\,|\, X_{it},\alpha_{i0},\theta_0)$ where $X_{it}$ has finite support. In this case, we have under appropriate conditions: $$\underset{=h_i(x)}{\underbrace{\frac{\sum_{t=1}^T\boldsymbol{1}\{X_{it}=x\}h(Y_{it},X_{it})}{\sum_{t=1}^T\boldsymbol{1}\{X_{it}=x\}}}}=\underset{=\varphi(x,\xi_{i0})}{\underbrace{\mathbb{E}_{X_{it}=x,\xi_{i0}}[h(Y_{it},X_{it})]}}+o_p\left(1\right),$$ where $h_i(x)$ is only defined when $\sum_{t=1}^T\boldsymbol{1}\{X_{it}=x\}\neq 0$, and, importantly, $\varphi(x,\xi_{i0})$ does not depend on $\nu_{i0}$. In the supplemental material we discuss implementation, and we report simulation results in a probit model with binary covariates. We find that using conditional moments can enhance the performance of GFE in such settings. We leave the analysis of conditional moments in the presence of continuous covariates to future work.
In this paper, we analyze some properties of two-step grouped fixed-effects (GFE) methods in settings where population heterogeneity is not discrete. Our framework relies on two main assumptions: low-dimensional individual heterogeneity, and the availability of moments to approximate the latent types. In many economic models, individual types are low-dimensional. By taking advantage of this feature, GFE can allow for flexible forms of heterogeneity across individuals and over time.
GFE methods are of interest in various applied settings. In a previous version of this paper, we used two-step GFE to estimate a dynamic structural model of location choice in the spirit of Kennan and Walker (2011), and we analyzed the performance of the discrete estimator of Bonhomme et al. (2019) for matched employer-employee data in the presence of continuous firm heterogeneity. Other potential applications include nonlinear factor models, nonparametric and semi-parametric panel data models such as quantile regression with individual effects, and network models.