Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.
Online appendix to “Latent group structure in linear panel data models with endogenous regressors”
\begingroup
\setcounter{footnote}{0}
\thispagestyle{plain}
\renewcommand\@fnsymbol\c@footnote{\@fnsymbol\c@footnote}
\if@twocolumn
\ifnum \col@number=\@ne
\@maketitle
\else
\twocolumn[\@maketitle]
\fi
\else
\global\@topnum\z@
\@maketitle
\fi
\thispagestyle{plain}\@thanks
\endgroup
\global\let\@thanks\@empty
abstractThis paper concerns the estimation of linear panel data models with endogenous regressors and a latent group structure in the coefficients. We consider instrumental variables estimation of the group-specific coefficient vector. We show that direct application of the Kmeans algorithm to the generalized method of moments objective function does not yield unique estimates. We newly develop and theoretically justify two-stage estimation methods that apply the Kmeans algorithm to a regression of the dependent variable on predicted values of the endogenous regressors. The results of Monte Carlo simulations demonstrate that two-stage estimation with the first stage modeled using a latent group structure achieves good classification accuracy, even if the true first-stage regression is fully heterogeneous. We apply our estimation methods to revisiting the relationship between income and democracy.
Keywords: latent group structure, Kmeans, instrumental variable, generalized method of moments, panel data
JEL Classification: C23, C26, C38
mtchideinmaintoc\section{Introduction}
This paper considers the estimation of linear panel data models with endogenous regressors where the coefficients exhibit a grouped pattern of heterogeneity. Instrumental variables (IVs) estimation becomes nontrivial because of the latent group structure. We examine various approaches for estimating such models. We particularly examine when these methods can provide consistent group membership and coefficient estimators.
Endogeneity is an important concern in empirical research in economics and other social sciences. It is also relevant for panel data analysis. Regressors are endogenous when they are correlated with the error term. This is important as least squares methods do not yield consistent estimators in the presence of endogenous regressors. IVs are employed to solve this issue. A valid IV is correlated with the endogenous regressors, but uncorrelated with the error term in the model. IV estimation is a standard tool in econometrics and is discussed in many textbooks.
In addition, heterogeneity across units has recently gained attention in the econometrics literature. In our setting, the coefficients exhibit a latent group structure.
An advantage of panel data is the ability to investigate heterogeneity across units because a panel data set contains multiple observations per unit. However, allowing complete heterogeneity may not be practical in many applications. One reason is that there may not be many observations per unit. However, even if there are many observations per unit to enable us to estimate a model separately for each unit, it would be difficult to summarize the features of heterogeneity across many units if each unit has its own estimate. A grouped pattern of heterogeneity is a useful modeling device for understanding heterogeneity. We divide the set of units into a small number of groups and assume that units in each group are homogeneous, but two units belonging to two different groups are different. Such a structure facilitates estimation and interpretation.
We thus examine how a model with endogenous regressors, and a grouped pattern of heterogeneity, can be estimated. Despite the importance of endogeneity and the recent popularity of models with a latent group structure, the literature lacks a comprehensive analysis. The present paper fills this gap. In particular, we show that na\"ively combining IV methods with Kmeans MacQueen1965, arguably the best-known clustering method, does not yield a uniquely defined estimator. We then consider alternative approaches, including those newly proposed here, that yield consistent estimators.
We first demonstrate that it is not possible to determine a unique minimizer of the generalized method of moments (GMM) objective function when both the coefficient vector and the group membership structure are unknown. That is, na\"ive extensions of the Kmeans procedure, or the group fixed effects (GFE) estimator of BonhommeManresa2015 for regression models, fail to yield unique estimates, and the estimators lack bases for their statistical and asymptotic analyses. Specifically, we show that for any hypothetical group membership structure, there may exist a coefficient vector that sets the GMM objective function to zero, at least asymptotically. This result is in stark contrast to the case of models without endogeneity where a least squares objective function can be employed to estimate the group memberships and coefficients simultaneously.
We develop two-stage least squares (2SLS) estimation to consistently estimate the group memberships and consider various specifications for the first-stage regression. This procedure appears novel in the literature on latent group structures.
We then consider other approaches under which the estimation of the group membership structure and the coefficients are separated.
In the 2SLS estimation approach, we estimate a first-stage regression that describes the relationship between the endogenous regressors and IVs and then apply the GFE estimator of BonhommeManresa2015 to a second-stage regression of the dependent variable on the predicted values of the endogenous regressors from the first stage. We consider three modeling frameworks for the first-stage regression. The simplest version is a homogeneous regression model.
The second modeling framework allows the first-stage regression to exhibit a grouped pattern of heterogeneity. The first-stage regression is estimated by GFE, and the predicted values are constructed based on estimated group-specific coefficients. The group structures in the first and second stages are allowed to be different. The third modeling framework is to allow complete heterogeneity across units. The first-stage regression is estimated for each unit. It turns out that this approach coincides with minimizing the sum of the GMM objective functions for each unit with a specific choice of weighting matrices.\footnote{The estimators in SuShiPhillips2016 and Mehrabani2022 are based on this objective function, but they consider penalized estimations. In the context of generalized estimating equation, itosugasawa2023 consider a similar estimator.} We prove the consistency of these 2SLS estimators of both the coefficient vector and the group membership structure.
The other two procedures provide consistent group membership estimators under some additional conditions but do not directly deliver a consistent coefficient estimator. The first procedure, referred to as the “ignoring endogeneity approach,” is to apply the GFE estimator to the model for the dependent variable and the endogenous regressors while ignoring the endogeneity. This provides a consistent group membership estimator if the grouped pattern in the linear projection of the dependent variable on the endogenous regressors is identical to that in the true coefficient vector. While this assumption may be debatable, this approach is suggested, for example, in BonhommeManresa2015Supp. Once an estimate of the group membership structure is obtained, the coefficient vector is estimated by GMM for each group. The second procedure, referred to as the “reduced-form approach,” is to consider the reduced-form equation of the dependent variable on the IVs. This approach requires that the group pattern in the reduced-form equation is identical to that for the coefficients of interest. To our best knowledge, this approach has not hitherto been considered in the literature. The coefficients are estimated by group-by-group GMM estimations after the group membership structure is estimated.
We also consider several extensions. First, we consider models with unit-specific fixed effects. Here, we merely need to apply the estimation procedures to the within or first-differencing transformed models.
We use dynamic linear panel data models for illustration.
Second, we consider models with time-varying group-specific fixed effects. Because the dimension of the coefficient vectors now grow as the data becomes longer, a different theoretical analysis is required.
We illustrate the advantages of the 2SLS estimation approach in these extensions.
We compare the performance of our estimators using simulations.
Performance is measured using the average classification accuracy and the average deviation from the true coefficient vector.
Four data-generating processes (DGPs) are considered, differentiated by the first-stage modeling.
Every estimator performs well when predicted to be so by asymptotic theory, confirming the usefulness of the asymptotic analyses.
Among the estimators, the 2SLS estimators, allowing heterogeneity in the first stage, show good performance under all DGPs. Interestingly, even in the DGPs where each unit has different first-stage coefficients, the 2SLS estimators, assuming a first-stage group structure, provide comparable performance to the estimator allowing complete heterogeneity.
The ignoring endogeneity approach shows excellent performance when the group structures of the estimating model and the structural equation coincide. However, when the correlation between the first- and second-stage errors is high and varies across units in addition to the complete heterogeneity in the first stage, the heterogeneity pattern in the correlation contaminates the group structure, and the ignoring endogeneity approach performs poorly. In such cases, the 2SLS estimator provides reliable estimates, and imposing a group structure in the first stage improves the estimation.
Lastly, we apply our estimators to revisit the empirical analysis in AcemogluJohnsonRobinsonYared08, which examines the relationship between income growth and democracy.
The authors use dynamic linear models to estimate the causal effect of income on democracy.
They employ the savings rate and trade-weighted world income as IVs to address the endogeneity of income growth because of reverse causality.
We extend these models and allow the coefficients to differ across groups.
The number of groups is set to two. The IV is trade-weighted world income.
When only the coefficient of income growth is allowed to exhibit a group structure, all the procedures produce compatible estimates, which are positive and mostly statistically significant. Re-running a separate IV regression for each group does not change the results qualitatively.
Overall, larger groups tend to demonstrate a greater effect of income. However, when we introduce group-specific fixed effects additionally, the procedures yield estimates that differ to some extent---particularly when including the lagged democracy in the model. The coefficient estimates are different across (second-stage) groups when the number of groups in the first stage is small, but they become similar as we increase the number of groups in the first stage. This result underscores the importance of modeling first-stage heterogeneity
in correctly eliciting the heterogeneity pattern in the structural parameters.
\subsection*{Related literature}
Latent group structure has been a popular topic in the recent econometrics literature. An (admittedly incomplete) list of papers on this topic includes Sun2015, HahnMoon2010, LinNg2012, BonhommeManresa2015, SuShiPhillips2016, AndoBai2016, VogtLinton2016, WangPhillipsSu2018, Mehrabani2022, and mugnier2022simple. They primarily focus on models with exogenous regressors. The present paper considers cases with endogenous regressors and their IV estimation.
While IV estimation for models with a latent group structure has been examined in the existing literature, discussion of these models is typically included only in extensions or appendices of main papers. See, for example, BonhommeManresa2015Supp, SuShiPhillips2016, Mehrabani2022. An application is found in Czarnowske2022, which uses the method in SuShiPhillips2016 to estimate heterogeneous production functions. Our aim is to provide a comprehensive account of an application of group heterogeneity to IV estimation. In many economic applications, endogeneity is a valid concern, and IV estimation is required. The present paper expands the scope of the heterogeneity analysis with a latent group structure in the existing economics literature.
We also develop a new estimation approach, namely the 2SLS approach. This approach has several advantages over existing methods. In particular, it enables us to consider various extensions. For example, it can accommodate time-varying parameters while allowing general heterogeneity patterns in the first stage. We note that controlling for time-varying group-specific intercepts is the motivation of BonhommeManresa2015. Generic clustering approaches such as that by YuGuVolgushev2022 cannot be applied here because it is impossible to obtain preliminary estimates from unit-by-unit estimation in the presence of time-varying parameters. It is thus necessary to develop estimation methods tailored to handle time-varying parameters.
The proposed estimator can exploit the heterogeneity in the first stage. adadiegushen2023 also consider first-stage equations with group heterogeneity but assume that the group structure is known. wiemann2023 applies the Kmeans algorithm to cluster categorical instruments to improve the efficiency of the estimator. For our part, we allow latent group structure in the first stage but require panel data.
Regarding proof techniques and estimation ideas, our paper relates to the structural break literature. There is a rich body of research on structural break detection using time series data. The proofs of the negative results for the na\"ive approaches follow the argument of HallHanBoldea2012. They document the failure of the GMM procedure to detect structural breaks in time series analysis. We consider different econometric situations, and the proofs are not identical. Nonetheless, their mathematical insights are useful in deriving our results. Furthermore, some of the estimation methods we consider have their origins in the structural break literature. The 2SLS approach to detect structural breaks is examined in HallHanBoldea2012 and PerronYamamoto2014. The approach based on biased estimation (by ignoring endogeneity) is considered in PerronYamamoto2015, QianSu2014. The present paper demonstrates that although the results in the structural break literature are at first very different in scope from ours, they are useful in investigating econometric problems for models with group patterns of heterogeneity.
While some research papers consider combining structural breaks with a grouped pattern of heterogeneity, such as OkuiWang2022, LumsdaineOkuiWang2022, they have very different scopes from the present paper. They consider the estimation of break points that affect group-specific coefficient vectors and/or group membership structures. The present paper does not consider structural breaks; we merely utilize their theoretical and mathematical insights for models with endogenous regressors.
\paragraph{Organization of the paper}
Section (ref) introduces a panel data model with endogenous regressors and a latent group structure. Section (ref) demonstrates that the na\"ive approaches of minimizing the GMM objective function with respect to the coefficient vector and the group membership structure fail to provide unique estimates. Section (ref) introduces various estimation procedures that provide consistent group membership estimators. Section (ref) considers various extensions. The first involves models with individual fixed effects, and the second comprises dynamic panel data models. The third extension concerns models with time-varying group-specific coefficients. The results of the Monte Carlo simulations are discussed in Section (ref). We illustrate the estimation procedures by an empirical example in Section (ref). Section (ref) concludes. The mathematical proofs are included in the Appendix.
\section{Model}
We consider a linear panel data model with possibly endogenous regressors. The coefficients are heterogeneous across observational units, and the heterogeneity exhibits a grouped pattern. IVs that are uncorrelated with the error term and relevant to the endogenous regressors are available.
Suppose that we observe panel data $(y_{it}, x_{it}', z_{it}')$, $i= 1, \dots, N$ and $t= 1, \dots, T$, where $y_{it}$ is a scalar dependent variable, $x_{it}$ is a vector of regressors, and $z_{it}$ is a vector of IVs. Elements of $x_{it}$ and $z_{it}$ may overlap. The index $i$ refers to an observational unit, and $t$ denotes a time period. Let $d$ and $m$ be the dimension of $x_{it}$ and $z_{it}$, respectively.
We consider the following linear model with group-specific coefficients:
\begin{align}
y_{it} = x_{it}'\beta_i^0 + u_{it},
\end{align}
where $\beta_i^0$ is the true unit-specific coefficient, which we explain in more detail below, and $u_{it}$ is an error term.
We suspect that $x_{it}$ may be endogenous, i.e., $x_{it}$ may be correlated with $u_{it}$. IVs, $z_{it}$, are exogenous in the sense that they are uncorrelated with the error term $u_{it}$ and are also relevant such that $z_{it}$ and $x_{it}$ are correlated.
The departure from the standard IV regression model is that we allow the coefficients to be heterogeneous across $i$. Specifically, we consider a grouped pattern of heterogeneity. Observational units are divided into $G$ groups. Let $\mathbb{G}\equiv\{1,\dots,G\}$ denote the set of groups. Units in the same group share the same value of the coefficient vector. Two units in different groups have different values of the coefficient. Unit $i$ belongs to $g_i^0 \in \mathbb{G}$. Group memberships $\{g_i^0 \}_{i=1}^N$ are unobservable and need to be inferred from data. We manipulate the notation, and a value of the coefficient for group $g$ is denoted as $\beta_g$. The true coefficient for $i$ is
\begin{align*}
\beta_i^0 = \beta_{g_i^0}^0.
\end{align*}
A grouped pattern of heterogeneity is a useful modeling device.
This is because allowing complete heterogeneity may not be attractive in practice because information about the coefficient of a given unit can then be obtained only from that specific unit. By grouping units, we can combine the information from multiple units in a group to obtain more precise coefficient estimates. Grouping methods also enhance the interpretability of the estimated coefficients because the dimension of $\{\beta_g\}_{g\in \mathbb{G}}$ is $dG$ and is not large when $G$ is small.
We are interested in estimating $\beta^0 \equiv \{ \beta_{g}^0 \}_{g\in \mathbb{G}} \in \mathbb{R}^{dG}$ and $\gamma^0 \equiv \{ g_i^0 \}_{i=1}^N \in \mathbb{G}^N$. In contrast to standard IV estimation, the group membership structure $\gamma^0$ is unknown, which makes the estimation nontrivial. In the next section, we argue that na\"ive applications of standard GMM estimation may not yield unique estimates. Various alternative estimation procedures are introduced and examined in Section (ref).
\section{Na\"ive approaches do not work}
This section considers the estimation procedures that minimize the GMM objective function with respect to the coefficient vector and the group membership structure. We argue that such na\"ive approaches do not provide unique estimates. It is easy to gauge the problem when the dimension of the regressor vector and that of the IVs are equal. A conceptually similar yet technically more involved argument demonstrates that this approach fails to yield a unique minimizer of the objective function asymptotically, even when the dimension of the IVs is larger than that of the regressors.
\subsection{Na\"ive procedures}
Consider estimating $\beta^0$ and $\gamma^0$ by minimizing the GMM objective function. Let $\beta \in \mathbb{R}^{dG}$ and $\gamma = \{g_i \}_{i=1}^N \in \mathbb{G}^N$ denote a value of the coefficient parameter and a group membership assignment structure, respectively. A na\"ive GMM estimator is
\begin{align}
(\hat \beta , \hat \gamma ) = \arg \min_{\beta, \gamma} \biggl( \sum_{i=1}^N \sum_{t=1}^T z_{it} (y_{it} - x_{it}'\beta_{g_i} ) \biggr)' \hat W \biggl( \sum_{i=1}^N \sum_{t=1}^T z_{it} (y_{it} - x_{it}'\beta_{g_i} ) \biggr),
\end{align}
where $\hat W$ is an $m \times m$ positive definite matrix possibly depending on the data. Consider also estimating $\beta_g$ for each $g$ by minimizing the group-specific GMM objective functions:
\begin{align}
(\hat \beta , \hat \gamma ) = \arg \min_{\beta, \gamma} \sum_{g\in \mathbb{G}} \biggl( \sum_{g_i =g} \sum_{t=1}^T z_{it} (y_{it} - x_{it}'\beta_g ) \biggr)' \hat W_g \biggl( \sum_{g_i=g} \sum_{t=1}^T z_{it} (y_{it} - x_{it}'\beta_g ) \biggr),
\end{align}
where $\hat W_g$ is an $m \times m$ positive definite matrix which depends on $g$ and possibly on the data. When $\gamma^0$ is known, solving these minimization problems under $\gamma = \gamma^0$ provides a standard GMM estimator of $\beta^0$. Thus, its consistency and asymptotic normality can be shown under conditions established in the literature. The problem is that $\gamma$ is unknown and needs to be estimated together with $\beta$.
We demonstrate the failure of these estimators: that is, they are not consistent and may not even be uniquely determined.
A common motivation behind these approaches is an expectation (which turns out to be wrong) that
\begin{align}
(\hat{\beta},\hat{\gamma})\rightarrow_p &\arg\min_{\beta,\gamma} \operatorname*{plim}_{N,T\rightarrow\infty} \biggl( \sum_{i=1}^N \sum_{t=1}^T z_{it} (y_{it} - x_{it}'\beta_{g_i} ) \biggr)' \hat W \biggl( \sum_{i=1}^N \sum_{t=1}^T z_{it} (y_{it} - x_{it}'\beta_{g_i} ) \biggr)
\intertext{and}
(\beta^0,\gamma^0)= &\arg\min_{\beta,\gamma} \operatorname*{plim}_{N,T\rightarrow\infty}\biggl(\sum_{i=1}^N \sum_{t=1}^T z_{it} (y_{it} - x_{it}'\beta_{g_i} ) \biggr)' \hat W \biggl( \sum_{i=1}^N \sum_{t=1}^T z_{it} (y_{it} - x_{it}'\beta_{g_i} ) \biggr)
\end{align}
would hold. This is a typical argument for the asymptotic analysis of an extremum estimator: the minimizer of a sample objective function converges to the minimizer of the limit of the objective function, and the latter is the true value of the parameter. We now show that one of these two conditions might not be true. When the dimension of the regressors and that of the IVs are equal, $(\hat{\beta},\hat{\gamma})$ cannot be uniquely defined (Theorem (ref)), which renders (ref) unrealizable. When the dimension of the IVs is larger than that of the regressors, the probability limits of the objective functions may have zeros at parameter values other than $(\beta^0,\gamma^0)$ (Theorem (ref) and (ref)), which violates (ref) because the true parameter value is not the unique minimizer.
\subsection{Just-identified cases}
We first consider the case of $m=d$. We refer to this as a “just-identified” case because, under a known group membership structure, solving (ref) is the standard IV/GMM estimation under just-identification. However, when the group membership structure is unknown and needs to be estimated jointly with the coefficient vector, the na\"ive approaches fail and do not provide a unique solution to the minimization problem.
Here, we show that for any group membership structure $\gamma$, a value of $\beta$ exists that sets the objective functions to zero.
Take any $\gamma = \{g_i\}_{i=1}^N$ and consider
\begin{align}
\hat \beta_g (\gamma) \equiv \biggl( \sum_{g_i=g} \sum_{t=1}^T z_{it} x_{it}' \biggr)^{-1} \sum_{g_i=g} \sum_{t=1}^T z_{it} y_{it},
\end{align}
for each $g\in \mathbb{G}$. This is the IV estimator based on the observations assigned to group $g$ under $\gamma$.
It is easy to see that
\begin{align*}
\sum_{g_i=g} \sum_{t=1}^T z_{it} (y_{it} - x_{it}' \hat \beta_{g} (\gamma)) &=0
\intertext{for all $g$, which implies}
\sum_{i=1}^N \sum_{t=1}^T z_{it} (y_{it} - x_{it}' \hat \beta_{g_i} (\gamma)) &=0.
\end{align*}
Both of the objective functions in (ref) and (ref) are zero under $\{ \hat \beta_g (\gamma) \}_{g\in \mathbb{G}}$ for any $\gamma$.
The value of $(\beta, \gamma)$ that minimizes the objective functions is not unique, and
the estimators cannot be determined uniquely. Consequently, the estimators (ref) and (ref) are not consistent. The discussion so far is summarized in the following theorem.
\begin{theorem}
Suppose that $m=d$. Let $\hat \beta_g (\gamma)$ satisfy
\begin{align*}
\sum_{g_i=g} \sum_{t=1}^T z_{it} x_{it}' \hat \beta_g (\gamma) = \sum_{g_i=g} \sum_{t=1}^T z_{it} y_{it}.
\end{align*}
Then, for any $\gamma \in \mathbb{G}^N$,
\begin{align*}
\biggl( \sum_{i=1}^N \sum_{t=1}^T z_{it} (y_{it} - x_{it}' \hat \beta_{g_i} (\gamma ) \biggr)'
&\hat W \biggl( \sum_{i=1}^N \sum_{t=1}^T z_{it} (y_{it} - x_{it}' \hat\beta_{g_i} (\gamma ) \biggr) =0
\intertext{and}
\sum_{g\in \mathbb{G}} \biggl( \sum_{g_i =g} \sum_{t=1}^T z_{it} (y_{it} - x_{it}'\hat \beta_g (\gamma ) )\biggr)'
&\hat W_g \biggl( \sum_{g_i=g} \sum_{t=1}^T z_{it} (y_{it} - x_{it}'\hat \beta_g (\gamma ) ) \biggr) =0
\end{align*}
hold. Note that if $ \sum_{g_i=g} \sum_{t=1}^T z_{it} x_{it}'$ is invertible, $\hat \beta_g (\gamma)$ is given in (ref).
\end{theorem}
\subsection{Over-identified cases}
This section considers the cases when $m>d$. We refer to these as “over-identified” cases because if the group membership structure were known in advance, solving (ref) reduces to standard IV/GMM estimation under over-identification. Note that in the case of over-identification, no parameter value sets the objective functions to zero in finite samples in general.
We argue that the objective functions for the na\"ive approaches do not have a unique minimizer asymptotically.
For simplicity, we consider cases with $G=2$ only. Considering cases with $G>2$ makes the analysis more complicated but adds nothing conceptually. Set $\mathbb{G} = \{1, 2\}$.
We use double asymptotics: both $N$ and $T$ tend to infinity at the same time, that is, $N, T\rightarrow \infty$. This represents the following situation. Let $T\equiv T(N)$ be an increasing function of $N$ such that $T (N) \to \infty$ as $N \to \infty$. We understand that $\operatorname*{plim}_{N, T\to \infty} $ gives the limit of a sequence indexed by $N$ and $T$ under $N\to \infty$ with any choice of $T (\cdot)$. If there are conditions on the relative magnitude of $N$ and $T$, then we consider a class of $T(\cdot)$ that is compatible with those conditions.
We introduce several notations that are needed to represent the limit of the objective functions.
For $g\in\mathbb G$, let
\begin{align*}
\lambda_g^0 \equiv \lim_{N \to \infty} \frac{1}{N} \sum_{i=1}^N \mathbf{1} \{g_i^0 =g\}
\end{align*}
where $\mathbf{1}\{\cdot\}$ is the indicator function that returns one if the argument is true and zero otherwise.
Thus, $\lambda_1^0$ and $\lambda_2^0$ are the (asymptotic) fractions of units belonging to groups 1 and 2, respectively.
Note that $\lambda_1^0 + \lambda_2^0 =1$.
We also define the asymptotic fractions of units that are assigned to their true groups under a group membership structure $\gamma$.
Here, we manipulate the notation and let $\gamma \equiv \{ g_i \}_{i=1}^{\infty}\in\mathbb G^\mathbb N$, where $\mathbb{N}$ is the set of natural numbers. Note that, in finite samples, the first $N$ elements of $\gamma$ are relevant.
For $g\in\mathbb G$, let
\begin{align*}
\lambda_{gg} (\gamma) \equiv \lim_{N \to \infty} \frac{1}{N} \sum_{i=1}^N\mathbf{1} \{g_i^0 =g\} \mathbf{1} \{g_i =g\}
\end{align*}
The (asymptotic) fraction of units that belong to group 1 and are assigned to group 1 under $\gamma$ is $\lambda_{11} (\gamma)$. Similarly, $\lambda_{22} (\gamma)$ is the fraction of units that belong to group 2 and are assigned to group 2. We note that $\lambda_{11} (\gamma^0) = \lambda_1^0$ and $\lambda_{22} (\gamma^0) = \lambda_2^0$.
Next, we define the matrices that represent the relationship between the IVs and the regressors for each group.
For $g\in\mathbb G$, let
\begin{align*}
M_g \equiv \operatorname*{plim}_{N, T\to \infty} \frac{1}{\sum_{i=1}^n \mathbf{1} \{g_i^0 =g\}} \sum_{i=1}^n \frac{1}{T} \sum_{t=1}^T \mathbf{1} \{g_i^0=g\} z_{it} x_{it}'
\end{align*}
Roughly speaking, $M_1$ and $M_2$ may be interpreted as the expected values of $z_{it}x_{it}'$ for groups 1 and 2, respectively.
In addition, for each $(g,\tilde g)\in\mathbb G^2$, let
\begin{align*}
M_{g\tilde g}(\gamma) \equiv& \operatorname*{plim}_{N, T\to \infty} \frac{1}{\sum_{i=1}^n \mathbf {1}\{g_i^0 =g\}\mathbf{1}\{g_i=\tilde g\}} \sum_{i=1}^n \frac{1}{T} \sum_{t=1}^T \mathbf{1} \{g_i^0=g\}\mathbf{1}\{g_i=\tilde g\}z_{it} x_{it}'.
\end{align*}
Similarly, we can understand $M_{g\tilde g}(\gamma)$ as the conditional expected values of $z_{it}x_{it}'$ for the units in group $g$ that are assigned to group $\tilde g$ under $\gamma$.
Lastly, we define the vectors that measure the exogeneity of the IVs for each assigned group. Specifically, for each $g\in\mathbb G$, let
\begin{align*}
L_g(\gamma)\equiv \operatorname*{plim}_{N,T\rightarrow\infty}\frac{1}{\sum_{i=1}^N\mathbf{1}\{g_i=g\}}\sum_{i=1}^N\frac{1}{T}\sum_{t=1}^T\mathbf{1}\{g_i=g\}z_{it}u_{it}.
\end{align*}
We now make the following sets of assumptions:
\begin{assumption}
\begin{enumerate}[label=\normalfont(\alph*)]
• The limits $\lambda_1^0$ and $\lambda_2^0$ are well-defined.
• The limits $M_1$ and $M_2$ are well-defined and finite.
\end{enumerate}
\end{assumption}
\begin{assumption}
\begin{enumerate}[label=\normalfont(\alph*),ref=(\alph*)]
• The limits $\lambda_{11} (\gamma)$ and $\lambda_{22}(\gamma)$ are well-defined.
• For each $(g,\tilde g)\in\mathbb G^2$, the limit $M_{g\tilde g}(\gamma)$ is well-defined and finite. Also, for each $g\in\mathbb G$, $M_{g1}(\gamma)=M_{g2}(\gamma)$.
• For each $g\in\mathbb G$, the limit $L_g(\gamma)$ is well-defined and finite.
\end{enumerate}
\end{assumption}
In Assumption (ref), we restrict the set of possible group assignments $\gamma$. That $\lambda_{11} (\gamma)$ and $\lambda_{22}(\gamma)$ are well-defined is a weak condition. The condition $M_{g1}(\gamma)=M_{g2}(\gamma)$ for both $g=1$ and $g=2$ indicates that the group assignment is not correlated with the strength of the IVs within each group. The condition is satisfied when the first-stage relationship between $x_{it}$ and $z_{it}$ is homogeneous. We note that the first-stage relationship is allowed to be heterogeneous, but then we consider group assignments that are not correlated with the first-stage relationship. Importantly, there remain infinitely many group assignments even after Assumption (ref) is imposed.
The following theorem demonstrates that for any group membership assignment satisfying certain conditions, there exists a value of the coefficient vector that sets the limit of the objective function in (ref) zero.
\begin{theorem}
Suppose that Assumption (ref) holds. Assume that $\operatorname*{plim}_{N,T\to \infty} \hat W = W$, where $W$ is a symmetric positive definite matrix, and that $\operatorname*{plim}_{N,T \to \infty} (NT)^{-1}\sum_{i=1}^N\sum_{t=1}^T z_{it} u_{it}=0$. Then, if $(M_2-M_1)(\beta_2^0 - \beta_1^0)=0$,
\begin{align}
\operatorname*{plim}_{N,T\to \infty} \biggl( \frac{1}{NT} \sum_{i=1}^N \sum_{t=1}^T z_{it} (y_{it} - x_{it}'\beta_{g_i} ) \biggr)' \hat W \biggl( \frac{1}{NT} \sum_{i=1}^N \sum_{t=1}^T z_{it} (y_{it} - x_{it}'\beta_{g_i} ) \biggr)
\end{align}
is zero at
\begin{align*}
(\beta_1(\gamma),\beta_2 (\gamma)) &\equiv
\biggl(\frac{\lambda_{11} (\gamma ) \beta_1^0 + (\lambda_2^0 - \lambda_{22} (\gamma ) ) \beta_2^0}{\lambda_{11} (\gamma ) + \lambda_2^0 - \lambda_{22} (\gamma)},\frac{\lambda_{22} (\gamma ) \beta_2^0 + (\lambda_1^0 - \lambda_{11} (\gamma ) ) \beta_1^0}{\lambda_{22} (\gamma ) + \lambda_1^0 - \lambda_{11} (\gamma)}\biggr)
\end{align*}
for any $\gamma\in\mathbb G^{\mathbb N}$ satisfying Assumptions (ref)(ref)--(ref).
\end{theorem}
A similar result holds for the objective function in (ref), as shown in the following theorem.
\begin{theorem}
Suppose that Assumption (ref) holds. Assume that for each $g \in \mathbb{G}$, $\operatorname*{plim}_{N,T\to \infty} \hat W_g = W_g$, where $W_g$ is a symmetric positive definite matrix. Then, if $(M_2-M_1)(\beta_2^0 - \beta_1^0)=0$,
\begin{align*}
\operatorname*{plim}_{N,T\to \infty} \sum_{g\in \mathbb{G}} \biggl( \frac{1}{NT} \sum_{g_i =g} \sum_{t=1}^T z_{it} (y_{it} - x_{it}'\beta_g ) \biggr)' \hat W_g \biggl( \frac{1}{NT} \sum_{g_i=g} \sum_{t=1}^T z_{it} (y_{it} - x_{it}'\beta_g ) \biggr)
\end{align*}
is zero at
\begin{align*}
(\beta_1 (\gamma), \beta_2 (\gamma)) \equiv \biggl(\frac{\lambda_{11} (\gamma ) \beta_1^0 + (\lambda_2^0 - \lambda_{22} (\gamma) ) \beta_2^0}{\lambda_{11} (\gamma ) + \lambda_2^0 - \lambda_{22} (\gamma)},
\frac{\lambda_{22} (\gamma ) \beta_2^0 + (\lambda_1^0 - \lambda_{11} (\gamma) ) \beta_1^0}{\lambda_{22} (\gamma ) + \lambda_1^0 - \lambda_{11} (\gamma)}\biggr)
\end{align*}
for any $\gamma\in\mathbb G^{\mathbb N}$ that satisfies Assumption (ref) and $L_1(\gamma)=L_2(\gamma)=0$.
\end{theorem}
These theorems indicate that there are infinitely many values of the parameters that can set the limits of the objective functions to zero. Note that the true value also satisfies this property, so the moment condition is correct. However, the moment condition does not provide sufficient information to identify the parameters. Accordingly, the estimators are not expected to be consistent.
The corresponding proofs are included in the Appendix. These proofs broadly follow the argument in HallHanBoldea2012. Their article discusses a failure of break detection with GMM in time series contexts. While the situations are different, we can utilize their mathematical insights in the present context.
We impose exogeneity conditions. In Theorem (ref), the IVs and the error term are assumed to be uncorrelated at least asymptotically. Theorem (ref) imposes a different condition that $L_1(\gamma)=L_2(\gamma) =0$. These conditions are satisfied, for example, when the exogeneity holds for each unit: $\operatorname*{plim}_{T \to \infty} T^{-1}\sum_{t=1}^T z_{it} u_{it}=0$ uniformly over $i$.
The key condition is $(M_2-M_1)(\beta_2^0 - \beta_1^0)=0$, namely that $\beta_2^0 - \beta_1^0$ is in the null space of $M_2-M_1$. In practice, it is difficult, if not impossible, to know whether $(M_2-M_1)(\beta_2^0 - \beta_1^0)=0$ in advance because the group structure is latent and the estimation of $M_2$ and $M_1$ requires knowledge of the true group membership structure. As $\beta_2^0$ and $\beta_1^0$ are two of the parameters that we want to estimate, we should always take into account the possibility that this failure may occur in any application.
Indeed, it is not difficult to imagine situations in which $(M_2-M_1)(\beta_2^0 - \beta_1^0)=0$ holds.\footnote{Obviously, $(M_2-M_1)(\beta_2^0 - \beta_1^0)=0$ occurs when $\beta_2^0=\beta_1^0$. However, in such a case, there is no group heterogeneity in the coefficients. It is natural that we cannot uniquely identify a group membership structure.}
For example, if $M_2=M_1$, then $(M_2-M_1)(\beta_2^0 - \beta_1^0)=0$ holds regardless of the value of $\beta^0$. This case implies that the first stage does not have a group structure, i.e., when we can postulate the first stage $x_{it}' = z_{it}'\Pi + v_{it}$ with a common coefficient matrix $\Pi$ across $i$, where $v_{it}$ is an error term that is uncorrelated with $z_{it}$.
Even when $M_2 \neq M_1$, $(M_2-M_1)(\beta_2^0 - \beta_1^0)=0$ may occur.
For example, consider a case in which there are two regressors and three IVs and the DGP satisfies the following:
\begin{align*}
M_2 = \begin{pmatrix}
1 & 0 \\
1 & 0 \\
0 & 1
\end{pmatrix}, \quad
M_1 = \begin{pmatrix}
0 & 1 \\
0 & 1 \\
1 & 0
\end{pmatrix}, \quad
\beta_2 = \begin{pmatrix}
1 \\
1
\end{pmatrix}, and
\beta_1 = \begin{pmatrix}
0 \\
0
\end{pmatrix}.
\end{align*}
In Group 2, both regressors have coefficients equal to one, while in Group 1, they do not affect the dependent variable. In Group 2, the first and second IVs are related to the first regressor but not the second, and the third IV is related to the second regressor, but not the first. The relationship is alternated in Group 1. In this case, $(M_2-M_1)(\beta_2^0 - \beta_1^0)=0$ holds.
\begin{remark}[Computational problem]
Even if we were perfectly certain that $(M_2-M_1) (\beta_2^0 - \beta_1^0) \neq 0$, both estimation approaches discussed above encounter a computational problem. Indeed, it is not clear whether an efficient way to solve the minimization problems exists. For example, consider the following Kmeans type algorithm:
\begin{enumerate}
• Set $\gamma^{(0)}$ and $s=1$.
• Update the coefficient: For all $g\in\mathbb{G}$,
\begin{align*}
\hat \beta_g^{(s)} = \arg \min_{\beta_g}\biggl( \sum_{g_i^{(s-1)} =g} \sum_{t=1}^T z_{it} (y_{it} - x_{it}'\beta_g ) \biggr)' \hat W_g \biggl( \sum_{g_i^{(s-1)}=g} \sum_{t=1}^T z_{it} (y_{it} - x_{it}'\beta_g ) \biggr).
\end{align*}
• Update the group membership estimate:
\begin{align*}
\gamma^{(s)} = \arg \min_{\gamma} \sum_{g\in \mathbb{G}} \biggl( \sum_{g_i =g} \sum_{t=1}^T z_{it} (y_{it} - x_{it}'\hat \beta_g^{(s)} ) \biggr)' \hat W_g \biggl( \sum_{g_i=g} \sum_{t=1}^T z_{it} (y_{it} - x_{it}'\hat \beta_g^{(s)} ) \biggr).
\end{align*}
• $s=s+1$.
• Repeat 2--4 until convergence.
\end{enumerate}
This procedure has a problem at Step 3 because it is not clear how to solve the minimization problem quickly. The only exception is the case of just-identification, where, as observed above, the algorithm would converge in the first step. But the minimizer then depends on the initial group membership assignment and fails to be a consistent estimator.
\end{remark}
\begin{remark}
SuShiPhillips2016 report an observation related to our findings. They note that they fail to establish an asymptotic theory for the fully pooled criterion, which corresponds to the objective function in (ref) in our case. However, they do not define it as an identification failure, nor do they provide any conditions under which it occurs. Another related observation is given by Mehrabani2022 who states that “...because group membership of individual units is unknown, we cannot apply the usual GMM objective function here,” but they do not explain the source of the problem.
\end{remark}
\section{Alternative estimation strategies}
In this section, we propose alternative approaches that produce consistent estimators. They are two-step procedures. We first fit the endogenous regressors using IVs. Then, we apply the Kmeans (i.e., GFE) estimator to the structural model (ref) whose endogenous variables are substituted with the corresponding fitted values from the first-stage regression.
We consider three different approaches to model the first-stage regression:
(i) The first is to assume that the first-stage coefficient is homogeneous across units:
\begin{align*}
x_{it} = \Pi^{0\prime} z_{it} + v_{it},
\end{align*}
where $\Pi^0$ is a $m\times d$ coefficient matrix, and $v_{it}$ is a first-stage error term.
(ii) The second is to assume that the first-stage coefficients exhibit a grouped pattern of heterogeneity $\kappa^0\equiv \{ k_i^0 \}_{i=1}^N\in\mathbb K^N=\{1,\dots,K\}^N$:
\begin{align}
x_{it} = \Pi_{k_i^0}^{0\prime}z_{it} + v_{it}.
\end{align}
We allow $\gamma^0$ and $\kappa^0$ (also $G$ and $K$) to differ. $\Pi_{k}^{0}$ is a $m\times d$ coefficient matrix for group $k$. $K=1$ corresponds to the homogeneous first stage mentioned above.
(iii) Lastly, we consider the model that allows full heterogeneity across units, that is,
\begin{align}
x_{it} = \Pi_i^{0\prime}z_{it} + v_{it},
\end{align}
where $\Pi_i^{0}$ is the $d\times m$ unit-specific first-stage coefficient parameter matrix.
The relative merit of this modeling is robustness to misspecification, though it may cause some efficiency loss because of the large number of parameters in the first stage.
We develop estimation procedures tailored to each of the first-stage models we postulate. In the following subsections, we elaborate on the estimation procedures and show that they yield consistent estimators. The proofs for the theorems are in the Appendix.
They broadly follow the argument given in BonhommeManresa2015,BonhommeManresa2015Supp.
\subsection{Homogeneous first stage}
We first consider the model in which the first-stage coefficient is homogeneous. In this case, the two-stage estimation runs the first-stage regression of $x_{it}$ and $z_{it}$ and then applies the GFE in the second stage using the fitted values from the first stage:
\begin{enumerate}[label=(\arabic*)]
• Regress $x_{it}$ on $z_{it}$ and obtain
\begin{align*}
\hat\Pi \equiv \biggl(\frac{1}{NT}\sum_{i=1}^N\sum_{t=1}^Tz_{it}z_{it}'\biggr)^{-1}\frac{1}{NT}\sum_{i=1}^N\sum_{t=1}^Tz_{it}x_{it}'.
\end{align*}
Construct $\hat x_{it} \equiv \hat \Pi' z_{it}$.
• Apply the GFE estimation to the structural model (ref) whose endogenous variables $x_{it}$ are substituted with the corresponding fitted values $\hat x_{it}$:
\begin{align*}
\min_{\beta\in \mathcal B^G, \gamma\in\mathbb G^N}\ \frac{1}{NT} \sum_{i=1}^N\sum_{t=1}^T(y_{it}-\hat x_{it}'\beta_{g_i})^2.
\end{align*}
The minimizer $(\hat\beta,\hat\gamma)\equiv( \{ \hat\beta_g \}_{g=1}^G, \{\hat g_i \}_{i=1}^N)$ of this problem is the final estimate.
\end{enumerate}
Because this approach is a special case of that based on (ref), the asymptotic properties are examined in more general settings, and we do not mention them specifically for this approach.
Abusing terminology, we specifically refer to this procedure as “2SLS” estimation.
However, this approach may encounter a weak instrument problem AndrewsStockSun2019 in the presence of heterogeneity, even when IVs are relevant for each unit. The first-stage coefficient is a weighted average of heterogeneous coefficients and may not be sufficiently away from zero even when the coefficient for each unit is large. For example, if half of the units have positive first-stage coefficients and the other half has negative coefficients, then the first-stage coefficient in the homogeneous regression may be close to zero. We thus consider approaches that explicitly account for possible heterogeneity in the first stage.
\subsection{First stage with group structure}
Next, we consider the first-stage model (ref) where the coefficients exhibit a group structure.
The two-step estimation procedure proceeds in the following way, which, as a whole, we call the “two-stage GFE” (TGFE) estimation:
\begin{enumerate}[label=(\arabic*)]
• Apply the GFE estimation to first-stage model (ref):
\begin{align*}
\min_{\Pi\in\mathbf \Pi^K,\kappa\in\mathbb K^N}\ \frac{1}{NT}\sum_{i=1}^N\sum_{t=1}^T\|x_{it}-\Pi_{k_i}'z_{it}\|^2,
\end{align*}
where $\mathbf \Pi\subseteq \mathbb R^{md}$ denotes the common parameter space of the group-specific first-stage coefficient matrices $\Pi_k$'s.
Let $(\hat\Pi^\text{\normalfont tgfe},\hat\kappa^\text{\normalfont tgfe})\equiv (\{\hat\Pi_k^\text{\normalfont tgfe}\}_{k=1}^K,\{\hat k_i^\text{\normalfont tgfe}\}_{i=1}^N)$ be the minimizer of this problem.
Construct the fitted values $\hat x_{it}^\text{\normalfont tgfe}\equiv \hat \Pi_{\hat k_i^\text{\normalfont tgfe}}^{\text{\normalfont tgfe}\prime}z_{it}$.
• Apply the GFE estimation to the structural model (ref) whose endogenous variables $x_{it}$ are substituted with the fitted values $\hat x_{it}^\text{\normalfont tgfe}$:
\begin{align*}
\min_{\beta\in \mathcal B^G, \gamma\in\mathbb G^N}\ \frac{1}{NT} \sum_{i=1}^N\sum_{t=1}^T(y_{it}-\hat x_{it}^{\normalfont tgfe\prime}\beta_{g_i})^2.
\end{align*}
The minimizer $(\hat\beta^\text{\normalfont tgfe},\hat\gamma^\text{\normalfont tgfe})\equiv (\{\hat\beta_g^\text{\normalfont tgfe} \}_{g=1}^G, \{\hat g_i^\text{\normalfont tgfe} \}_{i=1}^N)$ of this problem is the final estimate.
\end{enumerate}
Separate GFE estimations are conducted in the first and second steps. By doing so, the procedure allows the group heterogeneity in the first-stage regression to be different from that in the coefficients of interest.
We examine the asymptotic properties of the TGFE estimator. We introduce a set of assumptions.
The superscript “fs” that appears in our notation indicates that it is associated with the first stage. Similarly, the superscript “ss” denotes that it is related to the second stage.
We make assumptions similar to those in BonhommeManresa2015,BonhommeManresa2015Supp.
Below, $\|\cdot\|$ denotes the Frobenius norm, and $\|\cdot\|_2$ the spectral norm.
\begin{assumption}
For some constant $M>0$,
\begin{enumerate}[label=(\alph*),ref=(\alph*)]
• Let $\mathbf \Pi$ and $\mathcal B$ be compact sets in $\mathbb R^{m\times d}$ and $\mathbb R^d$ respectively. For all $k\in\mathbb K$, $\Pi_k^0\in\mathbf \Pi$, and for all $g\in\mathbb G$, $\beta_g^0\in\mathcal B$.
• $\operatorname*{\mathbb E}[(NT)^{-1}\sum_{i=1}^N\|\sum_{t=1}^Tz_{it}v_{it}'\|^2]\leq M$.
• $\operatorname*{\mathbb E}[(NT)^{-1}\sum_{i=1}^N\|\sum_{t=1}^Tz_{it}u_{it}'\|^2]\leq M$.
• $\operatorname*{\mathbb E}[(NT)^{-1}\sum_{i=1}^N\|\sum_{t=1}^Tz_{it}z_{it}'\|_2] \leq M$.
• There exists $\underline \rho^\text{\normalfont ss}$, which depends on data, such that $\min_{\gamma\in\mathbb G^N}\max_{\tilde g\in\mathbb G}\rho^\text{\normalfont ss}(\gamma,g,\tilde g)\geq \underline \rho^\text{\normalfont ss} $, for all $g\in\mathbb G$, where $\rho^\text{\normalfont ss}(\gamma,g,\tilde g)$ is the smallest eigenvalue of
\begin{align*}
M^\normalfont ss(\gamma,g,\tilde g)
\equiv &\frac{1}{N}\sum_{i=1}^N\mathbf 1\{g_i^0=g\}\mathbf 1 \{g_i=\tilde g\}\frac{1}{T}\sum_{t=1}^T(\Pi_{k_i^0}^{0\prime}z_{it})(\Pi_{k_i^0}^{0\prime}z_{it})'
\end{align*}
and $\underline \rho^\text{\normalfont ss}\rightarrow \rho^{\text{\normalfont ss}*}>0$ as $N$ and $T$ tend to infinity.
• There exists $c^\text{\normalfont ss}>0$ such that
\begin{align*}
\operatorname*{plim}_{N,T\rightarrow \infty}\frac{1}{N}\sum_{i=1}^N\min_{g\neq g}D_{i,g,\tilde g}^\normalfont ss >c^\normalfont ss,
\end{align*}
where $D_{i,g,\tilde g}^\text{\normalfont ss} \equiv T^{-1}\sum_{t=1}^T((z_{it}'\Pi_{k_i^0}^0)(\beta_g^0-\beta_{\tilde g}^0))^2$.
\end{enumerate}
\end{assumption}
The compactness of parameter spaces in Assumption (ref)(ref) is standard in econometrics.
Assumptions (ref)(ref)--(ref) impose weak dependence of $z_{it}v_{it}'$ and $z_{it}u_{it}$, respectively.
Note also that Assumption (ref)(ref) corresponds to the exogeneity of IVs.
To see why, suppose that $(z_{it}u_{it})_{t=1}^T$ is i.i.d. across $i$ and $T^{-1}\sum_{t=1}^T\operatorname*{\mathbb E}[z_{it}u_{it}]=a_T$ for some constant $a_T\in\mathbb R^m$.
Then, because $M/T\geq (NT^2)^{-1}\sum_{i=1}^N\sum_{t=1}^T\sum_{s=1}^T \mathbb E[z_{it}z_{is}'u_{it}u_{is}] = \|a_T\|^2$, $a_T$ should shrink to zero as $T$ tends to infinity.
Assumption (ref)(ref) requires that the second-order moment of $z_{it}$ does not explode as $N$ and $T$ tend to infinity.
Assumption (ref)(ref) excludes multicollinearity among the population fitted values $\Pi_{k_i^0}^{0\prime}z_{it}$.
Assumption (ref)(ref) is called the “group separation condition” in the literature.
It requires that $\beta_g^0$ and $\beta_{\tilde g}^0$ have sufficiently different values.
Note that the IVs need to be relevant for Assumption (ref)(ref) to hold. It is violated if $\Pi_k^0=0$.
Under these assumptions, we establish the consistency of $\hat\beta^\text{\normalfont tgfe}$. We show that its Hausdorff distance from the true parameter $\beta^0$, which is defined as:
\begin{align*}
d_H(\beta,\beta^0)
\equiv \max \bigg\{\max_{g\in\mathbb G}\min_{\tilde g\in\mathbb G}\|\beta_g - \beta_{\tilde g}^0\|,\
\max_{\tilde g\in\mathbb G}\min_{g\in\mathbb G} \|\beta_g - \beta_{\tilde g}^0\|\bigg\},
\end{align*}
converges to zero in probability as $N$ and $T$ diverge.
\begin{theorem}
Suppose that Assumption (ref) holds for the two-stage model comprising equations (ref) and (ref). Then, as $N$ and $T$ tend to infinity,
\begin{align*}
d_H(\hat\beta^\normalfont tgfe,\beta^0)\rightarrow_p 0.
\end{align*}
\end{theorem}
Next, we proceed to the consistency of the estimated group memberships. The notion of consistency here refers to the fact that the probability of exact classification converges to one as $N$ and $T$ tend to infinity. Recall that the first and second stages exhibit different grouped patterns of heterogeneity. Correct classification in the first stage leads to correct classification in the second. We thus first state conditions for the consistency of the group membership estimator in the first stage.
We impose the following set of assumptions. They are used to show the consistency of $\hat\Pi^\text{\normalfont tgfe}$, which is a prerequisite of the consistency of $\hat \kappa$.
\begin{assumption}
\begin{enumerate}[label=(\alph*),ref=(\alph*)]
• There exists $\underline \rho^\text{\normalfont fs}$, which depends on data, such that for all $k\in\mathbb K$, $\min_{\gamma\in\mathbb G^N}\max_{\tilde k\in\mathbb K}\rho^\text{\normalfont fs}(\kappa,k,\tilde k)\geq \underline \rho^\text{\normalfont fs}$, where $\rho^\text{\normalfont fs}(\gamma,k,\tilde k)$ is the smallest eigenvalue of
\begin{align*}
M^\text{\normalfont fs}(\gamma,g,\tilde g)\equiv \frac{1}{N}\sum_{i=1}^N\mathbf 1\{k_i^0=k\}\mathbf 1 \{k_i=\tilde k\}\frac{1}{T}\sum_{t=1}^Tz_{it}z_{it}'
\end{align*}
and $\underline \rho^\text{\normalfont fs}\rightarrow \rho^{\text{\normalfont fs}*}>0$ as $N$ and $T$ tend to infinity.
• There exists $c^\text{\normalfont fs}>0$ such that
\begin{align*}
\operatorname*{plim}_{N,T\rightarrow \infty}\frac{1}{N}\sum_{i=1}^N\min_{k\neq \tilde k}D_{i,k,\tilde k}^\text{\normalfont fs} >c^\text{\normalfont fs}
\end{align*}
where $D_{i,k,\tilde k}^\text{\normalfont fs} \equiv T^{-1}\sum_{t=1}^T(z_{it}'(\Pi_{k}^0-\Pi_{\tilde k}^0))^2$.
\end{enumerate}
\end{assumption}
Assumption (ref)(ref) excludes the multicollinearity between IVs and requires that each group have sufficiently many members in the first stage.
Assumption (ref)(ref) is a first-stage version of the group separation condition.
In addition, impose the following set of assumptions for the consistency of $\hat\kappa^\text{\normalfont tgfe}$.
\begin{assumption}
\begin{enumerate}[label=(\alph*),ref=(\alph*)]
• $\lim_{T\rightarrow\infty}\min_{i\in\{1,\dots,N\}}\operatorname*{\mathbb E}[\min_{k\neq \tilde k}D_{i,k,\tilde k}^\text{\normalfont fs}]>c^\text{\normalfont fs}$.
• There exists a constant $M_z>0$ such that as $N$ and $T$ tend to infinity, for all $\delta>0$,
\begin{align*}
\max_{i\in\{1,\dots,N\}}\operatorname{Pr}\Biggl(\biggl\|\frac{1}{T}\sum_{t=1}^Tz_{it}z_{it}'\biggr\|_2>M_z\Biggr) &= o(T^{-\delta})
\end{align*}
• For every positive constant $c$, as $N$ and $T$ tend to infinity, for all $\delta>0$,
\begin{align*}
\max_{i\in\{1,\dots,N\}} \operatorname{Pr}\Biggl(\Biggl\|\frac{1}{T}\sum_{t=1}^T z_{it}v_{it}'\Biggr\|>c\Biggr) &= o(T^{-\delta})
\end{align*}
• There exist positive constants $a^\text{\normalfont fs}$ and $d^\text{\normalfont fs}$, and a sequence $\alpha^\text{\normalfont fs}[t]\leq \exp(-a^\text{\normalfont fs} t^{d^\text{\normalfont fs}})$ such that, for all $i=1,\dots,N$ and $(k,\tilde k)\in\mathbb K^2$ such that $k\neq \tilde k$,
$\{(\Pi_k^0 - \Pi_{\tilde k}^0)'z_{it}\}_t$ is a strongly mixing process with mixing coefficients $\alpha^\text{\normalfont fs}[t]$.
• There exist positive constants $b_z^\text{\normalfont fs}$ and $d_z^\text{\normalfont fs}$ such that, for all $c>0$, $\operatorname{Pr}(\|(\Pi_k^0 - \Pi_{\tilde k}^0)'z_{it}\|>c)\leq \exp(1-(c/b_z^\text{\normalfont fs})^{d_z^\text{\normalfont fs}})$.
\end{enumerate}
\end{assumption}
Assumptions (ref)(ref)--(ref) are satisfied, for example, if $(z_{it},v_{it})$ is Gaussian and independent over time. They indicate that the tails of the distributions of the averages of $||z_{it}||^2$ and $z_{it} v_{it}'$ vanish faster than any polynomial rate. For example, if they vanish at an exponential rate, the assumptions are satisfied.
Assumption (ref)(ref) restricts the serial dependence of $(\Pi_k^0-\Pi_{\tilde k}^0)'z_{it}$.
Assumption (ref)(ref) is related to its tail property, and requires that the probability mass on the tail decays at an exponential rate. We note that Assumptions (ref)(ref)--(ref) can be satisfied under conditions on the serial dependence and the tail of the distribution of $z_{it}$ and $z_{it} v_{it}'$ similar to Assumptions (ref)(ref) and Assumption (ref)(ref).
With these sets of assumptions, $\hat\kappa^\text{\normalfont tgfe}$ is shown to be consistent. The details are in the Appendix.
We now discuss the consistency of $\hat\gamma^\text{\normalfont tgfe}$.
The following set of assumptions is imposed.
\begin{assumption}
\begin{enumerate}[label=(\alph*),ref=(\alph*)]
• $\lim_{T\rightarrow\infty}\min_{i\in\{1,\dots,N\}}\operatorname*{\mathbb E}[\min_{g\neq \tilde g}D_{i,g,\tilde g}^\text{\normalfont ss}]>c^\text{\normalfont ss}$.
• For every positive constant $c$, as $N$ and $T$ tend to infinity, for all $\delta>0$,
\begin{align*}
\max_{i\in\{1,\dots,N\}} \operatorname{Pr}\Biggl(\Biggl\|\frac{1}{T}\sum_{t=1}^T z_{it}u_{it}\Biggr\|>c\Biggr) = o(T^{-\delta})
\end{align*}
• There exist positive constants $a^\text{\normalfont ss}$ and $d^\text{\normalfont ss}$, and a sequence $\alpha^\text{\normalfont ss}[t]\leq \exp(-a^\text{\normalfont ss} t^{d^\text{\normalfont ss}})$ such that for all $i=1,\dots,N$ and $(g,\tilde g)\in\mathbb G^2$ such that $g\neq \tilde g$, $\{ z_{it}'\Pi_{k_i^0}^0(\beta_g^0-\beta_{\tilde g}^0)\}_t$ is a strong mixing process with mixing coefficients $\alpha^\text{\normalfont ss}[t]$.
• There exist positive constants $b_x^\text{\normalfont ss}$ and $d_x^\text{\normalfont ss}$ such that, for all $c>0$, $\operatorname{Pr}(|z_{it}'\Pi_{k_i^0}^0(\beta_g^0-\beta_{\tilde g}^0)|>c)\leq \exp(1-(c/b_x^\text{\normalfont ss})^{d_x^\text{\normalfont ss}})$.
\end{enumerate}
\end{assumption}
Discussions analogous to those for Assumption (ref) apply here.
The following theorem states that the group membership estimators for each are uniformly consistent.
\begin{theorem}
Suppose that Assumptions (ref)--(ref) hold for the two-stage model comprising equations (ref) and (ref). Then, as $N$ and $T$ tend to infinity, for all $\delta>0$,
\begin{align*}
\operatorname{Pr}\biggl(\max_{i\in\{1,\dots,N\}}|\hat g_i^\text{\normalfont tgfe} - g_i^0|>0\biggr)=o(1) + o(NT^{-\delta}).
\end{align*}
\end{theorem}
Note that the required condition on the relative size of $N$ and $T$ is $NT^{-\delta} \to 0$, which is weak because $\delta$ can be as large as a researcher wants. It is violated if, for example, $T=\log N$, but is satisfied when $T$ is of geometric order $N$.
\subsection{Unit-specific first stage}
Finally, we consider the first-stage model (ref) that allows full heterogeneity across units.
We describe the following two-step estimation procedure, which we hereinafter call the “unit-specific first-stage GFE” (UGFE) estimation:
\begin{enumerate}[label=(\arabic*)]
• Apply least squares estimation to the first-stage model (ref):
\begin{align*}
\min_{(\Pi_i)_{i=1}^N\in\mathbf \Pi^N}\frac{1}{NT}\sum_{i=1}^N\sum_{t=1}^T\|x_{it}-\Pi_i'z_{it}\|^2,
\end{align*}
where $\mathbf \Pi$ denotes the common parameter space of the unit-specific first-stage coefficient matrices $\Pi_i$'s.
Let $\{\hat\Pi_i^\text{\normalfont ugfe}\}_{i=1}^N$ be the minimizer of this problem.
Construct the fitted values $\hat x_{it}^\text{\normalfont ugfe} \equiv \hat\Pi_i^{\text{\normalfont ugfe}\prime}z_{it}$.
Note that because $ \hat\Pi_i^\text{\normalfont ugfe} = \arg \min_{\Pi_i\in\mathbf \Pi}T^{-1}\sum_{t=1}^T\|x_{it}-\Pi_i'z_{it}\|^2$,
a closed-form solution is available (when $\mathbf \Pi$ is sufficiently large):
\begin{align*}
\hat\Pi_i^\text{\normalfont ugfe}
= \Biggl(\frac{1}{T}\sum_{t=1}^Tz_{it}z_{it}'\Biggr)^{-1} \frac{1}{T}\sum_{t=1}^Tz_{it}x_{it}'.
\end{align*}
• Apply GFE estimation to the structural model (ref) whose endogenous variables $x_{it}$ are substituted with the corresponding fitted values $\hat x_{it}^\text{\normalfont ugfe}$:
\begin{align*}
\min_{\beta\in\mathcal B^G,\gamma\in\mathbb G^N}\frac{1}{NT}\sum_{i=1}^N\sum_{t=1}^T(y_{it}-\hat x_{it}^{\text{\normalfont ugfe}\prime}\beta_{g_i})^2
\end{align*}
The minimizer $(\hat\beta^\text{\normalfont ugfe},\hat\gamma^\text{\normalfont ugfe}) \equiv (\{ \hat\beta_g^\text{\normalfont ugfe} \}_{g=1}^G,\{\hat g_i^\text{\normalfont ugfe} \}_{i=1}^N)$ of this problem is our final estimate.
\end{enumerate}
The derivation of the asymptotic properties of the UGFE estimator is like that of the TGFE estimator.
Indeed, the following two assumptions are very similar to Assumptions (ref) and (ref)--(ref), respectively, except that $\Pi_k^0$ and $\Pi_{k_i^0}^0$ are now replaced by $\Pi_i^0$.
We first demonstrate the consistency of $\hat\beta^\text{\normalfont ugfe}$ under the following set of assumptions.
\begin{assumption}
\begin{enumerate}[label=(\alph*),ref=(\alph*)]
• Let $\mathbf \Pi$ be a compact set in $\mathbb R^{m\times d}$. For all $i\in\{1,\dots,N\}$, $\Pi_i^0\in\mathbf \Pi$.
• There exists $\underline \rho^\text{\normalfont ss}$, which depends on data, such that $\min_{\gamma\in\mathbb G^N}\max_{\tilde g\in\mathbb G}\rho^\text{\normalfont ss}(\gamma,g,\tilde g)\geq \underline \rho^\text{\normalfont ss} $, for all $g\in\mathbb G$, where $\rho^\text{\normalfont ss}(\gamma,g,\tilde g)$ is the smallest eigenvalue of
\begin{align*}
M^\text{\normalfont ss}(\gamma,g,\tilde g)
\equiv &\frac{1}{N}\sum_{i=1}^N\mathbf 1\{g_i^0=g\}\mathbf 1 \{g_i=\tilde g\}\frac{1}{T}\sum_{t=1}^T(\Pi_i^{0\prime}z_{it})(\Pi_i^{0\prime}z_{it})'
\end{align*}
and $\underline \rho^\text{\normalfont ss}\rightarrow \rho^{\text{\normalfont ss}*}>0$ as $N$ and $T$ tend to infinity.
• There exists $c^\text{\normalfont ss}>0$ such that
\begin{align*}
\operatorname*{plim}_{N,T\rightarrow \infty}\frac{1}{N}\sum_{i=1}^N\min_{g\neq \tilde g}D_{i,g,\tilde g}^\text{\normalfont ss} >c^\text{\normalfont ss}
\end{align*}
where $D_{i,g,\tilde g}^\text{\normalfont ss} \equiv T^{-1}\sum_{t=1}^T(z_{it}'\Pi_i^0(\beta_g^0-\beta_{\tilde g}^0))^2$.
\end{enumerate}
\end{assumption}
\begin{theorem}
Suppose that Assumptions (ref)(ref)--(ref) and (ref) hold for the two-stage model comprising equations (ref) and (ref). Then, as $N$ and $T$ tend to infinity,
\begin{align*}
d_H(\hat\beta^\text{\normalfont ugfe},\beta^0)\rightarrow_p 0.
\end{align*}
\end{theorem}
Next comes the consistency of $\hat\gamma^\text{\normalfont ugfe}$ under the following additional set of assumptions.
\begin{assumption}
\begin{enumerate}[label=(\alph*),ref=(ref)(\alph*)]
• $\lim_{T\rightarrow\infty}\min_{i\in\{1,\dots,N\}}\operatorname*{\mathbb E}[\min_{g\neq \tilde g}D_{i,g,\tilde g}^\text{\normalfont ss}]>c^\text{\normalfont ss}$.
• There exist positive constants $a^\text{\normalfont ss}$ and $d^\text{\normalfont ss}$, and a sequence $\alpha^\text{\normalfont ss}[t]\leq \exp(-a^\text{\normalfont ss} t^{d^\text{\normalfont ss}})$ such that for all $i=1,\dots,N$ and $(g,\tilde g)\in\mathbb G^2$ such that $g\neq \tilde g$, $\{z_{it}'\Pi_i^0(\beta_g^0-\beta_{\tilde g}^0\}_t$ is a strong mixing process with mixing coefficients $\alpha^\text{\normalfont ss}[t]$.
• There exist positive constants $b_x^\text{\normalfont ss}$ and $d_x^\text{\normalfont ss}$ such that, for all $c>0$, $\operatorname{Pr}(|z_{it}'\Pi_i^0(\beta_g^0-\beta_{\tilde g}^0)|>c)\leq \exp(1-(c/b_x^\text{\normalfont ss})^{d_x^\text{\normalfont ss}})$.
\end{enumerate}
\end{assumption}
\begin{theorem}
Suppose that Assumptions (ref)(ref)--(ref), (ref)(ref)--(ref), (ref)(ref), and (ref)--(ref) hold for the two-stage model comprising equations (ref) and (ref). Then, as $N$ and $T$ tend to infinity, for all $\delta>0$,
\begin{align*}
\operatorname{Pr}\biggl(\max_{i\in\{1,\dots,N\}}|\hat g_i^\text{\normalfont ugfe} - g_i^0|>0\biggr)=o(1) + o(NT^{-\delta}).
\end{align*}
\end{theorem}
\begin{remark}
The estimator can be rewritten as
\begin{align*}
(\hat\beta^\text{\normalfont ugfe},\hat\gamma^\text{\normalfont ugfe}) = \arg \min_{\beta\in\mathcal B^G, \gamma\in\mathbb G^N} \sum_{i=1}^N\biggl(\sum_{t=1}^T z_{it}(y_{it} - x_{it}'\beta_{g_i})\biggr)\biggl(\sum_{t=1}^T z_{it}z_{it}'\biggr)^{-1}\biggl(\sum_{t=1}^T z_{it}(y_{it} - x_{it}'\beta_{g_i})\biggr),
\end{align*}
which implies that we can run the UGFE estimation in a single stage. Note that this objective function is the unpenalized version of the PGMM objective function in SuShiPhillips2016 and Mehrabani2022 with the weighting matrix for unit $i$ set to $(\sum_{t=1}^T z_{it}z_{it}')^{-1}$. itosugasawa2023 consider a similar objective function for generalized estimating equation models.
\end{remark}
\subsection{Choice of number of groups}
The number of latent groups may not be known a priori, and methods to choose it are required.
We propose data-driven procedures that can select the correct number of groups in the limit.
Our selection procedures follow PesaranSmith1994. We use the residuals from the second-stage estimation. Note that the second-stage residuals are different from the residuals of the structural equation.
We start with the TGFE estimation. Both the first and second stages require choosing the number of groups, and the two stages require different procedures.
The information criteria for each stage are
\begin{align*}
\operatorname{IC}^{\text{\normalfont tgfe},\text{\normalfont fs}}(K)
\equiv& \frac{1}{NT}\sum_{i=1}^N\sum_{t=1}^T\|x_{it}-\hat\Pi_{\hat k_i^{\text{\normalfont tgfe}\operatorname{(K)}}}^{\text{\normalfont tgfe}\operatorname{(K)}\prime}z_{it}\|^2 + \frac{c^{\text{\normalfont tgfe},\text{\normalfont fs}}(K,N,T)}{NT} \text{ and} \\
\operatorname{IC}^{\text{\normalfont tgfe},\text{\normalfont ss}}(G,K) \equiv& \frac{1}{NT}\sum_{i=1}^N\sum_{t=1}^T(y_{it}-(z_{it}'\hat\Pi_{\hat k_i^{\text{\normalfont tgfe}\operatorname{(K)}}}^{\text{\normalfont tgfe}\operatorname{(K)}})\hat\beta_{\hat g_i^{\text{\normalfont tgfe}\operatorname{(G,K)}}}^{\text{\normalfont tgfe}\operatorname{(G,K)}})^2
+ \frac{c^{\text{\normalfont tgfe},\text{\normalfont ss}}(G,N,T)}{NT},
\end{align*}
where the superscripts $\operatorname{(K)}$, $\operatorname{(G)}$, and $\operatorname{(G,K)}$ denote that the given estimate is obtained assuming that the numbers of groups in the first and second stages are $K$ and $G$ respectively. $c^{\text{\normalfont tgfe},\text{\normalfont fs}}$ and $c^{\text{\normalfont tgfe},\text{\normalfont ss}}$ are functions having the stated arguments.
Following Mallows1973, the literature suggests using $c^{\text{\normalfont tgfe},\text{\normalfont fs}}(K,N,T)=\hat{\sigma^{\text{\normalfont tgfe},\text{\normalfont fs}}}^2(mdK+N)\log(NT)$ and $c^{\text{\normalfont tgfe},\text{\normalfont ss}}(K,N,T)=\hat{\sigma^{\text{\normalfont tgfe},\text{\normalfont ss}}}^2(dG+N)\log(NT)$, where $\hat{\sigma^{\text{\normalfont tgfe},\text{\normalfont fs}}}^2\equiv (NT)^{-1}\sum_{i=1}^N\sum_{t=1}^T\|x_{it}-\hat\Pi_{\hat k_i^{\text{\normalfont tgfe}\operatorname{(K^{\max})}}}^{\text{\normalfont tgfe}\operatorname{(K^{\max})}\prime}z_{it}\|^2$ and $\smash{\hat{\sigma^{\text{\normalfont tgfe},\text{\normalfont ss}}}^2\equiv (NT)^{-1}\sum_{i=1}^N\sum_{t=1}^T(y_{it}-(z_{it}'\hat\Pi_{\hat k_i^{\text{\normalfont tgfe}\operatorname{(K^{\max})}}}^{\text{\normalfont tgfe}\operatorname{(K^{\max})}})\hat\beta_{\hat g_i^{\text{\normalfont tgfe}\operatorname{(G^{\max},K^{\max})}}}^{\text{\normalfont tgfe}\operatorname{(G^{\max},K^{\max})}})^2}$.
The information criterion for the first stage is standard. The first term is the sum of squared residuals, and the second is a penalty term. The information criterion for the second stage relies on the residuals from the second-stage GFE estimation, which are different from the residuals from the structural equation: $y_{it}-x_{it}\hat\beta_{\hat g_i^{\text{\normalfont tgfe}\operatorname{(G^{\max},K^{\max})}}}^{\text{\normalfont tgfe}\operatorname{(G^{\max},K^{\max})}}$. Let
\begin{align*}
\hat K^\text{\normalfont tgfe} \equiv \arg \min_{K \in \{1,\dots,K^{\max}\}}\operatorname{IC}^{\text{\normalfont tgfe},\text{\normalfont fs}}(K) \text{ and }
\hat G^\text{\normalfont tgfe}(K) \equiv \arg \min_{G\in\{1,\dots,G^{\max}\}}\operatorname{IC}^{\text{\normalfont tgfe},\text{\normalfont ss}}(G,K)
\end{align*}
be the corresponding minimizers. The estimated number of groups is then $\hat G^\text{\normalfont tgfe} \equiv \hat G^\text{\normalfont tgfe}(\hat K^\text{\normalfont tgfe})$.
Next, we consider the UGFE estimation. Define:
\begin{align*}
\operatorname{IC}^{\text{\normalfont ugfe}}(G) \equiv \frac{1}{NT}\sum_{i=1}^N\sum_{t=1}^T(y_{it}-(z_{it}'\hat\Pi_i^\text{\normalfont ugfe})\hat\beta_{\hat g_i^{\text{\normalfont ugfe}\operatorname{(G)}}}^{\text{\normalfont ugfe}\operatorname{(G)}})^2
+ \frac{c^{\text{\normalfont ugfe},\text{\normalfont ss}}(G,N,T)}{NT}.
\end{align*}
The estimated number of groups is the minimizer $\hat G^\text{\normalfont ugfe}\equiv \arg \min_{G\in\{1,\dots,G^{\max}\}}\operatorname{IC}^{\text{\normalfont ugfe}}(G)$ of this criterion function.
As before, we can employ $c^{\text{\normalfont ugfe},\text{\normalfont ss}}(G,N,T)=\hat{\sigma^{\text{\normalfont ugfe},\text{\normalfont ss}}}^2(dG+N)\log(NT)$, where $\hat{\sigma^{\text{\normalfont ugfe},\text{\normalfont ss}}}^2\equiv (NT)^{-1}\sum_{i=1}^N\sum_{t=1}^T(y_{it}-(z_{it}'\hat\Pi_i^\text{\normalfont ugfe})\hat\beta_{\hat g_i^{\text{\normalfont ugfe}\operatorname{(G^{\max})}}}^{\text{\normalfont ugfe}\operatorname{(G^{\max})}})^2$.
\subsection{Other possible procedures}
We introduce two additional estimation strategies. They are simple and easy to implement. However, they only yield consistent estimators under possibly strong additional assumptions.
\paragraph{Ignoring endogeneity.} The first strategy is to apply GFE estimation to the structural model (ref) ignoring the presence of endogeneity between $x_{it}$ and $u_{it}$. This approach directly solves:
\begin{align*}
\min_{\beta\in\mathcal B^G,\gamma\in\mathbb G^N}\frac{1}{NT}\sum_{i=1}^N\sum_{t=1}^T(y_{it} - x_{it}'\beta_{g_i})^2.
\end{align*}
Let $(\hat\beta^\text{\normalfont ig},\hat\gamma^\text{\normalfont ig})$ be the minimizer of this problem.
Even if $\hat\beta^\text{\normalfont ig}$ suffers from endogeneity bias, $\hat\gamma^\text{\normalfont ig}$ is still consistent if the coefficient of the linear projection of $y_{it}$ on $x_{it}$ for each $i$---$\operatorname*{plim}_{T\to \infty} (\sum_{t=1}^T x_{it} x_{it}')^{-1} \sum_{t=1}^T x_{it} y_{it}$---is also grouped by $\gamma^0$.
This condition is satisfied, for example, when $(x_{it},u_{it})$ are i.i.d. across $i$ and $t$.
Then, the coefficient of the linear projection is $\beta_{g_i^0}^0 + (\mathbb E[x_{it}x_{it}'])^{-1} \mathbb E[x_{it}u_{it}]$, which exhibits the same grouped pattern as $\beta_{g_i^0}^0$.
Once the condition is assumed, we can rely on the existing results in the literature for exogenous cases such as those in BonhommeManresa2015.
That is, we can invoke their assumptions to show the consistency of $\hat\gamma^\text{\normalfont ig}$.
However, justifying the required condition in any given empirical situation would be difficult. The relevant assumptions should be made concerning the model of linear projection, not to the original model. It is unclear whether an economic theory can support the conditions on linear projection. There are cases in which the relevant condition is violated.
For example, if the degree of endogeneity, namely $\mathbb E[x_{it}u_{it}]$, varies with a distinct grouped pattern from $\beta_{g_i^0}^0$, $\hat\gamma^\text{\normalfont ig}$ may not be consistent.
This approach is considered in BonhommeManresa2015Supp in the framework of dynamic linear models.
Furthermore, a conceptually similar strategy has been employed for break detection in the time series literature. PerronYamamoto2015 and QianSu2014, among others, consider the detection of break points in models with endogenous regressors.
The authors first estimate the break point while ignoring endogeneity and then, given that break point, estimate regime-specific model parameters.
\paragraph{Reduced form.} The second strategy examines the reduced form, namely, the relationship between the dependent variable and the IVs. This solves:
\begin{align*}
\min_{\xi\in\Xi^G,\lambda\in\mathbb H^N}\frac{1}{NT}\sum_{i=1}^N\sum_{t=1}^T(y_{it} - z_{it}'\xi_{l_i})^2,
\end{align*}
where $\xi\equiv \{\xi_g \}_{g=1}^G\subseteq \Xi^G$ and $\lambda\equiv \{l_i\}_{i=1}^N\in \mathbb G^N$.
Let $(\hat\xi^\text{\normalfont rf},\hat\lambda^\text{\normalfont rf})$ be the minimizer of this problem. Coefficient estimates are obtained by IV estimation for each group given by $\hat\lambda^\text{\normalfont rf}$.
If the reduced form and the structural model share the same group structure, $\hat\lambda^\text{\normalfont rf}$ is consistent for $\gamma^0$.
To examine the situations in which consistency holds, consider the first-stage model (ref). The reduced form is:
\begin{align*}
y_{it} = z_{it}'(\Pi_{k_i^0}^0\beta_{g_i^0}^0) + (v_{it}'\beta_{g_i^0}^0 + u_{it}).
\end{align*}
When $\Pi_{k_i^0}^0\beta_{g_i^0}^0$ exhibit the same grouped pattern as $\beta_{g_i^0}^0$, the consistency of $\hat\lambda^\text{\normalfont rf}$ is guaranteed by BonhommeManresa2015Supp.
This condition is satisfied if, for example, the first-stage model has no group structure, that is, $\Pi_k^0=\Pi^0$.
However, when the first-stage model and the structural model exhibit different group structures, the reduced form may have a different group structure, and the condition is violated.
\section{Extensions}
This section discusses two extensions of our baseline model.
We begin with models with unit-specific fixed effects.
In this case, we simply need to consider the within or first-differencing transformed models and apply the estimation procedures discussed above to transformed models.
As an important example, we discuss dynamic linear panel data models.
Then, we consider models with time-varying group-specific fixed effects. This extension is important because previous results concerning IV estimation did not handle this case, yet it is practically relevant.
We establish a new theory to accommodate the situation that the dimension of the coefficient vector increases as $T$ increases.
As in the previous section, we separate our theoretical results based on the model we assume for the first-stage regression.
\subsection{Unit-specific fixed effects}
The first extension adds unit-specific fixed effects to our baseline model.
Specifically, we consider the model
\begin{align}
y_{it} = x_{it}'\beta_{g_i^0}^0 + \eta_i + u_{it},
\end{align}
where $\{\eta_i\}_{i=1}^N$ are unit-specific fixed effects and are allowed to be correlated with the IVs, $z_{it}$.
The estimation starts with eliminating the unit-specific fixed effects using the within or first-differencing transformation, each of which transforms the model into
\begin{align*}
\ddot y_{it} &= \ddot x_{it}' \beta_{g_i^0}^0 + \ddot u_{it}
\intertext{or}
\Delta y_{it} &= \Delta x_{it}' \beta_{g_i^0}^0 + \Delta u_{it},
\end{align*}
where, for a random variable or vector $w_{it}$, $\ddot w_{it} \equiv w_{it} - \sum_{t=1}^T w_{it}/T$ and $\Delta w_{it} \equiv w_{it} - w_{it-1}$.
The coefficient vectors $\{\beta_g^0\}_{g=1}^G$ and the group membership structure $\{g_i^0\}_{i=1}^N$ are then estimated by applying a method discussed in Section (ref) to the transformed model.
Note that we need to impose the relevant assumptions on the transformed variables instead of the original ones in level.
For example, in order to use the TGFE estimator, Assumptions (ref)--(ref) should be applied for $(\ddot x_{it},\ddot u_{it})$ or $(\Delta x_{it},\Delta u_{it})$, not for $(x_{it},u_{it})$.
In particular, the exogeneity of IVs, $z_{it}$, needs to hold with respect to transformed error terms ($\ddot u_{it}$ or $\Delta u_{it}$).
This requirement implies that the choice of the transformation depends on the IVs, $z_{it}$, or alternatively, that the choice of instruments depends on a transformation.
The consistency of the TGFE estimator is then secured by Theorems (ref)--(ref).
\paragraph{Dynamic linear panel data models.}
As an illustration, we examine a panel autoregressive model with lag order 1 (panel AR(1) model):
\begin{align}
y_{it} = \beta_{g_i^0}^0 y_{i,t-1} + \eta_i + u_{it},
\end{align}
where $\{\beta_g^0\}_{g=1}^G$ are the group-specific AR(1) coefficients.
The arguments that follow can be readily generalized for models with more lags.
Applying our procedures to this model encounters an additional challenge. The number of IVs differs across time periods. The first-differencing transformation transforms model (ref) into
\begin{align*}
\Delta y_{it} = \beta_{g_i^0}^0\Delta y_{it-1} + \Delta u_{it}.
\end{align*}
Because $\Delta y_{it-1}$ and $\Delta u_{it}$ are systematically correlated, the literature \parencites{AndersonCheng1981,AndersonCheng1982,ArellanoBond91,HoltzEakinNeweyRosen88} suggests using various combination of lags $y_{it-2},\dots,y_{i0}, \Delta y_{it-2},\dots,\Delta y_{i1}$ as IVs for $\Delta y_{it-1}$.
No issue arises if one uses the same number of lags for each period.
The estimation methods in Section (ref) will suffice.
However, if we intend to vary the number for efficiency gain, the direct application of our methods becomes infeasible, since they are based on the fixed dimension of the IVs.
Nevertheless, a simple solution is available.
The solution we propose is to combine IVs linearly.
Let $z_{it}$ be an $m_t$ dimensional IV vector for $\Delta y_{it-1}$.
We suggest using linearly transformed $\Gamma_t'z_{it}$, where $\Gamma_t$ is an $m_t \times d$ matrix that possibly depends on data, as new IVs.
Because their dimension is fixed at $d$, our methods are then executable.
A candidate for $\Gamma_t$ is
\begin{align*}
\biggl(\frac{1}{N}\sum_{i=1}^N z_{it} z_{it}'\biggr)^{-1}\frac{1}{N}\sum_{i=1}^N z_{it}\Delta y_{it-1},
\end{align*}
which is obtained by running the cross-sectional regression of $\Delta y_{it-1}$ on $z_{it}$ at period $t$.\footnote{Wooldridge05,Wooldridge10 shows that, in the absence of a latent group structure, that is, when $G=1$, the estimation using instruments $(\sum_{i=1}^N z_{it}\Delta y_{it-1})'(\sum_{i=1}^N z_{it} z_{it}')^{-1}z_{it}$ is equivalent to GMM estimation using all the available moment restrictions, that is, $\mathbb E(z_{i1}\Delta u_{i1})=0, \dots,\mathbb E(z_{iT}\Delta u_{iT})=0$.}
\subsection{Time-varying group-specific fixed effects}
The second extension adds time-varying group-specific effects to our baseline model:
\begin{align}
y_{it} = x_{it}'\beta_{g_i^0}^0 + \alpha_{g_i^0 t}^0 + u_{it}.
\end{align}
where $\{\alpha_{gt}^0\}_{t=1}^T$ represents the time profile shared by the units in group $g$.
The main motivation of including $\alpha_{g_i^0 t}^0$ is to avoid omitted variable bias caused by time-varying unobserved confounders. BonhommeManresa2015 and mugnier2022simple consider models with time-varying group-specific effects. Furthermore, we allow the coefficients on $x_{it}$ to exhibit a grouped pattern of heterogeneity.
The first-stage models (ref) and (ref) are adjusted to
\begin{align}
x_{it}=& \Pi_{k_i^0}^{0\prime}z_{it} + \mu_{k_i^0 t}^0 +v_{it}
\intertext{and}
x_{it}=& \Pi_i^{0\prime}z_{it} + \mu_{k_i^0 t}^0 + v_{it}
\end{align}
respectively. Note that, in model (ref), time-varying effects are kept group-specific.
If they were unit-specific, then the first-stage coefficients $\{\Pi_i^0\}_{i=1}^N$ would not be identified.
\subsubsection{First stage with group structure}
We start by considering the first-stage model (ref).
We define the following two-step procedure, which we also call the “two-stage GFE” (TGFE) estimation:
\begin{enumerate}[label=(\arabic*)]
• Apply the GFE estimation to the first-stage model (ref):
\begin{align*}
\min_{\Pi\in\mathbf \Pi^K,\mu\in\mathcal M^{KT},\kappa\in\mathbb K^N}\ \frac{1}{NT}\sum_{i=1}^N\sum_{t=1}^T\|x_{it}-\Pi_{k_i}'z_{it}-\mu_{k_i t}\|^2,
\end{align*}
where $\mu\equiv (\mu_k)_{k=1}^K$, $\mu_k\equiv (\mu_{kt})_{t=1}^T$, and $\mathcal M$ denotes the common parameter space of $\mu_{kt}$'s.
Let $(\hat\Pi^\text{\normalfont tgfe},\hat\mu^\text{\normalfont tgfe},\hat\kappa^\text{\normalfont tgfe})\equiv (\{\hat\Pi_k^\text{\normalfont tgfe}\}_{k=1}^K,\{\hat\mu_k^\text{\normalfont tgfe}\}_{k=1}^K,\{\hat k_i^\text{\normalfont tgfe}\}_{i=1}^N)$ be the minimizer of this problem.
Construct the fitted values $\hat x_{it}^\text{\normalfont tgfe}\equiv \hat \Pi_{\hat k_i^\text{\normalfont tgfe}}^{\text{\normalfont tgfe}\prime}z_{it} + \hat\mu_{\hat k_i^\text{\normalfont tgfe}}^\text{\normalfont tgfe}$.
• Apply the GFE estimation to the structural model (ref) whose endogenous variables $x_{it}$ are substituted with the fitted values $\hat x_{it}^\text{\normalfont tgfe}$:
\begin{align*}
\min_{\beta\in \mathcal B^G,\alpha\in\mathcal A^{GT},\gamma\in\mathbb G^N}\ \frac{1}{NT} \sum_{i=1}^N\sum_{t=1}^T(y_{it}-\hat x_{it}^{\text{\normalfont tgfe}\prime}\beta_{g_i}-\alpha_{g_i t})^2,
\end{align*}
where $\alpha\equiv \{(\alpha_{gt})_{t=1}^T\}_{g=1}^G$ and $\mathcal A$ denotes the common parameter space of $\alpha_{gt}$'s.
The minimizer $(\hat\beta^\text{\normalfont tgfe},\hat\alpha^\text{\normalfont tgfe},\hat\gamma^\text{\normalfont tgfe})\equiv (\{\hat\beta_g^\text{\normalfont tgfe} \}_{g=1}^G, \{(\hat\alpha_{gt})_{t=1}^T\}_{g=1}^G,\{\hat g_i^\text{\normalfont tgfe} \}_{i=1}^N)$ of this problem becomes the final estimate.
\end{enumerate}
Next, we study its asymptotic properties.
We introduce the following additional set of assumptions to Assumption (ref).
\begin{assumption}
For some constant $M>0$,
\begin{enumerate}[label=(\alph*),ref=(\alph*)]
• Let $\mathcal M$ and $\mathcal A$ be compact sets in $\mathbb R^d$ and $\mathbb R$, respectively. For all $k\in\mathbb K$, $\mu_{kt}^0\in\mathcal M$, and for all $g\in\mathbb G$, $\alpha_{gt}^0\in\mathcal A$,
• $(NT)^{-1}\sum_{i=1}^N\sum_{j=1}^N\sum_{t=1}^T|\operatorname*{\mathbb E}[u_{it}u_{jt}]|\leq M$.
• $|(N^2T)^{-1}\sum_{i=1}^N\sum_{j=1}^N\sum_{t=1}^T\sum_{s=1}^T\operatorname*{Cov}[u_{it}u_{jt},u_{is}u_{js}]|\leq M$
• $(NT)^{-1}\sum_{i=1}^N\sum_{j=1}^N\sum_{t=1}^T|\operatorname*{\mathbb E}[v_{it}'v_{jt}]|\leq M$.
• $|(N^2T)^{-1}\sum_{i=1}^N\sum_{j=1}^N\sum_{t=1}^T\sum_{s=1}^T\operatorname*{Cov}[v_{it}'v_{jt},v_{is}'v_{js}]|\leq M$
• There exists $\underline \rho^\text{\normalfont ss}$, which depends on data, such that $\min_{\gamma\in\mathbb G^N}\max_{\tilde g\in\mathbb G}\rho^\text{\normalfont ss}(\gamma,g,\tilde g)\geq \underline \rho^\text{\normalfont ss} $, for all $g\in\mathbb G$, where $\rho^\text{\normalfont ss}(\gamma,g,\tilde g)$ is the smallest eigenvalue of
\begin{align*}
&M^\text{\normalfont ss}(\gamma,g,\tilde g)\\
\equiv& \frac{1}{N}\sum_{i=1}^N\mathbf 1\{g_i^0=g\}\mathbf 1 \{g_i=\tilde g\}\\
\times &
\begin{pmatrix}
\frac{1}{T}\sum_{t=1}^T(\Pi_{k_i^0}^{0\prime}z_{it} + \mu_{k_i^0t}^0)(\Pi_{k_i^0}^{0\prime}z_{it} + \mu_{k_i^0t}^0)' & \frac{1}{\sqrt{T}}(\Pi_{k_i^0}^{0\prime}z_{i1} + \mu_{k_i^01}^0) & \cdots & \frac{1}{\sqrt{T}}(\Pi_{k_i^0}^{0\prime}z_{iT} + \mu_{k_i^0T}^0)\\
\frac{1}{\sqrt{T}}(\Pi_{k_i^0}^{0\prime}z_{i1} + \mu_{k_i^01}^0)' & 1 & \cdots & 0 \\
\vdots & 0 & \ddots & 0 \\
\frac{1}{\sqrt{T}}(\Pi_{k_i^0}^{0\prime}z_{iT} + \mu_{k_i^0T}^0)' & 0 & \cdots & 1 \\
\end{pmatrix}
\end{align*}
and $\underline \rho^\text{\normalfont ss}\rightarrow \rho^{\text{\normalfont ss}*}>0$ as $N$ and $T$ tend to infinity.
• There exists $c^\text{\normalfont ss}>0$ such that
\begin{align*}
\operatorname*{plim}_{N,T\rightarrow \infty}\frac{1}{N}\sum_{i=1}^N\min_{g\neq \tilde g}D_{i,g,\tilde g}^\text{\normalfont ss} >c^\text{\normalfont ss}
\end{align*}
where $D_{i,g,\tilde g}^\text{\normalfont ss} \equiv T^{-1}\sum_{t=1}^T((z_{it}'\Pi_{k_i^0}^0+\mu_{k_i^0 t}^0)(\beta_g^0-\beta_{\tilde g}^0) + (\alpha_{g t}^0 - \alpha_{\tilde g t}^0))^2$.
\end{enumerate}
\end{assumption}
In Assumption (ref)(ref), we require the support of the group-specific time fixed effects to be compact.
Assumptions (ref)(ref) and (ref)(ref) restrict the degree of cross-sectional dependence of the first- and second-stage errors,
and Assumptions (ref)(ref) and (ref)(ref) bear on their weak dependence along the time dimension.
With the aid of these assumptions, we establish the consistency of $(\hat\beta^\text{\normalfont tgfe},\hat\alpha^\text{\normalfont tgfe})$. Define:
\begin{align*}
d_H(\alpha,\alpha^0) \equiv \max\biggl\{\max_{g\in\mathbb G}\min_{\tilde g\in\mathbb G}\frac{1}{T}\sum_{t=1}^T\|\alpha_{gt}-\alpha_{\tilde gt}^0\|^2,\max_{\tilde g\in\mathbb G}\min_{g\in\mathbb G}\frac{1}{T}\sum_{t=1}^T\|\alpha_{gt}-\alpha_{\tilde gt}^0\|^2\biggr\}.
\end{align*}
\begin{theorem}
Suppose that Assumptions (ref)(ref)--(ref) and (ref) hold for the two-stage model comprising equations (ref) and (ref). Then, as $N$ and $T$ tend to infinity,
\begin{align*}
d_H(\hat\beta^\text{\normalfont tgfe},\beta^0) \rightarrow_p 0
\text{ and }
d_H(\hat\alpha^\text{\normalfont tgfe},\alpha^0) \rightarrow_p 0.
\end{align*}
\end{theorem}
Our next result concerns the uniform consistency of $\hat\gamma^\text{\normalfont tgfe}$, for which the following assumptions are employed.
\begin{assumption}
\begin{enumerate}[label=(\alph*),ref=(\alph*)]
• There exists $\underline \rho^\text{\normalfont fs}$, which depends on data, such that for all $k\in\mathbb K$, $\min_{\gamma\in\mathbb G^N}\max_{\tilde k\in\mathbb K}\rho^\text{\normalfont fs}(\kappa,k,\tilde k)\geq \underline \rho^\text{\normalfont fs}$, where $\rho^\text{\normalfont fs}(\gamma,k,\tilde k)$ is the smallest eigenvalue of
\begin{align*}
M^\text{\normalfont fs}(\kappa,k,\tilde k)
\equiv &\frac{1}{N}\sum_{i=1}^N\mathbf 1\{k_i^0=k\}\mathbf 1 \{k_i=\tilde k\}
\begin{pmatrix}
\frac{1}{T}\sum_{t=1}^Tz_{it}z_{it}' & \frac{1}{\sqrt{T}}z_{i1} & \cdots & \frac{1}{\sqrt{T}}z_{iT}\\
\frac{1}{\sqrt{T}}z_{i1}' & 1 & \cdots & 0 \\
\vdots & 0 & \ddots & 0 \\
\frac{1}{\sqrt{T}}z_{iT}' & 0 & \cdots & 1 \\
\end{pmatrix}
\end{align*}
and $\underline \rho^\text{\normalfont fs}\rightarrow \rho^{\text{\normalfont fs}*}>0$ as $N$ and $T$ tend to infinity.
• There exists $c^\text{\normalfont fs}>0$ such that
\begin{align*}
\operatorname*{plim}_{N,T\rightarrow \infty}\frac{1}{N}\sum_{i=1}^N\min_{k\neq \tilde k}D_{i,k,\tilde k}^\text{\normalfont fs} >c^\text{\normalfont fs}
\end{align*}
where $D_{i,k,\tilde k}^\text{\normalfont fs} \equiv T^{-1}\sum_{t=1}^T\|(\Pi_{k}^0-\Pi_{\tilde k}^0)'z_{it} + (\mu_{k t}^0 - \mu_{\tilde k t}^0)\|^2$.
\end{enumerate}
\end{assumption}
\begin{assumption}
\begin{enumerate}[label=(\alph*),ref=(\alph*)]
• $\lim_{T\rightarrow\infty}\min_{i\in\{1,\dots,N\}}\operatorname*{\mathbb E}[\min_{k\neq \tilde k}D_{i,k,\tilde k}^\text{\normalfont fs}]>c^\text{\normalfont fs}$.
• There exist positive constants $a^\text{\normalfont fs}$ and $d^\text{\normalfont fs}$, and a sequence $\alpha^\text{\normalfont fs}[t]\leq \exp(-a^\text{\normalfont fs} t^{d^\text{\normalfont fs}})$ such that, for all $i\in\{1,\dots,N\}$ and $(k,\tilde k)\in\mathbb K^2$ such that $k\neq \tilde k$, $\{v_{it}\}_t$, $\{\mu_{k_i^0t}^0v_{it}'\}_t$, $\{v_{it}'(\mu_{k t}^0 - \mu_{\tilde kt}^0)\}_t$, and $\{(\Pi_k^0 - \Pi_{\tilde k}^0)'z_{it} + (\mu_{k t}^0 - \mu_{\tilde k t}^0)\}_t$ are strongly mixing processes with mixing coefficients $\alpha^\text{\normalfont fs}[t]$. Moreover, $\operatorname*{\mathbb E}[\mu_{k_i^0t}^0v_{it}']=0$ and $\operatorname*{\mathbb E}[v_{it}'(\mu_{k t}^0 - \mu_{\tilde kt}^0)]=0$.
• There exist positive constants $b_v^\text{\normalfont fs}$ and $b_z^\text{\normalfont fs}$, and $d_v^\text{\normalfont fs}$ and $d_{z}^\text{\normalfont fs}$ such that, for all $c>0$,
$\operatorname{Pr}(|v_{it}|>c)\leq \exp(1-(c/b_v^\text{\normalfont fs})^{d_v^\text{\normalfont fs}})$
and $\operatorname{Pr}(\|(\Pi_k^0 - \Pi_{\tilde k}^0)'z_{it} + (\mu_{k t}^0 - \mu_{\tilde k t}^0)\|>c)\leq \exp(1-(c/b_z^\text{\normalfont fs})^{d_{z}^\text{\normalfont fs}})$.
\end{enumerate}
\end{assumption}
\begin{assumption}
\begin{enumerate}[label=(\alph*),ref=(\alph*)]
• $\lim_{T\rightarrow\infty}\min_{i\in\{1,\dots,N\}}\operatorname*{\mathbb E}[\min_{k\neq \tilde k}D_{i,k,\tilde k}^\text{\normalfont ss}]>c^\text{\normalfont ss}$.
• There exist positive constants $a^\text{\normalfont ss}$ and $d^\text{\normalfont ss}$, and a sequence $\alpha^\text{\normalfont ss}[t]\leq \exp(-a^\text{\normalfont ss} t^{d^\text{\normalfont ss}})$ such that, for all $i\in\{1,\dots,N\}$ and $(g,\tilde g)\in\mathbb G^2$ such that $g\neq \tilde g$,
$\{u_{it}\}_t$, $\{\mu_{k_i^0t}^0u_{it}\}_t$,
$\{v_{it}(\alpha_{g t}^0 - \alpha_{\tilde g t}^0)\}_t$,
$\{u_{it}(\alpha_{g t}^0 - \alpha_{\tilde g t}^0)\}_t$, and
$\{(z_{it}'\Pi_{k_i^0 t}^0+\mu_{k_i^0 t}^0)(\beta_g^0-\beta_{\tilde g}^0) + (\alpha_{g t}^0 - \alpha_{\tilde g t}^0)\}_t$ are strongly mixing with mixing coefficients $\alpha^\text{\normalfont ss}[t]$.
Moreover, $\operatorname*{\mathbb E}[\mu_{k_i^0t}^0u_{it}]=0$, $\operatorname*{\mathbb E}[v_{it}(\alpha_{g t}^0 - \alpha_{\tilde g t}^0)]=0$, and $\operatorname*{\mathbb E}[u_{it}(\alpha_{g t}^0 - \alpha_{\tilde g t}^0)]=0$.
• There exist positive constants $b_{u}^\text{\normalfont ss}$ and $b_x^\text{\normalfont ss}$, and $d_{u}^\text{\normalfont ss}$ and $d_x^\text{\normalfont ss}$ such that, for all $c>0$,
$\operatorname{Pr}(|u_{it}|>c)\leq \exp(1-(c/b_u^\text{\normalfont ss})^{d_u^\text{\normalfont ss}})$ and
$\operatorname{Pr}(|(z_{it}'\Pi_{k_i^0}^0 + \mu_{k_i^0 t}^0)(\beta_g^0-\beta_{\tilde g}^0) + (\alpha_{g t}^0 - \alpha_{\tilde g t}^0)|>c)\leq \exp(1-(c/b_x^\text{\normalfont ss})^{d_x^\text{\normalfont ss}})$.
\end{enumerate}
\end{assumption}
We obtain the same rate of uniform convergence as Theorem (ref).
\begin{theorem}
Suppose that Assumptions (ref)(ref)--(ref), (ref)(ref)--(ref), (ref)(ref), and (ref)--(ref) hold for the two-stage model comprising equations (ref) and (ref). Then, as $N$ and $T$ tend to infinity, for all $\delta>0$,
\begin{align*}
\operatorname{Pr}\biggl(\max_{i\in\{1,\dots,N\}}|\hat g_i^\text{\normalfont tgfe} - g_i^0|>0\biggr)=o(1) + o(NT^{-\delta}).
\end{align*}
\end{theorem}
\subsubsection{Unit-specific first stage}
Now we consider the first-stage model (ref), where each unit can have different first-stage linear projection coefficients.
The estimation procedure proceeds in two steps, which is called the “unit-specific first-stage GFE” (UGFE) estimation as before:
\begin{enumerate}[label=(\arabic*)]
• Apply the least squares estimation to the first-stage model (ref):
\begin{align*}
\min_{(\Pi_i)_{i=1}^N\in\mathbf \Pi^N, \mu\in\mathcal M^{KT},\kappa\in\mathbb K^N}\frac{1}{NT}\sum_{i=1}^N\sum_{t=1}^T\|x_{it}-\Pi_i'z_{it}-\mu_{k_i t}\|^2.
\end{align*}
Let $(\hat\Pi^\text{\normalfont ugfe},\hat\mu^\text{\normalfont ugfe},\hat\kappa^\text{\normalfont ugfe})\equiv (\{\hat\Pi_i^\text{\normalfont ugfe}\}_{i=1}^N,\{\hat\mu_k^\text{\normalfont ugfe}\}_{k=1}^K,\{\hat k_i^\text{\normalfont ugfe}\}_{i=1}^N)$ denote the minimizer of this problem.
Construct the fitted values $\hat x_{it}^\text{\normalfont ugfe} \equiv \hat\Pi_i^{\text{\normalfont ugfe}\prime}z_{it} + \hat\mu_{\hat k_i^\text{\normalfont ugfe}}^\text{\normalfont ugfe}$.
• Apply the GFE estimation to the structural model (ref) whose endogenous variables $x_{it}$ are substituted with the corresponding fitted values $\hat x_{it}^\text{\normalfont ugfe}$:
\begin{align*}
\min_{\beta\in\mathcal B^G,\alpha\in\mathcal A^{GT},\gamma\in\mathbb G^N}\frac{1}{NT}\sum_{i=1}^N\sum_{t=1}^T(y_{it}-\hat x_{it}^{\text{\normalfont ugfe}\prime}\beta_{g_i}-\alpha_{g_i t})^2.
\end{align*}
The minimizer $(\hat\beta^\text{\normalfont ugfe},\hat\alpha^\text{\normalfont ugfe},\hat\gamma^\text{\normalfont ugfe}) \equiv (\{\hat\beta_g^\text{\normalfont ugfe} \}_{g=1}^G,\{(\hat\alpha_{gt}^\text{\normalfont ugfe})_{t=1}^T\}_{g=1}^G, \{\hat g_i^\text{\normalfont ugfe}\}_{i=1}^N)$ of this problem is the final estimate.
\end{enumerate}
The following assumptions are needed for the consistency of $(\hat\beta^\text{\normalfont ugfe},\hat\alpha^\text{\normalfont ugfe})$.
\begin{assumption}
\begin{enumerate}[label=(\alph*),ref=(\alph*)]
• There exists $\underline \rho^\text{\normalfont ss}$, which depends on data, such that for all $g\in\mathbb G$, $\min_{\gamma\in\mathbb G^N}\max_{\tilde g\in\mathbb G}\rho^\text{\normalfont ss}(\gamma,g,\tilde g)\geq \underline \rho^\text{\normalfont ss}$, where $\rho^\text{\normalfont ss}(\gamma,g,\tilde g)$ is the smallest eigenvalue of
\begin{align*}
&M^\text{\normalfont ss}(\gamma,g,\tilde g)\\
\equiv &\frac{1}{N}\sum_{i=1}^N\mathbf 1\{g_i^0=g\}\mathbf 1 \{g_i=\tilde g\}\\
\times &
\begin{pmatrix}
\frac{1}{T}\sum_{t=1}^T(\Pi_i^{0\prime}z_{it} + \mu_{k_i^0t}^0)(\Pi_i^{0\prime}z_{it} + \mu_{k_i^0t}^0)' & \frac{1}{\sqrt{T}}(\Pi_i^{0\prime}z_{i1} + \mu_{k_i^01}^0) & \cdots & \frac{1}{\sqrt{T}}(\Pi_i^{0\prime}z_{iT} + \mu_{k_i^0T}^0)\\
\frac{1}{\sqrt{T}}(\Pi_i^{0\prime}z_{i1} + \mu_{k_i^01}^0)' & 1 & \cdots & 0 \\
\vdots & 0 & \ddots & 0 \\
\frac{1}{\sqrt{T}}(\Pi_i^{0\prime}z_{iT} + \mu_{k_i^0T}^0)' & 0 & \cdots & 1 \\
\end{pmatrix}
\end{align*}
and $\underline \rho^\text{\normalfont ss}\rightarrow \rho^{\text{\normalfont ss}*}>0$ as $N$ and $T$ tend to infinity.
• There exists $c^\text{\normalfont ss}>0$ such that
\begin{align*}
\operatorname*{plim}_{N,T\rightarrow \infty}\frac{1}{N}\sum_{i=1}^ND_{i,g,\tilde g}^\text{\normalfont ss} >c^\text{\normalfont ss},
\end{align*}
where $D_{i,g,\tilde g}^\text{\normalfont ss} \equiv T^{-1}\sum_{t=1}^T((z_{it}'\Pi_i^0+\mu_{k_i^0 t}^0)(\beta_g^0-\beta_{\tilde g}^0) + (\alpha_{g t}^0 - \alpha_{\tilde g t}^0))^2$.
\end{enumerate}
\end{assumption}
Our result then follows:
\begin{theorem} Suppose that Assumptions (ref)(ref)--(ref), (ref)(ref), (ref)(ref)--(ref), and (ref) hold for the two-stage model comprising equations (ref) and (ref). Then, as $N$ and $T$ tend to infinity,
\begin{align*}
d_H(\hat\beta^\text{\normalfont ugfe},\beta^0)\rightarrow_p 0
\text{ and }
d_H(\hat\alpha^\text{\normalfont ugfe},\alpha^0) \rightarrow_p 0.
\end{align*}
\end{theorem}
The next three assumptions are needed for the uniform consistency of $\hat\gamma^\text{\normalfont ugfe}$.
\begin{assumption}
\begin{enumerate}[label=(\alph*),ref=(\alph*)]
• There exists $\underline \rho^\text{\normalfont fs}$, which depends on data, such that for all $k\in\mathbb K$, $\min_{\kappa\in\mathbb K^N}\max_{\tilde k\in\mathbb K}\rho(\gamma,k,\tilde k) \geq \underline \rho^\text{\normalfont fs} $, where $\rho(\gamma,k,\tilde k)$ is the smallest eigenvalue of
\begin{align*}
M_i^\text{\normalfont ss}(\kappa,k,\tilde k)
\equiv & \frac{1}{N}\sum_{i=1}^N\mathbf 1\{k_i^0=k\}\mathbf 1\{k_i=\tilde k\}I_T
\end{align*}
and $\underline \rho^\text{\normalfont fs}\rightarrow \rho^{\text{\normalfont fs}*}>0$ as $N$ and $T$ tend to infinity.
• There exists $c^\text{\normalfont fs}>0$ such that
\begin{align*}
\lim_{T\rightarrow\infty} \operatorname*{\mathbb E}[\min_{k\neq \tilde k}D_{k,\tilde k}^\text{\normalfont fs}] >c^\text{\normalfont fs},
\end{align*}
where $D_{k,\tilde k}^\text{\normalfont fs} \equiv T^{-1}\sum_{t=1}^T\|\mu_{k t}^0 - \mu_{\tilde k t}^0\|^2$.
\end{enumerate}
\end{assumption}
\begin{assumption}
\begin{enumerate}[label=(\alph*),ref=(\alph*)]
• There exist positive constants $a^\text{\normalfont fs}$ and $d^\text{\normalfont fs}$, and a sequence $\alpha^\text{\normalfont fs}[t]\leq \exp(-a^\text{\normalfont fs} t^{d^\text{\normalfont fs}})$ such that, for all $i\in\{1,\dots,N\}$ and $(k,\tilde k)\in\mathbb K^2$ such that $k\neq \tilde k$ and
$\{v_{it} \}_t$, $\{v_{it}\mu_{k_i^0t}^{0\prime}\}_t$, $\{v_{it}'(\mu_{k t}^0 - \mu_{\tilde kt}^0)\}_t$ are strongly mixing processes with mixing coefficients $\alpha^\text{\normalfont fs}[t]$. Moreover, $\operatorname*{\mathbb E}[v_{it}\mu_{k_i^0t}^{0\prime}]=0$ and $\operatorname*{\mathbb E}[v_{it}'(\mu_{k t}^0 - \mu_{\tilde kt}^0)]=0$.
• There exist positive constants $b_v^\text{\normalfont fs}$ and $d_v^\text{\normalfont fs}$ such that, for all $c>0$,
$\operatorname{Pr}(\|v_{it}\|>c)\leq \exp(1-(c/b_v^\text{\normalfont fs})^{d_v^\text{\normalfont fs}})$.
\end{enumerate}
\end{assumption}
\begin{assumption}
\begin{enumerate}[label=(\alph*),ref=(\alph*)]
• $\lim_{T\rightarrow\infty}\min_{i\in\{1,\dots,N\}}\operatorname*{\mathbb E}[\min_{g\neq \tilde g}D_{i,g,\tilde g}^\text{\normalfont ss}]>c^\text{\normalfont ss}$.
• There exist positive constants $a^\text{\normalfont ss}$ and $d^\text{\normalfont ss}$, and a sequence $\alpha^\text{\normalfont ss}[t]\leq \exp(-a^\text{\normalfont ss} t^{d^\text{\normalfont ss}})$ such that, for all $i\in\{1,\dots,N\}$ and $(g,\tilde g)\in\mathbb G^2$ such that $g\neq \tilde g$,
$\{u_{it}\}_t$,
$\{u_{it}\mu_{k_i^0t}^0\}_t$,
$\{v_{it}(\alpha_{g t}^0 - \alpha_{\tilde g t}^0)\}_t$,
$\{u_{it}(\alpha_{g t}^0 - \alpha_{\tilde g t}^0)\}_t$, and
$\{(z_{it}'\Pi_i^0+\mu_{k_i^0 t}^0)(\beta_g^0-\beta_{\tilde g}^0) + (\alpha_{g t}^0 - \alpha_{\tilde g t}^0)\}_t$ are strongly mixing processes with mixing coefficients $\alpha^\text{\normalfont ss}[t]$.
Moreover, $\operatorname*{\mathbb E}[u_{it}\mu_{k_i^0t}^0]=0$, $\operatorname*{\mathbb E}[v_{it}(\alpha_{g t}^0 - \alpha_{\tilde g t}^0)]=0$, and $\operatorname*{\mathbb E}[u_{it}(\alpha_{g t}^0 - \alpha_{\tilde g t}^0)]=0$.
• There exist positive constants $b_{u}^\text{\normalfont ss}$ and $b_x^\text{\normalfont ss}$, and $d_{u}^\text{\normalfont ss}$ and $d_x^\text{\normalfont ss}$ such that, for all $c>0$,
$\operatorname{Pr}(|u_{it}|>c)\leq \exp(1-(c/b_u^\text{\normalfont ss})^{d_u^\text{\normalfont ss}})$ and
$\operatorname{Pr}(|(z_{it}'\Pi_i^0 + \mu_{k_i^0 t}^0)(\beta_g^0-\beta_{\tilde g}^0) + (\alpha_{g t}^0 - \alpha_{\tilde g t}^0)|>c)\leq \exp(1-(c/b_x^\text{\normalfont ss})^{d_x^\text{\normalfont ss}})$.
\end{enumerate}
\end{assumption}
We arrive at the same rate of uniform convergence as Theorem (ref).
\begin{theorem} Suppose that Assumptions (ref)(ref)--(ref), (ref)(ref)--(ref), (ref)(ref), (ref)(ref), (ref)(ref)--(ref), (ref)--(ref) hold for the two-stage model comprising equations (ref) and (ref). Then, as $N$ and $T$ tend to infinity,
\begin{align*}
\operatorname{Pr}\biggl(\max_{i\in\{1,\dots,N\}}|\hat g_i^\text{\normalfont ugfe} - g_i^0|>0\biggr) = o(1) + o(NT^{-\delta}).
\end{align*}
\end{theorem}
\section{Monte Carlo simulations}
We evaluate the finite sample performance of our estimators and confirm their consistency using simulations under various specifications.
Let $d=m=1$ and $\mathbb G = \{1,2\}$.
The structural and first-stage equations are then streamlined into $y_{it} = \beta_{g_i^0}^0 x_{it} + u_{it}$ and $x_{it} = \Pi_i^0 z_{it} + v_{it}$, where $\beta_{g_i^0}^0$ and $\Pi_i^0$ are now scalars.
\subsection{Data-generating process}
We consider four DGPs that are differentiated by their first-stage modeling.
We start by specifying the distribution of the underlying variables $(z_{it},v_{it},u_{it})$.
The error terms $(v_{it},u_{it})$ are generated from a jointly normal distribution whose mean is 0, respective variances are $\sigma^2$ and 1, and correlation is $\rho_i$. Note that the correlation can be unit-specific.
An IV, $z_{it}$, is drawn independently of $(v_{it},u_{it})$ and follows a normal distribution with mean 0 and variance $\sigma^2$.
$(z_{it},v_{it},u_{it})$ are drawn independently across $i$ and $t$.
Summing up, the underlying variables are generated by:
\begin{align*}
\begin{pmatrix}
z_{it} \\ v_{it} \\ u_{it}
\end{pmatrix}
\sim_{i.i.d.} \mathcal N\left(\begin{pmatrix}
0 \\ 0 \\ 0
\end{pmatrix}
, \begin{pmatrix}
\sigma^2 & 0 & 0 \\
0 & \sigma^2 & \sigma \cdot \rho_i \\
0 & \sigma \cdot \rho_i & 1
\end{pmatrix}\right),
\end{align*}
for each $t= 1,\dots, T$, and are independently drawn across $i$, where the distribution of $\rho_i$ varies across the DGPs.
Next, we specify the structural equation.
Every DGP shares the same grouped pattern of the coefficient parameter $\beta_i^0$, which is defined as follows:
\begin{align*}
\beta_i^0 = \begin{cases}
\ 1 &\text{if $i=1,\dots, N/2$}\\
\ -1 &\text{if $i=N/2+1,\dots,N$}
\end{cases}.
\end{align*}
Namely, $g_i^0=1$ for $i=1,\dots,N/2$ and $g_i^0=2$ for $i=N/2+1,\dots,N$.
We list the first-stage specifications of each DGP:
\begin{description}
• $\Pi_i^0=1$ for all $i$. $\rho_i=\rho$ for all $i$.
• $\Pi_i^0=1$ for odd $i$ and $\Pi_i^0=-1$ for even $i$. $\rho_i=\rho$ for all $i$.
• $\Pi_i^0\sim_{i.i.d}\text{Unif}_{[0.5,1.5]}$ for $i=1,\dots, N/2$ and $\Pi_i^0\sim_{i.i.d}\text{Unif}_{[-1.5,-0.5]}$ for $i=N/2+ 1,\dots, N$. $\rho_i=\rho$ for all $i$.
• $\Pi_i^0=1$ for all $i$. $\rho_i=\rho$ for $i=1,\dots, N/2$ and $\rho_i=-\rho$ for $i=N/2+ 1,\dots, N$,
\end{description}
where $\rho=-0.5$. We set $\sigma \in\{0.5,0.75\}$.
DGP 1 is our baseline. There, the first stage is homogeneous across units.
In DGP 2, the first stage exhibits a group structure that is orthogonal to that of the structural equation.
In DGP 3, each unit has a different first stage.
We note that while the IV is relevant for each $i$, the cross-sectional correlation between the IV and the regressor is zero in DGPs 2--3.
In DGP 4, the correlation between the two error terms shows a different grouped pattern from that of the structural equation. In this specification, the linear projection of $y_{it}$ on $x_{it}$ exhibits a different group structure from that of the structural equation.
\subsection{Comparison of performance}
We investigate the finite sample performance of our procedures under these DGPs.
We use the Rand index (RI) Rand1971 to measure their classification accuracy.
The RI is a commonly used measure for evaluating the similarity between two classifications.
The formal definition is as follows. We take two group membership structures: $\gamma_1=\{g_i^1\}_{i=1}^N$ and $\gamma_2=\{g_i^2\}_{i=1}^N$. Let $N_{11}$ be the number of pairs $(i,j)$ such that $g_i^1=g_j^1$ and $g_i^2=g_j^2$, $N_{10}$ the number of pairs $(i,j)$ such that $g_i^1=g_j^1$ and $g_i^2\neq g_j^2$, $N_{01}$ the number of pairs $(i,j)$ such that $g_i^1\neq g_j^1$ and $g_i^2= g_j^2$, and $N_{00}$ the number of pairs $(i,j)$ such that $g_i^1\neq g_j^1$ and $g_i^2\neq g_j^2$.
The RI between these two classifications is then calculated by $(N_{11}+N_{00})/(N_{11}+N_{10}+N_{01}+N_{00})$.
The average (over Monte Carlo replications) RI between the true group structure $\gamma^0$ and those estimated from each procedure measures its classification performance.
\begin{table}[hp]
\begin{center}
\resizebox{\textwidth}{!}{
\begin{threeparttable}
\caption{Average Rand index}
\begin{tabular}{lllcccccccccc}
\toprule
\multicolumn{3}{c} & \multicolumn{5}{c}{$\sigma=0.5$} & \multicolumn{5}{c}{$\sigma=0.75$} \\
\cmidrule(l){4-8} \cmidrule(l){9-13}
& $N$ & $T$ & IG & TSLS & TGFE$_2$ & TGFE$_{N/4}$ & UGFE & IG & TSLS & TGFE$_2$ & TGFE$_{N/4}$ & UGFE \\
\midrule
\midrule\multirow{6}{*}{\textbf{DGP 1}} & 50 & 5 & 0.856 & 0.714 & 0.717 & 0.718 & 0.717 & 0.953 & 0.799 & 0.805 & 0.810 & 0.814 \\
& 50 & 10 & 0.956 & 0.819 & 0.824 & 0.824 & 0.825 & 0.994 & 0.902 & 0.902 & 0.907 & 0.907 \\
& 50 & 20 & 0.998 & 0.936 & 0.935 & 0.934 & 0.934 & 1.000 & 0.977 & 0.978 & 0.977 & 0.977 \\
& 100 & 5 & 0.861 & 0.713 & 0.718 & 0.722 & 0.723 & 0.943 & 0.789 & 0.795 & 0.801 & 0.802 \\
& 100 & 10 & 0.959 & 0.823 & 0.828 & 0.830 & 0.830 & 0.995 & 0.903 & 0.904 & 0.907 & 0.907 \\
& 100 & 20 & 0.997 & 0.930 & 0.930 & 0.929 & 0.929 & 1.000 & 0.981 & 0.980 & 0.981 & 0.981 \\
\midrule\multirow{6}{*}{\textbf{DGP 2}} & 50 & 5 & 0.840 & 0.496 & 0.701 & 0.709 & 0.710 & 0.948 & 0.493 & 0.795 & 0.800 & 0.799 \\
& 50 & 10 & 0.959 & 0.493 & 0.840 & 0.847 & 0.845 & 0.996 & 0.491 & 0.911 & 0.911 & 0.910 \\
& 50 & 20 & 0.996 & 0.491 & 0.945 & 0.944 & 0.944 & 1.000 & 0.490 & 0.981 & 0.981 & 0.982 \\
& 100 & 5 & 0.853 & 0.498 & 0.708 & 0.720 & 0.721 & 0.951 & 0.497 & 0.805 & 0.811 & 0.811 \\
& 100 & 10 & 0.962 & 0.496 & 0.833 & 0.837 & 0.837 & 0.996 & 0.496 & 0.907 & 0.910 & 0.910 \\
& 100 & 20 & 0.996 & 0.496 & 0.931 & 0.933 & 0.933 & 1.000 & 0.495 & 0.980 & 0.980 & 0.980 \\
\midrule\multirow{6}{*}{\textbf{DGP 3}} & 50 & 5 & 0.862 & 0.505 & 0.699 & 0.707 & 0.708 & 0.951 & 0.501 & 0.798 & 0.802 & 0.803 \\
& 50 & 10 & 0.945 & 0.499 & 0.806 & 0.814 & 0.816 & 0.991 & 0.498 & 0.879 & 0.886 & 0.887 \\
& 50 & 20 & 0.992 & 0.500 & 0.904 & 0.907 & 0.906 & 1.000 & 0.500 & 0.955 & 0.957 & 0.958 \\
& 100 & 5 & 0.848 & 0.502 & 0.700 & 0.707 & 0.707 & 0.947 & 0.499 & 0.788 & 0.792 & 0.792 \\
& 100 & 10 & 0.948 & 0.502 & 0.794 & 0.804 & 0.803 & 0.991 & 0.500 & 0.888 & 0.892 & 0.893 \\
& 100 & 20 & 0.992 & 0.502 & 0.905 & 0.910 & 0.910 & 1.000 & 0.500 & 0.952 & 0.956 & 0.956 \\
\midrule\multirow{6}{*}{\textbf{DGP 4}} & 50 & 5 & 0.548 & 0.831 & 0.828 & 0.772 & 0.766 & 0.793 & 0.945 & 0.947 & 0.885 & 0.885 \\
& 50 & 10 & 0.613 & 0.954 & 0.953 & 0.942 & 0.940 & 0.921 & 0.992 & 0.994 & 0.983 & 0.983 \\
& 50 & 20 & 0.700 & 0.995 & 0.995 & 0.994 & 0.994 & 0.983 & 1.000 & 1.000 & 1.000 & 1.000 \\
& 100 & 5 & 0.555 & 0.842 & 0.838 & 0.786 & 0.787 & 0.806 & 0.951 & 0.950 & 0.891 & 0.891 \\
& 100 & 10 & 0.605 & 0.954 & 0.955 & 0.943 & 0.942 & 0.928 & 0.994 & 0.994 & 0.985 & 0.984 \\
& 100 & 20 & 0.703 & 0.996 & 0.995 & 0.995 & 0.995 & 0.988 & 1.000 & 1.000 & 1.000 & 1.000 \\
\bottomrule
\end{tabular}
\begin{tablenotes}
• Note: The value in each cell denotes the mean of the Rand index across 100 simulations, with 100 starting group memberships randomly generated for each estimate.
\end{tablenotes}
\end{threeparttable}
}
\end{center}
\end{table}
Table (ref) displays the average RI of our estimators under the above DGPs.
For TGFE, we consider two specifications of first-stage heterogeneity, $K=2$ and $K=N/4$. These are denoted as TGFE$_{2}$ and TGFE$_{N/4}$, respectively.
We omit the reduced-form approach because it yields estimates identical to the 2SLS estimates in just-identified cases.\footnote{The second stage in the TSLS estimation minimizes $\sum_{i=1}^N \sum_{t=1}^T (y_{it} - z_{it}' (\hat \Pi ' \beta_{g_i}))^2$. In just-identified cases, $\hat \Pi$ is invertible when the IVs are relevant. Thus, setting $\hat \Pi' \beta_{g_i} =\xi_{l_i}$ makes the minimization problem equivalent to minimizing $\sum_{i=1}^N \sum_{t=1}^T (y_{it} - z_{it}'\xi_{l_i})^2$.}
Every estimator performs well when predicted to be so by asymptotic theory.
Overall, an increase in $N$ does not contribute much to classification accuracy as opposed to $T$.
The 2SLS estimator displays high classification performance in DGPs 1 and 4, both of which assume a homogeneous first stage. In particular, in DGP 4, 2SLS far outperforms the other competing estimators when $T$ is small.
However, it suffers from the weak instrument problem in DGPs 2--3, which
indicates the importance of modeling the first stage as exhibiting a group structure.
The TGFE and UGFE estimators perform well in all specifications.
A notable observation is that the TGFE$_2$ estimator performs similarly to the UGFE estimator even under DGP 3, where the first stage exhibits full heterogeneity.
The result indicates that the TGFE estimators may be employed as an approximation tool for the UGFE estimator.
Indeed, the TGFE$_{N/4}$ estimator performs very similarly to the UGFE estimator in all cases.
The IG estimator performs very well in DGPs 1--3, but poorly in DGP 4, especially when $\sigma=0.5$ even after we increase $N$ and $T$.
Intuitively, this is because the true group structure $\gamma^0$ and the grouped pattern $\{\rho_i\}_{i=1}^N$ in the correlation between the two error terms are tangled up, and the IG estimator is unable to disentangle them.
\begin{table}[hp]
\begin{center}
\resizebox{0.9\textwidth}{!}{
\begin{threeparttable}
\caption{Average Hausdorff distance (Pre-estimates)}
\begin{tabular}{lllcccccccc}
\toprule
\multicolumn{3}{c} & \multicolumn{4}{c}{$\sigma=0.5$} & \multicolumn{4}{c}{$\sigma=0.75$} \\
\cmidrule(l){4-7} \cmidrule(l){8-11}
& $N$ & $T$ & TSLS & TGFE$_2$ & TGFE$_{N/4}$ & UGFE & TSLS & TGFE$_2$ & TGFE$_{N/4}$ & UGFE \\
\midrule
\midrule\multirow{6}{*}{\textbf{DGP 1}} & 50 & 5 & 0.438 & 0.334 & 0.328 & 0.326 & 0.266 & 0.200 & 0.190 & 0.190 \\
& 50 & 10 & 0.251 & 0.213 & 0.212 & 0.212 & 0.137 & 0.127 & 0.120 & 0.120 \\
& 50 & 20 & 0.129 & 0.130 & 0.132 & 0.132 & 0.081 & 0.078 & 0.079 & 0.080 \\
& 100 & 5 & 0.394 & 0.276 & 0.267 & 0.266 & 0.240 & 0.174 & 0.180 & 0.181 \\
& 100 & 10 & 0.240 & 0.198 & 0.192 & 0.192 & 0.103 & 0.101 & 0.103 & 0.103 \\
& 100 & 20 & 0.101 & 0.105 & 0.107 & 0.107 & 0.047 & 0.052 & 0.054 & 0.054 \\
\midrule\multirow{6}{*}{\textbf{DGP 2}} & 50 & 5 & 115.135 & 0.445 & 0.306 & 0.301 & 61.235 & 0.257 & 0.198 & 0.201 \\
& 50 & 10 & 73.606 & 0.211 & 0.180 & 0.182 & 41.249 & 0.149 & 0.130 & 0.131 \\
& 50 & 20 & 51.571 & 0.118 & 0.116 & 0.116 & 86.899 & 0.075 & 0.071 & 0.071 \\
& 100 & 5 & 91.446 & 0.425 & 0.285 & 0.287 & 69.831 & 0.230 & 0.178 & 0.179 \\
& 100 & 10 & 109.139 & 0.202 & 0.176 & 0.176 & 185.538 & 0.125 & 0.115 & 0.115 \\
& 100 & 20 & 93.052 & 0.098 & 0.097 & 0.099 & 544.140 & 0.053 & 0.058 & 0.058 \\
\midrule\multirow{6}{*}{\textbf{DGP 3}} & 50 & 5 & 73.604 & 0.463 & 0.315 & 0.310 & 161.554 & 0.253 & 0.171 & 0.171 \\
& 50 & 10 & 138.895 & 0.232 & 0.176 & 0.177 & 106.551 & 0.152 & 0.115 & 0.116 \\
& 50 & 20 & 83.750 & 0.145 & 0.119 & 0.119 & 67.040 & 0.086 & 0.073 & 0.074 \\
& 100 & 5 & 87.780 & 0.419 & 0.260 & 0.260 & 168.426 & 0.252 & 0.163 & 0.164 \\
& 100 & 10 & 111.934 & 0.264 & 0.169 & 0.169 & 111.968 & 0.118 & 0.086 & 0.087 \\
& 100 & 20 & 169.544 & 0.120 & 0.093 & 0.093 & 139.661 & 0.070 & 0.061 & 0.061 \\
\midrule\multirow{6}{*}{\textbf{DGP 4}} & 50 & 5 & 0.185 & 0.250 & 0.319 & 0.319 & 0.115 & 0.171 & 0.215 & 0.218 \\
& 50 & 10 & 0.134 & 0.153 & 0.189 & 0.190 & 0.079 & 0.108 & 0.139 & 0.139 \\
& 50 & 20 & 0.092 & 0.107 & 0.125 & 0.126 & 0.054 & 0.074 & 0.088 & 0.089 \\
& 100 & 5 & 0.145 & 0.185 & 0.260 & 0.262 & 0.081 & 0.157 & 0.209 & 0.210 \\
& 100 & 10 & 0.089 & 0.126 & 0.178 & 0.179 & 0.058 & 0.088 & 0.117 & 0.118 \\
& 100 & 20 & 0.064 & 0.087 & 0.109 & 0.110 & 0.039 & 0.054 & 0.068 & 0.068 \\
\bottomrule
\end{tabular}
\begin{tablenotes}
• Note: The value in each cell denotes the mean of the Hausdorff distance across 100 simulations, with 100 starting group memberships randomly generated for each estimate.
\end{tablenotes}
\end{threeparttable}
}
\end{center}
\end{table}
Table (ref) presents the average Hausdorff distance between the estimates and the true coefficient parameter $\beta^0$.
The IG estimation is omitted because it only provides estimates for the group memberships.
The results overall follow those in Table (ref).
As an estimator attains higher classification accuracy, its estimate also approaches the true coefficient value.
\begin{table}[hp]
\begin{center}
\resizebox{\textwidth}{!}{
\begin{threeparttable}
\caption{Average Hausdorff distance (Post-estimates)}
\begin{tabular}{lllcccccccccc}
\toprule
\multicolumn{3}{c} & \multicolumn{5}{c}{$\sigma=0.5$} & \multicolumn{5}{c}{$\sigma=0.75$} \\
\cmidrule(l){4-8} \cmidrule(l){9-13}
& $N$ & $T$ & IG & 2SLS & TGFE$_2$ & TGFE$_{N/4}$ & UGFE & IG & 2SLS & TGFE$_2$ & TGFE$_{N/4}$ & UGFE \\
\midrule
\midrule\multirow{6}{*}{\textbf{DGP 1}} & 50 & 5 & 0.204 & 0.319 & 0.312 & 0.281 & 0.279 & 0.146 & 0.187 & 0.180 & 0.159 & 0.157 \\
& 50 & 10 & 0.138 & 0.187 & 0.179 & 0.184 & 0.183 & 0.083 & 0.097 & 0.095 & 0.094 & 0.094 \\
& 50 & 20 & 0.100 & 0.116 & 0.113 & 0.114 & 0.114 & 0.068 & 0.068 & 0.068 & 0.068 & 0.068 \\
& 100 & 5 & 0.156 & 0.283 & 0.275 & 0.225 & 0.223 & 0.106 & 0.157 & 0.152 & 0.134 & 0.134 \\
& 100 & 10 & 0.103 & 0.174 & 0.165 & 0.161 & 0.161 & 0.065 & 0.076 & 0.074 & 0.073 & 0.073 \\
& 100 & 20 & 0.075 & 0.085 & 0.084 & 0.084 & 0.084 & 0.043 & 0.041 & 0.041 & 0.041 & 0.041 \\
\midrule\multirow{6}{*}{\textbf{DGP 2}} & 50 & 5 & 7.825 & 27.009 & 5.349 & 4.829 & 5.070 & 4.324 & 45.869 & 11.318 & 4.829 & 5.160 \\
& 50 & 10 & 3.518 & 52.256 & 6.476 & 5.172 & 9.523 & 3.921 & 83.225 & 4.600 & 14.287 & 14.309 \\
& 50 & 20 & 39.294 & 64.611 & 44.333 & 44.502 & 44.502 & 4.887 & 2411.873 & 4.852 & 4.691 & 4.695 \\
& 100 & 5 & 47.885 & 157.277 & 7.541 & 10.174 & 10.802 & 3.549 & 253.861 & 10.258 & 4.781 & 4.906 \\
& 100 & 10 & 7.073 & 161.123 & 6.398 & 6.171 & 4.338 & 32.893 & 78.121 & 3.476 & 6.383 & 6.372 \\
& 100 & 20 & 9.156 & 210.235 & 5.259 & 5.533 & 5.485 & 21.470 & 255.334 & 20.574 & 20.745 & 20.733 \\
\midrule\multirow{6}{*}{\textbf{DGP 3}} & 50 & 5 & 0.249 & 41.228 & 0.408 & 0.367 & 0.359 & 0.136 & 47.176 & 0.172 & 0.164 & 0.160 \\
& 50 & 10 & 0.141 & 63.842 & 0.197 & 0.186 & 0.185 & 0.088 & 55.036 & 0.105 & 0.102 & 0.101 \\
& 50 & 20 & 0.104 & 51.546 & 0.117 & 0.117 & 0.117 & 0.069 & 35.139 & 0.073 & 0.073 & 0.073 \\
& 100 & 5 & 0.193 & 263.935 & 0.337 & 0.316 & 0.317 & 0.100 & 178.029 & 0.161 & 0.153 & 0.153 \\
& 100 & 10 & 0.110 & 55.298 & 0.206 & 0.195 & 0.195 & 0.066 & 202.919 & 0.080 & 0.078 & 0.078 \\
& 100 & 20 & 0.075 & 57.106 & 0.101 & 0.099 & 0.099 & 0.049 & 84.641 & 0.052 & 0.051 & 0.051 \\
\midrule\multirow{6}{*}{\textbf{DGP 4}} & 50 & 5 & 0.431 & 0.213 & 0.232 & 0.229 & 0.225 & 0.191 & 0.146 & 0.150 & 0.144 & 0.145 \\
& 50 & 10 & 0.384 & 0.155 & 0.158 & 0.154 & 0.152 & 0.115 & 0.098 & 0.097 & 0.098 & 0.098 \\
& 50 & 20 & 0.299 & 0.108 & 0.108 & 0.107 & 0.107 & 0.072 & 0.070 & 0.070 & 0.070 & 0.070 \\
& 100 & 5 & 0.374 & 0.167 & 0.177 & 0.159 & 0.157 & 0.150 & 0.098 & 0.101 & 0.110 & 0.110 \\
& 100 & 10 & 0.371 & 0.104 & 0.108 & 0.111 & 0.110 & 0.084 & 0.069 & 0.069 & 0.071 & 0.071 \\
& 100 & 20 & 0.288 & 0.075 & 0.076 & 0.076 & 0.076 & 0.050 & 0.048 & 0.048 & 0.048 & 0.048 \\
\bottomrule
\end{tabular}
\begin{tablenotes}
• Note: The value in each cell denotes the mean of the Hausdorff distance across 100 simulations, with 100 starting group memberships randomly generated for each estimate.
\end{tablenotes}
\end{threeparttable}
}
\end{center}
\end{table}
Table (ref) includes estimates obtained by conducting separate IV estimations for each estimated group.
Again, the estimates largely correspond to those in Table (ref).
The estimators achieve small biases in situations under which the estimators attain high classification accuracy.
An exception is DGP 2, where every estimator, including the TGFE and UGFE estimators, suffers from the weak instrument problem.
This is because these estimates obtained ex-post do not utilize the first-stage information.
This result exemplifies the importance of taking into account possible heterogeneity in the first stage, which constitutes the distinct characteristic of our procedures.
\section{Empirical example}
In this section, we apply our estimators to revisit the relationship between democracy and income studied by AcemogluJohnsonRobinsonYared08.
The authors use linear dynamic models and find that once we control for the country fixed effects, the positive effect of income on the level of democracy disappears.
However, because their model assumes that the effect of income growth is homogeneous across countries, the literature has pointed out that we might reach different conclusions when the potential heterogeneity is considered.
For example, in their comment, CervellatiJungSundeVischer14 report opposing estimation results for former colonies and noncolonies.
Meanwhile, by using GFE estimation, BonhommeManresa2015Supp show that the effect varies across groups.
We contribute to this literature by reporting new results on heterogeneity in the effect of income using our estimators.
The key deviation from BonhommeManresa2015Supp is that we take into account the potential endogeneity of lagged income growth due to the reverse causality.
AcemogluJohnsonRobinsonYared08 manage this issue by employing the savings rate and trade-weighted world income as instruments for income growth.
By contrast, BonhommeManresa2015Supp implicitly assume that the potential correlation between the income growth and unobserved factors is entirely subsumed by the group-specific fixed effects.
Here, our interest is in the consequences of simultaneously allowing the endogeneity of income and the heterogeneity in its coefficient.
We consider various models of the relationship between the level of democracy and income growth.
Let $d_{it}$ and $y_{it-1}$ denote the level of democracy measured by the Freedom House score and the lagged value of log income per capita, respectively.
Let $\mathbb G =\{1,2\}$, and denote the group-specific fixed effects by $\alpha_{g}$.
The $u_{it}$ in each equation is the error term.
\begin{description}
• $d_{it}=\beta_{2,g_i} y_{it-1} + \alpha + u_{it}$
• $d_{it}=\beta_{2,g_i} y_{it-1} + \alpha_{g_i} + u_{it}$
• $d_{it}=\beta_1 d_{it-1} + \beta_{2,g_i} y_{it-1} + \alpha + u_{it}$
• $d_{it}=\beta_1 d_{it-1} + \beta_{2,g_i} y_{it-1} + \alpha_{g_i} + u_{it}$
\end{description}
Models 1a and 1b are static in the sense that they do not contain the lagged term for the level of democracy.
They differ by whether the group-specific fixed effects are included.
Models 2a and 2b are dynamic, where the term $d_{it-1}$ captures the mean-reverting dynamics of democracy.
In Model 2a, only the coefficient of income growth exhibits a grouped pattern, while in Model 2b the constant also does so.
As in CervellatiJungSundeVischer14, we employ trade-weighted world income as an instrument.
Similar to the simulation section, we report two types of estimates for $\beta_{2,g_i}$, which captures the effect of income growth on democracy. The first ones, which we call the “pre-estimates,” are directly generated from our procedures, and the second ones, which we call the “post-estimates,” are from separate IV regressions post-run for each estimated group.
\begin{table}[htp]
\begin{center}
\resizebox{0.8\columnwidth}{!}{
\begin{threeparttable}
\caption{Estimation results for the models}
\begin{tabular}{llccccc}
\toprule
Model & Method & Pre-estimates & SE & Group size & Post-estimates & SE \\
\midrule
\midrule\multirow{1}{*}{\textbf{No groups}} & \multirow{1}{*}{2SLS} & & & & 0.112 & 0.126 \\
\midrule\multirow{10}{*}{\textbf{Model 1a}} & \multirow{2}{*}{IG} & & & 47 & 0.180 & 0.036 \\
& & & & 32 & 0.134 & 0.040 \\
& \multirow{2}{*}{2SLS} & 0.166 & 0.083 & 37 & 0.150 & 0.053 \\
& & 0.232 & 0.083 & 42 & 0.192 & 0.045 \\
& \multirow{2}{*}{TGFE$_2$} & 0.170 & 0.022 & 35 & 0.138 & 0.054 \\
& & 0.220 & 0.019 & 44 & 0.180 & 0.048 \\
& \multirow{2}{*}{TGFE$_{N/4}$} & 0.210 & 0.012 & 48 & 0.156 & 0.036 \\
& & 0.170 & 0.015 & 31 & 0.108 & 0.038 \\
& \multirow{2}{*}{UGFE} & 0.218 & 0.012 & 52 & 0.200 & 0.036 \\
& & 0.178 & 0.015 & 27 & 0.155 & 0.039 \\
\midrule\multirow{10}{*}{\textbf{Model 1b}} & \multirow{2}{*}{IG} & & & 47 & 0.276 & 0.102 \\
& & & & 32 & 0.092 & 0.028 \\
& \multirow{2}{*}{2SLS} & 0.141 & 0.069 & 34 & 0.113 & 0.037 \\
& & 0.632 & 0.148 & 45 & 0.445 & 0.234 \\
& \multirow{2}{*}{TGFE$_2$} & 0.236 & 0.026 & 36 & 0.109 & 0.050 \\
& & 0.167 & 0.020 & 43 & 0.303 & 0.267 \\
& \multirow{2}{*}{TGFE$_{N/4}$} & 0.181 & 0.026 & 33 & 0.082 & 0.038 \\
& & 0.197 & 0.015 & 46 & 0.261 & 0.078 \\
& \multirow{2}{*}{UGFE} & 0.187 & 0.029 & 27 & 0.101 & 0.031 \\
& & 0.215 & 0.014 & 52 & 0.271 & 0.071 \\
\bottomrule
\end{tabular}
\begin{tabular}{llccccc}
\toprule
Model & Method & Pre-estimates & SE & Group size & Post-estimates & SE \\
\midrule
\midrule\multirow{1}{*}{\textbf{No groups}} & \multirow{1}{*}{2SLS} & & & & -0.006 & 0.053 \\
\midrule\multirow{10}{*}{\textbf{Model 2a}} & \multirow{2}{*}{IG} & & & 27 & 0.098 & 0.046 \\
& & & & 52 & 0.125 & 0.047 \\
& \multirow{2}{*}{2SLS} & 0.113 & 0.050 & 45 & 0.109 & 0.039 \\
& & 0.083 & 0.049 & 34 & 0.081 & 0.040 \\
& \multirow{2}{*}{TGFE$_2$} & 0.088 & 0.017 & 48 & 0.125 & 0.041 \\
& & 0.059 & 0.017 & 31 & 0.096 & 0.041 \\
& \multirow{2}{*}{TGFE$_{N/4}$} & 0.083 & 0.012 & 27 & 0.098 & 0.046 \\
& & 0.107 & 0.012 & 52 & 0.125 & 0.047 \\
& \multirow{2}{*}{UGFE} & 0.088 & 0.011 & 27 & 0.098 & 0.046 \\
& & 0.112 & 0.011 & 52 & 0.125 & 0.047 \\
\midrule\multirow{10}{*}{\textbf{Model 2b}} & \multirow{2}{*}{IG} & & & 27 & 0.089 & 0.049 \\
& & & & 52 & 0.219 & 0.089 \\
& \multirow{2}{*}{2SLS} & 0.326 & 0.061 & 54 & 0.337 & 0.146 \\
& & 0.116 & 0.049 & 25 & 0.103 & 0.051 \\
& \multirow{2}{*}{TGFE$_2$} & 0.099 & 0.019 & 48 & 0.222 & 0.171 \\
& & 0.033 & 0.023 & 31 & 0.059 & 0.046 \\
& \multirow{2}{*}{TGFE$_{N/4}$} & 0.105 & 0.013 & 52 & 0.219 & 0.089 \\
& & 0.091 & 0.019 & 27 & 0.089 & 0.049 \\
& \multirow{2}{*}{UGFE} & 0.122 & 0.012 & 54 & 0.213 & 0.077 \\
& & 0.121 & 0.020 & 25 & 0.093 & 0.052 \\
\bottomrule
\end{tabular}
\begin{tablenotes}
• Note: 10,000 starting group memberships are randomly generated for each estimate. Standard errors clustered at country level.
\end{tablenotes}
\end{threeparttable}
}
\end{center}
\end{table}
Table (ref) presents our estimation results. In Models 1a and 2a, the pre- and post-estimates are positive and mostly statistically significant. They are similar to each other, while the pre-estimates tend to show smaller standard errors. The results are overall consistent across procedures. The group with more units displays larger estimates---around 0.4 to 0.6 in Model 1a and around 0.3 in Model 2a.
Likewise, in Models 1b and 2b, the estimates are positive and statistically significant. However, in Model 1b, some pre- and post-estimates are markedly different from each other. Their difference is not prevalent in Model 2b, but there, the pre-estimates show some variation by procedure. For instance, the 2SLS estimator reports a large difference between groups---around 0.2---whereas the UGFE estimator reports a negligible difference. This illustrates the situation where modeling the first-stage heterogeneity requires careful attention.
\section{Conclusion}
This paper considers the estimation of linear panel data models with endogenous regressors and a latent group structure in the coefficients.
In particular, it provides a relatively comprehensive view of possible estimation methods and clarifies the conditions under which the estimators work.
We first argue that a na\"ive approach of minimizing the GMM objective function should not be used to estimate group membership parameters because at least asymptotically, the minimizer of the objective function is not uniquely determined.
Then, we consider various alternative estimation strategies.
The first three procedures are 2SLS estimations which differ in the first-stage modeling.
The 2SLS estimator is based on the homogeneity of the first-stage coefficients.
The TGFE estimator assumes that the first-stage coefficients exhibit a group structure.
The UGFE estimator allows complete heterogeneity across units.
We provide conditions under which these procedures yield consistent estimates for the true coefficient and group memberships.
Furthermore, for the 2SLS estimators, we provide data-driven procedures that select the number of groups.
The next two approaches are more heuristic.
The ignoring endogeneity approach estimates the true memberships by applying the GFE estimation directly to the endogenous variables.
The reduced-form approach applies the GFE estimation directly to the instrument.
The conditions that ensure their consistency can be obtained from the literature, but it is difficult to test or verify them in practice.
The results from the Monte Carlo simulations suggest that estimating the group memberships using the TGFE/UGFE estimators performs well in terms of classification performance; the IG estimator also shows good performance, but it might perform poorly in cases where the first- and second-stage errors are highly correlated with heterogeneous correlation coefficients.
\printbibliography[notcategory={onlyappx}]