EconBase
← Back to paper

Convergence rate of estimators of clustered panel models with misclassification

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

21,208 characters · 4 sections · 8 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Convergence rate of estimators of clustered panel models with misclassification

\onehalfspacing

abstractWe study kmeans clustering estimation of panel data models with a latent group structure and $N$ units and $T$ time periods under long panel asymptotics. We show that the group-specific coefficients can be estimated at the parametric root $NT$ rate even if error variances diverge as $T \to \infty$ and some units are asymptotically misclassified. This limit case approximates empirically relevant settings and is not covered by existing asymptotic results. Keywords: Panel data, latent grouped structure, clustering, kmeans, convergence rate, misclassification. JEL codes: C23, C33, C38 Mathematical Subjects Classification (2010): 62H30, 62H12 Declarations of interest: none

Introduction

Panel models can account for unobserved heterogeneity by dividing units into a finite number of latent groups and allowing a unit's coefficients to be group-specific bonhomme2015grouped,su2016identifying, vogt2015classification, wang2016homogeneity, okuiwang2018. Estimators of such models simultaneously estimate group memberships and group-specific coefficients. For example, bonhomme2015grouped propose a kmeans-type estimator and su2016identifying propose the CLasso estimator that is based on solving a penalized regression program. These two and other related estimators are justified under a long panel asymptotic framework that sends both the number of units $N$ and the number of time periods $T$ to infinity. Existing theoretical results show that coefficients that are group-specific and time invariant can be estimated at a root $NT$ rate, i.e., at the parametric rate. In this paper we show that the parametric rate can be obtained even if some units have a positive probability of being misclassified in the limit. This limit case is highly relevant in practice since it is common to misclassify at least some units in empirical applications BonhommeDiscretize. However, existing results do not apply in such settings.

Existing asymptotic results for linear panel models assume that the variance of the error term is universally bounded. From this assumption, it can then be shown that group memberships can be estimated uniformly consistently, i.e., the probability of misclassifying one or more units vanishes as $N, T \to \infty$. This implies that the rate at which group-specific coefficients can be estimated is the same as under a known group structure and is therefore equal to the parametric rate.

However, the assumption of a universal bound on the variance of the error term may not reflect real circumstances. It implies that the asymptotic limit as $T \to \infty$ prescribes that, for each unit, the level of statistical noise is negligible when compared to the number of observed time periods. This is not characteristic of typical empirical applications. The number of observed time periods is often rather small and, at least for some units, statistical noise plays an important role in determining the outcome.

In this paper, we extend previous theoretical results to a heteroscedastic setting in which units are endowed with unit-specific error variances $\sigma_1^2, \dotsc, \sigma_N^2$. A unit $i$ with small $\sigma_i$ is easy to classify, whereas a unit $i$ with large $\sigma_i$ is difficult to classify. The individual error variances may depend on $N$ and $T$ and may diverge as $T \to \infty$. We expect our asymptotic framework to be a more faithful approximation of the finite sample behavior of the estimators than the conventional framework.

For kmeans-estimation, we show that uniform consistency of group memberships holds provided that the unit-specific error variances do not diverge too fast. Units $i$ for which $\sigma_i$ diverges too fast are potentially misclassified in the limit. However, if the proportion of such potentially misclassified units is sufficiently small then it is still possible to estimate the group-specific coefficients at a root $NT$ rate.

pollard1981strong,pollard1982central,bonhomme2015grouped consider panel models with fixed $T$ and estimate cluster-specific coefficients. They show that the cluster-specific coefficients converge to a pseudo-true value at rate root $N$ even though units are misclassified in the limit with positive probability. Their setting and results are distinct from ours. We consider long panel asymptotics under which true rather than pseudo-true cluster-specific coefficients can be identified and estimated at a root $NT$ rate.

We prove our results for a simple linear panel model with group-specific intercepts and focus on estimation by least squares (equivalent to kmeans). By focusing on this simple model we are able to derive our results under interpretable and intuitive conditions on the structure of heteroscedasticity. While we think that our argument can be extended to more general regression models with group-specific coefficients, we believe that such an exercise would impose more involved assumptions and would not be as instructive about the mechanisms that allow root $NT$-consistency to arise despite of diverging error variances and possibly misclassified units.

bonhomme2015grouped conduct a simulation experiment that is calibrated to their empirical application. They find that the group-specific coefficients are estimated precisely, even though it is likely that one or more units are misclassified. Existing theoretical results about the rate of consistency of the group-specific coefficients cannot explain this phenomenon as they do not apply in the presence of misclassification. We fill this gap in the literature by showing that uniform consistency is sufficient but not necessary for precise estimation of the group-specific coefficients.

Setting

The units $i = 1, \dotsc, N$ are partitioned into $G$ groups. The set of all groups is $\mathbb{G} = \{1, \dotsc, G\}$ and unit $i$ belongs to group $g_i^0 \in \mathbb{G}$. For units in group $g \in \mathbb{G}$ the mean outcome in each period is given by $\mu_g$. At time $t = 1, \dotsc, T$ we observe the scalar outcome $y_{it}$ generated by

align*[align* omitted — 54 chars of source]

where $v_{it}$ is a noise term with variance one. Let $\Gamma$ denote the space of possible group assignments $\mathbf{g} = (g_1, \dotsc, g_N)$ and let $\mathcal{M}$ denote the space of possible group-specific means $\boldsymbol{\mu} = (\mu_1, \dotsc, \mu_G)$. The true group assignment $\mathbf{g}^0 \in \Gamma$ and the true group-specific mean $\boldsymbol{\mu}^0 \in \mathcal{M}$ are unknown parameters and are estimated.

We consider kmeans-type estimation as suggested in bonhomme2015grouped. The objective function for estimation is defined on $\Gamma \times \mathcal{M}$ and is given by

align*[align* omitted — 138 chars of source]

The estimator is defined as $(\hat{\boldsymbol{\mu}}, \hat{\mathbf{g}})= \arg \min_{ \boldsymbol{\mu} \in \mathcal{M}, \boldsymbol{g} \in \Gamma} Q_{N, T}(\mathbf{g}, \boldsymbol{\mu})$. In practice, the estimator is computed by the iterative kmeans procedure. We start with an initial group membership structure $\mathbf{g}^{(0)}$ and then iterate $\boldsymbol{\mu}$ and $\mathbf{g}$ such that the $s$-th iteration sets $\boldsymbol{\mu}^{(s)} = \arg \min_{ \boldsymbol{\mu} \in \mathcal{M}} Q_{N, T}(\mathbf{g}^{(s-1)}, \boldsymbol{\mu})$ and $\mathbf{g}^{(s)} = \arg \min_{ \boldsymbol{g} \in \Gamma} Q_{N, T}(\mathbf{g}, \boldsymbol{\mu}^{(s)})$ until convergence. Since the iteration may converge to a local minimum we re-start the procedure from many initial values for $\mathbf{g}$.

Main results

We consider asymptotic sequences under which $N, T \to \infty$ and

align[align omitted — 93 chars of source]

We treat $(\sigma_1, \dotsc, \sigma_N)$ and $\mathbf{g}^0$ as unobserved deterministic parameters.

We first state sufficient conditions for consistent estimation of $\boldsymbol\mu^0$.

assumption\begin{enumerate}[label=\roman*)] • $\{v_{it}\}_{t=1}^T$ is an independent sequence with $\mathbb{E} v_{it} = 0$ and $\mathbb{E} v_{it}^2 = 1$. • The average error variance satisfies $N^{-1} \sum_{i = 1}^N \sigma_i^2 = o (T)$. • There is a bounded set $\mathcal{M} \subset{\mathbb{R}}^G$ such that $\boldsymbol \mu^0 \in \mathcal{M}$. • There is a positive constant $M_{G}$ such that \[ \min_{g \in \mathbb{G}} \min_{h \in \mathbb{G} \setminus \{g\}} \left\lvert \mu_g^0 - \mu^0_h \right\rvert > M_G. \] • For all $g \in \mathbb{G}$, $ N^{-1} \sum_{i = 1}^N 1 (g^0_i = g) \geq q_{\min} $. \end{enumerate}

Part (ref) imposes independence of the error term over time. Using this assumption we obtain asymptotic results under simply conditions on between-unit heteroscedasticity. The assumption can be relaxed to allow for weak serial correlation at the expense of conditions on heteroscedasticity that are more difficult to interpret. Part (ref) states that the average error variance increases at a slower rate than $T$. This assumption ensures that, as $T \to \infty$, the additional information from observing more time periods is not undone by an increased noisiness of the signal. Part (ref) is a standard regularity assumption. Part (ref) requires that the group-specific means are distinct (group separation). Part (ref) ensures that the effective sample size that can be used to estimate the group-specific mean grows at the same asymptotic rate for all groups.

Assumption (ref) does not restrict cross-sectional dependence. Assumption (ref) below limits the amount of cross-sectional dependence and is required for our result on $NT$-convergence of the group-specific parameters, but not any of our intermediate results.

The grouped model is invariant to a relabeling of the groups and the vector of group-specific means $\boldsymbol\mu^0$ is therefore only identified up to a re-ordering of its components. The following result states that the identified set is consistently estimated.

lemma[Consistency of group-specific means] Suppose that Assumption (ref) holds. Then, there is a (possibly random) permutation function $\pi: \mathbb{G} \to \mathbb{G}$ such that for all $\epsilon > 0$ \begin{align*} \lim_{N, T \to \infty} P \left( \max_{g \in \mathbb{G}} \left\lvert \hat{\mu}_{\pi (g)} - \mu_g^0 \right\rvert > \epsilon \right) = 0. \end{align*}

Similarly to related results in the literature bonhomme2015grouped, proving this result does not require establishing that group memberships are consistently estimated for all units. In Theorem (ref) below, we strengthen the result to root $NT$ convergence under weaker assumptions on heteroscedasticity than are commonly assumed in the literature.

The subsets of units for which we can guarantee that group memberships are uniformly consistently estimated is given by

align[align omitted — 151 chars of source]

For the units in $\mathcal{I}_{N,T}$ the error variances are allowed to diverge but only at rate $\sqrt{T/\log N}$. Controlling the rate of divergence is necessary to ensure that observing additional time periods adds enough information to estimate group memberships precisely. What rates of divergence are permissible is determined by bounds on the tail of the error distribution. The error term of our panel model is given by $\sigma_i v_{it}$. We assume that $v_{it}$ is a sub-exponential random variable. Under this assumption, new observations add information at the usual parametric rate root $T$ and the price of uniformity is root $\log N$.

assumption[Sub-exponential errors] There are positive constants $\nu, \alpha$ such that \begin{align*} \max_{1 \leq i \leq N} \max_{1 \leq t \leq T} \mathbb{E} \exp(\lambda \left\lvert v_{it} \right\rvert) \leq \exp\left(\frac{\lambda^2 \nu^2}{2}\right) \qquad for all $\lambda > 0$ such that $\lambda < \frac{1}{\alpha}$. \end{align*}

In addition to errors that are Gaussian and sub-Gaussian (conditional on $\sigma_i$) this assumption allows also for certain “fat-tailed” distributions such as Poisson or chi-squared. It is possible to relax this assumption and allow for distributions with even heavier tails, but only at the expense of a different rate condition in (ref) that is more difficult to state and to interpret. In our setting, misclassification can occur even for moderate realizations of $v_{it}$ if $\sigma_i$ is sufficiently large. Therefore, misclassification does not hinge on heavy tails of $v_{it}$ and is not ruled out or limited by Assumption (ref).

The following lemma states that group membership is estimated consistently uniformly over all units in $\mathcal{I}_{N,T}$.

lemmaSuppose that Assumptions (ref) and (ref) hold. Then, there exists a (possibly random) permutation function $\pi: \mathbb{G} \to \mathbb{G}$ such that \begin{align*} \lim_{N, T \to \infty} P \left( \sup_{i \in \mathcal{I}_{N, T}} \left\lvert \pi(\hat{g}_i) - g_i^0 \right\rvert > 0 \right) \to 0. \end{align*}

This lemma extends existing results in the literature that are derived under the assumption that $\max_{1 \leq i \leq N} \sigma_i^2$ is bounded in which case $\mathcal{I}_{N,T} = \{1, \dotsc, N\}$ eventually. Lemma (ref) shows that uniform consistency over all units can be obtained even if the error variance $\sigma_i^2$ diverges for some or all units. In this case, all unit-specific error variances must diverge at most at the rate given in (ref) and the average error variance must diverge at most at the rate given in Assumption (ref)(ref).

We study the asymptotic behavior of $\hat{\boldsymbol{\mu}}$ without requiring that all units are contained in $\mathcal{I}_{N,T}$ and therefore guaranteed to be estimated consistently. The idea of Theorem (ref) below is that units that are not in $\mathcal{I}_{N,T}$ do not affect the asymptotic distribution provided that there are sufficiently few of them.

Let $\mathcal{I}_{N, T}^\mathsf{c} = \{1, \dotsc, N\} \setminus \mathcal{I}_{N, T}$ and write $\# A$ to denote the cardinality of a set $A$. We assume

align[align omitted — 253 chars of source]

Existing theoretical results cover only settings under which no units are potentially misclassified in the asymptotic limit, i.e., $\# \mathcal{I}_{N,T}^{\mathsf{c}} = 0$. In this case (ref) is trivially satisfied. Our result allows $\# \mathcal{I}_{N,T}^{\mathsf{c}} \neq 0$ provided that the proportion of possibly misclassified units ${\# \mathcal{I}_{N,T}^{\mathsf{c}}}/{N}$ vanishes at a sufficiently fast rate. The rate in the first component of the max ensures that units in $\mathcal{I}_{N,T}^{\mathsf{c}}$ asymptotically do not affect the mean of $\hat{\boldsymbol{\mu}}$. The rate of the second component in the max ensures that units in $\mathcal{I}_{N,T}^{\mathsf{c}}$ asymptotically do not affect the variance of $\hat{\boldsymbol{\mu}}$. By (ref), the second component satisfies

align*[align* omitted — 173 chars of source]

This shows that the first component can dominate the second component at most at a root $\log N$ rate. Therefore, replacing the max in (ref) by the second component gives a good approximation (up to order root $\log N$) of the required rate condition.

To state the assumption for asymptotic normality of $\hat{\mu}_g$, $g \in \mathbb{G}$, let $\mathcal{I}_{N, T}(g) = \left\{i \in \mathcal{I}_{N, T} : g_i^0 = g \right\}$ and

gather*[gather* omitted — 192 chars of source]
assumption\begin{enumerate}[label = \roman*)] • Condition (ref) is satisfied. • For each $g \in \mathbb{G}$ there are positive constants $\delta_g$ and $q_g$ such that $N_g/N \to q_g$ and \begin{align*} \frac{1}{\tilde{N}_g} \sum_{i \in \mathcal{I}_{N, T} (g)} \sigma_i^2 + \frac{1}{\tilde{N}_g} \sum_{\substack{i, j \in \mathcal{I}_{N, T} (g) \\i \neq j}} \sigma_i \sigma_j \operatorname{cov}(v_{i1}, v_{j1}) \to \delta_g. \end{align*} • We have \begin{align*} \frac{1}{\# \mathcal{I}_{N,T}} \sum_{i \in \mathcal{I}_{N,T}} \sigma_i^2 = O (\sqrt{T}) \quad and \quad \frac{1}{\# \mathcal{I}_{N,T}} \sum_{i \in \mathcal{I}_{N,T}} \sigma_i^4 = O (N T). \end{align*} • In addition, \begin{align*} \sum_{\substack{i, j, k \in \mathcal{I}_{N, T}\\ \{i\} \cap \{j\} \cap \{k\} = \emptyset}} \sigma_i \sigma_j \sigma_k \mathbb{E} [v_{i1}^2 v_{j1} v_{k1}] = & O (N^2 T), \\ \sum_{\substack{i, j, k, \ell \in \mathcal{I}_{N, T} \\ \{i\} \cap \{j\} \cap \{k\} \cap \{\ell\}= \emptyset}} \sigma_i \sigma_j \sigma_k \sigma_\ell \mathbb{E} [v_{i1} v_{j1} v_{k1} v_{\ell 1}] = & O (N^2 T). \end{align*} \end{enumerate}

Part (ref) ensures that the asymptotic variance of $\hat{\mu}_g$ converges. Part (ref) imposes two conditions on the rate of divergence of the $L_2$ and the $L_4$ norm of $\{\sigma_i : i \in \mathcal{I}_{N,T}\}$. Under cross-sectional independence the first condition is implied by (ref). The second condition is satisfied if $ N \log^2 N /T \to \infty. $ Part (ref) limits the amount of cross-sectional dependence.

The following theorem guarantees root $NT$-consistency and asymptotic normality of $\hat{\mu}_g$.

theoremSuppose that Assumptions (ref)--(ref) hold. Then, for $g \in \mathcal{\mathbb{G}}$ as $N,T \to \infty$ \begin{align*} \sqrt{N T} \left( \hat{\mu}_{\pi(g)} - \mu_g^0 \right) \overset{d}{\longrightarrow} \mathcal{N} (0, q_g^{-1} \delta_g) . \end{align*}

This result shows that root $NT$-consistency can be obtained even if some units are potentially misclassified in the limit. In addition, the error variance for the units that are consistently estimated need not be bounded. For root $NT$-consistency we require a stronger assumption on the average error variance than for the result on consistent estimation of group memberships in Lemma (ref). Assumption (ref)(ref) implies that the average error variance is bounded. In contrast, Lemma (ref) allows the average error variance to diverge at a controlled rate.

Conclusion

We have shown that uniformly consistent estimation of group memberships is not a necessary condition of root $NT$ estimation of time invariant group-specific parameters. The simple model with group-specific intercepts served our purpose of providing an example of a grouped panel model in which a root $NT$ rate can be obtained even under misclassification in the limit. We are confident that similar results can be obtained for general linear panel regression, albeit under more involved conditions that may not be straightforward to interpret. We leave such extensions to future research. For scenarios where the amount of misclassification permitted by our assumption (ref) is exceeded by only a sufficiently small margin, our proofs suggest that it is possible to obtain a convergence rate that is slower than root $NT$ but faster than root $N$. This suggests a negative relationship between the difficulty of classifying individual units and the precision of the estimator of the vector of group-specific coefficients.