Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
89,812 characters · 15 sections · 58 citation commands
Confidence Set for Group Membership
\onehalfspacing
Clustering units into discrete groups is one of the oldest problems in statistics pearson1896mathematical. It has received interest in the recent econometric literature on grouped panel models lin2012estimation, bonhomme2015grouped, sarafidis2015partially, ando2016panel,vogt2015classification, su2016identifying, wang2016homogeneity, vogt2017clustering, lu2017determining, GuQuantileClustering, liu2020identification, wang2021identifying, mammen2022estimation,mehrabani2022estimation,mugnier2022simple,chetverikov2022spectral,yu2022group, mugnier2023nonlinear.
In grouped panel models, a data-driven clustering algorithm is used to estimate a latent group structure. As statistical procedures, clustering algorithms suffer from sampling errors and produce a noisy version of the true group structure. The existing literature gives little guidance on how to assess clustering uncertainty in a given application. Inferential theory for grouped panel models has focused on group characteristics, but is underdeveloped for assessing the uncertainty about individual group memberships mclachlan2004finite.
In this paper, we quantify the statistical uncertainty about the true group memberships in a grouped panel model. As far as we know, we are the first to propose and justify a rigorous frequentist method to evaluate clustering uncertainty.
In grouped panel models, individual regression curves are heterogeneous and exhibit a grouped pattern. All units that belong to the same group face the same regression curve. Group memberships are unobserved and estimated by a clustering algorithm. Clustering uncertainty means that the algorithm may misclassify some units and assign them an incorrect regression curve.
We propose a confidence set for group membership that quantifies clustering uncertainty jointly for all units in the panel. For a panel of $N$ units, an element of a joint confidence set is an $N$-dimensional vector that specifies a group assignment for every unit. Our confidence set gathers all $N$-dimensional vectors of group assignments that are “not ruled out by the data” and is guaranteed to contain the vector of the true group memberships with a pre-specified probability, say 95%.
The latent groups in a grouped panel model have no natural labels and can only be identified up to a permutation. For notational convenience, we write our confidence set using an arbitrary ordering of the groups. We interpret this as a shorthand for linking units to regression curves. For example, suppose that there are a “group 1” with a slope coefficient of $0.2$ and a “group 2” with a slope coefficient of $0.8$. If our confidence set rules out that unit $i$ belongs to “group 1” then we take this to mean that unit $i$ does not face a slope coefficient of $0.2$. This interpretation does not depend on the ordering of the groups. Similarly, if our confidence set determines that units $i$ and $j$ belong to different groups, then we take this to mean that they face different slope coefficients.
This interpretation of a confidence set presumes that the data are rich enough to recover the group-specific coefficients. If the data do not provide any clue about the group-specific coefficients, then we are not able to statistically examine group memberships either. On the other hand, if the data are known to be “very rich,” and group memberships are guaranteed to be estimated correctly, then our confidence set is not needed.
Settings between these two extreme scenarios are relevant in practice. In Monte Carlo experiments calibrated to their empirical application, bonhomme2015grouped find that units are frequently misclassified, whereas group-specific coefficients are estimated precisely (see Table S.III in their supplemental appendix). dzemskiokui2021convergence provide a theoretical framework to explain this observation.\footnote{ For mathematical convenience, the theoretical analysis of clustered panel models often proceeds under assumptions that rule out any misclassification in the asymptotic limit bonhomme2015grouped,vogt2015classification.} They assume that unit $i$ faces an error term with unit-specific variance $\sigma_i^2$. Units with small $\sigma_i$ are classified reliably. Units with large $\sigma_i$ are potentially misclassified. dzemskiokui2021convergence show that the group-specific coefficients can be estimated consistently if the proportion of potentially misclassified units is sufficiently small. This is the main setting that we have in mind for applications of our confidence set. It is less restrictive than assuming, as is typically done in the literature, that both group-specific coefficients and all group assignments can be reliably estimated.
Our empirical application illustrates how our confidence set provides new economic insights. We follow wang2019heterogeneous who estimate a panel model that allows the effect of a minimum wage on unemployment to vary between US states, depending on the assignment of each state to one of four latent groups. This group assignment is potentially estimated with error. We use our confidence set to identify, up to a small pre-specified error probability, states without clustering uncertainty. In terms of the framework discussed above, these are states with low values of $\sigma_i$. For these states, we can identify the state-specific effects of the minimum wage.
Our confidence set can also be used to enhance a plot of the estimated groups by adding information about clustering uncertainty. We illustrate this in our empirical application. Providing a plot of the estimated groups is standard practice.\footnote{For example, see Figure 2 in wang2019heterogeneous, Figure 2 in bonhomme2015grouped and Figure 6 in wang2016homogeneity.} This is true even in applications where the group structure is considered merely a nuisance parameter.\footnote{For example, time-varying unobserved heterogeneity can be controlled by imposing a latent group structure with group-specific time fixed-effects. In this context, the heterogeneity in the fixed-effect is a nuisance parameter similar to the interacted fixed effects in, for example, moonweidner2019.}
An alternative to our frequentist approach is Bayesian inference. Fully parametric grouped-panel models are finite mixtures models that can be estimated by the EM algorithm dempster1977maximum. The E-step of the EM algorithm computes unit-wise posterior probabilities for group membership. These are valid if the units in the panel are independently drawn from the assumed parametric distribution. Our frequentist approach is more general. We do not assume a parametric distribution of the error term and show that our approach is valid for error distributions in a broad nonparametric class. We also allow for cross-sectional dependence, non-random patterns of heteroscedasticity, and non-random group assignments. Another advantage of our procedure is that it allows for joint inference on the $N$-dimensional vector of all group memberships, whereas unit-wise posterior probabilities only address uncertainty about the group membership of a single unit.
Quantifying the uncertainty about the true group structure is a high-dimensional inference problem. The $N$-dimensional vector of true group memberships is high-dimensional since its size grows as $N \to \infty$. To construct a confidence set for this high-dimensional parameter, we invert a test of the many moment inequalities that characterize group memberships. Testing many moment inequalities is the problem considered in chernozhukov2013testing. Their test is based on a single test statistic, whereas ours combines many simultaneous group membership tests for individual units. The advantage of our approach is that it can be inverted without running a computationally infeasible exhaustive search over the space of all partitions.
Our confidence set is valid in the presence of weak-dependence and serial correlation, whereas chernozhukov2013testing assume independence over time. We account for serial correlation by using a heteroskedasticity--and--autocorrelation--robust (HAC) variance estimator Andrews91,NeweyWest87 when constructing the unit-wise test statistics.
The asymptotic analysis of our procedure accounts for the high-dimensional nature of our setting and allows for weak time-dependence. We built on chang2022central who provide results for HAC estimators and a high-dimensional central limit theorem for dependent data. Our setting requires extending their approach in different ways. Specifically, we use a Nasarov-type inequality chernozhukov2016central to control for the effect of parameter estimation. Moreover, we develop a regularization scheme for the HAC estimates. This regularization scheme allows us to control the estimation error in our data-driven critical values by using a comparison bound based on li2002normal.
We provide several extensions of our method. First, we suggest alternative critical values. These are slightly conservative but much easier to compute than our benchmark critical values. Second, we show that HAC estimation is not needed if there is no serial correlation. In this case, test statistics based on the usual variance estimators provide valid confidence sets and are easier to implement than those based on the HAC estimator. Third, we propose a two-step method, called unit selection, to shrink the cardinality of the confidence set when there are many units for which clustering uncertainty is low. The two-step procedure identifies and discards such units. We then construct a confidence set for the remaining units. The method accounts for errors in the unit selection and provides a valid confidence set. This additional error control can inflate the confidence set if unit selection does not eliminate sufficiently many units.
The remainder of this paper is organized as follows. Section (ref) introduces the grouped panel model. Section (ref) defines our confidence set for group membership. Section (ref) proves the asymptotic validity of our confidence sets. Section (ref) discusses the extensions of our method. Section (ref) provides an empirical application. Section (ref) presents Monte Carlo simulations that investigate the validity and power of our confidence set based on simulation designs inspired by our empirical application. Section (ref) concludes.
An R package implementing the methods proposed in this paper is included in the replication package that is published together with this article.
We observe panel data $(y_{it}, w_{it}', x_{it}')'$ for units $i = 1, \dotsc, N$ and time periods $t = 1, \dotsc, T$, where $y_{it}$ is a scalar dependent variable and $w_{it}$ and $x_{it}$ are covariate vectors. Unit $i$ belongs to group $g_i^0 \in \mathbb{G} = \{1, \dotsc, G\}$. Group memberships are unobserved. The data are generated from the model
where $v_{it}$ is a noise term with variance one and is potentially serially correlated, and $\sigma_i$ is a latent heteroscedasticity parameter. The slope coefficient $\theta^w$ on $w_{it}$ is common to all units. The slope coefficient on $x_{it}$ is group-specific and given by $\theta_g$ for units $i$ belonging to group $g \in \mathbb{G}$.
We assume that the regressors $(w_{it}', x_{it}')'$ are uncorrelated with the contemporaneous error term $\sigma_i v_{it}$,
This assumption does not rule out predetermined regressors such as lagged dependent variables.
Different estimation strategies for estimating the common and group-specific coefficients ($\theta^w$ and $\theta_g$) have been proposed in the literature. For example, bonhomme2015grouped estimate slope coefficients $\hat{\theta}^w, \hat{\theta}_1, \dotsc, \hat{\theta}_G$ and group memberships $\hat{g}_1, \dotsc, \hat{g}_N$ simultaneously by solving the least-squares problem
via the kmeans algorithm. The choice of squared loss is justified under the orthogonality condition (ref). The estimators in su2016identifying, wang2016homogeneity augment a squared loss function by a penalization scheme that imposes the grouped structure.
We assume that the number of groups $G$ is either pre-specified or consistently estimated. Consistent estimates of $G$ can be obtained, for example, by using the information criteria proposed by bonhomme2015grouped or su2016identifying or by employing the testing procedure in lu2017determining.
This section describes our confidence set for group membership. First, we define the formal requirements for an asymptotically valid confidence set for group membership. Second, we show that each group allocation corresponds to a set of moment inequalities and that a confidence set can be obtained by inverting a test of these inequalities. Finally, we introduce the test statistic and critical values.
A joint confidence set of group membership at confidence level $1 - \alpha$ is a random set $\widehat{C}_{\alpha}$ of vectors in $\mathbb{G}^{N}$ that satisfies
where $\mathbb{P}_N$ is a class of data-generating processes. If we observe $ \{g_i\}_{1 \leq i \leq N} \in \widehat{C}_{\alpha}$, then the group structure that assigns unit $i = 1, \dotsc, N$ to group $g_i$ is not ruled out by the data at confidence level $1 - \alpha$. Inequality (ref) ensures that the confidence set is asymptotically valid in the sense that it rules out the population partition at most with probability $\alpha$.
We impose uniform validity over sequences $P_N$ on $\mathbb{P}_N$. Changing the data-generating process along the asymptotic sequence allows for $\sigma_i$ that diverge as $N \to \infty$ for some units $i$, rendering these units potentially misclassified in the limit. Data-generating processes that are constant in $N$ cannot model asymptotic clustering uncertainty dzemskiokui2021convergence.
We construct our joint confidence set by combining unit-wise marginal confidence sets. This approach is computationally simple and can be tabulated and visualized easily. The marginal confidence set for unit $i$ is computed by inverting a test for group membership
where $\hat{g}_i$ denotes an estimator of the group membership of unit $i$, $\widehat{T}_i(g)$ is a test statistic and $\hat{c}_{\alpha, N, i}$ is a unit-specific and data-dependent critical value. Test statistic and critical value are defined in Sections (ref) and (ref), respectively.
By explicitly adding $\hat{g}_i$ to the marginal confidence set, we guarantee that the joint confidence set is never empty and can always be interpreted as containing the estimated group structure padded by a margin of error. For the typical unit, inverting the test already includes the unit's estimated group membership in its marginal confidence set.\footnote{In our simulations, we find that the probability of the test rejecting the estimated group membership to be very close to, but not equal to, zero (see Supplemental Material \ref*{sec:testing ghat}).}
Our joint confidence set is given by the Cartesian product of the unit-wise confidence sets:
We use Bonferroni correction to control dependence between units and compute each unit-wise marginal confidence set at a nominal level of $1 - \alpha/N$.
In principle, it is possible to construct a joint confidence set directly without first computing unit-wise confidence sets. This can be accomplished by inverting a joint test for group membership. Testing group memberships for all groups simultaneously avoids the possible power loss from Bonferroni correction. However, inverting the test to obtain the confidence set requires testing all $G^N$ possible groupings. This task is computationally infeasible unless $N$ is very small. In contrast, our approach carries out only $G \times N$ tests and is feasible even if $N$ is large.
An additional advantage of Bonferroni correction is that it produces confidence sets that are easy to report and to interpret. The joint confidence set can be fully described by reporting the marginal unit-wise confidence sets without enumerating all $N$-dimensional vectors contained in $\widehat{C}_{\alpha}$. Interpreting a potentially large collection of such high-dimensional vectors would be challenging.
The Bonferroni correction renders our confidence set robust to any kind of cross-sectional dependence. This correction is only minimally conservative if the unit-wise confidence sets are approximately independent, i.e., if
where $p_\Delta$ is a small number. Approximate independence holds if units are cross-sectionally independent and $N$ and $T$ are large enough to estimate the group-specific coefficients precisely.
For example, suppose that $N \geq 8$ and $1 - \alpha = 0.9$ and that conditions (ref) and (ref) hold with $p_\Delta \approx 0$. Theorem (ref) implies that the Bonferroni correction inflates the joint coverage probability by only about $0.5$-$0.55\%$.
Our approach to testing the group membership hypothesis $H_0: g_i^0 = g$ is based on
The first two terms on the right-hand side are squared residuals representing the fit of assigning unit $i$ to group $g$ and the fit of assigning unit $i$ to group $h$, respectively. The third term applies moment re-centering and ensures that $d_{it}(g, h)$ has mean zero under the null hypothesis. This can be seen by re-writing $d_{it}(g, h)$ as
The first term on the right-hand side has mean zero under the orthogonality assumption (ref). If the null hypothesis is true, i.e., if $g_i^0 = g$, then the second term vanishes and $\sum_{t = 1}^T \mathbb{E} [d_{it} (g, h)] = 0$ for all $h \neq g$.
If $\sum_{t = 1}^T \mathbb{E}[x_{it} x_{it}']$ has full rank and the null hypothesis is false, then $\sum_{t=1}^T \mathbb{E} \left[ d_{it}(g, h) \right] > 0$ for some $h \neq g$. The strict inequality holds because for $h = g_i^0 \in \mathbb{G}\setminus \{g\}$ the second term in $d_{it}(g, h)$ is a quadratic form with strictly positive mean.
In summary, testing $H_0: g_i^0 = g$ is equivalent to testing
against
This is a one-sided significance test for a vector of moments.
Our test statistic $\widehat{T}_i (g)$ for unit $i=1, \dotsc, N$ is the maximum of $(G-1)$ statistics $\widehat{D}_i(g, h)$ that test a hypothesized group membership $g$ against alternative group assignments $h \neq g$:
$\widehat{D}_i (g, h)$ tests group $g$ against group $h$ based on the restriction $\sum_{t=1}^T \mathbb{E}[d_{it}(g, h)]=0$ from the previous section and is equal to the $t$-statistic
where $\hat{d}_{it}(g, h)$ is a sample counterpart of $d_{it} (g,h)$ that replaces the true slope coefficients $\theta^w$ and $\theta_g$ by their estimated values $\hat{\theta}^w$ and $\hat{\theta}_g$,
and $\widehat{\Xi}_i (g,h,h)$ is an estimator of the long-run variance of $d_{it}(g, h)$.
The long-run variance estimator is kernel-based as in Andrews91 and NeweyWest87. In order to define the estimator, let $\widehat{H}_{ij} (g, h, h')$ denote the sample covariance of order $j$ between $\hat{d}_{it}(g, h)$ and $\hat{d}_{it} (g, h')$,
where $\bar{\hat{d}}_{i} (g, h) = \frac{1}{T}\sum_{t = 1}^T \hat{d}_{it} (g, h)$. The long-run variance-covariance estimator is given by
where $K(\cdot)$ is a kernel function and $\kappa_N$ is a bandwidth parameter. In our simulation studies and empirical application, we use the quadratic spectral (QS) kernel Andrews91 given by
We select a data-driven bandwidth $\hat{\kappa}_N$ by using the following algorithm adapted from chang2022central:
The critical value $\hat{c}_{\alpha, N, i}(g)$ is computed from the multivariate $t$-distribution (MVT) in $(G-1)$ dimensions.\footnote{ Using the multivariate $t$-distribution instead of a Gaussian distribution improves the finite sample performance of our confidence set if $T$ is small. } In the case of two groups ($G=2$), this distribution is equal to Student's $t$-distribution and our critical value is given by
where $t_{T - 1}^{-1}(p)$ denotes the $p$-quantile of Student's $t$-distribution with $(T - 1)$ degrees of freedom. This critical value is straightforward to evaluate in most statistical software packages.
For $G \geq 3$, the computation of the critical value is more involved and requires the estimation of unit-specific correlation matrixes. In Section (ref), we discuss a conservative approximation of the critical value that is easy to implement and independent of the data.
The critical value is given by
where $t_{\max, \Omega, T-1}$ denotes the distribution function of the maximal entry of a centered random vector with multivariate $t$-distribution with scale matrix $\Omega$ and $(T-1)$ degrees of freedom, $\widehat{\Omega}_i (g)$ is an estimator of the correlation matrix of the moment inequalities, $\rho$ is a regularization function and $\epsilon_N$ is a regularization parameter.\footnote{ The distribution function of the multivariate $t$-distribution can be efficiently approximated by modern algorithms genz1992numerical. Implementations exist for Stata grayling2016mvtnorm and R GenzRPackage.}
To define the estimated correlation matrix $\widehat{\Omega}_i (g)$ for given $g$ and $i$, map $j, j' = 1, \dotsc, G-1$ to $h, h' \in \mathbb{G}$ such that $h$ and $h'$ give the $j$th and $j'$th element of the vector $\mathbb{G} / \{g\}$, respectively.\footnote{More formally, $h = j - \mathbf{1}_{\{j > g\}}$ and $h' = j' - \mathbf{1}_{\{j' > g\}}$, where $\mathbf{1}_{\{\cdot \}}$ is the indicator function.} $\widehat{\Omega}_i (g)$ is given by the $(G - 1) \times (G - 1)$ matrix with entry $(j, j')$ equal to
For a $(G-1) \times (G-1)$ correlation matrix $\Omega$, the regularization function $\rho$ is defined as
where $I_{G-1}$ is the identity matrix in $\mathbb{R}^{G-1}$, $\operatorname{diag}(A)$, for a matrix $A$, returns a diagonal matrix of the same dimension as $A$ with the diagonal entries equal to the diagonal entries of $A$, and
We set the regularization parameter equal to $\epsilon_N = 0.01$. For robustness, we also conduct simulations with other values of $\epsilon_N$ and find that the exact choice of $\epsilon_N$ is not crucial to the validity of our method. Regularization is a technical tool needed to prove the asymptotic validity of our confidence set.
Our regularization scheme bounds pairwise correlations away from one but may output a singular matrix. A related regularization scheme in andrews2012inference bounds the resulting matrix away from singularity.
Our asymptotic framework is of the long-panel variety and takes both the number of units $N$ and the number of time periods $T$ to infinity. In particular, we consider asymptotic sequences in which $T= T(N)$, where $T(\cdot)$ is increasing, but its exact form is unspecified except for conditions given in the statement of the theorems. In many panel data sets, the number of units far exceeds the number of time periods. We replicate this feature along the asymptotic sequence by allowing $N$ to diverge at a much faster rate than $T$.
We consider a sequence $\mathbb{P}_N$ of classes of probability measures. All our theoretical results hold uniformly over the sequence $\mathbb{P}_N$. For a probability measure $P$, let $\mathbb{E}_P$ denote the expectation operator that integrates with respect to measure $P$. The parameters $\theta^w$, $\{\theta_g\}_{g \in \mathbb{G}}$ and $\sigma_i$ depend potentially on $N$. For notational convenience, we keep this dependence implicit.
The number of latent groups $G$ is fixed and does not depend on $N$.
To state our assumptions, we define the matrix $\Omega_i (g_i^0)$ which provides the population counterpart to $\widehat{\Omega}_i (g_i^0)$. For unit $i$, define the population long-run covariance of the averages of $d_{it} (g_i^0, h)$ and $d_{it}(g_i^0, h')$ as
where
For $i = 1, \dotsc, N$, $\Omega_i (g_i^0)$ is the $(G-1) \times (G-1)$ correlation matrix with entries
where the convention relating $(j, j')$ to $(h, h')$ is defined in footnote (ref).
We now state our assumptions for the asymptotic validity of our joint confidence set.
Assumption (ref).(ref) requires the regressors to be uncorrelated with the contemporaneous error term. It is a much weaker exogeneity assumption than, for example, strict exogeneity and allows for a rich set of regressors, including lagged dependent variables. Assumption (ref).(ref) requires the number of groups to be consistently estimated. This condition is weak and can be guaranteed by using an appropriate procedure for choosing the number of groups lu2017determining,vogt2017clustering.
Assumption (ref).(ref) requires the estimators $\hat{\theta}^w$ and $\hat{\theta}_g$ to be consistent for $\theta^w$ and $\theta_g$, respectively, at a rate that vanishes as fast as or faster than $r_{\theta, N}$. Theorem (ref) below require $r_{\theta, N}$ to vanish faster than $(T \log N)^{-1/2}$. Therefore, $\hat{\theta}^w$ and $\hat{\theta}_g$ have to converge faster than $T^{-1/2}$. Since $T^{-1/2}$ is the rate obtained by estimators based on time-series regression within units, it is important to choose estimators that also exploit cross-sectional variation such as the kmeans estimator bonhomme2015grouped or the estimators in su2016identifying,wang2016homogeneity. These estimators are known to be $\sqrt{NT}$-consistent under assumptions that rule out misclassification in the limit. dzemskiokui2021convergence study consistency of the estimated coefficients under weaker assumptions. They distinguish between units with $\sigma_i \prec \sqrt{T/\log N}$ that can be classified reliably and noisy units with $\sigma_i \succeq \sqrt{T/\log N}$ that are potentially misclassified in the limit. They show that the kmeans estimator is $\sqrt{NT}$-consistent if the proportion of noisy units is sufficiently small. In Section (ref), we complement this theoretical result by simulating designs where kmeans misclassifies units but still recovers group-specific coefficients at a rate that is faster than $\sqrt{T}$.
The full-rank condition in Assumption (ref).(ref) ensures that the denominator of $\widehat{D}_{i} (g,h)$ is not too close to zero. The term in the minimum eigenvalue function is the population long-run covariance matrix of $v_{it}x_{it}$. The assumption restricts $v_{it}$ (which has variance one) but does not limit the magnitude of the error term $\sigma_i v_{it}$ in our panel model (ref). It allows conditional heteroskedasticity of $v_{it}$ given $x_{it}$.
Assumption (ref).(ref) maintains that groups are unique in the sense that there are not two groups that share the same coefficient values. The minimal distance between any two groups is measured by $\iota_N$ and is allowed to vanish asymptotically, provided that it satisfies additional rate conditions stated below. Vanishing group separation is an asymptotic modeling device to study settings where groups are distinct but difficult to distinguish. Most existing results that establish asymptotic properties of estimators of the group-specific coefficient assume strict group separation, i.e., that $\iota_N$ is bounded away from zero. This makes it difficult to verify Assumption (ref).(ref) if $\iota_N \to 0$. This difficulty can be overcome by using our result for kmeans estimation in Supplemental Material \ref*{appendix:weak separation}, which gives a consistency rate under vanishing group separation.
Assumptions (ref).(ref)-(ref).(ref) restrict the distribution of the time series $\{x_{it}', w_{it}', v_{it}\}_{t = 1}^T$. Assumption (ref).(ref) imposes exponential decay of the tails of the marginal distributions. Assumption (ref).(ref) restricts the time-series dependence of the data by imposing exponential decay of the mixing coefficients. This assumption rules out processes with long memory. Assumption (ref).(ref) requires the time series to be stationary. We use the stationarity assumption primarily to show that certain long-run variances are bounded. It can be replaced by other conditions that bound the long-run variances.
Assumption (ref).(ref) rules out that the correlation matrix $\Omega_i (g_i^0)$ contains entries that are too close to negative one. This assumption does not rule out singularity of $\Omega_i (g_i^0)$. Singularity occurs mechanically in our setting whenever $G > p + 1$, where $p$ is the dimension of $x_{it}$.\footnote{Our testing approach is designed to be able to handle singular correlation matrices. In contrast, e.g., the quasi-likelihood ratio statistic used in kudo1963multivariate is not defined for singular correlation matrices.} No restrictions are placed on positive correlations, which is important in settings with vanishing group separation, where groups $h$ and $h'$ have similar coefficients and hence highly positively correlated moment inequalities.\footnote{In Supplemental Material \ref*{appendix:weak separation}, we develop a theoretical framework based on local alternatives to study settings with very similar groups.} Our regularization approach controls positive correlations that are close to one.
We now introduce the last assumption that restricts the choice of kernel function. The validity of this assumption is under complete control of the researcher and does not depend on the underlying data. It is satisfied by the QS kernel Andrews91.
The following theorem establishes the validity of our joint confidence set. In the limit, it covers the true group membership at least with pre-specified probability $1- \alpha$.
In addition to the abovementioned assumptions, this theorem introduces some rate conditions. The first condition restricts the relative magnitudes of $T$ and $N$. It still accommodates both “short panels” where $T$ is small relative to $N$ and “long panels” where $T$ is large relative to $N$. The second condition requires the regularization parameter $\epsilon_N$ to vanish at a sufficiently slow rate. The third rate condition controls the rate at which the bandwidth sequence $\kappa_N$ diverges. This rate condition is automatically satisfied for the QS kernel if the bandwidth is chosen by the procedure described in Section (ref). Finally, condition (ref) restricts the rate of convergence of the estimators $\hat{\theta}^w$ and $\hat{\theta}_g$.
We now discuss the latter condition in the context of two examples. For both examples, assume that $\sigma_i$ is bounded away from zero uniformly over units $i$. For the first example, groups are well-separated, i.e., $\iota_N$ is a positive constant and does not depend on $N$. Existing results for clustering with well-separated groups suggest $r_{\theta, N} = (NT)^{-1/2} \zeta_N$ with $\zeta_N \to \infty$ bonhomme2015grouped,su2016identifying,wang2016homogeneity.\footnote{The sequence $\zeta_N$ can go to infinity at any slow rate.} Under this convergence rate, condition (ref) is equivalent to $\log N / N = o(1)$ and hence trivially satisfied. For the second example, group separation is vanishing with $\iota_N = T^{-1/2 + e}$ for $0 < e < 1/2$. In Theorem \ref*{thm:weak-sep_gamma-beta-w} in Supplemental Material \ref*{sec:weak-sep_theory}, we show $r_{\theta, N} = (NT)^{-1/2} \zeta_N$, under some technical conditions. Then, condition (ref) becomes $T^{1 - 2e} \log N / N = o(1)$. This condition restricts the relative magnitudes of $N$ and $T$. In particular, the weaker group separation is, i.e., the closer $e$ is to zero, the larger $N$ has to be relative to $T$. It is sufficient that $N \geq T^{\delta_1}$ with $\delta_1 > 1 - 2e$.
A caveat to the calculations in the previous paragraph is that the results for the rate of consistency of the kmeans estimator rely on assumptions that rule out misclassification in the limit. In the discussion following Theorem \ref*{thm:weak-sep_gamma-beta-w} in Supplemental Material \ref*{sec:weak-sep_theory}, we indicate how this limitation can be overcome based on the approach in dzemskiokui2021convergence, but leave a formal proof to future research.
This section discusses several extensions of our procedure. We first demonstrate the possibility of simplifying the procedure at the cost of power and/or losing robustness against serial correlation. We also propose a two-step method to increase the power of our confidence set.
The implementation of our confidence set can be greatly simplified by using different critical values that are slightly conservative but can be computed without estimating and regularizing a covariance matrix. We call these the SNS critical values, borrowing a term from chernozhukov2013testing who propose similar critical values for a high-dimensional testing problem and justify them using the theory of self-normalized sums (SNS).
For $G=2$, the SNS critical values are identical to our critical values. For $G \geq 3$, the SNS critical values are an upper bound to our critical values and will always yield a weakly larger confidence set. The SNS critical values are given by
with $t_{T - 1}^{-1}(p)$ defined in Section (ref). The factor $(G - 1)$ carries out a Bonferroni correction to account for the $(G-1)$ moment inequalities that are simultaneously tested for unit $i$. The critical values do not depend on $i$ nor $g$; identical values can be used for all $i$ and $g$. Let $\widehat{C}_\alpha^{\text{SNS}}$ denote the confidence set computed by applying our procedure with SNS critical values. The following result establishes the asymptotic validity of this confidence set.
Another simplification of our procedure is possible in the absence of serial correlation. Suppose that, for each unit $i$, the time series $\{x_{it} v_{it}\}_{t=1}^T$ is serially uncorrelated.
Under this assumption
can be consistently estimated by $\widehat{\Xi}_i (g, h, h')$ by setting
Then $\widehat{\Xi}_i (g, h, h')$ takes the simple form
and the test statistic is given by
Critical values are computed by $\hat{c}_{\alpha, N,i} (g)$ with the simplified version of $\widehat{\Omega}_i (g)$. The following result states the conditions for the validity of the simplified procedure. In contrast to Theorem (ref), stationarity is not required.
Combining the test statistic (ref) under no serial correlation with SNS critical values yields a particularly simple procedure that can be implemented with minimal programming effort.
For units that are very easy to classify (i.e., have very small $\sigma_i$), we estimate the true group memberships with probability strictly larger than $1 - \alpha / N$. This renders our confidence set conservative. Using the information provided by the units that are easy to classify, we may be able to shrink the marginal confidence sets for the units that are difficult to classify (i.e., have large $\sigma_i$). This idea inspires our two-step procedure that we call unit selection.
The key part of our two-step procedure is the detection of units that are easy to classify. For these units, we report singleton marginal confidence sets. We then compute our joint confidence on the sub-sample of remaining units. We slightly adjust the nominal level of the confidence set to control for classification error in unit selection. Unit selection can increase the power of the confidence set because it carries out fewer simultaneous tests than our one-step procedure. For example, if unit selection eliminates $N/3$ units, the resulting confidence set is based on Bonferroni correction to adjust for $2N/3$ rather than $N$ simultaneous tests.
For our two-step procedure we assume that $\hat{g}_i$ minimizes squared loss, i.e., we assume
where
This requirement is automatically satisfied if the grouped panel model is estimated by kmeans clustering bonhomme2015grouped.
Our algorithm for unit selection identifies a unit as easy to classify if it satisfies two conditions that we call moment selection and hypothesis selection.\footnote{ The term “moment selection” is borrowed from the literature on testing moment inequalities chernozhukov2013testing,andrews2010inference,andrews2012inference,romano2014practical,canay2016practical. In this literature, moment selection reduces the power loss from possibly slack moment inequalities by identifying inequalities that are “obviously” satisfied. Our use is different. We use moment selection to reduce the power loss from running many simultaneous tests by identifying units that are “obviously” correctly classified. } A unit $i$ satisfies the moment selection criterion if we detect substantial slackness in inequality (ref) for all $h \neq \hat{g}_i$. A unit $i$ satisfies the hypothesis selection criterion if all group memberships $h \neq \hat{g}_i$ are rejected.
The unit selection procedure is parameterized by $\beta \in [0,\alpha/3)$. The larger $\beta$, the more unit selection is carried out. Let
where $\widehat{\Xi}^U_{i}(g, h, h)$ is defined as $\widehat{\Xi}_{i}(g, h, h)$ with $\hat{d}_{it}^U$ replacing $\hat{d}_{it}$. $\widehat{D}_i^U (g,h)$ is a counterpart to $\widehat{D}_{i}(g, h)$ that does not adjust for the mean under the null hypothesis.
The first step of our two-step procedure carries out moment selection by computing the set
for $g \in \mathbb{G}$ and $i = 1, \dotsc, N$. This set gives the selected moment inequalities for the hypothesis $H_0: g_i^0 = g$. For units $i$ for which $\widehat{M}_i (g)$ is empty, we have strong evidence that $g_i^0 = g$. These units satisfy the moment selection criterion for elimination in the first step. Condition (ref) ensures that $\widehat{M}_i(g)$ is never empty for $g \neq \hat{g}_i$. This property ensures that moment selection does not eliminate misclassified units.
The second step of our two-step procedure is given by the following algorithm that carries out hypothesis selection:
Step 2.A initializes the algorithm by designating all possible group assignments as hypotheses that have to be tested. Step 2.B counts the number $\widehat{N}(s)$ of units that are not easy to classify. Unit $i$ is easy to classify if and only if $\widehat{M}_i(\hat g_i)$ is empty (moment selection) and $H_i (s) = \{\hat g_i\}$ (hypothesis selection). Step 2.C carries out hypothesis selection with critical values adjusted for $\widehat{N}(s)$ simultaneous tests of group membership. $H_i(s+1)$ gives a preliminary marginal confidence set for unit $i$ after $s$ iterations of hypothesis selection. Step 2.D checks the convergence of the algorithm.\footnote{The algorithm always converges since $\widehat{N} (s)$ and the cardinality of $H_i (s)$ are decreasing in $s$.} If the algorithm has converged after $s = s^*$ iterations, the final joint confidence set is given by
The second-step confidence set is calculated at nominal confidence level $1 - \alpha + 2 \beta$ (see the definition of $H_i(s+1)$ in Step 2.C). The adjustment by $2 \beta$ represents the cost of unit selection and controls for two possible errors at the first step. The first error is estimating an incorrect group membership for a unit that is easy to classify “in population”. The second error is erroneously declaring a unit as easy to classify.
Unit selection increases the power of the confidence set if its benefits (decreasing the number of units at the second step) outweigh the cost (adjustment of nominal level at the second step). If sufficiently many units can be eliminated, then $\widehat{C}_{\mathrm{sel},\alpha, \beta}$ is more powerful (“smaller”) than the corresponding one-step confidence set $\widehat{C}_{\alpha}$. If too few units are eliminated, then a two-step confidence set can be slightly more conservative (“larger”) than the corresponding one-step confidence set.
We establish the validity of the two-step procedure under the assumption of no serial correlation (Assumption (ref) above) and the following additional assumption.
The first part of this assumption is standard. The second part imposes an orthogonality condition on triples of time periods. It is stronger than the assumption of no serial correlation between pairs of time periods but weaker than independence across time. We now state the formal result for the asymptotic validity of the two-step procedure.
In this section, we revisit the work by wang2019heterogeneous to illustrate our procedures. wang2019heterogeneous estimate a grouped panel model to study heterogeneous effects of a minimum wage in the restaurant sector. Their analysis builds on dube2010minimum who employ a similar panel model but do not allow for effect heterogeneity. We assess the clustering uncertainty of the group memberships estimated in wang2019heterogeneous by computing confidence sets for group memberships.
We use the panel data described in dube2010minimum. It contains quarterly data for 1380 US counties, ranging from the first quarter of 1990 to the second quarter of 2006. The grouped panel model estimated in wang2019heterogeneous is given by
for state $i = 1, \dotsc, 51$, county $c = 1, \dotsc, n_i$ and time period $t = 1, \dotsc, T = 66$, where $n_i$ is the number of counties in state $i$. The variable $\texttt{emp}_{ict}$ gives employment in the restaurant sector, $\texttt{mw}_{ict}$ gives the minimum wage and $\texttt{emp}_{ict}^{\text{TOT}}$ gives total employment in all sectors. Finally, $\phi_c$ is a county fixed effect, $\tau_t$ is a time fixed effect, and $\sigma_i v_{ict}$ is an idiosyncratic error term.
su2016identifying propose an information criterion to select the number of groups $G$ in a grouped-panel regression and prove its consistency. wang2019heterogeneous apply this criterion to the data and select $G = 4$. Table (ref) gives their CLasso su2016identifying estimates for the slope coefficients for the four groups.
There are two groups with estimated positive effects of the minimum wage on employment and two groups with negative effects (see $\theta_{g,1} $ in the table).
Based on these estimates for the slope coefficients, we estimate group memberships by running one update step of the kmeans algorithm. This updating step ensures that the estimated group memberships satisfy inequality (ref) and produces estimated group memberships that are identical to the CLasso estimates in all but six states.
Estimated group memberships are displayed in Figure (ref) and reported in Table \ref*{tab:supp:cs_1step} in the Supplemental Material. Note that Figure (ref) is slightly different from Figure 2 in wang2019heterogeneous due to the additional kmeans step.
To translate the panel model to our framework, we identify states (subscript $i$) as the cross-sectional dimension and counties and quarters (subscripts $c$ and $t$) jointly as the second (“time”) dimension. We compute our confidence sets based on the fixed-effect transformation
Here, $\widetilde{\texttt{lemp}}_{ict} = \log (\texttt{lemp}_{ict}) - \sum_{t'= 1}^T \log (\texttt{lemp}_{ict'}) /T$ and $\widetilde{\texttt{lmw}}_{ict}$, $\widetilde{\texttt{lpop}}_{ict}$, and $\widetilde{\texttt{lemp}}_{ict}^{\texttt{TOT}}$ are defined similarly. Moreover, $\tilde{\tau}_t = \tau_t - \sum_{t' = 1}^T \tau_{t'} /T $ and $\tilde{v}_{ict} = v_{ict} - \sum_{t' = 1}^T v_{ict'} /T$. Remark (ref) above heuristically justifies applying our procedure to fixed-effect-transformed data.
Our second panel dimension comprises both cross-sectional variation between counties and time-series variation between different quarters. To adapt our definition of the estimated long-run variance to this setting, we redefine $\widehat{H}_{ij}$ in (ref) as
where $\bar{\hat{d}}_{ic} = \sum_{t = 1}^T \hat{d}_{ict}/T$ and $\hat{d}_{ict}$ is defined as in Section (ref) with the double index $ct$ replacing $t$. We use the following modified algorithm for bandwidth selection:
This algorithm yields .
We compute the joint-confidence set at level $1 - \alpha = 0.95$. For the regularization parameter, we set $\epsilon_N = 0.01$. Our results are robust to different choices of $\epsilon_N$.\footnote{For $\epsilon_N = 0, 0.05$, we obtain a confidence set that differs only in the group assignments for Kentucky. For $\epsilon_N = 0.01, 0.05$, group $g=2$ is contained in the marginal confidence set for Kentucky ($p$-values $0.0662$ and $0.0610$). For $\epsilon = 0$, it is not ($p$-value $0.0372$).}
The full joint confidence set is reported in Table \ref*{tab:supp:cs_1step} in the Supplemental Material. Marginal confidence sets for a subset of states are given in Table (ref). The cardinality of the state-wise marginal confidence sets is four for three states, three for seven states, two for twelve states, and one for twenty-nine states.
Based on our joint confidence set, we compute $p$-values for the significance of the estimated group memberships. We say that the estimated group membership $\hat{g}_i$ for state $i$ is significant at level $\alpha$ if the marginal confidence set for state $i$ with confidence level $1-\alpha$ contains only $\hat{g}_i$. Up to a joint failure probability of at most $\alpha$, significantly estimated group memberships reveal the true group membership and cannot be attributed to estimation error. The $p$-value for significance of $\hat{g}_i$ is the smallest value of $\alpha$ such that $\hat{g}_i$ is significant at level $\alpha$. The $p$-values for a subset of states are reported in Table (ref), and results for all states are reported in Table \ref*{tab:supp:cs_1step} in the Supplemental Material.
Table (ref) demonstrates that, for our sample, the SNS critical values from Section (ref) yield slightly larger confidence sets than our baseline procedure.
Displaying a visual representation of the estimated clusters as we do in Figure (ref) is standard practice even in applications where the clustering structure is a nuisance parameter and not of interest in its own right. Visual inspection of the clusters is meant to confirm their economic plausibility and serves as an informal test of model specification. Based on our confidence set, such an informal analysis can be complemented by a graphical representation of clustering uncertainty as illustrated in Figure (ref).
The upper panel in Figure (ref) represents clustering uncertainty by shading US states according to the $p$-values of their estimated group memberships. The lower panel shades states by the cardinality of their marginal confidence sets ($1 - \alpha = 0.95$). The figure suggests a low degree of clustering uncertainty. It would suggest a large degree of clustering uncertainty if the maps were shaded mostly in a dark hue. Large clustering uncertainty indicates that the clustering algorithm may be overfitting on the sample, rather than picking up structural heterogeneity.
Our two-step procedure with parameter $\beta = \alpha / 5 = 0.01$ eliminates only one unit in the first step and does not improve the final confidence set or $p$-values. In Supplemental Material \ref*{app:application 2step}, we apply the two-step procedure under the implausible assumption of no serial correlation which eliminates more units and lowers the final $p$-values.
Our simulation designs are based on our empirical application. The data-generating process is given by
for $i = 1, \dotsc, N$ and $t = 1, \dotsc, T$. For $g = 1, \dotsc, 4$, the coefficients $\theta_{g, 1}$, $\theta_{g, 2}$ and $\theta_{g, 3}$ are set equal to the estimated coefficients in Table (ref).
To generate $x_{it} = (\widetilde{\texttt{lmw}}_{it}, \widetilde{\texttt{lpop}}_{it}, \widetilde{\texttt{lemp}}_{it}^{\text{TOT}})$ that exhibit similar patterns of serial correlation as the time series in our empirical application, we estimate a $\text{VAR}(1)$ model of $x_{it}$ for each county $c$ in our data sample and save the estimated coefficient $\hat{\Gamma}_c$ and the empirical residuals $\hat{e}_{ct}$. We then average over all county-wise coefficients to obtain a single $\text{VAR}(1)$ coefficient matrix $\bar{\Gamma}$. To generate a time series $(x_{it})_{t = 1}^T$, we sample the initial value $x_{i0}$ from the covariates observed in our data and then iteratively generate $x_{i,t+1} = \bar{\Gamma} x_{it} + e_t$, where $e_t$ is a randomly drawn empirical residual $\hat{e}_{ct}$.
Each unit $i$ is assigned to one of the four groups with equal probability and assigned a heteroscedasticity parameter $\sigma_i$. For $\sigma = 0.1, 0.2$, we draw $\sigma_i = \sigma \chi^2(4)/4$, where $\chi^2 (df)$ is a random draw from a $\chi^2$-distribution with $df$ degrees of freedom.
For $\rho = 0, 0.25, 0.5$, we set $v_{it} = \sqrt{1 - \rho^2} \tilde{v}_{it}$ where $\tilde{v}_{it}$ follows an $AR(1)$ process with autoregressive parameter $\rho$ and standard normal innovations and $\tilde{v}_{i0}$ is drawn from the stationary distribution of $(\tilde{v}_{it})_{t = 1}^T$. To introduce relevant serial correlation, we require both $v_{it}$ and $x_{it}$ to be serially correlated (cf. Section (ref)). Therefore, setting $\rho = 0$ turns off the serial correlation (but not all temporal dependence) in the moment inequalities even though $x_{it}$ is still serially correlated.
We simulate panels of size $N = 50, 100, 200$ and $T = 60, 120$. The lower values of these ranges are in the ballpark of the number of states (51 states) and time periods (66 quarters) observed in our application. We set the regularization parameter for covariance estimation to $\epsilon_N = 0.01$. In Supplemental Material \ref*{appendix:epsilon_N}, we demonstrate that our results are robust to alternative specifications of $\epsilon_N$. We simulate both the oracle confidence set, for which we take the true values of the group-specific coefficients as given, and the confidence set with group-specific coefficients estimated by the kmeans estimator. We steer the kmeans algorithm towards recovering the population labels of the groups by using the true group memberships as initial values. The simulation results are based on 500 replications and reported in Table (ref).\footnote{Our Monte Carlo simulations were carried out on computing resources of the Swedish National Infrastructure for Computing (SNIC) funded by Swedish Research Council grant 2018-05973.}
We first discuss the empirical coverage probability of our confidence set with kmeans estimates and a data-driven bandwidth (coverage\textrightarrow kmeans\textrightarrow hac in Table (ref)). The joint confidence set has coverage probability of at least $1 - \alpha = 95\%$ in most cases. This is in line with our theoretical result on the asymptotic validity of the confidence set. When $T=60$, our confidence set under-covers for some designs. In these designs, we compute a large confidence set. This result indicates that our confidence set can detect that statistical uncertainty is large even when it is under-covering. In the designs with moderate serial correlation ($\rho = 0.25$), under-coverage can be alleviated by increasing the number of cross-sectional observations, keeping the number of time series observations constant. In the designs with large serial correlation ($\rho = 0.5$), improving coverage requires increasing the number of time-series observations. This result may be explained by the fact that our algorithm for bandwidth selection mechanically chooses a small bandwidth when the time series is short. With a small bandwidth, our procedure cannot control for temporal dependence over many time periods.
In some of our designs, our confidence set is conservative, with an empirical coverage probability of up to close to one. Since group membership is a discrete parameter, this is not necessarily a sign that the confidence set is underpowered. To see this, consider the na\"ive confidence set that contains only the vector of estimated group memberships. This is the smallest possible confidence set. As $T$ increases, true group memberships are revealed with increasing probability, and even the na\"ive confidence set can become conservative eventually (see Section (ref)). The coverage probability of the na\"ive confidence set (coverage\textrightarrow kmeans\textrightarrow na\"ive in Table (ref)) gives the probability that data-driven clustering recovers the true group structure. In our designs, this probability varies between 0% and 83%.
We now turn to the power of our confidence set with kmeans estimates and a data-driven bandwidth. The simulated average cardinality of a unit-wise marginal confidence set is reported in Table (ref) under cardinality\textrightarrow kmeans\textrightarrow hac. Increasing $\sigma$ or $\rho$ makes the data more noisy as time-series shocks become larger or more persistent. As is expected, the power of our test decreases, and the size of our confidence set increases as the data become more noisy.
In the less noisy designs, the confidence set is highly informative. For example, if $\rho = 0$ and $\sigma = 0.1$, the confidence set rules out 2-3 out of 4 possible group assignments for most units. Our confidence set is only slightly larger than the na\"ive set (i.e., the set containing only the vector of estimated group memberships) in the designs where the na\"ive set performs well. The na\"ive set still under-covers, and some enlargement improves the coverage.
In the noisiest designs, our confidence set becomes quite uninformative and assigns the trivial marginal confidence set (all four possible groups) to most units.
Our asymptotic result is valid under a high-level assumption about the rate at which the estimator $\hat{\theta} = \{\hat{\theta}_g'\}'_{g \in \mathbb{G}}$ converges to the true group-specific coefficients $\theta = \{\theta_g'\}'_{g \in \mathbb{G}}$. This rate may be affected by misclassification. In this section, we study the effect of model estimation in finite samples, focusing on misclassification driven by the noise variance $\sigma^2$. Additional results for settings where model estimation is potentially affected by weak group separation are reported in Supplemental Material \ref*{appendix:weak separation}.
The column “$\hat{\theta} - \theta$” in Table (ref) quantifies estimation error in $\hat{\theta}$ and reports the simulated value of
where $\lVert \cdot \rVert_2$ is the $L_2$-norm. The simulated estimation error is not negligible, varying between 7% and 64% of the magnitude of the true coefficient vector. In all our designs, increasing $T$ keeping $N$ fixed or increasing $N$ keeping $T$ fixed decreases estimation error. This shows that, even though states are misclassified, the kmeans estimator can exploit both time-series and cross-sectional variation. This is necessary for fulfilling the rate condition imposed in our asymptotic result.
As a more direct test of the effect of parameter estimation, we compare the confidence set using the true coefficients (oracle\textrightarrow hac in Table (ref)) to the confidence set using kmeans estimates (kmeans\textrightarrow hac in Table (ref)). In many designs, both confidence sets are valid with a coverage probability of at least 95%. In most designs where the oracle confidence set has coverage of at least 95%, the confidence set with kmeans estimates is valid or undercovers only slightly (1-3%). For all designs, increasing either $N$ or $T$ decreases the difference in coverage probability between the oracle confidence set and the confidence set using the kmeans estimates. This simulation evidence aligns with our theoretical argument that the asymptotic effect of parameter estimation is not of first order.
Next, we compare our benchmark procedure based on the heteroskedasticity-and-autocorrelation-robust (HAC) variance estimator (“hac” in results table) and the alternative procedure discussed in Section (ref) with variance estimation that is robust to heteroscedasticity but not auto-correlation (“h” in results table).
In the designs with no serial correlation ($\rho = 0$), the confidence set using the heteroscedasticity-robust variance estimator is valid. It covers the true group structure at least with probability $1 - \alpha = 95\%$ provided that the sample size is large enough. In these designs, choosing the heteroscedasticity-robust estimator over the HAC estimator increases the power of the confidence set. If $T = 120$, the power gain is modest. If $T = 60$, the power gain can be quite large. For example, in the design with $\rho = 0$, $\sigma = 0.1$, $N = 200$, and $T = 60$, the average cardinality of a unit-wise confidence set is $1.95$ for the HAC estimator and $1.34$ for the heteroscedasticity-robust estimator. This power gain points to the inherent difficulty of estimating long-run variances for a high-dimensional number of relatively short time series.
For the settings with serial correlation ($\rho = 0.25, 0.5$), the confidence set using the variance estimator that is only robust to heteroscedasticity is not valid. In the designs with a lot of serial correlation ($\rho = 0.5$), it under-covers severely. For example, for $\rho = 0.5$, $\sigma = 0.2$, $N = 200$, $T = 120$, its coverage probability is only 48%, even though the sample size is fairly large. Our benchmark confidence set based on the HAC estimator set has correct coverage in this design.
An interpretation of our confidence set is that it distinguishes low-noise units for which the clustering algorithm reveals the true group memberships from noisy units for which group membership is uncertain. Our model and simulation designs parameterize the unit-specific noise-level by the latent parameter $\sigma_i$. Our confidence set picks up the latent heteroscedasticity and assigns, on average, a small marginal confidence set for a unit $i$ with small $\sigma_i$ and a large marginal confidence set for a unit $i$ with large $\sigma_i$. This property is demonstrated in Table (ref), where we report the average cardinality of the marginal confidence set for units with different magnitudes of $\sigma_i$. For example, in the first reported design ($\rho = 0$, $\sigma = 0.1$, $N = 50$, $T = 60$), the average cardinality for the units in the lowest quintile of the distribution of $\sigma_i$ is $1.31$. The average cardinality for the units in the highest quintile of the distribution of $\sigma_i$ is $2.07$.
We have constructed a confidence set for group membership for grouped panel models with time-invariant group-specific regression curves. Our confidence set can be easily tabulated and visualized as demonstrated in our empirical application. Empirical researchers can use our confidence set to quantify the statistical uncertainty about the true group structure in a grouped panel model.
Our method extends naturally to other models with a latent group structure. We construct our confidence set based on a test for the best fit in a least-squares sense. This idea can be adapted to a wide variety of settings with possibly different notions of what constitutes a best fit. For example, it may be possible to compute a confidence set for group membership in non-linear likelihood-based models \parencites{liu2020identification, wang2021identifying} by testing group membership based on the fit measured by the log-likelihood function. This and other extensions of our method require new and non-trivial theoretical work and are interesting avenues for future research.
Our approach can be generalized to models with time-varying coefficients $\theta_g = \theta_{g, t}$, including models with time-varying intercepts as in bonhomme2015grouped. We studied this extension in a previous version of this paper dzemskiokui2018confidence. Our present focus on time-invariant coefficients is motivated by the literature su2016identifying,wang2016homogeneity,vogt2015classification and theoretical considerations. The conditions for establishing the validity of our method for models with time-varying coefficients are less transparent and substantially more restrictive. Intuitively, time-varying coefficients are identified purely from cross-sectional variation and are, therefore, estimated at a slower rate than time-invariant coefficients. Therefore, stronger rate conditions requiring at least $T \log N / N \to 0$ are needed to control estimation error in the time-varying coefficients.
Our confidence set is tailored to address research questions where unit identities are relevant. For example, in our application, we can identify (at a pre-specified confidence level) a set of US states that exhibit a positive effect of the minimum wage on employment. Identifying such states is relevant for conducting further research or implementing targeted policies. Our method cannot be applied directly to assess the effect of misclassification on statistics that average over units without regard to unit labels. One example of such a statistic is the estimator of the group-specific slope coefficients. Developing a theory tailored to averages is an interesting research agenda that is complementary to our work.
\printbibliography