Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
150,000 characters · 36 sections · 148 citation commands
Incorporating Prior Knowledge of Latent Group Structure in Panel Data Models
{5pt} {5pt}
JEL CLASSIFICATION: C11, C14, C23, E31
KEY WORDS: Grouped Heterogeneity; Bayesian Nonparametrics; Dirichlet Process Prior; Density Forecast; Inflation Rate Forecasting; Democracy and Development.
\thispagestyle{empty} \setcounter{page}{0}
Numerous studies have examined and demonstrated the important role of panel data models in empirical research throughout the social and business sciences, as the availability of panel data has increased. Using fixed-effects, panel data permits researchers to model unobserved heterogeneity across individuals, firms, regions, and countries as well as possible structural changes over time. As individual heterogeneity is often empirically relevant, fixed-effects are an objective of interest in numerous empirical studies. For example, teachers' fixed-effects are viewed as a measure of teacher quality in the literature on teacher valued-added rockoff2004, chetty2014; the heterogeneous coefficients are crucial for panel data forecasting liu2020, pesaran2022. In practice, however, researchers may face a short panel where $N$ is large and $T$ is short and fixed. When applying the least squares estimator for the fixed-effects, a significant number of noisy estimates are produced. To alleviate this issue, a popular and parsimonious assumption that has recently been used is to introduce a group pattern into the individual coefficients, so that units within each group have identical coefficients (bonhomme2015, su2016, bonhomme2022).
To recover the group pattern, we essentially face a clustering problem, e.g., dividing $N$ units into several unknown groups. All existing methods for estimating group heterogeneity solve a clustering problem by assuming that units are exchangeable and treating all units equally a priori. In a cross-country application of evaluating the impact of climate change on economic growth hsiang2016, henseler2019,kahn2021, countries in different climatic zones are assumed to have equal probabilities of being grouped together. The assumption of exchangeability might not be reasonable since correlations are common between observations at proximal locations and researchers could have knowledge of the underlying group structure based on theories or empirical findings. For instance, Sweden and Finland, which share a border, an economic structure, and weather conditions, may have a higher chance of being in the same group than African countries. In such a scenario, it is preferable to use additional information to break the exchangeability between countries to facilitate grouping as opposed to clustering based solely on observations in the sample. The availability of this information drives us to formalize such prior knowledge, which we wish to leverage to improve model performance.
In this paper, we focus on the group heterogeneity in the linear panel data model and develop a nonparametric Bayesian framework to incorporate prior knowledge of groups, which is considered additional information that does not enter the likelihood function. The prior knowledge aids in clustering units into groups and sharpens the inference of group-specific parameters, particularly when units are not well-separated.
The whole framework is built on the nonparametric Bayesian method, where we do not impose a restriction on the number of groups, and model selection is not required. The baseline model is a linear panel data model with an unknown group pattern in fixed-effects, slope coefficients, and cross-sectional error variances. We estimate the model using the stick-breaking representation sethuraman1994 of the Dirichlet process (DP) ferguson1973, ferguson1974 prior, a standard prior in nonparametric Bayesian inference. In this framework, the number of groups is considered a random variable and is subject to posterior inference. The number of groups and group membership are estimated together with the heterogeneous coefficients. Moreover, since the DP prior implicitly defines a prior distribution on the group partitionings, the posterior analysis takes the uncertainty of the latent group structure into account.
The derivation of the proposed prior starts with summarizing prior knowledge in the form of pairwise constraints, which describe a bilateral relationship between any two units. Inspired by the work of wagstaff2000, we consider two types of constraints: positive-link and negative-link constraints, representing the preference of assigning two units to the same group or distinct groups. Instead of imposing these constraints dogmatically, each constraint is given a level of accuracy that shows how confident the researchers are in their choice. There is a hyperparameter that controls the overall strength of the prior knowledge: a small value partially recovers the exchangeability assumption on units, whereas a large value confines the prior distribution of group partitioning around group structure based on prior knowledge. We choose the optimal value for the hyperparameter by maximizing the marginal data density. Summarizing prior knowledge in the form of pairwise constraints is practical and flexible since it eliminates the need to predetermine the number of groups and focuses on the bilateral relationships within any subset of units.
The aforementioned pairwise constraints are used to modify the standard DP prior. In particular, the pairwise constraints are combined with the prior distribution of the group partitioning, shrinking the distribution toward my prior knowledge. We refer to the estimator using the proposed prior as the Bayesian group fixed-effects (BGFE) estimator.
We derive a posterior sampling algorithm for the framework with the modified DP prior. Adopting conjugate priors on group-specific coefficients allows for drawing directly from posteriors using a computationally efficient Gibbs sampler. With the newly proposed prior, it can be shown that, compared to the framework that uses a standard DP prior, all that is needed to implement pairwise constraints is a simple modification to the posterior of the group indices.
The pairwise constraint-based framework is closely related and applicable to other models where group structure plays a role. Although we concentrate primarily on the panel data model, the DP prior with pairwise constraints applies to models without the time dimension, such as the standard clustering problem and the estimation of heterogeneous treatment effects. The framework is also applicable to estimating panel VARs holland1983, which involves multiple dependent variables. The group structure is used to overcome overparameterization and overfitting issues by clustering the VAR coefficients into groups, and pairwise constraints add additional information to the highly parameterized model. Moreover, the proposed Gibbs sampler with pairwise constraints is connected to the KMeans-type algorithm, motivating a frequentist's counterpart of our estimator with a fixed $K$. Essentially, the assignment step in the Pairwise Constrained-KMeans algorithm basu2004, a constrained version of the KMeans algorithm macqueen1967, is remarkably similar to the step of drawing a group membership indicator from its posterior. The same exact equivalence can be achieved by applying small-variance asymptotics to the posterior densities under certain conditions. To obtain the frequentist's analog of our pairwise constrained Bayesian estimators, one can utilize the same approach in BM with the Pairwise Constrained-KMeans algorithm.
We compare the performance of the BGFE estimator to alternative estimators using simulated data. The Monte Carlo simulation demonstrates that the BGFE estimator generates more accurate estimates of the group-specific parameters and the number of groups than the BGFE estimator without including any constraints. The improved performance is mostly attributable to the precise group structure estimation. The BGFE estimator clearly dominates the estimators that omit the group structure by assuming homogeneity or full heterogeneity. We also evaluate the performance of one-step ahead point, set, and density forecasts. Unsurprisingly, the accurate estimates translate into the predictive power of the underlying model; the BGFE estimator outperforms the rest of the estimators.
We apply the proposed method to two empirical applications. An application to forecasting the inflation of the U.S. CPI sub-indices demonstrates that the suggested predictor yields more accurate density predictions. The better forecasting performance is mostly attributable to three key characteristics: the nonparametric Bayesian prior, prior belief on group structure, and grouped cross-sectional heteroskedasticity. In a second application, we revisit the relationship between a country's income and its democratic transition. This question was originally studied by acemoglu2008, who demonstrate that the positive income effect on democracy disappears if country fixed effects are introduced into the model. The proposed framework recovers a group structure with a moderate number of groups. Each group has a clear and distinct path to democracy. In addition, we identify heterogeneous income effects on democracy and, contrary to the initial findings, show that a positive income effect persists in some groups of countries, though quantitatively small.
Literature. This paper relates to the econometric literature on clustering in panel data models. Early contributions include sun2005 and buchinsky2005. hahn2010 provide economic and theoretical foundations for fixed effects with a finite support. Most recent works focus on linear\footnote{See wang2021, bonhomme2022, among others, for procedures to identify latent group structures in nonlinear panel data models.} panel data models with discrete unobserved group heterogeneity. lin2012 and sarafidis2015 apply the KMeans algorithm to identify the unobserved group structure of slope coefficients. bonhomme2015 also use the KMeans algorithm to recover the group pattern, but they assume group structure in the additive fixed effects. bonhomme2022 modify this method and split the procedure into two steps. They first classify individuals into groups using KMeans algorithm and then estimate the coefficients. ando2016 improved on BM's approach by allowing for group structure among the interactive fixed effects. The underlying factor structure in the interactive fixed effects is the key to forming groups. su2016 develop a new variant of Lasso to shrink individual slope coefficients to unknown group-specific coefficients. This method is then extended by su2018 and su2019. freeman2022 consider two-way grouped fixed effects that allow for different group patterns in time and cross-sectional dimensions. okui2021 and lumsdaine2022 identify structure breaks in parameters along with grouped patterns. From the Bayesian perspective, kim2019, zhang2020, and liu2022 adopt the Dirichlet process prior to estimate grouped heterogeneous intercepts in linear panel data models in the semiparametric Bayesian framework. moon2023 incorporate a version of a spike and slab prior to recover one core group of units. Alternative methods, such as binary segmentation wang2018 and assumptions, such as multiple latent groups structure cheng2019, cytrynbaum2020 have also been explored to flourish group heterogeneity literature.
Our work concerns prior knowledge. bonhomme2015's grouped fixed-effects (GFE) estimator is able to include prior knowledge, but it is plagued by practical issues to some extent. They add a collection of individual group probabilities as a penalty term in the objective function, which is a $N$ by $K$ matrix describing the probability of assigning each unit to all potential groups. This additional penalty term balances the respective weights attached to prior and data information in estimation. The main challenge is providing the set of individual group probabilities for each potential value of $K$ as the underlying KMeans algorithm requires model selection. It is rather cumbersome to assess these probabilities for each possible $K$ and to adjust for changes in reallocating probabilities across $K$.
None but aguilar2022 explore heterogeneous error variance, and they extend BM's GFE estimator to allow for group-specific error variances. They modify the objective function to avoid the singularity issue in pseudo-likelihood. Despite the fact that their work paves the way for identifying groups in the error variance, their framework is not yet ready to satisfactorily incorporate prior knowledge because they face the same issue as BM. Building on these works, we investigate the value of prior knowledge of group structure.
This paper also relates to the literature of constraint-based semi-supervised clustering in statistics and computer science. Pairwise constraints have been widely implemented in numerous models and have been shown to improve clustering performance. In the past two decades, various pairwise constrained KMeans algorithms using prior information have been suggested wagstaff2001, basu2002, basu2004, bilenko2004, davidson2005, pelleg2007, yoder2017. Prior information is also introduced in the model-based method. shental2003 develop a framework to incorporate prior information for the density estimation with Gaussian mixture models. The Dirichlet process mixture model with pairwise constraints has been discussed in vlachos2008, vlachos2009, orbanz2008, vlachos2010, ross2013. lu2004, lu2007 and lu2007_2 assume the knowledge on constraints is incomplete and penalize the constraints in accordance with their weights. law2004 extents shental2003 to allow for soft constraints in the mixture model by adding another layer of latent variables for the group label. nelson2007 propose a new framework that samples pairwise constraints given a set of probabilities related to the weights of constraints.
Our paper is closely related to paganin2021, who address a similar problem using a novel Bayesian framework. Their proposed method shrinks the prior distribution of group partitioning toward a full target group structure, which is an initial clustering of all units provided by experts. This is demanding since not every application can have a full target group structure, as their birth defect epidemiology study did. Our framework circumvents this problem by using pairwise constraints, which are flexibly assigned to any two units. In addition, the induced shrinkage of their framework is produced by the distance function defined by Variation of Information meilua2007. It can be demonstrated that a partition can readily become caught in local modes, preventing it from ever shrinking toward the prior partition. The use of pairwise relationships in this paper circumvents this issue as well. By fixing the group indices of other pairs, our framework makes sure that the partition with a specific pair that fits our prior belief has a higher prior probability than the partition with a pair that goes against our prior belief.
Outline. In section (ref), we present the specification of the dynamic panel data model with group pattern in slope coefficients and error variances and provide details on nonparametric Bayesian priors without prior knowledge, which are then extended to accommodate soft pairwise constraints. Section (ref) focuses on the posterior analysis, where the posterior sampling algorithm is provided. We also highlight the posterior estimate of group structure and discuss the connection to constrained KMeans models. We briefly discuss the extensions of the baseline model in section (ref). In section (ref), we present empirical analysis in which we forecast the inflation rate of the U.S. CPI sub-indices and estimate the country's income effect on its democracy. Finally, we conclude in section (ref). Monte Carlo simulations, additional empirical results, and proofs are relegated to the appendix.
We begin our analysis by setting up a linear panel data model with group heterogeneity in intercepts, slope coefficients, and cross-sectional innovation variance. We then elaborate a nonparametric Bayesian prior for the unknown parameters that takes prior beliefs in the group pattern into account. We briefly highlight several key concepts of a standard nonparametric Bayesian prior as our proposed prior inherits some of its properties.
We consider a panel with observations for cross-sectional units $i=1, \ldots, N$ in periods $t=1, \ldots, T$. Given the panel data set $(y_{it}, x'_{it})$, a basic linear panel data model with grouped heterogeneous slope coefficients and grouped heteroskedasticity takes the following form:
where $x_{i t}$ are a $p \times 1$ vector of covariates, which may contain intercept, lagged $y_{it}$, other informative covariates. $\alpha_{g_{i}}$ denote the group-specific slope coefficients (including intercepts). $\sigma^2_{g_{i}}$ are the group-specific variance. $g_{i} \in \{1,..., K \}$ is the latent group index with an unknown number of groups $K$. $\varepsilon_{i t}$ are the idiosyncratic errors that are independent across $i$ and $t$ conditional on $g_i$. They feature zero mean and grouped heteroskedasticity $\sigma_{g_{i}}^{2}$, with cross-sectional homoskedasticity being a special case where $\sigma_{g_{i} }^{2}=\sigma^{2}$. This setting leads to a heterogeneous panel with group pattern modeled through both $\alpha_{g_{i}}$ and $\sigma_{g_{i}}^{2}$.
It is convenient to reformulate the model in ((ref)) in matrix form by stacking all observations for unit $i$:
where $\boldsymbol{y_{i}} = \left[y_{i 1}, y_{i 2}, \ldots, y_{i T}\right]^{\prime}$, $ \boldsymbol{x_i} = \left[x_{i 1}, x_{i 2}, \ldots, x_{i T}\right]^{\prime}$, $\boldsymbol{\varepsilon_{i}} = \left[\varepsilon_{i 1}, \varepsilon_{i 2}, \ldots, \varepsilon_{i T}\right]^{\prime}$, and $G = \left[ g_{1}, \ldots, g_{N}\right]$ is a vector of group indices.
Group structure is the key element in our approach. It can be either represented as a vector of group indices $G$ describing to which group each unit belongs or as a collection of disjoint blocks $\mathcal{B} = \left\{B_{1}, B_{2}, \ldots, B_{K}\right\}$ induced by $G$, where $B_{k}$ contains all the units in the $k$-th group and $K$ is the number of groups in the sample of size $N$. $|B_k|$ denotes the cardinality of the set $B_k$ with $\sum_{k=1}^K |B_k| = N$.
Following sun2005, lin2012 and BM, we assume that the composition of groups does not change over time. In addition, for any group $k \neq k'$, we assume that they have different slope coefficients, e.g., $\alpha_{k} \neq \alpha_{k'}$, and no single unit can simultaneously belong to these two groups: $B_k \bigcap B_{k'} = \emptyset$. Note that these assumptions are used to simplified the prior construction and are not necessary to incorporate prior knowledge. As we show in Section (ref), both assumptions can be relaxed by using slightly different priors.
The primary objective of this paper is to estimate the group-specific slope coefficients $\alpha_{g_i}$, group-specific variance $\sigma_{g_{i}}^{2}$, group membership $G$ as well as the unknown number of groups $K$ using full sample and prior knowledge of the group structure. Given estimates of group-specific coefficients, we are able to offer the point, set, and density forecasts of $y_{i t+h}$ for each unit $i$. Throughout this paper, we will concentrate on the one-step ahead forecast where $h=1$. For multiple-step forecasting, the procedure can be extended by iterating $y_{i T+h}$ in accordance with ((ref)) given the estimates of parameters or estimating the model in the style of direct forecasting. The method proposed in this paper is applicable beyond forecasting. In certain applications, the heterogeneous parameters themselves are the objects of interest. For example, the technique developed here can be adapted to infer group-specific heterogeneous treatment effects.
We propose a nonparametric Bayesian prior for the unknown parameters with prior beliefs on the group pattern. Figure (ref) provides a preview of the procedure for introducing prior knowledge into the model. We propose to use pairwise constraints to summarize researchers' prior knowledge, with each constraint accompanied by a hyperparameter $W$ indicating the researchers' levels of confidence in their choice. The $W$ is then incorporated directly in the prior distribution of the group partition $G$, which is induced from a standard nonparametric Bayesian prior, yielding a new prior. We will elaborate the details throughout this subsection and highlight the clustering properties of the underlying nonparametric Bayesian priors in Section (ref).
The derivation of the proposed prior starts from summarizing prior knowledge in the form of pairwise constraints, which describe a bilateral relationship between any two units. Inspired by the literature on semi-supervised learning wagstaff2000,\footnote{We essentially follow the same idea of the pairwise constraints in wagstaff2000. To better demonstrate the beliefs on constraints, we use different names: positive-link and negative-link, rather than must-link and cannot-link. } we consider two types of pairwise constraints: (1) positive-link (PL) constraints, $\mathcal{P}$, and (2) negative-link (NL) constraints, $\mathcal{N}$. A positive-link constraint specifies that two units are more likely to be assigned to the same group, whereas a negative-link constraint indicates that the units are prone to be assigned to different groups.
Instead of imposing these constraints dogmatically, the constraint between units $i$ and $j$ is given a hyperparameter $W_{ij}$ which describes how confident the researchers are in their choice for different types of constraints. $W_{ij}$ is continuously valued on the real line, as depicted in Figure (ref). On the one hand, the sign of $W_{ij}$ specifies the constraint type, with a positive (negative) value indicating a PL (NL) constraint between $i$ and $j$. On the other hand, the absolute value of $W_{ij}$ reflects the strength of the prior belief. We become increasingly confident in our prior belief on units $i$ and $j$ as $|W_{ij}| \to \infty$. If $|W_{ij}| = \infty$, we essentially impose the constraint, which is known as a hard PL/NL constraint. Otherwise, it's a soft PL/NL constraint with a nonzero and finite $W_{ij}$. $W_{ij} = 0$ if there is no prior belief in units $i$ and $j$.
We assume the weight $W_{ij}$ is a logit function of two user-defined hyperparameters\footnote{This parametric form is related to the penalized probabilistic clustering proposed by lu2004, lu2007_2. See detailed discussion in Appendix (ref).}, accuracy $\psi_{i j}$ and type $T_{ij}$:
Accuracy, $\psi_{i j} \in [0.5, 1)$, describes the user-specified probability of assigning a constraint for unit $i$ and $j$ being correct given our prior preference. Specifically, $\psi_{i j} = 1$ implies the constraint between $i$ and $j$ must be imposed since we confident that it is accurate, while specifying $\psi_{i j} = 0.5$ is equivalent to a random guess or no information is provided. $\psi_{i j} $ is bounded below by 0.5, following the assumption that leaving the pair unrestricted is more rational than setting a less likely constraint. The type of constraints is denoted by $T_{i j}$. $T_{i j} = 1$ if unit $i$ and $j$ are specified to be positive-linked, and $T_{i j} = -1$ for a NL constraints. If the pair $(i,j)$ doesn't involve any constraint, we assume $T_{i j} = 0$.
To incorporate these constraints into the prior, we propose modifying the exchangeable partition probability function (EPPF) or the prior distribution of group indices, $p(G)$, of the baseline Dirichlet process, which we will highlight in the Section (ref). The resulting group partition will receive a strictly higher (lower) probability if it is (in)consistent with pairwise constraints. As a result, the induced prior on the group indices $G$ directly depend on the characteristics of user-specific pairwise constraints and is able to increase or decrease the likelihood of a certain $G$.
In the presence of soft constraints, we modify the EPPF by multiplying a function of characteristics of constraints,
where $\psi_{i j}/ (1 - \psi_{i j} )$ is the prior odds for the constraint between unit $i$ and $j$, $\delta_{i j}(G)$ is a transformed Kronecker delta function such that
and $c$ is a positive number that controls the overall strength of prior belief. For $c \rightarrow 0$, $p(G | \psi, T)$ corresponds to the baseline EPPF $p(G)$, while for $c \rightarrow \infty$, $p\left(G = G^* | \psi, T \right) \rightarrow 1$, where $G^*$ satisfies all pairwise constraints.
With the definition of $W_{ij}$, we rewrite the partition probability function defined in ((ref)) in terms of $W_{ij}$ to ease notation,
where
is a normalization constant and we will use the prior $p(G | W)$ hereinafter. In practice, we will first specify $(T_{ij}, \psi_{ij}) = (\text{type, accuracy})$ for the constraint between unit $i$ and $j$ and then construct the corresponding weight $W_{i j}$ via the equation ((ref)).
The function $\pi(\psi, T | G)$ in Equation ((ref)) is crucial in shifting the prior probability of $G$. By design, $T_{i j}\delta_{ij} = 1$ when the constraint between $i$ and $j$ is met in a group partitioning defined by $G$. The prior probability for $G$ is therefore increased since $\left[\psi_{i j}/ (1 - \psi_{i j} )\right]^c > 1$. Similarly, if a group partitioning $G$ violates the constraint between $i$ and $j$, then $T_{i j}\delta_{ij} = -1$ and the prior probability for $G$ drops due to $\left[\psi_{i j}/ (1 - \psi_{i j} )\right]^{-c} < 1$. Therefore, with $\pi(\psi, T | G)$, the resulting group partition is shrunk toward our prior knowledge without imposing any constraint.
To fix ideas, consider a simplified scenario with $N = 2$ units where there are at most two groups. For illustrative purposes, we set the concentration parameter to $a = 1$ so that $\Pr (g_1 = g_2) = \Pr (g_1 \ne g_2) = 0.5$\footnote{antoniak1974 provides analytical formulas for probabilities of more general events with larger $N$. In this example, $\Pr (g_1 = g_2)$ = $\frac{1}{a+1}$.} if no constraint exists. When $N = 2$, listing all partitions $G$ is possible, $G \in \{(1,1), (1,2), (2,1), (2,2)\}$, and we can calculate the probabilities for each $G$ using ((ref)). As a result, we are able to derive the probability of units 1 and 2 belonging to the same or different groups, i.e., analytical formulae for $\Pr(g_1 = g_2)$ and $\Pr(g_1 \ne g_2)$, which neatly demonstrate the effect of $\psi$ and $c$ on group partitioning.
It is straightforward to show $\Pr(g_1 = g_2)$ as a function of $\psi_{12}, T_{12}$, and $c$:
Figure (ref) traces out the equation ((ref)) for a range of $c$ values. The left panel (a) displays the curve for a PL constraint. Firstly, observe that when $c = 0$, $\Pr(g_1 = g_2)$ remains unchanged at $0.5$ regardless of the value of $\psi$. This is the situation in which $c$ eliminates the constraint's effect on the prior. Next, given a particular $c$, $\Pr(g_1 = g_2)$ increases in $\psi$, which means that a stronger soft PL constraint between units 1 and 2 leads to higher chance of assigning both units to the same group. When $\psi$ is fixed, increasing $c$ easily results in a higher $\Pr(g_1 = g_2)$, indicating that a larger $c$ value magnifies the effect of the PL constraint. In contrast, panel (b) depicts the curve with a NL constraint. $\psi_{12}$ and $c$ clearly have the opposite effect on $Pr (g_1 = g_2)$: $\Pr(g_1 = g_2)$ drops significantly as $\psi_{12}$ or $c$ increases. Notably, even with a large $\psi_{12}$, the soft constraint framework maintains the possibility of breaching the constraint, which is another important feature that preserves the chance of correctly assigning group indices even if the constraint is erroneous.
In the general case where numerous PL and NL constraints are enforced, $c$ concurrently affects all constraints. In other words, the value of $c$ determines the overall “strength" of the prior belief of $G$. If the prior belief is coherent with the real group partition, it would be preferable to have a large $c$ to intensify the effect on constraints, allowing prior information to take precedence over data information, and vice versa.
In reality, it is practical to establish soft pairwise constraints based on existing information on group, even if it is not the genuine group partitioning. In the empirical analysis, for instance, we use the official expenditure categories of CPI sub-indices to construct soft pairwise constraints. When information on group partitioning is insufficient, especially when the number of units is large, these official expenditure categories may serve as a trustworthy starting point. Before formalizing the idea, we first introduce the prior similarity matrix $\underline{\pi}^S$ which is a $N \times N$ symmetric matrix describing the prior probability of any two units belonging to the same group, i.e., $\underline{\pi}^S_{ij} = \Pr(g_i = g_j)$ conditional on all hyperparameters in the prior.
The general idea is to derive soft pairwise constraints using the existing information on a preliminary group partitioning $\underline{G}$, which is allowed to involve only a subset of units. We start with the type of constraints $T_{ij}$ between any two units. Given the preliminary group structure, such as expenditure categories, we specify PL constraints for all pairs of units within the same group and NL constraints for all pairs of units from different groups. This means that we believe the preliminary group structure is correct a priori. Despite the fact that more elaborate and subtle constraints might be implemented, this rough specification is usually a great starting point.
The accuracy $\psi_{ij}$ for constraints is then specified. When our prior knowledge is limited or the number of units is large, we cannot specify $\psi_{ij}$ for all pairs with solid knowledge of them. Instead, one desirable yet simple choice is to assume $\psi_{ij}$ again based on preliminary group partitioning $\underline{G}$. More specifically, all units in the same group are positive-linked with identical $\psi^{PL}_{ij} $, i.e., for units $i$ and $j$ from the group $\underline{g}_i = \underline{g}_j = \underline{g}$, we have $\psi^{PL}_{ij} = c_{\underline{g}}$. Units from different groups are assumed to be negative-linked with identical $\psi^{NL}_{ij}$, i.e., for units $i$ and $j$ from distinct groups, we assume $\psi^{NL}_{ij} = c_{\underline{g}_i\underline{g}_j}$ and $c_{\underline{g}_i\underline{g}_j} = c_{\underline{g}_j\underline{g}_i}$. Following this strategy, $\psi_{ij}$ depends solely on $\underline{G}$ and hence two units from the same group would have identical soft pairwise constraints with other units. Notice that the number of possible distinct $\psi_{ij}$ reduces from $N(N-1)/2$ to $\widebar{K}(\widebar{K}+1)/2$, where $\widebar{K}$ is the number of groups in $\underline{G}$ and $\widebar{K} \ll N$.
This framework permits no prior belief in certain units. If at least one unit in a pair $(i,j)$ is not included in $\underline{G}$, we assume that this pair of units is free of constraints and we set $T_{ij}$ to 0 or $\psi_{ij}$ to 0.5 in the prior. Note that the absence of a constraint does not ensure that the units $i$ and $j$ are completely unrelated. Instead, if both $i$ and $j$ are involved in constraints with a third unit $k$, or are connected through a series of $l$ constraints, $i \leftrightarrow k_1 \leftrightarrow \cdots \leftrightarrow k_l \leftrightarrow j$, then the prior probability of $i$ and $j$ belonging to the same group differs from the prior probability without any constraints. If we wish to prevent two units from linking a priori, they must not be subject to any constraints with the remaining units.
The aforementioned specification strategy induces a block prior similarity matrix, i.e., for an unit $i$, $\underline{\pi}^S_{ij} = \underline{\pi}^S_{ik}$ if $\underline{g}_j = \underline{g}_k$. Intuitively, if two units have identical soft pairwise constraints and hence posit an identical relationship with all other units, they are equivalent and exchangeable. As a result, these units should have an equal prior probability of sharing the same group index with any other units. More formally,
Theorem (ref) echos the concept of stochastic equivalence nowicki2001 in stochastic block model\footnote{For a more comprehensive review of the stochastic block model, see lee2019.} (SBM) holland1983. In less technical terms, for nodes $p$ and $q$ in the same group, $p$ has the same (and independent) probability of connecting with node $r$, as $q$ does. Interestingly, this relationship is not coincidental. The prior draw of group membership with the aforementioned specification of $T_{ij}$ and $\psi_{ij}$ can be viewed as a simulation of a simple SBM. In a simple SBM, there are two essential components: a vector of group memberships and a block matrix, each element of which represents the edge probability of two nodes, given their group memberships. In our case, the preliminary group structure serves as the group membership in SBM. The DP prior and the weight (or $T_{ij}$ and $\psi_{ij}$) of each constraint induce a prior similarity probability comparable to the block matrix.
Let's consider an example of 90 units. The preliminary group structure divides units into 3 groups, with groups 1, 2 and 3 containing 25, 30 and 35 units, respectively. Figure (ref) shows the prior similarity matrix, which is based on the aforementioned specification strategy, so it becomes a block matrix with equal entries in each of nine blocks. Units within the same group are stochastically equivalent, as their prior probabilities of being grouped not only with each other but also with units from other groups are the same. As a result, the similarity probability of each pair depends solely on their preliminary membership (and $\psi$).
bonhomme2015 incorporates prior knowledge of group membership by adding a penalty term to the objective function. They assume that prior information is in the form of probabilities which describe the prior probability of unit $i$ belonging to group $k$ with at most $K$ groups as $\omega^{(K)}_{ik}$. Consequently, the estimated group index is given by:
where $C>0$ is a hyperparameter need to be tuned further and $K$ is the predetermined number of groups.
The penalty determines the weights assigned to prior and data information in estimation. Due to the fact that $K$ is frequently unknown in advance, this method requires model selection to determine the ideal number of groups. Assume we have $n_K$ alternative options for $K$. We have a $N \times K$ matrix $\omega^{(K)} = \left \{\omega^{(K)}_{ik} \right\}$ for prior information for a given $K$. As a result, in order to pick a model, we must therefore provide $n_K$ sets of prior probability matrix $\omega^{(K)}$, which is cumbersome and inconvenient. For instance, if $K$ has values ranging from $3,5,10$ and $N = 200$, there are 3,600 entries for $\omega$, none of which can be missing or undefined. In addition, the information criteria may be unreliable in the finite-sample results, necessitating further care when selecting an appropriate variant for empirical application.
Summarizing prior knowledge through pairwise constraints is often more practical than the penalty function approach in BM and solves the aforementioned practical issues. Pairwise relationships can be derived intuitively from researchers' input without requiring in-depth knowledge of the underlying groups; researchers do not need to fix the number of groups $K$ or group membership a priori. One only needs to focus on a pair of units each time and specify the preference of assigning them to the same or different groups. Moreover, since the pairwise constraints are incorporated into the DP prior, which implicitly defines a prior distribution on the group partitions, model selection is not required and the posterior analysis also takes the uncertainty of the latent group structure into account.
paganin2021 offer a statistical framework for including concrete prior knowledge on the partition. Their proposed method aims to shrink the prior distribution towards a complete prior clustering structure, which is an initial clustering involving all units provided by experts. Specifically, they suggest a prior on group partition that is proportional to a baseline EPPF of the DP prior multiplied by a penalization term,
with $m>0$ is a penalization parameter, $d\left(G, G_{0}\right)$ is a suitable distance measuring how far $G$ is from $G_{0}$, and $p(G)$ indicates the same EPPF as in ((ref)). Because of the penalization term, the resulting group indices $G$ shrink toward the initial target $G_0$.
This framework is parsimonious and easy to implement, but it comes with a cost. The method is incapable of coping with an initial clustering in a subset of the units under study or multiple plausible prior partitions; otherwise, the distance measure is not well-defined. In addition, the authors suggest utilizing Variation of Information meilua2007 as the distance measure. It can be shown that the resulting partition can easily become trapped in local modes, leading the partition to never shrink toward $G_{0}$. They also argue that other available distance methods have flaws. As a result, the penalty term does not function as anticipated.
Our proposed framework with pairwise constraints is more flexible than adopting a target partition. Actually, the target partition in paganin2021 can be viewed as a special case of the pairwise constants, in which every unit must be involved in at least one PL constraint. Our framework could manage partitions involving arbitrary subsets of the units by tactically specifying the bilateral relationships. Most importantly, when the group indices of other pairs are fixed, our framework ensures that the partition containing a specific pair that is consistent with our prior belief receives a strictly greater prior probability than the partition that is inconsistent with our prior belief. This guarantees that the generated $G$ shrinks in the direction of our prior belief.
The baseline model contains the parameters listed below: $(\alpha, \sigma^2, \pi, \xi, a, \phi)$. We rely mostly on nonparametric Bayesian models.\footnote{For a more comprehensive review of the nonparametric Bayesian literature, see ghosal2017 and mueller2018.} Bayesian nonparametric models have emerged as rigorous and principled paradigms to bypass the model selection problem in parametric models by introducing a nonparametric prior distribution on the unknown parameters. The prior assumes that a collection of $\alpha$ and $\sigma^2$ is drawn from the Dirichlet process prior.\footnote{There have been some empirical works that use Dirichlet process model with panel data. Dirichlet process mixture prior is specified for either the distribution of the innovations hirano2002 or intercepts fisher2022.} $\pi$ is a vector of mixture probabilities in Dirichlet process that is produced by the stick-breaking approach with stick length $\xi$. $a$ is the concentration parameter in the Dirichlet process, whereas $\phi$ is a collection of hyperparameters in the base measure $B_{0}$. We consider prior distributions in the partially separable form,\footnote{The joint prior includes $\xi$ but not $\pi$. Because the stick-breaking formulation of $\xi$ is a deterministic transformation of $\xi$, knowing $\xi$ is identical to knowing $\pi$.}
We tentatively focus on the random coefficients model where, conditional on $G$, $\alpha_{g_{i}}$ and $\sigma^2_{g_i}$ are independent to the conditional set that includes initial value of each unit $y_{i0}$, the initial values of predetermined variables, and the whole history of exogenous variables. The assumption guarantees that $\alpha_{g_{i}}$ and $\sigma_{g_i}$ can be sampled separately and simplifies the inference of the underlying distribution of $\alpha_{g_{i}}$ and $\sigma_{g_i}$ to an unconditional density estimation problem, therefore lowering computational complexity. The joint distribution of heterogeneous parameters as a function of the conditioning variables can then be modeled to extend the model to the correlated random coefficient model, which is briefed in Section (ref). A full explanation and derivation for the correlated random coefficient model are provided in the online appendix.
In the nonparametric Bayesian literature, the Dirichlet Process (DP) prior ferguson1973, ferguson1974, sethuraman1994 is a canonical choice, notable for its capacity to construct group structure and accommodate an infinite number of possible group components.\footnote{See Appendix (ref) for a brief overview of the DP and Appendix (ref) for its clustering properties.} The DP mixture is also known as a “infinite" mixture model due to the fact that the data indicate a finite number of components, but fresh data can uncover previously undetected components neal2000. When the model is estimated, it chooses automatically an appropriate subset of groups to characterize any finite data set. Therefore, there is no need to determine the “proper" number of groups.
The DP prior can be written as an infinite mixture of point mass with the probability mass function:
where $\delta_{x}$ denotes the Dirac-delta function concentrated at $x$ and $B_{0}$ is the base distribution. We adopt an Independent Normal Inverse-Gamma (INIG) distribution for the base distribution $B_0$:
with a set of hyperparameters $\phi = \left( \mu_{\alpha}, \Sigma_{\alpha}, \frac{\nu_{\sigma}}{2}, \frac{\delta_{\sigma}}{2} \right)$.
The group probabilities $\pi_k$ are constructed by an infinite-dimensional stick-breaking process sethuraman1994 governed by the concentration parameter $a$,
where stick lengths $\xi_k$ are independent random variables drawn from the beta distribution,\footnote{Recall that a $Bata(m, n)$ distribution is supported on the unit interval and has mean $m/(m+n)$.} $Beta(1, a)$. The group probability will be random but still satisfy $\sum_{k=1}^{\infty} \pi_{k}=1$ almost surely.
Equation ((ref)) is essential to understanding how the DP prior controls the number of groups. The building of group probabilities is compared to the breaking of a stick of unit length sequentially, in which the length of each break is assigned to the current value of $\pi_k$. As the number of groups increases, the probability created by the stochastic process decreases because the remaining stick becomes shorter with each break. In practice, the number of groups does not increase as fast as $N$ due to the characteristic of the stick-breaking process that leads the group probability to soon approach zero.
Although in principle we do not restrict the maximum number of groups and allow the number to rise as $N$ increases, a finite number of instances will only occupy a finite number of $K$ components. The concentration parameter $a$ in the prior of $\xi_k$ determines the degree of discretization -- the complexity of the mixture and, consequently, $K$, as also revealed in ((ref)). As $a \to 0$, the realizations are all concentrated at a single value, however as $a \to \infty$, the realizations become continuous-valued as its based distribution. Specifically, antoniak1974 derives the relationship between $a$ and the number of unique groups,
that is, the expected number of unique groups is increasing in both $a$ and the number of units $N$.
escobar1995 highlights the importance of specifying $a$ when imposing prior smoothness on an unknown density and demonstrates that the number of estimated groups under a DP prior is sensitive to $a$. This suggests that a data-driven estimate of $a$ is more reasonable. Moreover, ascolani2022 emphasizes the importance of introducing a prior for $a$ as it is crucial for learning the true number of groups as $N$ increases and hence establishing the posterior consistency. We define a gamma hyperprior for $a$ and update it based on the observed data in order to alter the smoothness level. This step generates a posterior estimate of $a$, which indirectly determines the number of groups $K$ without reestimating the models with different group sizes. Essentially, this represents “automated" model selection.
Collectively, we specify a DP prior for $\left(\alpha_{i}, \sigma^2_{i}\right)$. The DP prior is a mixture of an infinite number of possible point masses, which can be constructed through the stick-breaking process. The discretization of the underlying distribution is governed by the concentration parameter $a$. With a hyerprior on $a$, we permit the data to determine the number of groups $K$ present in the data, which can expand unboundedly along with the data.
In a formal Bayesian formulation, a prior distribution is specified to partition $\mathcal{B}$ with associated indices $G$. Despite the fact that DP prior does not specify this prior distribution explicitly, we can characterize it using the exchangeable partition probability function (EPPF) pitman1995. As we briefly mentioned in the last subsection, the EPPF plays a significant role in connecting the prior belief on group structure to the DP prior, which is included as part of our proposed prior distribution in Equation ((ref)).
The EPPF characterizes the distribution of a partition $\mathcal{B} = \left\{B_{1}, B_{2}, \ldots, B_{K}\right\}$ induced by $G$. As the generic Dirichlet process assumes units are exchangeable, any permutation has no effect on the joint probability distribution of $G$; hence, the EPPF is determined entirely by the number of groups and the size of each group. pitman1995 demonstrates that the EPPF of the Dirichlet process has the closed form,
where $a$ is the concentration parameter and $\Gamma(x) = (x-1)!$ denotes the Gamma function. Note that the partition $\mathcal{B}$ is conceived as a random object and hence the group number $K$ is not predetermined, but rather is a function of $G$, $K = K(G)$.
sethuraman1994 and pitman1996 constructively show that group indices/partitions can be drawn from the EPPF for DP using the stick-breaking process defined in ((ref)). As a result, the EPPF does not explicitly appear in the posterior analysis in the current setting so long as the priors for the stick lengths are included.
As suggested by chamberlain1980, allowing the individual effects to be correlated with the initial condition can eliminate the omitted variable bias. This subsection presents the first attempt to introduce dependence between grouped effects and covariates, under the presence of group structure in both heterogeneous slope coefficients and cross-sectional variance. The underlying theorems, such as posterior consistency, and the performance of the framework are left for future study.
We primarily follow the proposed framework in liu2022 and utilize Mixtures of Gaussian Linear Regressions (MGLRx) for the group-specific parameters. MGLRx prior is discussed in pati2013 and can be viewed as a Dirichlet Process Mixture (DPM) prior that takes the dependence of covariates into account. Notice that the correlated random coefficients model requires a DPM-type prior for $\alpha_i$ and $\sigma_i^2$. This is because $\alpha_i$ and $\sigma_i^2$ are assumed to be correlated with covariates of each individual, and as such, they are not identical within a group.
Following liu2022, we first transform $\sigma_i^2$ and define $l_i=\log \frac{\bar{\sigma}^2\left(\sigma_i^2-\underline{\sigma}^2\right)}{\bar{\sigma}^2-\sigma_i^2}$, where $\underline{\sigma}^2\left(\bar{\sigma}^2\right)$ is some small (large) positive number. This transformation simplifies the prior for $\sigma_i^2$, which is now dependent on covariates, and ensures that a similar prior structures can be applied to both $\alpha_i$ and $l_i$.
In the correlated random coefficients model, the DPM prior for $\alpha_{i}$ or $\sigma_i^2$ is an infinite mixture of Normal densities with the probability density function:
where $w_{i0} = [1, y_{i0}, x_{i,0:T}]'$ is the conditioning set at period 0, which includes initial value of each unit $y_{i0}$, the initial values of predetermined variables, and the whole history of exogenous variables. Notice that $\alpha_{i}$ and $\sigma_i^2$ share the same set of group probabilities $\pi_k (w_{i0})$.
Similar but not identical to the DP prior, it is the component parameters $\left(\mu^{\alpha}_{k}, \Omega^{\alpha}_{k} \right)$ or $\left(\mu^{\sigma}_{k}, \Omega^{\sigma}_{k} \right)$ that are directly drawn from the base distribution $G_0$. $G_0$ is assumed to be a conjugate Matricvariate-Normal-Inverse-Wishart distribution.
The group probabilities are now characterized by a probit stick-breaking process rodriguez2011,
where the stochastic function $\zeta_{k}$ is drawn from a Gaussian process, $\zeta_{k} \sim G P\left(0, V_{k}\right)$ for $k=1,2, \cdots$. The Gaussian process is assumed to have zero mean and the covariance function $V_{k}$. defined as follows,
where $\tau_{v} \sim IG(\frac{\nu_{v}}{2}, \frac{\nu_{v}}{2})$ and $A_k$ has its own hyperprior, see details in pati2013.
This section describes the procedure for analyzing posterior distributions for the baseline model described in ((ref)) with the priors specified in Section (ref). The joint posterior distribution of model parameters is
where $ p(Y | X, \alpha, \sigma^2, G)$ is the likelihood function given by equation ((ref)) for an i.i.d. model conditional on group indices $G$, and $p(W | G)$ is the additional term of pairwise constraints with the form $p(W | G) = \prod_{i=1}^{N} \prod_{j=1}^{N}\exp\left(c W_{ij} \delta_{ij}\right)$.
Draws from the joint posterior distribution can be obtained by using blocked Gibbs sampling. The algorithm is derived from ishwaran2001 and walker2007. Due to the use of a finite-dimensional prior and truncation, the method described in ishwaran2001 cannot truly address our demand for estimating the number of groups without a predetermined value or upper bound. We employ the slice sampler walker2007, which is the exact block Gibbs sampler for the posterior computation in infinite-dimensional Dirichlet process models, modifying the block Gibbs sampler of ishwaran2001 to avoid truncation approximations. walker2007 augments the posterior distribution with a set of auxiliary variables consisting of i.i.d. standard uniform random variables, i.e., $u_i \stackrel{iid}{\sim} U(0,1)$ for $i = 1,2,..,N$. The augmented posterior is then represented as
where $\prod_{i} \mathbf{1}(u_i \le \pi_{g_i})$ is substituted for $p(G | \Xi)$ in the equation ((ref)).
To roughly see how slice sampling works, recall that the group probabilities are constructed in a sequential manner, following a stick-breaking procedure. The leftover of the stick after each break gets smaller and smaller. Given the finite number of units, we can always find the smallest $K^*$ such that for all groups $k \ge K^*$, the minimum of $u_i$ among all units is larger than $\pi_k$, which is bounded above by the length of the leftover after $k$ breaks, $1-\sum_{j = 1}^{k} \pi_j$. More formally,
As a result, all units receive strictly zero probability of being assigned to any group $k = K^* + 1, K^* + 2,...,N$ since the indicator function $\mathbf{1}(u_i \le \pi_{k})$ is zero.
There are two advantages to incorporating the auxiliary variable $u$ into the model. First and foremost, $u$ directly determines the largest possible number of groups in each sampling iteration. This reduces the support of $G$ and $\boldsymbol{\Xi}$ to a finite space, enabling us to solve a problem of finite dimensions without truncation. Furthermore, $u$ have no effect on the joint posterior of other parameters because the original posterior can be restored by integrating out $u_i$ for $i = 1,2,...,N$.
The Gibbs sampler are used to simulate the joint posterior distribution of $\left(\alpha, \sigma^2, \Xi, u, a, G\right)$. We break this vector into blocks and sequentially sampling for each block conditional on the current draws for the other parameters and the data. The full conditional distributions for each block are easily derived using the conjugate priors specified in Section (ref).
For the group-specific parameters, we directly draw samples from their posterior densities as we adapt conjugate priors. The posterior inference with respect to $(\alpha, \sigma^2)$ becomes standard once we condition on the latent group indices $G$. It is essentially a Bayesian panel data regression for each group. The conditional posterior for the stick length $\Xi$ is a beta distribution given $G$, and hence direct sampling is possible.
We follow walker2007 to derive the posterior of auxiliary variable $u$. As $u$ are standard uniformly distributed, the posterior is a uniform distribution defined on $\left(0, \pi_{g_i}\right)$, conditional on the group probabilities and group indices. In terms of the concentration parameter $a$, we use a 2-step procedure proposed by escobar1995. Following their approach, we first draw a latent variable $\eta$ from $Beta(a+1, N)$. Then, given $\eta$ and number of groups $K^a$ in the current iteration, we directly draw $a$ from a mixture of two Gamma distribution.
It is worth noting that the steps for implementing the DP prior with or without soft pairwise constraints are the same for all parameter besides the group indices $G$. This is due to the fact that soft pairwise constraints only affect other parameters through the group indices. It is handy to sample group indices with soft pairwise constraints conditional on other parameters. The posterior probability of assigning unit $i$ to group $k$ includes additional term $p(W_i | G) = \prod_{j \ne i, g_j = k} \exp \left( 2 c W_{i j} \delta_{i j}\right) $ to rewards (penalizes) the abidance (violation) of constraints,
where $K^*$ is the maximal number of groups after generating potential new group-specific slope coefficients and variance. We then draw the group index for unit $i$ from a multinomial distribution:
Algorithm (ref) below summarizes the algorithm for the proposed Gibbs sampling. For illustrative purposes, we focus primarily on the posterior densities of major parameters and omit details on step ((ref)). In short, step ((ref)) creates potential groups by sampling new $(\alpha_k, \sigma_k^2)$ from the prior if the latest $K^*$ based on newly drawn $u_i$ and $\pi_k$ is larger than previous $K^*$, which indicate the current iteration permits more groups. This Detailed derivations and explanation of each step are provided in Appendix (ref).
In contrast to popular algorithms such as agglomerative hierarchical clustering or the KMeans algorithm, which return a single clustering solution, Bayesian nonparametric models provide a posterior over the entire space of partitions, enabling the assessment of statistical properties, such as the uncertainty on the number of groups.
However, when the group structure is part of the major conclusion of an empirical analysis, the point estimate of group structure becomes crucial. wade2018 discuss in detail an appropriate point estimate of the group partitioning based on the posterior draws. From the decision theory, the point estimate $G^{*}$ minimizes the posterior expected loss,
where the loss function $ L(G, \widehat{G})$ is the variation of information by meilua2007, which measures the amount of information lost and gained in changing from partition $G$ to $\widehat{G}$.\footnote{Another possible loss function is the 0-1 loss function $L(G, \widehat{G}) = \mathbf{1} (G = \widehat{G})$, which leads to $G^*$ being the posterior mode. This loss function is undesirable since it ignores similarity between two partitions. For instance, a partition that deviates from the truth in the allocation of only one unit is penalized the same as a partition that deviates from the truth in the allocation of numerous units. Furthermore, it is generally recognized that the mode may not accurately reflect the distribution's center. } The Variation of Information is based on the Shannon entropy $H(\cdot)$, and can be computed as
where $\log$ denotes $\log$ base 2, $\lambda_j=\left|B_j\right|$ is the cardinality of the group $j$, and $\lambda_{j l}^{\wedge}$ the size of blocks of the intersection $G \wedge \widehat{G}$ and hence the number of indices in block $j$ under partition $G$ and block $l$ under $\widehat{G}$.
wade2018 show that the optimal group partitioning can be identified based on the posterior similarity matrix,
where $P \left(g_{j}=g_{i} | Y,X,W\right)$ is the $(i,j)$ entry of the posterior similarity matrix. We refer to wade2018 for additional properties and empirical evaluations.
The procedure of Gibbs sampling with soft constraints in Algorithm (ref) is closely related to constrained clustering in the computer science literature. In this parallel literature, constrained clustering refers to the process of introducing prior knowledge to guide a clustering algorithm. For a subset of the data, the prior knowledge takes the form of constraints that supplement the information derived from the data via a distance metric. As we shall see below, under several simplifying assumptions, our framework could be reduced to a deterministic method for estimating group heterogeneity using a constrained KMeans algorithm. Though this deterministic method may address the practical issues in BM, it only works for certain restricted models and hence is not as general as our proposed framework.
We start with a brief review of the Pairwise Constrained KMeans (PC-KMeans) clustering algorithm by basu2004, which is a well-known clustering algorithm in the field of semi-supervised machine learning. It's a pairwise constrained variant of the standard KMeans algorithm in which an augmented objective function is used in the assignment step. Given a collection of observations $\left(y_{1}, y_{2}, \ldots, y_{N}\right)$, a set of positive-link constraints $\mathcal{P}$, a set of negative-link constraints $\mathcal{N}$, the cost of violating constraints $w = \{w^p_{i j}, w^n_{i j} \}$ and the number of groups $K$, the PC-KMeans algorithm divides $N$ observations into $K$ groups (the assignment step) so as to minimize the following objective function,
where $\mu_{k}$ is the centroid of group $k$, i.e., $\mu_{k}=\frac{1}{\left|B_k\right|} \sum_{i \in B_k} y_i$, $B_k$ is the set of units assigned to group $k$, and $\left|B_k\right|$ is the size of group $k$. The first part is the objective function for the conventional KMeans algorithm, while the second part accounts for the incurred cost of violating either PL constraints ($w^p_{i j}$) or NL constraints ($w^n_{i j}$).
Similar to KMeans, PC-KMeans alternates between reassigning units to groups and recomputing the means. In the assignment step, it determines a disjoint $K$ partitioning that minimizes ((ref)). Then the update step of the algorithm recalculates centroids of observations assigned to each cluster and updates $\mu_{k}$ for all $k$.
By applying asymptotics to the variance of distributions within the model, we demonstrate linkages between the posterior sampler of our constrained BGFE estimator and KMeans-type algorithms in Theorem (ref). We investigate small-variance asymptotics for posterior densities, motivated by the asymptotic connection between the Gibbs sampling algorithm for the Dirichlet process mixture model and KMeans kulis2011, and demonstrate that the Gibbs sampling algorithm for the CBG estimator with soft constraints encompasses the constrained clustering algorithm PC-KMean in the limit.
We return to the world of grouped fixed-effects models. In fact, the clustering algorithm is essential for BM and bonhomme2022, who use the KMeans algorithm to reveal the group pattern in the fixed-effects. With the theorem described above, it motivates a constrained version of BM's GFE estimator. We show that it is straightforward to incorporate prior knowledge in the form of soft paired restrictions into the GFE estimator. The soft pairwise constrained grouped fixed-effects (SPC-GFE) estimator is defined as the solution to the following minimization problem given the number of groups $K$: {
}where the minimum is taken over all possible partitions $G$ of the $N$ units into $K$ groups, common parameters $\theta$, and group-specific time effects $\alpha$. $w^p_{i j}$ and $w^n_{i j}$ are the user-specified costs on PL and NL constraints.
For given values of $\theta$ and $\alpha$, the optimal group assignment for each individual unit is {
}where we essentially apply the PC-KMeans algorithm to get the group partition. The SPC-GFE estimator of $(\theta, \alpha)$ in ((ref)) can then be written as
where $\widehat{g}_{i} = \widehat{g}_{i}(\theta, \alpha)$ is given by ((ref)). $\theta$ and $\alpha$ are computed using an OLS regression that controls for interactions of group indices and time dummies. The SPC-GFE estimate of $g_{i}$ is then simply $\widehat{g}_{i}(\widehat{\theta}, \widehat{\alpha})$.
We apply our panel forecasting methods to the following two empirical applications: inflation of the U.S. CPI sub-indices and the income effect on democracy. The first application focuses mostly on predictive performance, whereas the second application focuses primarily on parameter estimation and group structure.
To accommodate richer assumptions on model, we use variants of the baseline model in Equation ((ref)) in this section, either by adding common regressors or allowing for time-variation in the intercept. We use the conjugate prior for all parameters, see details in Appendix (ref).
In both applications, the prior knowledge of the latent group structure or the pre-grouping structure covers all units. In the first application, CPI sub-indices can be clustered by expenditure category, whereas countries in the second application may be grouped according to their geographic regions. We build positive-link and negative-link constraints given the prior knowledge: all units within the same group are presumed to be positive-linked, while units from different categories are believed to be negative-linked. In terms of the accuracy of constraints, $\psi^{PL}_{ij}$ and $\psi^{NL}_{ij}$ are equal for all constraints with the same type, following the strategy described in Section (ref). In the applications below, we fix $\psi^{PL}_{ij} = 0.65$ and $\psi^{NL}_{ij} = 0.55$ to reflect the belief that PL constraints (attracting forces) play slightly more important role than the NL constraints (repelling forces) and NL constraints cannot be ignored. Finally, we construct weights $W_{ij}$ using ((ref)). Notice that these assumptions on prior and hyperparameters are an example to showcase how the proposed framework works with real data. In practice, we may specify constraints for a subset of units with different levels of weights, either in a data-driven manner (for instance, highly correlated units may fall into the same group with a high level of confidence) or in a model-based manner (i.e., $W$ is a function of covariates).
Given that the dimension of the space of group partitions increases exponentially with the number of units $N$, attention must be given while selecting $c$ across analyses with different $N$. As suggested by paganin2021, calibrating the modified prior is computationally intensive. We are facing a trade-off between investing time to get the prior “exactly right" and letting the constant $c$ be an estimated model parameter. As such, we propose to find the optimal $c$ that maximizes marginal data density using grid search.
In the Monte Carlo simulation, the value of $c$ is fixed for simplicity, but in the empirical applications, $c$ is determined by marginal data density (MDD). We calculate MDD using the harmonic mean estimator newton1994, which defined as
given a sample $\theta^{(j)}$ from the posterior $p(\theta | y)$. The simplicity of the harmonic mean estimator is its main advantage over other more specialized techniques. It uses only within-model posterior samples and likelihood evaluations, which are often available anyway as part of posterior sampling. We finally choose the optimal value for $c$ that maximizes MDD.
We consider six estimators in the section. The first three estimators are our proposed Bayesian grouped fixed-effects (BGFE) estimator with different assumptions on cross-sectional variance and pairwise constraints. The last three estimator ignore the group structure.
In the first application, we focus on inflation forecasting. For the most recent advances in this topic, faust2013 provide a comprehensive overview of a large set of traditional and recently developed forecasting methods. Among many candidate methods, we choose the AR model as the benchmark and exclusively include it as an alternative estimator in this exercise. This is because the AR is relatively hard to beat and, notably, other popular methods, such as the Atkeson–Ohanian version random walk model atkeson2001, UCSV stock2007, and TVP-VAR primiceri2005, generally do as reasonably well as the AR model, according to faust2013.
We generate one-step ahead forecasts of $y_{i, T+1}$ for $i = 1,...,N$ conditional on the history of observations
and newly available variables $x_{i T+1}$ at $T+1$.
The posterior predictive distribution for unit $i$ is given by
where $\Theta$ is a vector of parameters $\Theta = \left( \alpha_{g_i}, \sigma^2_{g_i}, g_i \right)$. This density is the posterior expectation of the following function:
which is invariant to relabeling the components of the mixture and $K(G)$ is the number of groups in $G$. Given $S$ posterior draws, the posterior predictive distribution estimated from the MCMC draws is
where
We can therefore draw samples from $\hat{p}(y_{i T+1} | Y,X)$ by simulating ((ref)) forward conditional on the posterior draws of $\Theta $ and observations. Note that MCMC exhibits the true Bayesian predictive distribution, implicitly integrating over the entire underlying parameter space.
We evaluate the point forecasts via the real-time recursive out-of-sample Root Mean Squared Forecast Error (RMSFE) under the quadratic compound loss function averaged across units. Let $\hat{y}_{i T+1 | T} $ represent the predicted value conditional on the observed data up to period $T$, the loss function is written as
where $y_{i, T+1}$ is the realization at $T+1$ and $\hat{\varepsilon}_{i T+1 | T}$ denote the forecast error.
The optimal posterior forecast under quadratic loss function is obtain by minimizing the posterior risk,
This implies optimal posterior forecast is the posterior mean,
Conditional on posterior draws of parameters, the mean forecast can be approximated by the Monte Carlo averaging,
Finally, the RMSFE across units is given by
To compare the performance of density forecasts for various estimators, we report the average log predictive scores (LPS) to assess the performance of the density forecast from the view of the probability distribution function. As suggested in geweke2010, the LPS for a panel reads as,
where the expectation can be approximated using posterior draws:
The following results are also robust to other metrics such as the continuous ranked probability score matheson1976, hersbach2000.
Policymakers and market participants are very interested in the abilities to reliably predict the future disaggregated inflation rate. Central banks predict future inflation trends to justify interest rate decisions, control and maintain inflation around their targets. The Federal Reserve Board forecasts disaggregated price categories for short-term inflation forecasting bernanke2007. They rely primarily on the bottom-up approach that focuses on estimating and forecasting price behavior for the various categories of goods and services that make up the aggregate price index. Moreover, investors in fixed-income markets in the private sector wish to forecast future sectoral inflation in order to anticipate future trends in discounted real returns. Some private firms also need to predict specific inflation components in order to forecast price dynamics and reduce risks accordingly.
In this section, we demonstrate the use of constrained BGFE estimators with prior knowledge on the group pattern to forecast inflation rates for the sub-indices of U.S. Consumer Expenditure Index (CPI). We focus primarily on the one-step ahead point and density forecast. Due to space constraints, we only report the group partitioning for the most recent month in the main text.
{\bf Model:} We start by exploring the out-of-sample forecast performance of a simple, generic Phillips curve model. It is a panel autoregressive distributed lag (ADL) model with a group pattern in the intercept, coefficients, as well as error variance. The model is given by
where $y_{it}$ is year-over-year inflation rate, i.e., $y_{it} = \log(\text{price}_{it} / \text{price}_{it-12})$, and $u_t$ is the slack measure for the labor market, the unemployment gap. We fix $p$ at 3 because the benchmark AR model would have the best predictive performance.
{\bf Data:} We collect the sub-indices of CPI for all urban consumer (CPI-U) that include food and energy. The raw data is obtained from the U.S. Bureau of Labor Statistics (BLS), which is recorded on a monthly basis from January 1947 to August 2022. The CPI-U is a hierarchical composite index system that partitions all consumer goods and services into a hierarchy of increasingly detailed categories. It consists of eight major expenditure categories (1) Apparel; (2) Education and Communication; (3) Food and Beverages; (4) Housing; (5) Medical Care; (6) Recreation; (7) Transportation; (8) Other Goods and Services. Each category is composed of finer and finer sub-indexes until the most detailed levels or “leaves” are reached. This hierarchical structure can be represented as a tree structure, as shown in Figure (ref). It is important to note that the parent series and its child series may be highly correlated and readily form a group due to the fact that parent series are generated from child series. For instance, the Energy Services is expected to correlated with its child series Utility gas service and Electricity. Due to our focus on group structure, it is vital to eliminate all parent series in order to prevent not just double-counting but also dubious grouping results. More details regarding the data are provided in Appendix (ref).
{\bf Pre-grouping:} The official expenditure categories are used to build PL and NL constraints: all units within the same categories are presumed to be positive-linked, while units from different categories are believed to be negative-linked.
We focus on the CPI sub-indices after January 1990 for two reasons: (1) the number of sub-indices before 1990 was relatively small, diminishing the importance of the group structure; and (2) the consumption has been changed and more expenditure series were introduced in the 1990s as a result of the popularity of electronic products, food care, etc. After the elimination of all parent nodes, the unbalanced panel consists of 156 sub-indices in eight major expenditure categories. We employ rolling estimation windows of 48 months\footnote{The benchmark AR-he model scores the best overall performance with a window size of 48.} and require each estimation sample to be balanced, removing individual series lacking a complete set of observations in a given window. Finally, we generate 329 samples with the first forecast produced for April 1995.
We begin the empirical analysis by comparing the performance of point and density forecasts across 329 samples. Throughout the analysis, the AR-he estimator serves as the benchmark as it essentially assumes individual effects.
In Figure (ref), we present the frequency of each estimator with the lowest RMSFE in the panel (a) and the boxplot\footnote{The boundaries of the whiskers is based on the 1.5 IQR value. All other points outside the boundary of the whiskers are plotted as outliers in red crosses.} of the ratio of RMSFE relative to the AR-he estimator in the panel (b). First, the AR-he and AR-he-PC estimators, which rely only on an individual's own past data, are not competitive in point forecasts and perform considerably worse than the others. This implies that it is highly advantageous to explore cross-sectional information to improve point forecasts. Moreover, the BGRE-he-cstr estimator scores the highest frequency of being the best estimator despite the fact that BGFE-he-cstr, BGFE-he, BGFE-ho, and pooled OLS estimators all utilize cross-sectional information. Examining the box plot, we find that the BGFE-ho and pooled OLS estimators, which overlook heteroskedasticity, can achieve greater performance in some samples, but make poorer forecasts more often than the other estimators. BGFE-he-cstr and BGFE-he estimator, on the other hand, typically outperform the others and the benchmark in terms of median RMSFE and the ability to produce forecasts with the lowest RMSFE without also increasing the risk of generating the least accurate forecasts.
The revealing patterns of the density forecast are significantly distinct from those of the point forecast. Figure (ref) depicts the log predictive score (LPS) for density forecast. The most notable pattern from the panel (a) is that the BGFE-he-cstr estimator, which incorporates prior knowledge, is dominating and outperforms the rest in over 80% of the samples. It emerges as the apparent winner in this case. Furthermore, when generating density forecast, the BGFE-ho and pooled OLS are not as accurate as they are in point forecast: they never have the lowest LPS across samples. This also confirms that the heteroscedasticity\footnote{We provide more results in Section (ref) to explore the importance of heteroskedasticity in density forecast for the inflation.} is a well-known feature of the inflation time series clark2015. In the boxplot, we ignore BGFE-ho and pooled OLS and show the differences in LPS between the respective estimators and the AR-he estimator. As LPS differences represent percentage point differences, BGFE-he-cstr can provide density forecasts that are up to 22% more accurate compared to the benchmark model. Finally, despite the fact that the BGFE-he-cstr and BGFE-he estimators are mainly based on the same algorithm, the use of prior knowledge on group pattern further enhances the performance, resulting in the BGFE-he-cstr estimator having a lower LPS and scoring the best model with the highest frequency.
Next, we assess the value of adding prior information about groups by comparing the performance of the BGFE-he-cstr and BGFE-he estimators exclusively. The solid black line in Figure (ref) represents the ratio of RMSE between BGFE-he-cstr and BGFE-he. The periods during which BGFE-he-cstr, BGFE-he, and all other estimators achieve the lowest RMSE are indicated by pink, blue, and green shaded areas, respectively. Though the BGFE-he-cstr estimator is not always the best across samples, the prior information improves the performance of the Bayesian grouped estimator. The BGFE-he-cstr estimator performs better than the BGFE-he estimator in most samples, with an average improvement of 2%.
Adding prior information on groups substantially improves the accuracy of density forecasts. Figure (ref) shows the comparison between BGFE-he-cstr and BGFE-he in terms of the difference in LPS. We find the prior information valuable as BGFE-he-cstr outperforms BGFE-he in more than 98% of the samples. Clearly, the majority of the figure is covered by a pink background, showing that BGFE-he-cstr is typically the best choice. All of these facts demonstrate that adding prior informatio is favorable and essential, especially for density forecasting.
Having specified pairwise constraints across sub-indices, we provide a prior on $G$ that shrinks the group structure toward the eight expenditure categories with equal accuracy for all pairs inside each category. As Theorem (ref) suggested, our prior specification essentially assumes that the prior probability of any two units that come from the same expenditure category being in the same group is equal, and the prior group pattern is actually the expenditure category. We now examine the posterior of group structure to demonstrate how the distribution of $G$ gets updated by data. In order to accomplish this, we construct a posterior similarity matrix (PSM), whose $(i,j)$-element records the posterior probabilities of units $i$ and $j$ being in the same group. For illustrative purposes, we present the results for the last sample, in which we forecast CPI in August 2022. Figure (ref) depicts the PSM generated by BGFE-he-crst for the series in the categories of Food and Beverages and Transportation. A darker block indicates a higher posterior probability of being in one group. A common pattern emerges: even though some sub-indices are joined together frequently, as shown in the dark diagonal blocks, it is extremely unlikely that all series within the same category belong to the same group. Some series have relatively low or zero probabilities of being grouped together, as suggested by the white and gray off-diagonal blocks. This indicates that the group structure based on official expenditure categories is not optimal, which may result in inaccurate forecasting. Instead, our suggested framework uses information from both prior beliefs and data to reinvent the group pattern, leading to improved forecasting performance.
Finally, we restrict our analysis to the point estimate of group partition, i.e., the single grouping solution, rather than the posterior over the whole universe of partitions. Figure (ref) depicts the posterior point estimate of $G$ for the last sample ended in August 2022, derived using the approach described in Section (ref). Eight expenditure categories are divided into twelve groups of varied sizes. Two different forms of groups are generated based on the arrangement of their components. Groups 2, 3, 4, and 5 contain sub-indices from a variety of categories, with no clear dominance. In contrast, the majority of the series in groups 1 and 8, for example, belong to a certain category. Group 1 may refer to a Food group, whereas group 8 is a Transportation group. The detailed group 8 components are depicted in Figure (ref). There are seven sub-indices from Transportation, including car and truck rentals, gasoline (regular, midgrade, and premium), other motor fuels, airline fares, and ship fares, and one series from Housing - fuel oil (for residential heating). Clearly, all sub-indices share a common trend and have a close relationship with energy and oil prices, which have increased since the Pandemic. This is an example demonstrating that our proposed algorithm exploits cross-sectional information, not limited to our prior knowledge, and forms meaningful groups for forecasting.
We examine how the accuracy of pairwise constraints influences the point estimate of group partitioning. For demonstration purposes, we restrict our analysis to PL constraints solely by setting $\psi^{NL}_{ij} = 0.5$ and changing $\psi^{PL}_{ij}$. We do not select the constant $c$ in the setup since it would balance the impact of the PL restrictions with a different level of accuracy. We set $c$ to 0.5. Again, PL constraints are derived from the official expenditure categories, with the assumption that all units within the same category are positive-linked with equal probability of being in the same group.
Figure (ref) presents the point estimates of the group structure with two different levels of accuracy. The “weaker" PL constraints with $\psi^{PL}_{ij} = 0.55$, as shown in panel (a), demonstrate a limited influence of prior knowledge on the group structure. The eleven groups are composites of CPI sub-indices from the various categories, which is diverse from the official spending categories. Panel (b), on the other hand, illustrates the group structure with “stronger" PL constraints. By setting a high level of accuracy for PL constraints, such as $\psi^{PL}_{ij} = 0.95$, the prior knowledge dominates and pushes the group structure towards the official expenditure categories. As anticipated, panel (b) shows fewer groups, and the majority of CPI sub-indices within each group belong to the same category, bringing the group structure closer to that of the prior.
It is well-known in the literature that income per capita is strongly correlated with the level of democracy across countries. This strong empirical regularity is often known as "modernization theory" or Lipset hypothesis lipset1959. The theory claims a causal relation: democratic regimes are created and consolidated in affluent societies lipset1959, przeworski1995, barro1999, epstein2006.
In an influential paper, acemoglu2008 challenge the casual effect of countries' income on the level of democracy. They argue that it is essential to take into account other factors that affect both economic and political development simultaneously. Their analysis, based on panel data, indicates that the positive relationship between income and democracy disappears when fixed effects are included in the regression. They suggest that this finding is due to historical events, such as the end of feudalism, industrialization, or colonization, which have led countries to follow distinct paths of development. The fixed effects are meant to capture these persistent events. The finding is robust, as it holds for different measures of democracy, various econometric specifications, and additional covariates. Another study by bonhomme2015 uses a different econometric model but arrives at the same conclusion. Their analysis highlights the presence of diverse group-specific paths of democratization in the data, consistent with the observation that regime types and transitions tend to cluster in time and space gleditsch2006, ahlquist2012.
The seminal work of acemoglu2008 has been subject to critical scrutiny by several recent works, including moral2012, benhabib2013, and cervellati2014. moral2012 contend that a nonlinear relationship between income and democracy exists, even after accounting for country-specific effects. They show that a positive income-democracy relationship holds only in countries with low levels of income. benhabib2013 use panel estimation methods that adjust for the censoring of democracy measures at their upper and lower bounds and find that the positive relationship between income and democracy withstands the inclusion of country fixed effects. cervellati2014 extend the linear estimation framework of acemoglu2008 and unveil the presence of significant heterogeneity in the income effect on democracy across different subsamples. Specifically, they demonstrate that this effect exhibits an opposite sign for colonies and non-colonies, is substantially different from zero, and is of considerable magnitude. They also argue that the existence of a heterogeneous effect of income suggests that results from a linear framework, such as the finding of a zero effect, may lack robustness since they depend on the composition of the sample.
In this section, we contribute to the literature on the relationship between income and democracy by employing a novel grouped fixed-effects approach. Specifically, we expand on the econometric model proposed by bonhomme2015 to incorporate group structure not only in time fixed-effects, but also in slope coefficients and the variance of errors. This more complex model allows for a more detailed analysis of the heterogeneous effects of income on democracy across different groups of countries. To identify these groups, we incorporate prior knowledge about the latent group structure by clustering countries based on either geographic location or initial levels of democracy score. By leveraging this information, we are able to identify a moderate number of groups, each of which exhibits a distinct path to democracy.
Our results indicate that the effect of income on democracy is highly varied across countries, a finding which is consistent with previous research by cervellati2014. Furthermore, we find that the positive cumulative effect of income on democracy exists in groups of countries with a medium or relative low level of income, in line with the findings of moral2012. However, the effect could be relatively small for some groups, suggesting that other factors beyond income may also play important roles in democratic development.
{\bf Model:} To accommodate richer assumptions on models for the real-world applications, we extend the baseline model in Chapter 1 either by adding common regressors and allowing for time-variation in the fixed-effects. Time-variation are essential to this analysis as they capture highly persistent historical shocks. Following bonhomme2015, we introduce group-specific time patterns of heterogeneity $\alpha_{g_i t}$ and consider the following two specifications:
where $y_{it}$ is the democracy score of country $i$ at time $t$. The lagged value of democracy score $y_{it-1}$ is included to capture persistence in democracy and also potentially mean-reverting dynamics (i.e., the tendency of the democracy score to return to some equilibrium value for the country). The coefficient of main interest is $\beta_{g_i}$ and reflects the effect of the lagged value of $\log$ income per capita $x_{it-1}$ on democracy. In addition, $\alpha_{g_{i} t}$ denote a set of group-specific time fixed-effects; $\varepsilon_{i t}$ is an error term with grouped variance $\sigma_{g_i}^2$, capturing additional transitory shocks to democracy and other omitted factors. We use the conjugate prior for all parameters, see details in Appendix (ref).
Specification 2 in ((ref)) nests the linear dynamic panel data model in BM as a special case. If we assume homoskedasticity, it is the Equation (22) in BM. This specification enables us to reproduce BM's results and provide fresh insight into their framework. Specification 1 in ((ref)), on the other hand, generalizes specification 2 by introducing group-dependent slope coefficients. As we shall demonstrate in the following section, specification 1 yields a more refined group structure and provides a clearer view of the income effects.
{\bf Data:} We use the Freedom House (FH) Political Rights Index as a benchmark for measuring democracy. To standardize the index, we normalize it between 0 and 1, with higher scores indicating higher levels of democracy. FH assesses a country's political rights based on a checklist of questions, such as the presence of free and fair elections, the role of elected officials, the existence of competitive parties or other political groupings, the power of the opposition, and the extent of self-government or participation by minority groups. We measure countries' income using the logarithm of GDP per capita, which is adjusted for purchasing power parity (PPP) in 1996 prices, using data from the Penn World Tables 6.1 heston2002. Details on the data can be found in Section 1 of acemoglu2008, and all data in this section are from the replication files of bonhomme2015.\footnote{ \url{https://www.dropbox.com/s/ssjabvc2hxa5791/Bonhomme_Manresa_codes.zip?dl=0}} Our analysis is based on a five-year panel dataset that includes all independent countries since the postwar period, with observations taken every fifth year from 1970 to 2000. We chose this period for comparability with previous studies. The final dataset consists of a balanced panel of 89 countries.
{\bf Prior group structure:} We propose two prior grouping strategies as specified below and assume all units within the same prior group are presumed to be positive-linked, while units from different prior groups are believed to be negative-linked.
Figure (ref) presents the world maps with countries colored differently according to their respective groups. The panel on the left illustrates the geographic groups, while the panel on the right depicts the democratic groups. All gray nations/regions are excluded from the dataset. We concentrate primarily on the first pre-grouping strategy, as it needs no country-specific knowledge beyond geographic information. We then compare the results using different pre-grouping strategies in Section (ref).
{\bf Specification 1:} We begin with specification 1 where group-specific slope coefficients are allowed and new findings emerge. Table (ref) presents the posterior probability of the number of groups utilizing various estimators. BGFE-ho creates more than 5 groups in all posterior draws. Intriguingly, accounting for heteroskedasticity drastically reduces the number of groups, with BGFE-he identifying four groups. Adding pairwise constraints based on geographic information increases the number of groups to five, whereas six groups are expected in the prior.
The marginal data density (MDD) of each estimator in Table (ref) provides some insight on different models. Among all the estimators, the BGFE-ho estimator has the lowest MDD; it is even lower than that of specification 1. BGFE-he-cstr and BGFE-he, on the other hand, benefit from the introduction of group-specific slope coefficients, since both achieve substantially greater MDD than in specification 1. BGFE-he-cstr has the highest MDD since the pairwise constraints give direction on grouping and identify the ideal group structure, which BGFE-he cannot uncover without our prior knowledge.
We concentrate on the BGFE-he-cstr estimator and use the approach outlined in Section (ref) to identify the unique group partitioning $\widehat{G}$. The left panel of Figure (ref) presents the world map colored by $\widehat{G}$, while the right panel present the group-specific averages of democracy index over time. The estimated group structure $\widehat{G}$ features five distinct groups which we refer to as the “high-democracy", “low-democracy", “flawed-democracy", “late-transition" and “progressive-transition" group, respectively. With the exception of the “flawed-democracy" and “progressive-transition" group, the group-specific averages of the democracy index are comparable to those in BM for all other groups. BGFE-he-cstr does not identify the "early transition" group in comparison to BM but instead produces two new groups. Group 3 ("flawed-democracy") comprises primarily of relatively democratic but not the most democratic nations, including India, Sri Lanka, and Venezuela, among others. Group 5 ("progressive transition") contains 30 countries that have had a steady expansion of democracy, including Argentina, Greece, and Panama. Consequently, by incorporating group-specific slope coefficients, we recover a more refined group structure than that of BM.
Table (ref) presents the posterior mean and 90% credible set for each coefficient across all groups, with $G$ fixed at the point estimate $\widehat{G}$. The key feature of using the specification 1 is that we are able to see distinct (cumulative) income effects across groups as group-specific coefficients are allowed.
The effect of income on democracy is negligible for group 1 ("full-democracy") and group 4 ("late-transition") as the posterior means of $\beta$ are close to 0 and the associated credible intervals for $\hat{\beta}$ contain 0. Group 1, which we refer to as the "full-democracy" group, mostly contains high-income, high-democracy countries. It includes the United States, Canada, UK, most of European countries, Australia, and New Zealand, but also Costa Rica and Uruguay. These country kept their democracy index at the highest level throughout the sample, demonstrating that income has no effect on democracy. Group 4 is referred to as the "late-transition" group, which consists of Benin, Central African Republic, Mali, Malawi, Niger, and Romania. The transition to democracy for countries in group 4 was primarily driven by historical events in the 90s: Romania began a transition towards democracy after the 1989 Revolution; all other countries involved in the third wave of democratization in sub-Saharan Africa beginning in 1989. The impact of historical events is primarily captured by time fixed-effects, as the credible intervals for $\hat{\beta}$ and $\hat{\rho}$ well cover zero.
Group 2 ("low-democracy") and group 3 ("flawed-democracy") are two groups that have stable Freedom House scores. Lagged democracy is highly significant and indicates that there is a considerable degree of persistence in democracy. Log income per capita is also significant and illustrates the well-documented positive relationship between income and democracy. Though statistically significant, the effect of income is quantitatively small. For example, the coefficient of 0.055 for the group 3 implies that a 10 percent increase in GDP per capita is associated with an increase in the Freedom House score of 0.0055, which is very small. Group 2 includes low-democratic countries, China, Singapore, Iran, and a fraction of African countries, whose Freedom House scores remain relatively low throughout the sample. Countries in group 3, however, have Freedom House score stay in the relatively high level. The group covers countries with almost democratic but minor flaws in certain aspects, including India, Japan, South Korea, Finland, Sweden, and Portugal, among others. The cumulative income effects for these two groups are, however, different - it is negligible for group 2 (0.079) and modest for group 3 (0.244). Group 5 ("progressive-transition"), on the other hand, experiences a continuous increase in Freedom House score from a 0.35 in 1970 to 0.7 in 2020. It has the largest positive income coefficient, although the cumulative income effect is modest (0.156).
Figure (ref) depicts the historical average log income for each group of countries between 1970 and 2000. Except for the high income group (group 1) and low income group (group 4), all other three groups with medium or relatively low levels of income reveal positive cumulative income effect, as indicated in Table (ref). This finding is generally in line with the results of moral2012, who observe a positive effect in low-income nations. Notice that, their definition of low-income country is quite board, encompassing countries with a GDP per capita below the 80th percentile of the empirical cross-sectional density.. As a result, beside six counties in the group 4 that had the lowest income on average, other countries fit within this definition and confirm that the positive income effect is not prevalent in high-income countries.
{\bf Specification 2:} The results for specification 2 are reported in Appendix (ref). In short, the results are comparable to the key findings in BM. BGFE-ho in specification 1 is identical to the main model in BM; it produces eight groups, which is consistent with the upper bound on the number of groups in BM based on BIC. BGFE-he-cstr, on the other hand, is more preferable and has the highest marginal data density as shown in Table (ref). The point estimate of group partitioning based on BGFE-he-cstr consists of four groups that all have the similar pattern as BM's group structure. This justifies BM's subjective choice of four groups. Regarding the estimated coefficients, there is moderate persistence and a positive effect of income on democracy, but the cumulative effect of income is quantitatively small: $\beta / (1-\rho) = 0.08$.
All results presented thus far are based on pairwise constraints derived from spatial information. We now implement the alternative pre-grouping strategy based on the initial level of democracy. We stick with the BGFE-he-cstr estimator under specification 1.
Different pre-grouping strategies yields different estimates of group patterns. We consider the BGFE-he estimator with three prior group structures: geo-prior, dem-prior, and no prior knowledge, with point estimates of the group partition presented in panel (\subref{fig:app2_geo_prior}), (\subref{fig:app2_dem_prior}), and (\subref{fig:app2_no_prior}) of Figure (ref), respectively. Geo-prior and dem-prior produce comparable group structures, however certain nations are assigned to distinct groups.. They are encircled in the black dashed rectangle, including Portugal, Spain, Romania, Mali, Niger, Central African Republic, Benin, Malawi, and Jordan. Another country is South Korea. As depicted in panel (\subref{fig:app2_no_prior}), without any prior knowledge of groups, the group pattern is quite different, particularly for countries in Asia, Africa, and Latin America.
The difference in group patterns result in discrepancies in MDD, which are listed in Table (ref). The BGFE-he estimator with geo-prior has the highest MDD in both specifications, whereas the dem-prior is only informative in specification 1 in comparison to the BGFE-he estimator without prior knowledge.
Using the initial level of democracy as prior knowledge results in four groups, as indicated in the panel (\subref{fig:app2_dem_prior}). The dem-prior has two major impacts on the group structure comparing with the geo-prior. It combines the “late-transition" group (group 4 in geo-prior) with the “progressive-transition" group (group 5 in geo-prior) to form a bigger and boarder “progressive-transition" group. Additionally, Portugal and Spain are no longer categorized as “flawed-democracy" countries, but rather as “progressive-transition" group in a boarder sense. In terms of the posterior estimates, as shown in Table (ref), we observe similar value for the first three groups because they are merely subject to minor changes in the group structure. However, as “late-transition" and “progressive-transition" groups are agglomerated together under the dem-prior, the income effects of countries in the new group becomes larger, forcing some countries to exhibit strong and positive effects even though they are not under the geo-prior.
Within the domain of panel data models, the proposed constrained-based BGFE framework can be extended in multiple directions to allow for more subtle group structures or more covariates. In addition, the DP prior with soft pairwise constraints also applies to other related topics and models, such as clustering problems, heterogeneous treatment effects, and panel VARs.
Through the Dirichlet process defines a prior that possesses the clustering property and is flexible enough to incorporate pairwise constraints, the group structure itself is elementary. Aside from our prior belief on the group, the group structure, which is introduced in all $\alpha_{i}$ and $\sigma_i^2$, is entirely governed by the stick-breaking process defined in Equation ((ref)). The stick length $\xi_k$, on which we have a prior, is independent of any regressors or time. Consequently, each unit is associated with a single group, and the membership remains constant across time.
To create an even more flexible and richer group structure, we provide insight into three possible extensions, each of which requires a set of more distinctive nonparametric priors. (1) overlapping group and (2) time-varying group and (3) dependent group.
Overlapping group structures allow for multi-dimensional grouping. This is a natural extension without having to greatly modify the proposed DP prior. Following cheng2019, each of $\alpha_i$'s and $\sigma_i^2$ may have its own group structure and a separate Dirichlet process is specified to each of them. As a result, units simultaneously belong to multiple groups based on the heterogeneous effects among regressors or cross-sectional heteroskedasticity.
Time-varying group structures allow the membership of the group to change over time. We could replace the DP by variants of the hierarchical Dirichlet process Teh2006 to achieve this feature. In short, the hierarchical Dirichlet process (HDP), a nonparametric Bayesian approach to clustering grouped data, is now the foundation of the prior. The time dimension naturally divides the panel data into $T$ groups, and a Dirichlet process is assumed for each group, with all Dirichlet processes having the same base distribution, which is distributed according to a global base distribution. The HDP allows each group to have its own cluster, but most importantly, these clusters are shared across groups. This lays the groundwork for time-varying group structures, as it assumes that the number of clusters remains constant over time, while cluster memberships are subject to change. Variants of the HPD are then proposed to capture the time-persistence in group structures, including dynamic HDP ren2008 and sticky HDP fox2008, fox2011. A closely related area in the frequentists' methods is to identify structure breaks in parameters with grouped patterns, see okui2021, lumsdaine2022.
Dependent group structures allow the prior group probability to rely directly on a collection of characteristics. The dependence is introduced through a modification of the stick-breaking representation for DPs, where the group probabilities vary with the characteristics. rodriguez2011 introduced the probit-stick breaking (PSB) process where the Beta random variables are replaced by normally distributed random variables transformed using the standard normal CDF. The PSB is defined by,
where stochastic function $\zeta_k$ is drawn from Gaussian process $\zeta_k \sim G P\left(0, V_k\right)$ for $k=1,2, \cdots$ and $w_{i}$ is the set of characteristics that are informative to the latent group. Other forms of dependence are also available, see quintana2022 for a comprehensive review. A caveat of this approach is that analysis of group structure is confined to $w_{i}$ observed by the researcher. The approach requires researchers to know possible key characteristics, be able to observe them and ensure they are informative. In many cases, however, these characteristics might be hard to justify by researchers.
Although we concentrate on panel data model, our framework of the DP prior with soft pairwise constraints applies to other models where the group structure are crucial.
{\bf Gaussian Mixture Model}
If we ignore covariates and focus exclusively on group membership, we essentially face a classical clustering problem with an infinite-dimensional mixture model. A typical probabilistic model is the infinite Gaussian mixture model rasmussen1999, where the data itself is assumed to be drawn from a mixture of Gaussian components
where $\pi_k$ are the mixture weights. With soft pairwise constraints, observations are clustered in accordance with prior belief.
{\bf Heterogeneous Treatment Effects}
Following the potential outcomes framework of rubin1974, we posit the existence of potential outcomes $y_i(1)$ and $y_i(0)$ corresponding respectively to the response the $i$ th subject would have experienced with and without the treatment, and define the treatment effect at $x$ as
Previous works on Bayesian analysis for the treatment effects include chib2000, chib2002, chib2007a, chib2007b, heckman2014, etc. Another strand of literature propose to estimate ((ref)) by using several machine learning algorithms hill2011, athey2016, wager2018. These methods are built on the idea that researchers find the subsamples across which the effect of a treatment differs out of all possible subsamples on the basis of the values of $x_i$. Instead of trying to discover valid subsets of the data, shiraito2016 directly models the outcome as a function of the treatment and pre-treatment covariates $x_i$ and estimate of the distribution of conditional average treatment effects (CATE) across units by employing the Dirichlet process.
where $T$ is the binary treatment variable. Our approaches fit in this methods by adding side information on the treatment groups into the prior.
{\bf Panel VARs}
Panel VARs holtz1988, canova2013 has been widely used in macroeconomic analysis and policy evaluations to capture the interdependency across sectors, markets, and countries. Nevertheless, the large dimension of panel VARs typically makes the curse of dimensionality a severe problem. billio2019 propose nonparametric Bayesian priors that cluster the VAR coefficients and induce group-level shrinkage. Our paradigm with the DP prior with soft pairwise constraints is applicable to their method and injects prior information on groups into the underlying Granger causal networks.
Panel VARs have the same structure as VAR models, in the sense that all variables are assumed to be endogenous and interdependent, but a cross-sectional dimension is added to the representation. Thus, let $Y_t$ be the stacked version of $y_{i t}$, the vector of $J$ variables for each unit $i=1, \ldots, N$, i.e., $Y_t=\left(y_{1 t}^{\prime}, y_{2 t}^{\prime}, \ldots y_{N t}^{\prime}\right)^{\prime}$. Then a panel VAR is
where $u_{t}$ is a $J \times 1$ vector of idiosyncratic errors and $A_{0}$ and $A_j$ are $NJ \times NJ$ matrices of coefficients.
The main feature of billio2019 is to specify a prior that blends the DP prior with Lasso prior for each of $A_{0}$ and $A_j$, such that the VAR coefficients are either shrunk toward 0 or clustered at multiple non-zero locations. Our proposed DP prior with soft pairwise constraints, in the meantime, fit into their framework by replacing the original DP prior and permitting richer structure within each coefficient matrix. As the nonzero coefficients form Granger causal networks, equipping with soft pairwise constraints may result in a more plausible network by taking researchers' expertise into account.
This paper proposes a Bayesian framework for estimating and forecasting in panel data models when prior group knowledge is available and informative for the group pattern. We include prior knowledge in the form of soft pairwise constraints into the Dirichlet process prior. Then, an intuitive and coherent prior is presented. The constrained grouped estimator proposed examines both heteroskedasticity and heterogeneous slope coefficients to endogenously reveal group structure. Our framework immediately estimates the number of groups as opposed to relying on ex-post model selection, and the structure of pairwise restrictions circumvents the computational difficulties and limitations that afflict conventional approaches. In addition, when utilizing small-variance asymptotics, the suggested Gibbs sampler with pairwise constraint contains a clustering procedure comparable to that of the constrained KMeans algorithm. In Monte Carlo simulations, we demonstrate that constrained Bayesian grouped estimators outperform conventional estimators even in the presence of incorrect prior knowledge. Our empirical application to forecasting sub-indices of CPI inflation rates demonstrates that incorporating prior knowledge on the latent group structure yields more accurate density predictions. The better forecasting performance is mostly attributable to the key characteristics: nonparametric Bayesian prior and grouped cross-sectional variance. The method proposed in this paper is applicable beyond forecasting. In a second application, we revisit the relationship between a country's income and its democratic transition, where estimation of heterogeneous parameters is the object of interest. We recover a reasonable cluster pattern with a moderate number of groups and identify heterogeneous income effects on democracy.
The current work raises exciting questions for future research. It is desirable to investigate overlapping group structures, in which a unit might belong to many groups. This would allow us to increase the flexibility of a panel data model, potentially enhancing its predictive performance. Second, the assumption that an individual cannot change its group identity for the entire sample time can be amended, resulting in a specification that is even more flexible. Thirdly, our method is applicable to other econometric models, such as panel VARs with latent group structures in macro series.
\setstretch{1}
\setstretch{1.15}