Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
95,070 characters · 14 sections · 60 citation commands
Functional-Coefficient Quantile Regression for Panel Data with Latent Group Structure
\def\spacingset#1{ {#1}} \spacingset{1}
\if00 \fi
\if10 {
} \fi
\spacingset{1.2} {\it Keywords:} Cluster analysis; functional-coefficient models; incidental parameter; latent groups; local linear estimation; panel data; quantile regression
\spacingset{1.6}
\setcounter{equation}{0}
The quantile regression models and their estimation have received increasing attention since the seminal work by KB78. They have been widely applied in various disciplines including economics, finance, health science and social science. In contrast to classic mean regression, the quantile regression provides a more comprehensive picture in capturing the relationship between the response and explanatory variables, and serves as a robust alternative. Various parametric methods with theoretical treatment and empirical applications have been extensively studied for quantile regression, see Ko05 and KCHP17 for comprehensive reviews. Due to wide availability of panel/longitudinal data in many areas, it is natural to extend parametric linear quantile regression from independent data to more general panel data. Subject-specific individual effects are often incorporated in linear quantile panel models to reflect location shift effects on the quantile regression and describe heterogeneity over subjects. The number of these “incidental parameters" diverges as the number of subjects $N$ increases, affecting estimation accuracy of the common quantile regression coefficients (in particular when the number of observations per subject $T$ is fixed). Various quantile estimation and inferential techniques have been proposed in the literature Ko04, Ca11, KGM12, GLL13, GK16, GGV20 for large panel data, i.e., both $N$ and $T$ are large. However, the aforementioned literature relies on a pre-specified parametric linear model assumption, which may be too restrictive in quantile regression and is often rejected in practical data analysis. In this paper, we adopt a nonparametric panel modelling approach, allowing data to “speak for themselves" and thus providing more reliable numerical performance in quantile regression than the parametric one.
Nonparametric quantile regression estimation has been systematically studied in the literature for independent cross-sectional data or weakly dependent time series data YJ98, Ca02, YL04, CX08, LLR13, BCCF19, LLL21. In recent years, there have been some attempts to study nonparametric quantile regression for panel data with individual effects. For example, YL18 introduce a three-step nonparametric conditional quantile estimation method combining series approximation, first-order difference (to remove incidental parameters) and deconvolution; and Ch21 proposes two local-linear-based methods to estimate quantile partial effects and further derives the asymptotic distribution theory under the large panel setting. Among various nonparametric quantile regression models, the functional-coefficient model is one of the most commonly-used frameworks. It is a natural extension of the linear quantile regression model, avoiding the so-called “curse of dimensionality" problem in nonparametric estimation when the number of covariates is large. For the classic independent or time series data setting, the functional-coefficient quantile model and its generalised version have been extensively studied in the literature Ho04, Ki07, CX08, WZZ09, KLZ11, TSWZ13. In particular, the functional-coefficient quantile regression allows the dynamic quantile relationship to vary smoothly over a state variable, and is thus connected to the functional linear quantile regression K12, but the latter assumes the covariate takes a functional value and bases the estimation methodology on some dimension reduction techniques (such as the functional principal component analysis). For the panel data with subject-specific fixed effects, SH16 combine the series approximation and instrumental variable quantile regression CH06 to estimate functional coefficients under a large-$T$ framework, whereas CCF18 use a kernel weighted quasi-likelihood to estimate semiparametric functional-coefficient quantile regression under a fixed-$T$ framework.
A typical assumption imposed on nonparametric quantile regression for panel data in the existing papers is that the main nonparametric regression structure (after removing subject-specific location shift effects) is invariant over subjects, indicating that the dynamic relationship between the dependent and explanatory variables is the same for all subjects. However, such an assumption is often too restrictive in panel data studies when subjects involved have very different characteristics. An example is the house price data from UK local authority districts that we consider in Section (ref). Due to differences in their location, population, and the socio-economic backgrounds of their population, the effects of factors, such as population growth and personal income growth, on house price growth are very unlikely to be homogeneous. As a result, in this paper, we relax the homogenous panel model assumption, allowing the functional-coefficients in nonparametric quantile regression to vary over subjects.
However, for the heterogenous functional-coefficient panel quantile regression without imposing any structural restriction on the subject-specific coefficients, we can only rely on the sample information from an individual subject to estimate the subject-specific dynamic relationship, which leads to slow convergence of the estimated functional coefficients and unstable numerical performance of the estimates in finite samples. To address this problem, we assume that there exists a latent group structure on the subject-specific functional coefficients at each quantile level. If the underlying group structure is known or can be consistently estimated, more efficient coefficient estimates can be obtained by pooling information belonging to the same group. From an empirical perspective, some panel data studies, such as PS07 and HM10, have found group structures for the panel models they use. For the UK house price data in Section (ref), we identify two or five homogeneous groups (depending on the quantile level) for the 335 local authority districts. Such homogeneous groups may exist due to similarities in the characteristics of many districts (e.g., type of district, i.e., urban or rural, and socio-economic background of the majority of the population). Hence, it is both beneficial and reasonable to assume a group structure in some panel studies.
There has been increasing interest on estimating latent group structure in mean regression models for panel data in recent years. For example, KLZ16 use a binary segmentation technique to identify the latent group structure in linear regression models for panel data, whereas SSP16 introduce a penalised method via the so-called classifier-LASSO. VL17, VL20 propose a kernel-based classification of univariate nonparametric regression functions in panel data, which is further extended by Ch19 to estimate the group structure in time-varying coefficient panel data models. Other relevant developments can be found in BM15, AB17, SWJ19, LSZZ20, WS20 and LQZ21. In contrast, there is sparse literature on quantile regression models for panel data with latent group structures. CLP16 study IV panel quantile regression with group-specific coefficients defined as a linear regression with observable group-level covariates, but the “groups" in their paper are essentially the subjects in the context of this paper. ZZW19 propose an $L_1$-penalised estimation method to identify the group structure on the intercept in linear median regression; GV19 estimate linear quantile regression for panel data with a latent group structure on the subject-specific fixed effects; and ZWZ19 introduce an iterative algorithm using an idea similarly to the classic k-means clustering to estimate groups of units in panel data with heterogeneous slope coefficients. These estimation methods and algorithms rely on the parametric linear model assumption in quantile regression and cannot be directly applied to estimate the latent structure in nonparametric panel quantile regression.
In this paper, we aim to consistently estimate the group structure, the group number and the group-specific functional coefficients, all of which are allowed to vary over quantile levels. As there is no prior information on the latent groups, we start with a preliminary local linear quantile estimation of the subject-specific functional coefficients and the incidental parameter, only using the sample information from one subject. Based on the preliminary estimates of the functional coefficients, we compute the distance matrix between the subjects and subsequently use a classic agglomerative clustering algorithm to estimate the unknown group structure for the heterogenous functional coefficients. The resulting estimate is shown to be consistent (once the group number is pre-specified). Then, we introduce a simple ratio criterion to consistently estimate the group number. As the preliminary quantile estimates have rather slow convergence rates, we further propose a post-grouping local linear smoothing method to estimate the group-specific functional coefficients using the consistently estimated group structure, and derive the asymptotic normal distribution theory for the developed estimate with a convergence rate comparable to that in the literature. In the asymptotic analysis, we focus on the large panel setting with both $N$ and $T$ diverging to infinity. The panel observations are allowed to be temporally dependent and cross-sectionally correlated, relaxing the commonly-used cross-sectional independence restriction for panel quantile estimation KGM12, CCF18, Ch21.
We apply the proposed method to the house price data from UK local authority districts over the period Q1/1997--Q4/2016 and discover different group structures at different quantiles. At the lower quartile and median, we find more homogeneity in the effects of population and income growth on house price growth across districts, while at the upper quartile, more groups (i.e., five) are identified. By allowing the group structure to vary with the quantile level, we uncover a clearer picture about the relationship between population and income growth and house price growth across the distribution of house price growth.
The rest of the paper is organised as follows. Section (ref) introduces the model and latent group structure. Section (ref) describes the clustering algorithm and the ratio criterion for estimating the latent structure, and the post-grouping local linear quantile estimation. The technical assumptions and main asymptotic properties are provided in Section (ref). Section (ref) reports both the simulation and empirical studies. Section (ref) concludes the paper. Proofs of the main theorems and technical lemmas, extensions of the developed methods and theory and additional simulation studies are available in a supplement.
\setcounter{equation}{0}
Suppose that we collect the panel random observations $\left(Y_{it}, {\mathbf X}_{it}\right)$, $i=1,\cdots,N$, $t=1,\cdots,T$, and time series random observations $Z_t$, $t=1,\cdots,T$, where $Y_{it}$ and $Z_t$ are univariate and ${\mathbf X}_{it}$ is $d$-dimensional. Let $\alpha_i$ be a subject-specific effect which may be correlated with $X_{it}$ and $Z_t$. At a given quantile level $0<\tau<1$, the conditional quantile function for the $i$-th subject has the following functional-coefficient regression form:
where ${\boldsymbol\beta}_{\tau,i}(\cdot)$ is a $d$-dimensional vector of subject-specific functional coefficients. Both ${\boldsymbol\beta}_{\tau,i}(\cdot)$ and $\alpha_{\tau,i}$ are allowed to depend on the quantile level $\tau$. It is worth stressing that model ((ref)) is different from the random-coefficient quantile regression model KX06, see the discussion in Appendix C.1 of the supplement. The index variable $Z_t$ does not play any role in our model identification. The functional coefficients ${\boldsymbol\beta}_{\tau,i}(Z_t)$ capture smooth changes of the dynamic quantile relationship (over $Z_t$) between $Y_{it}$ and ${\mathbf X}_{it}$ at a fixed quantile level. Without loss of generality, we assume $Z_t$ has a compact support $[0,1]$. In practical applications, we may replace the random index variable $Z_t$ in ((ref)) by the fixed scaled time, $t/T$, or a variable $Z_{it}$ that varies over both $i$ and $t$, which would lead to the following functional-coefficient quantile regressions:
With slight modification, the methodology and theory to be developed in Sections (ref) and (ref) are still applicable to the above two model variants. Models in ((ref)) and ((ref)) can be seen as an extension of the functional-coefficient/time-varying panel data models studied by LCG11, Ch19, SWJ19 and PW22 from mean regression to quantile regression.
In this paper, we further assume that there exists a partition of the index set $\{1,2,,\cdots,N\}$, denoted by $\boldsymbol{{\cal G}}_\tau=\{{\cal G}_1^\tau,{\cal G}_2^\tau,\cdots,{\cal G}_{R_{\tau,0}}^\tau\}$ such that
where ${\boldsymbol\gamma}_{\tau,j}(\cdot)$ denotes a $d$-dimensional vector of group-specific functional coefficients that may also depend on $\tau$. Neither the group membership nor the group number is known a priori. Combining ((ref)) and ((ref)), we readily have that
Note that the total number of unknown functional coefficients in ((ref)) is $dR_{\tau,0}$, which is much smaller than $dN$, the number of heterogenous functional coefficients in ((ref)).
Two remarks are in order here. First, the group membership $\boldsymbol{\mathcal{G}}_\tau$ and the group number $R_{\tau,0}$ are allowed to vary over $\tau$, which implies that the latent group structure can change over quantile levels. This makes the proposed quantile regression model framework more flexible and applicable than the mean regression one for practical research. Second, although we assume a group structure on the functional coefficients, ${\boldsymbol\beta}_{\tau,i}(\cdot)$, no group structure is imposed on the individual effects $\alpha_{\tau,i}$. This means that the subjects belonging to the same group are still allowed some degree of heterogeneity, as represented by their individual specific effects, albeit having the same functional slope coefficients.
The main interest of this paper lies in estimation of $\boldsymbol{{\cal G}}_\tau$, $R_\tau$ and ${\boldsymbol\gamma}_{\tau,j}(\cdot)$, $j=1,\cdots,R_{\tau,0}$. For notational simplicity, we next write ${\boldsymbol\beta}_{\tau,i}(\cdot)={\boldsymbol\beta}_{i}(\cdot)=\left[\beta_{i,1}(\cdot),\cdots,\beta_{i,d}(\cdot)\right]^{^\intercal}$, ${\boldsymbol\gamma}_{\tau,j}(\cdot)={\boldsymbol\gamma}_{j}(\cdot)=\left[\gamma_{j,1}(\cdot),\cdots,\gamma_{j,d}(\cdot)\right]^{^\intercal}$, $\alpha_{\tau,i}=\alpha_i$, $R_{\tau,0}=R_0$ and $\boldsymbol{{\cal G}}_\tau=\boldsymbol{{\cal G}}=\{{\cal G}_1,{\cal G}_2,\cdots,{\cal G}_{R_{0}}\}$, suppressing their dependence on $\tau$.
\setcounter{equation}{0}
Note that model ((ref)) is a semiparametric functional-coefficient quantile model by treating $\alpha_i$ as an incidental parameter for fixed $i$. Assume that the unknown functional coefficients have continuous second-order derivatives. For $z\in [0,1]$, with the sample information from the $i$-th subject, we define
where $\rho_\tau(\cdot)$ is the quantile check function defined by $\rho_\tau(z)=z\left[\tau-I(z\leq0)\right]$ with $I({\cal A})$ being the indicator function of the event ${\cal A}$, $K_h(u)=K(u/h)$, $K(\cdot)$ is a kernel function and $h$ is a bandwidth. The local linear estimates $\widehat{\boldsymbol\beta}_i(z), \widehat{\boldsymbol\beta}_i^\prime(z), \widehat{\alpha}_i(z),\widehat{\alpha}_i^\prime(z)$ are obtained as the solution to minimise the objective function in ((ref)).
Let ${\boldsymbol\Delta}$ be an $N\times N$ distance matrix among the true functional coefficients ${\boldsymbol\beta}_{j}(\cdot)$, $j=1,\cdots,N$. The diagonal elements of ${\boldsymbol\Delta}$ are zeros, whereas the off-diagonal elements $\Delta(j,k)$, $1\leq j\neq k\leq d$, are defined by
where $f(\cdot)$ is the density function of $Z_t$ and $\Vert\cdot\Vert$ denotes the Euclidean norm. With the preliminary local linear quantile estimates, we have the following estimate of $\Delta(j,k)$:
With $\widehat\Delta(j,k)$, we obtain $\widehat{\boldsymbol\Delta}$, an $N\times N$ estimated distance matrix of ${\boldsymbol\Delta}$. The $(j,k)$-entry of $\widehat{\boldsymbol\Delta}$ is $\widehat\Delta(j,k)$ and the diagonal elements of $\widehat{\boldsymbol\Delta}$ are zeros. Using the estimated distance matrix, we may apply the agglomerative clustering method which has been widely used in the literature of cluster analysis ELLS11,RC12. Recently, such a method, combined with the kernel-based smoothing technique, has been applied to estimate the homogeneity/group structure in nonparametric mean regression models Ch19, VL20, CLWZ21. However, so far as we know, there is virtually no work on applying the kernel-based agglomerative clustering method to quantile regression models with a latent group structure. We next introduce the clustering algorithm when the group number is assumed to be $R$.
Let $\widehat{\cal G}_{1|R},\cdots, \widehat{\cal G}_{R |R}$ be the estimated clusters for a given group number $R$. If the true number $R_0$ is known a priori, we denote the estimated groups by $\widehat{\cal G}_r=\widehat{\cal G}_{r|R_0}$, $r=1, \cdots,R_0$, whose consistency property is given in Theorem (ref).
We next introduce a ratio criterion to consistently estimate the group number $R_0$. For a given number $R$, with the estimated groups $\widehat{\cal G}_{r|R}$ defined in Section (ref), we may pool the estimated functional coefficients $\widehat{\boldsymbol\beta}_j(\cdot)$, $j\in\widehat{\cal G}_{r|R}$, and obtain the following estimate: \[ \widehat{\boldsymbol\beta}_{r|R}(z)=\frac{1}{\left\vert\widehat{\cal G}_{r|R}\right\vert}\sum_{j\in\widehat{\cal G}_{r|R}}\widehat{\boldsymbol\beta}_j(z),\ \ r=1,\cdots,R, \] where $\vert {\cal A}\vert$ denotes the cardinality of a set ${\cal A}$. Then, we calculate the average deviation for $\widehat{\boldsymbol\beta}_j(\cdot)$ if the group number is assumed to be $R$:
It follows from Theorem (ref) that the functional-coefficient quantile panel regression model is either correctly- or over-fitted (with probability tending to one) when $R\geq R_0$, and ${\sf D}(R)$ is thus convergent to zero. On the other hand, the model is under-fitted when $R<R_0$, and at least two groups are falsely merged. Consequently ${\sf D}(R)$ is strictly larger than a positive constant (using Assumption (ref)(ii) in Section (ref)). Hence, it is sensible to determine $R_0$ via the following simple ratio criterion:
where $\overline{R}$ is a pre-specified positive integer larger than $R_0$, and we set ${\sf D}(1)/{\sf D}(0)=1$, ${\sf D}(R)=0$ if ${\sf D}(R)$ is smaller than $\omega_{NT}$, a threshold satisfying some mild restrictions, and define $0/0\equiv1$. A similar ratio criterion is also used by LY12 and AH13 to find the number of latent factors in approximate factor models, and by LRS20 to determine the dimension of dominant sub-space in functional time series. Theorem (ref) in Section (ref) below shows that $\widehat{R}$ is a consistent estimate of $R_0$. With $\widehat{R}$, we may extract the estimated groups $\widetilde{\boldsymbol{\mathcal G}}=\{\widetilde{\cal G}_1, \widetilde{\cal G}_2,\cdots, \widetilde{\cal G}_{\widehat{R}}\}$ by terminating the agglomerative clustering algorithm when $R=\widehat{R}$.
Note that the preliminary functional coefficient estimates defined in Section (ref) only make use of the sample information from one subject, resulting in a relatively slow uniform convergence rate, see Lemma A.2 in Appendix A of the supplement. The numerical performance of these estimates may be unstable in finite samples in particular when $T$ is not sufficiently large. With the consistent estimates of the group number and membership constructed in Sections (ref) and (ref), we next propose a post-grouping local linear quantile estimation method for the group-specific functional coefficients ${\boldsymbol\gamma}_j(\cdot)$, $j=1,\cdots,R_0$, improving the convergence rate of the preliminary functional coefficient estimates. From Corollary (ref) to be given in Section (ref), for any $j=1,\cdots,R_0$, there exists $1\leq j_\ast\leq \widehat{R}$ such that ${\sf P}\left({\cal G}_j=\widetilde{\cal G}_{j_\ast}\right)\rightarrow1$. Without loss of generality, we let $j_\ast=j$ in the rest of the section. Define the post-grouping local linear weighted objective function:
where $K_{h_1}(z)=K(z/h_1)$, $K(\cdot)$ is the kernel function and $h_1$ is a bandwidth which may be different from $h$ used in the preliminary local linear estimation. The post-grouping local linear estimates $\widetilde{\boldsymbol\gamma}_j(z), \widetilde{\boldsymbol\gamma}_j^{\prime}(z),\widetilde{\alpha}_i(z), \widetilde{\alpha}_i^\prime(z)$, $i\in\widetilde{\cal G}_j$, are obtained as a solution to minimise the objective function in ((ref)). As the fixed effects $\alpha_i$ are treated as “nuisance parameters", our primary interest lies in $\widetilde{\boldsymbol\gamma}_j(z)$ whose asymptotic distribution theory will be derived in Section (ref) below.
\setcounter{equation}{0}
For $i=1,\cdots,N$, we let
when $f(\cdot)$ is the density of $Z_t$ and $f_{ie}(\cdot |{\mathbf x},z)$ is the conditional density of $e_{it}=Y_{it}-{\mathbf X}_{it}^{^\intercal} {\boldsymbol\beta}_{i}(Z_t)-\alpha_i$ given ${\mathbf X}_{it}={\mathbf x}$ and $Z_t=z$. Assumptions (ref)--(ref) below are sufficient to prove the consistency properties for the estimated group membership and number.
\setcounter{remark}{0}
With the latent structure ((ref)), we write $e_{it}=Y_{it}-{\mathbf X}_{it}^{^\intercal}{\boldsymbol\gamma}_j(Z_t)-\alpha_i$ for $i\in{\cal G}_j$ and let \[b_{it}(z)={\mathbf X}_{it}^{^\intercal}\left[{\boldsymbol\beta}_i(Z_t)-{\boldsymbol\beta}_i(z)-{\boldsymbol\beta}_i^\prime(z)(Z_t-z)\right]={\mathbf X}_{it}^{^\intercal}\left[{\boldsymbol\gamma}_j(Z_t)-{\boldsymbol\gamma}_j(z)-{\boldsymbol\gamma}_j^\prime(z)(Z_t-z)\right].\] Re-write ${\boldsymbol\Omega}_{i}(z)$ defined in ((ref)) in the block-matrix form:
where $\omega_i^\alpha(z)$ is univariate and ${\boldsymbol\Omega}_{i}^\gamma(z)$ is a $d\times d$ matrix. For $j=1,\cdots,R_0$, we define
where $\eta_{it}(z)=\tau-I\left(e_{it}\leq -b_{it}(z)\right)$. To derive the asymptotic distribution theory for the post-grouping local linear quantile estimation, we need some additional conditions.
We start with the consistency property for the group membership estimate when $R_0$ is pre-specified.
\setcounter{theorem}{0}
The following theorem shows that the simple ratio criterion proposed in Section (ref) consistently estimates the group number $R_0$.
Combining Theorems (ref) and (ref), we readily have the following corollary on consistency of the group membership estimate when $R_0$ is unknown.
\setcounter{corollary}{0}
We finally turn to the asymptotic normal distribution theory for the post-grouping estimate $\widetilde{\boldsymbol\gamma}_j(z)$ of the group-specific functional coefficient. It follows from Corollary (ref) that there exists $1\leq j_\ast\leq \widehat{R}$ such that ${\sf P}\left({\cal G}_j=\widetilde{\cal G}_{j_\ast}\right)\rightarrow1$. Without loss of generality, we let $j_\ast=j$ as in Section (ref), and define $\mu_k=\int z^kK(z)dz$ and $\nu_k=\int z^kK^2(z)dz$ for $k=0,1,2,\cdots$.
\setcounter{equation}{0}
We consider two data generating processes (DGP) to examine the small-sample performance of the proposed methodologies in various scenarios.
{\bf DGP 1}.\ \ Data is generated via the following functional-coefficient quantile regression:
where $Z_t$, $t=1,\cdots T$, are independently drawn from the uniform distribution $U[0,1]$, ${\mathbf X}_{it}=(X_{it,1}, X_{it,2})^{^\intercal}$, $i=1,\cdots,N, t=1,\cdots, T$, are independently drawn from a bivariate normal distribution with zero means, unit variances, and a correlation coefficient of $1/2$, $\alpha_i=\left(\bar{X}_{i,1}^2+\bar{X}_{i,2}^2\right)/5$ with $\bar{X}_{i,k}=\frac{1}{T}\sum_{t=1}^TX_{it,k}$, and the idiosyncratic errors $e_{it}$ are independently generated from one of the following distributions: ${\sf N}(0,1)$, $t(5)$, and $0.4\big[\chi^2(3)-3\big]$, and are independent of $Z_t$ and ${\mathbf X}_{it}$. The results for the ${\sf N}(0,1)$ errors can be used as the benchmark for assessing how the proposed method performs when the errors are heavy tailed (with $t(5)$ distribution) or asymmetrically distributed (with $0.4\big[\chi^2(3)-3\big]$ distribution, where the scaling factor $0.4$ is used to give a comparable error variance to ${\sf N}(0,1)$).
As in Ch19 and SWJ19, the heterogenous functional coefficients ${\boldsymbol\beta}_{i}(\cdot)=\left[\beta_{i,1}(\cdot), \beta_{i,2}(\cdot)\right]^{^\intercal}$ satisfy the following group structure: \[ \beta_{i,1}(z)=\left\{
\right. \] and \[ \beta_{i,2}(z)=\left\{
\right. \] where $F(z;\xi,\eta)=1/(1+\exp[-(z-\xi)/\eta])$, ${\cal G}_1=\left\{1,2,\cdots,N_1\right\}$, ${\cal G}_2=\left\{N_1+1,\cdots,N_1+N_2\right\}$, and ${\cal G}_3=\left\{N_1+N_2+1,\cdots,N_1+N_2+N_3\right\}$ with $N_1=\lfloor 0.3N\rfloor $, $N_2=\lfloor 0.3N\rfloor $ and $N_3=N-N_1-N_2$.
{\bf DGP 2}.\ \ This DGP has the same model equation as ((ref)), but with quantile-dependent functional coefficients and group structure: \[ \beta_{i,1}(z)=\beta_{\tau,i,1}(z)=\left\{
\right. \] and \[ \beta_{i,2}(z)=\beta_{\tau,i,2}(z)=\left\{
\right. \] where ${\cal G}_k$ and $\gamma_{k,j}(z)$, $k=1,2,3$ and $j=1,2$, are the same as in DGP 1. Data for $\mathbf{X}_{it}$, $Z_t$, $\alpha_i$, and $e_{it}$ are generated in the same way as in DGP 1. Note that when $\tau=0.5$, ${\cal G}_1$ and ${\cal G}_2$ merge so there are only two groups, whereas when $\tau\neq0.5$, there are three groups.
The sample size is $N=50, 100$ and $T=50, 100$, and the number of replications is $M=200$. We consider three quantile levels, $\tau=0.25, 0.50$ and $0.75$, and use the Gaussian kernel, $K(u)=e^{-u^2/2}/\sqrt{2\pi}$, in the local linear estimation. The bandwidth is selected via the leave-one-out cross-validation method. To gauge the performance of the ratio criterion ((ref)) in estimating the number of groups, we set $\overline{R}=5$ and report the percentage of replications where each integer (between 1 and $\overline{R}$) is chosen. To understand the accuracy of the estimated groups (and their membership) from the agglomerative clustering algorithm in Section (ref), we consider two measures: purity and the normalised mutual information (NMI) of the estimated groups $\widetilde{\boldsymbol{\mathcal G}}=\left\{\widetilde{{\cal G}}_1, \cdots, \widetilde{{\cal G}}_{\widehat{R}}\right\}$ with the true groups $\boldsymbol{{\mathcal G}}=\left\{{\cal G}_1, \cdots, {\cal G}_{R}\right\}$, both of which are classic criteria of clustering quality and are defined respectively as $${\rm Purity}(\widetilde{\boldsymbol{\mathcal G}}, \boldsymbol{{\mathcal G}})=\frac{1}{N}\sum_{k=1}^{\widehat{R}}\max_{1\leq j\leq R_0}\left|\widetilde{{\cal G}}_k \cap {\cal G}_j\right|,\ \ \ \ \ \ {\rm NMI}(\widetilde{\boldsymbol{\mathcal G}}, \boldsymbol{\mathcal G})=2\cdot\frac{I(\widetilde{\boldsymbol{\mathcal G}}, \boldsymbol{\mathcal G})}{H(\widetilde{\boldsymbol{\mathcal G}})+H(\boldsymbol{\mathcal G})},$$ where $I(\widetilde{\boldsymbol{\mathcal G}}, \boldsymbol{\mathcal G})$ is the mutual information between $\widetilde{\boldsymbol{\mathcal G}}$ and $\boldsymbol{\mathcal G}$ defined as $$I(\widetilde{\boldsymbol{\mathcal G}}, \boldsymbol{\mathcal G})=\sum_{k=1}^{\widehat{R}}\sum_{j=1}^{R_0}\frac{\left|\widetilde{{\cal G}}_k \cap {\cal G}_j \right|}{N}\cdot\log_2\left(\frac{N\left |\widetilde{{\cal G}}_k \cap {\cal G}_j\right|}{\left|\widetilde{{\cal G}}_k\right| \cdot\left|{\cal G}_j \right|}\right),$$ $H(\widetilde{\boldsymbol{\mathcal G}})$ is the entropy of $\widetilde{\boldsymbol{\mathcal G}}$ defined as $$H(\widetilde{\boldsymbol{\mathcal G}})=-\sum_{k=1}^{\widehat{R}}\frac{\left|\widetilde{\cal G}_k\right|}{N}\log_2\left(\frac{\left|\widetilde{\cal G}_k\right|}{N}\right),$$ and $H(\boldsymbol{\mathcal G})$ is defined analogously. The closer the values of NMI and purity are to 1, the more accurate the estimated groups are to the true ones. We also look at the estimation accuracy of the functional coefficients for both the preliminary local linear quantile estimator defined in Section (ref) and the post-grouping local linear quantile estimator defined in Section (ref). For this we compute the average root mean squared errors (RMSE) defined as
for an estimate, $\widehat{\boldsymbol\beta}(\cdot)=\big(\widehat{\boldsymbol\beta}_1(\cdot),\cdots,\widehat{\boldsymbol\beta}_N(\cdot)\big)^{^\intercal}$, of the true functional coefficients ${\boldsymbol\beta}(\cdot)=\big({\boldsymbol\beta}_1(\cdot),\cdots,{\boldsymbol\beta}_N(\cdot)\big)^{^\intercal}$. As a benchmark, we also compute the RMSE of the oracle estimator, which assumes that the true group structure is known a priori and pools data belonging to each group to obtain group-specific estimates of the functional coefficients. The results for DGP 1 are reported in Table (ref) (for ${\sf N}(0,1)$ errors), Table (ref) (for $t(5)$ errors), and Table (ref) (for $0.4\big[\chi^2(3)-3\big]$ errors). Results for DGP 2 are reported in Tables (ref)--(ref).
Table (ref) shows that for DGP 1, when the errors follow the ${\sf N}(0,1)$ distribution, for the smallest sample size considered (i.e., $N=50$ and $T=50$), the ratio criterion ((ref)) picks the correct number of groups in about 75% of the replications at $\tau=0.50$ and about 60% of the replications at $\tau=0.25$ and 0.75. These percentages increase markedly as $T$ increases and usually as $N$ increases, but to a lesser extent. When $T=100$, the correct $R_0$ is chosen in 100% of the replications at $\tau=0.50$ and more than 85% of the replications at $\tau=0.25$ and 0.75. This is consistent with the theoretical result in Theorem (ref). On the other hand, the NMI values are above 0.77 and purities above 0.91 at all quantiles when $T=50$, and they increase to above 0.95 when $T$ increases to $100$, verifying the consistency results in Theorem (ref) and Corollary (ref) for the agglomerative clustering algorithm. The lower block of Table (ref) shows that by pooling data belonging to the same group in the estimation, the post-grouping local linear estimator cuts the RMSE of the preliminary local linear estimator by more than 30%, and its RMSE values are not far from those of the oracle estimator. When $T$ increases to $100$ and the groups are accurately estimated, the RMSEs of the post-grouping estimator are very close to the oracle estimator. Similar findings can be drawn from Tables (ref) and (ref), where the errors in DGP 1 follow the heavier tailed $t(5)$ distribution and the asymmetric $0.4\big[\chi^2(3)-3\big]$ distribution, respectively. The results for $t(5)$ errors are in general worse than those of ${\sf N}(0,1)$ and $0.4[\chi^2(3)-3]$ errors, which may be due to the fact that the $t(5)$ distribution has a larger variance than the other two. For $0.4[\chi^2(3)-3]$ errors, the results are better than those of ${\sf N}(0,1)$ at $\tau=0.25$, worse than ${\sf N}(0,1)$ at $\tau=0.75$ and comparable at median $\tau=0.50$. This may be because at $\tau=0.25$, the $0.4\big[\chi^2(3)-3\big]$ distribution has more data points than the ${\sf N}(0,1)$ distribution and hence, the preliminary functional coefficients estimation and latent group estimation are more accurate. At $\tau=0.75$, the reverse is true.
For DGP 2, the performance of the proposed methods is similar, despite the fact that the functional coefficients and the group structure are now quantile dependent. In Tables (ref)--(ref), we can observe the same evolving patterns when $T$ or $N$ increases as in Tables (ref)--(ref). This demonstrates the robustness of our methods to the different settings considered.
To further illustrate the applicability and usefulness of our methods, we next consider a panel house price growth model for UK local authority districts (LADs). Similarly to CSZ22, we use quarterly house price data over the period Q1/1997 -- Q4/2016\footnote{Due to data unavailability, we consider LADs only in England and Wales.}, downloaded from the UK Office of National Statistics (ONS) website: \url{https://www.ons.gov.uk/}. Growth rates in population and the nominal per capita personal income are used as the explanatory variables. Their data at the individual LAD level are available at the annual rate on the ONS website\footnote{Four LADs from England and Wales - Aylesbury Vale, Gloucester, Norwich, and Powys, are excluded due to their outlying values. This gives a total of 335 LADs for the subsequent analysis.}, and we construct quarterly data for the two variables using the interpolation method in Den71 and DC06. For the index variable, we use quarterly inflation rates, which are available at the country level at \url{https://www.bls.gov/cpi/data.htm}. This allows for interaction between inflation and the explanatory variables and for the effects of the explanatory variables on house prices to vary with inflation. Data for all the variables have been de-seasoned and de-trended\footnote{More data details can be found in CSZ22.}.
Assume the following panel quantile regression:
for $i=1,\cdots, 335$ and $t=2,\cdots, 80$, where $hp_{it}$ is the house price growth rate (in %) for the $i$-th LAD in the $t$-th quarter, $pop_{it}$ is the population growth (in %), $inc_{i,t-1}$ is the growth in per capita personal income (in %), $inf_t$ is the inflation rate (in %), and $\alpha_i$ is the fixed effect. Here $inc$ is lagged by one time period due to the likely lagged effect of income growth on house price.
While the model ((ref)) offers great flexibility by allowing the coefficients to vary across LADs and quantiles as well as with inflation, there may exist some homogeneity groups of LADs at each quantile level, where the coefficients are homogeneous within each group (while still varying with inflation) but heterogeneous across groups. That is, at a given quantile level $\tau$, there may exist a partition of the index set $\{1,2,,\cdots,335\}$, denoted as $\{{\cal G}_1,{\cal G}_2,\cdots,{\cal G}_{R_0}\}$, such that \[ {\boldsymbol\beta}_i(\cdot):=\left(
\right)={\boldsymbol\gamma}_j(\cdot):=\left(
\right)\ \ {\rm for}\ i\in{\cal G}_j,\ \ \ \ i=1,\cdots,335,\ j=1,\cdots,R_0.\notag \] We consider 3 quantile levels, i.e., $\tau=0.25, 0.50, 0.75$, and at each quantile, we then use the methods proposed in Section (ref) to estimate the number of latent groups as well as the group membership. The results are summarised in Table (ref). The post-grouping local linear estimates of the group-specific functional coefficients, together with their 95% confidence intervals, are plotted in Figures (ref)--(ref) for $\tau=0.25, 0.50, 0.75$, respectively. Construction of the confidence intervals is introduced in Appendix C.3 of the online supplement. The estimated groups at each quantile level, projected onto a choropleth map, are shown in Figures (ref)--(ref).
At both $\tau=0.25$ and $\tau=0.50$, two groups of LADs are identified, although the membership of the groups is not exactly the same\footnote{The membership of the two groups at $\tau=0.50$ is similar to that of the two groups at $\tau=0.25$ with a large number of overlapping member LADs (e.g. for Group 1, there are 45 overlapping LADs and for Group 2, there are 247 overlapping LADs).}. Furthermore, Figures (ref) and (ref) show that the coefficient functions for each group have similar patterns at these two quantiles, especially for that of population growth. The coefficients are mostly positive, consistent with the economic theory that growth in population and income leads to growth in demand for housing and hence, a rise in house price. At $\tau=0.25$, the effects of population and income growth in general decrease as inflation increases. This decreasing trend is more marked for the effect of population growth in Group 1 LADs (88% of which are non-metropolitan districts or unitary authorities), but much less so in Group 2 LADs (which include 88% of the London boroughs and 89% of the metropolitan districts). For example, when inflation rate is -0.5%, for each 1% increase in population growth, house price growth is increased by around 8%. But when inflation rate is 0.5%, a 1% increase in population growth leads only to a 2% increase in house price growth. At $\tau=0.50$, we observe a similar trend. We can also find, from Figures (ref) and (ref), that at $\tau=0.25$ and $\tau=0.50$, the effect of population growth on house price growth dominates that of income growth for Group 1 LADs, which are mainly non-metropolitan or unitary districts. At $\tau=0.75$, where five groups are identified, we see some different trends for different groups (see Figure (ref)). While most values of the coefficient functions are positive, there are negative values for some groups in some subintervals. For example, the coefficient function of population growth for Group 1 exhibits a downward sloping pattern with increasing inflation and is negative when inflation is above $-0.7\%$. A closer examination of the member LADs reveals that three out of the four LADs of this group are urban districts with major or minor conurbation, where immigration is more likely to occur. If population growth is driven by migration inflows, its effect on house price growth might be ambiguous and sometimes even negative Sa15. CSZ22 also find negative effects of population growth on house price growth in some LADs.
For comparison, we also conduct a grouping analysis based on the functional-coefficient mean regression for the same variables. The same ratio criterion and clustering algorithm are used for choosing the number of groups and estimating the group membership. The obtained results are similar to those from the quantile analysis at $\tau=0.50$: two groups are found with similar, although not exactly the same, membership to those from the quantile regression at $\tau=0.50$. Results from the mean regression are more susceptible to the influence of outliers. More detail about the mean regression grouping results can be found in Appendix D of the online supplement.
The above analysis reveals that at $\tau=0.25$ and $\tau=0.50$, there is more homogeneity in the effects of population and income growth on house price growth across LADs, and in general these effects are positive and decrease as inflation increases. At $\tau=0.75$, more heterogeneity is observed, and for some identified groups the effects of population and income growth are negative for some values of inflation. It also appears that at $\tau=0.75$, the effect of income growth is larger than at lower quantiles. This empirical application demonstrates the benefit of the quantile grouping analysis: it can shed more light on the impact of population and income growth on house price growth across LADs than a mean regression based grouping analysis does. The results may be useful for policy makers in developing more targeted policies for specific groups of districts at specific quantiles.
In this paper, we propose a general functional-coefficient quantile regression model for large panel data and assume a latent group structure on the heterogenous functional coefficients. An estimation methodology which combines preliminary functional coefficient estimates (ignoring the latent group structure), an agglomerative clustering algorithm and a simple ratio criterion is introduced to consistently estimate the group number and membership. Furthermore, a post-grouping local linear quantile regression method is used to estimate the group-specific functional coefficients, aiming to achieve faster convergence rates than the preliminary local linear estimator. To asymptotically remove the influence of nuisance parameters and derive an asymptotic normal distribution theory comparable to that in the literature such as KGM12, we impose a relatively restrictive condition on the divergence rate of the group size, but allow weak cross-sectional and temporal dependence for large panel observations. The simulation studies show that the proposed methods have reliable finite-sample performance. The empirical application to the UK house price data reveals that the latent structures vary over different quantile levels with more heterogeneity observed at the upper quartile.
The authors would like to thank an Associate Editor and two reviewers for the constructive comments, which helped to substantially improve the paper. The authors also thank Dr Chaowen Zheng for collecting the empirical data and preparing Figures (ref)-(ref). The first author is partially supported by Natural Science Foundation of Zhejiang Province (No. LY22A010006), National Social Science Foundation of China (No. 17BTJ027) and the Fundamental Research Funds for Provincial Universities of Zhejiang. The second author is partially supported by the Economic and Social Research Council in the UK (No. ES/T01573X/1). The third author is partially supported by the National Natural Science Foundation of China (No. 72033002). The usual disclaimer applies.
The supplemental document contains proofs of the main asymptotic theorems, technical lemmas with proofs, some extensions of the developed method and theory, and additional empirical result.
{
}
\if00 {
} \fi
\if10 {
} \fi
\spacingset{1.68}
In this supplement, we prove the main asymptotic theorems in Appendix A, prove some technical lemmas in Appendix B, apply the modified methodology to identify the latent group structure (uniformly over quantile levels) in linear panel quantile regression in Appendix C.1, discuss the methodology and theory for the time-varying coefficient panel quantile regression with the latent structure in Appendix C.2, construct the point-wise confidence intervals in Appendix C.3, and report the empirical result of functional-coefficient mean regression based grouping analysis in Appendix D.