Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
122,616 characters · 14 sections · 5 citation commands
Panel Data Quantile Regression with Grouped Fixed Effects
\pagestyle{myheadings} \markboth{\sc Panel Quantile regression with group fixed effects}{\sc Gu and Volgushev}
It is widely accepted in applied Econometrics that individual latent effects constitute an important feature of many economic applications. When panel data are available, a common approach is to incorporate latent structures in completely nonrestrictive way, i.e. the fixed effect approach. The fixed effect approach is attractive as it imposes minimal assumptions on the structure of the latent effects and on the correlation between the latent effects and the observed covariates and hence has become a very common empirical tool (see \citeasnoun{Hsiao} for a textbook treatment).
A major challenge of the fixed effects approach lies in the fact that it introduces a large number of parameters which grows linearly with the number of individuals. For a few specific models this can be avoided by differencing out individual effects and learning about the common parameter of interest. However, for most models, including quantile regression, this simple differencing method no longer exists. The literature contains various approaches that put additional structure on latent effects in order to reduce the number of parameters and obtain more interpretable models. One popular approach is to introduce some parametric distributional structure on the latent effects, see for example \citeasnoun{Mundlak78}, \citeasnoun{Chamberlain1982} and the correlated random effects literature. An alternative is to assume that the fixed effects have a group structure and hence only take a few distinct values which is the approach we take in this paper.
There is ample evidence from empirical studies that it is often reasonable to consider a number of homogeneous groups (clusters) within a heterogeneous population. This discrete approach was taken by \citeasnoun{HeckmanSinger82} for duration analysis of unemployment spells of a heterogeneous population of workers. \citeasnoun{BesterHansen} argue that in many applications individuals or firms are grouped naturally by some observable covariates such as classes, schools or industry codes. It is also widely accepted in the discrete choice model literature that individual agents are classified as a number of latent types (for instance \citeasnoun{Keane} among many others).
Estimating cluster structure has a long history in Statistics and Economics, and has generated a rich and mature literature. A general overview is given in \citeasnoun{Kaufman}. Among the many available clustering algorithms, the k-means algorithm (\citeasnoun{MacQueen}) is one of the most popular methods. It has been successfully utilized in many economic applications, for instance \citeasnoun{LinNg}, \citeasnoun{BM} and \citeasnoun{AndoBai}. Finite mixture models provide an alternative, likelihood based approach. In the latter, grouping is usually achieved by maximizing the likelihood of the observed data. \citeasnoun{Sun} builds a multinomial logistic regression model to infer the group pattern while nonparametric finite mixture models are considered in \citeasnoun{Allman} and \citeasnoun{Kasahara} among many others.
The focus of the present paper is on quantile regression for panel data with grouped individual heterogeneity. Panel data quantile regression has recently attracted a lot of attention, and there is a rich and growing literature that proposes various approaches to dealing with individual heterogeneity in this setting. In a pioneering contribution, \citeasnoun{Koenker2004} takes the fixed effect approach and introduces individual latent effects as location shifts. These individual effects are regularized through an $\ell_1$ penalty which shrinks them towards a common value. \citeasnoun{Lamarche} proposes an optimal way to choose the corresponding penalty parameter in order to optimize the asymptotic efficiency of the common parameters of the conditional quantile function, see \citeasnoun{harding2017} for an extension of this approach. Another line of work that focuses on estimating common parameters while putting no structure on individual effects includes \citeasnoun{Kato}, \citeasnoun{galvao2015} and \citeasnoun{galvao2016}. Alternative approaches have also emerged. \citeasnoun{AD} take a random effect view of these individual latent effects. They consider a correlated random-effect model in the spirit of \citeasnoun{Chamberlain1982} where the individual effects are modeled through a linear regression of some covariates. This is further developed in \citeasnoun{AB16} and \citeasnoun{CLP} where the conditional quantile function of the unobserved heterogeneity is modelled as a function of observable covariates.\footnote{For related literature on non-separable panel data models see \citeasnoun{Evdokimov}, \citeasnoun{Chernozhukov} and the references therein.}
Our contribution, which builds upon \citeasnoun{Koenker2004}, is a linear quantile regression method that accommodates grouped fixed effects. The advantages of our proposal over existing proposals are twofold. First, grouped fixed effects maintain the merit of unrestricted correlation between the latent effects and the observables and strike a good balance between the classical fixed effects approach and the other extreme which completely ignores latent heterogeneity. Second, in contrast to \citeasnoun{Koenker2004}, where the fixed effects are treated as nuisance parameters and are regularized to achieve a more efficient estimator for the global parameter, our method allows the researcher to learn the particular group structure of the latent effects together with common parameters of interest in the model. To the best of our knowledge, panel data quantile regression with grouped fixed effects has not been considered in the literature before. The only paper that goes in this direction is \citeasnoun{su2016}. While the general framework developed in this paper does include a version of quantile regression with smoothed quantile objective function, the theoretical analysis requires the smoothing parameter to be fixed. This results in a non-vanishing bias and hence does not correspond to quantile regression in a strict sense.
We do not assume any prior knowledge of the group structure and combine the quantile regression loss function with the recently proposed convex clustering penalty of \citeasnoun{Hocking}. The convex clustering method introduces a $\ell_1$-constraint on the pair-wise difference of the individual fixed effects, which tends to push the fixed effects into clusters. The number of clusters is controlled by a penalty parameter. The resulting optimization problem remains convex and can be solved in a fast and reliable fashion. Further modifications and a theoretical analysis of convex clustering were considered in \citeasnoun{zhu2014}, \citeasnoun{TanWitten} and \citeasnoun{Radchenko}. All of those authors combine $\ell_1$ penalties with the classical $\ell_2$ loss, and only consider clustering for cross-sectional data. Their theoretical results are not directly applicable to panel data or the non-smooth quantile loss function which is the main objective in this paper (all of the available theoretical results explicitly make use of the differentiability of the $\ell_2$ loss function in their proofs).
Our main theoretical contribution is to show consistency of the estimated grouping for a suitable range of penalty parameters when $n$ and $T$ tend to infinity jointly. We also propose a completely data-driven information criterion that facilitates the practical implementation of the method and prove its consistency for group selection as well as asymptotic normality of the resulting parameter estimators.
The remaining part of this paper is organized as follows. Section 2 contains a detailed description of the proposed methodology and provides details on its practical implementation. Assumptions and theoretical results are included in Section 3. Section 4 presents the convex optimization problem and its computational details. Monte Carlo simulation results are included in Section 5 where investigate the final sample behavior of the proposed methodology. In Section 6 we apply the method to an empirical application in studying the effect of the adoption of Right-to-Carry concealed weapon law on violent crime rate using a panel data of 51 U.S. states from 1977 - 2010. All proofs are collected in Section (ref) while additional simulation results and details for the empirical application are relegated to the Appendix.
Assume that for individuals $i=1,...,n$ we observe repeated measures $(X_{it},Y_{it})_{t=1,...,T}$ where $X_{it}$ denote covariates and $Y_{it}$ are responses.\footnote{Here, $T$ is assumed to be the same across individuals for notational simplicity. All results that follow can be extended to individual-specific values of $T_i$ as long as the ratio $(\max_{i=1,...,n} T_i)/(\min_{i=1,...,n} T_i)$ is uniformly bounded. In this case the theory goes through without changes if all instances of $T$ are replaced by $n^{-1} \sum_{i=1,...,n} T_i$} We shall maintain the assumption that data are i.i.d. within individuals and independent across individuals. The main object of interest in this paper is the conditional $\tau$-quantile function of $Y_{i1}$ given $X_{i1}$, which we will denote by $q_{i,\tau}$. We assume that $q_{i,\tau}$ is of the form \[ q_{i,\tau}(x) = \beta_0(\tau)^\top x + \alpha_{0i}(\tau), \quad i=1,....,n \] with individual fixed effects $\alpha_{0i}(\tau)$ taking only a finite number, say $K$, of different values, say $\alpha_{(01)}(\tau),...,\alpha_{(0K)}(\tau)$.\footnote{We follow \citeasnoun{Koenker2004} in treating the $\alpha_i$ as fixed parameters. An alternative approach which leads to equivalent results is to treat the $\alpha_i$ as random (with no restrictions placed on the dependence with $X_{it}$). In this case the model can be written as $Q_{Y_{it}|X_{it},\alpha_i(\tau)}(\tau) = \beta_0(\tau)^\top x + \alpha_{0i}(\tau)$; here $Q_{Y_{it}|X_{it},\alpha_i(\tau)}$ denotes the conditional quantile function of $Y_{it}$ given $(X_{it},\alpha_i(\tau))$ (see for instance \citeasnoun{Kato}, \citeasnoun{galvao2015} and \citeasnoun{galvao2016} for this interpretation). Both interpretations lead to the same asymptotic results.} We explicitly allow the group membership, and even the number of groups to be unknown and to depend on $\tau$ but will not stress this dependence in the notation for the sake of simplicity. Our main objective is to jointly estimate the number of groups, unknown group structure, and parameters $\alpha_{(01)},...,\alpha_{(0K)},\beta$ from the observations. To achieve this, we consider penalized estimators of the form
where\footnote{As pointed out by a Referee, one could also consider combining the objective functions corresponding to several quantiles as was done in \citeasnoun{Koenker2004} and force all coefficients $\alpha_i$ to be independent of $\tau$. This would result in efficiency gains if all $\alpha_i$ are purely location-shift effects but can introduce bias otherwise. We leave this extension for future research.} \[ \Theta(\alpha_1,...,\alpha_n,\beta) := \sum_{i,t} \rho_\tau(Y_{it} - X_{it}^\top \beta - \alpha_i) + \sum_{i\neq j} \lambda_{ij} |\alpha_i - \alpha_j|. \] Here $\rho_\tau$ denotes the usual 'check function' and the weights $\lambda_{i,j}$ are allowed to depend on $n,T$ and the data; one particular choice is discussed in below. The form of the penalty is motivated by the work of \citeasnoun{Hocking}. Intuitively, large values of $\lambda_{ij}$ will push different coefficients closer together and result in clustered structure of the estimators $\hat \alpha_i$. High-level conditions on the weights $\lambda_{i,j}$ which guarantee consistency of the resulting grouping procedure are provided in Theorem (ref).
There are various possible choices for the penalty parameters $\lambda_{i,j}$. We propose to use weights of the form
where $(\check\alpha_1,...,\check\alpha_n)$ are the fixed effects quantile regression estimators
of \citeasnoun{Kato}\footnote{As pointed out by a Referee, an alternative approach to obtain preliminary estimators for $\alpha_{i0}$ would be to run separate quantile regressions for each individual. This did not improve the performance of our procedure in the simulations that we tried.} and $\lambda$ is a tuning parameter. This form of weighting by preliminary estimators is motivated by the work of \citeasnoun{zou2006} on adaptive lasso. Intuitively, weighting by preliminary estimated distances tends to give smaller penalties to coefficients from different groups thus reducing some of the bias that is typically present in the classical lasso.
Given the developments above, it remains to find a value for the tuning parameter $\lambda$. The high-level results in Theorem (ref) together with findings in \citeasnoun{Kato} provide a theoretical range for those values (see the discussion following Theorem (ref) for additional details), but this range is not directly useful in practice since only rates and not constants are provided. Moreover, despite the fact that the weights $\check \lambda_{ij} := \lambda |\check \alpha_i - \check\alpha_j|^{-2}$ lead to asymptotically unbiased estimates, bias can still be a problem in finite samples. A typical approach in the literature to reduce bias which results from lasso-type penalties is to view the lasso problem solution as a candidate model (in our case, a candidate grouping of $\alpha_i$) and re-fit based on this candidate model (see \citeasnoun{belloni2009} or \citeasnoun{su2016} among many others).
To deal with bias issues and the choice of $\lambda$ in practice, we propose to combine the re-fitting idea with a simple information criterion which will simultaneously reduce the bias problem and provide a simple way to select a final model. A formal description of our approach is given in Algorithm (ref).
Theorem (ref) provides a formal justification of Algorithm (ref) under high-level conditions on $\hat C, p_{n,T}$. In particular we prove that the group structure is estimated consistently with probability tending to one and derive the asymptotic distribution of the resulting estimators $\hat \alpha_i^{IC}, \hat \beta^{IC}$. In order to make the proposed estimation procedure fully data-driven, we need to specify a choice for the tuning parameters $\hat C$ and $p_{n,T}$. In our simulations, we found that the following choices lead to good results\footnote{The exact constant $1/10$ in the factor $p_{n,T}$ does not matter asymptotically. The value $1/10$ was found to work well for a wide range of values of $n,T$ and for various models, details are provided in the Monte Carlo section (ref). There we also show that the impact of the precise form of the factor in $\hat p_{n,T}$ becomes less pronounced as $T$ increases}:
with \[ \hat s(\tau) := (\hat F^{-1}(\tau + h_{n,T}) - \hat F^{-1}(\tau - h_{n,T}))/(2h_{n,T}) \] where $\hat F(y) := \frac{1}{nT}\sum_{i,t} I\{Y_{it} - X_{it}^\top \check\beta - \check\alpha_{i}\leq y\}$ denotes the empirical cdf of the regression residuals from the fixed effects quantile regression estimator given in (ref), $\hat F^{-1}$ denotes the corresponding empirical quantile function, and $h_{n,T} \to 0$ is a bandwidth parameter (we use the Hall-Sheather rule in our simulations, see \citeasnoun{koenker2005}).
To motivate this particular choice of constant $\hat C$, observe the following expansion, which is derived in detail in the proof of Theorem (ref) \[ \sum_{i,t} \rho_\tau(Y_{it} - X_{it}^\top \check\beta - \check\alpha_{i}) - \rho_\tau(Y_{it} - X_{it}^\top \beta_0 - \alpha_{0i}) = - \sum_i \frac{\tau(1-\tau)}{2 \mathbb{E}[f_{Y_{i1}|X_{i1}}(q_{i,\tau}(X_{i1})|X_{i1})]}+ o_P(n). \] This shows that plugging in the estimated (by fixed effects quantile regression) instead of true errors underestimates the objective function evaluated at the residuals by roughly the first term on the right-hand side in the above expression. This term needs to be dominated by the penalty if we want to avoid selecting models that are too large, and so it is natural to scale the penalty by a constant which is proportional to $\sum_i 1/\mathbb{E}[f_{Y_{i1}|X_{i1}}(q_{i,\tau}(X_{i1})|X_{i1})]$ in order to ensure reasonable performance across different data generating processes. Under the simplifying assumption that $f_{Y_{i1}|X_{i1}}(q_{i,\tau}(X_{i1})|X_{i1}) =: f_\varepsilon(0)$ does not depend on $i, X_{i1}$, this term equals $n/f_\varepsilon(0)$, and under the same assumptions $\hat s$ provides a consistent estimator for the latter, see \citeasnoun{koenker2005}. Note that the sparsity term introduced here plays a similar role as the noise variance in classical information criteria such as AIC and BIC in least squares regression.
In this section we provide a theoretical analysis of the methodology proposed in Section (ref). We begin by stating an assumption on the true (but unknown) underlying group structure.
Assumption (C) implies that the individual fixed effects are grouped into $K$ distinct groups and that the group centers are separated. Note that the number of groups as well as group membership is allowed to differ across quantiles. For the sake of a concise notation, the dependence of the number of groups and group centers on $\tau$ will from now on be dropped unless there is risk of confusion. Note also that we require the number of groups to be fixed (i.e. independent of $n,T$ and non-random) and exogenous, i.e. independent of the covariates $X_{it}$.
Next we collect some technical assumptions on the data generating process. Define $Z_{it}^\top = (1, X_{it}^\top)$ and let $\mathcal{Z}$ denote the support of $Z_{it}$.
Assumptions (A1)-(A3) are fairly standard and routinely imposed in the quantile regression literature. Similar assumptions have been made, for instance in \citeasnoun{Kato} [see assumptions (B1)-(B3) in that paper].
To state our first main result define \[ \Lambda_D := \sup_{i \in I_k, j \in I_{k'}, k\neq k'} \lambda_{i,j}, \quad \Lambda_S := \inf_k \inf_{i,j \in I_k} \lambda_{i,j}. \] In words, $\Lambda_D$ corresponds to the largest penalty corresponding to the difference between two individual effects from different groups while $\Lambda_S$ describes the smallest penalty between two effects from the same group. Our first result provides high-level conditions on $\Lambda_S, \Lambda_D$ that guarantee asymptotically correct grouping.
Next we discuss the implications of this general result for the specific choice $\check \lambda_{i,j}$ given in (ref). Define \[ \check\Lambda_D := \sup_{i \in I_k, j \in I_{k'}, k\neq k'} \check\lambda_{i,j}, \quad \check\Lambda_S := \inf_k \inf_{i,j \in I_k} \check\lambda_{i,j}. \] From \citeasnoun{Kato} \footnote{more precisely, from the paragraph following equation (A.14) in the latter paper; note that this result is derived under the assumption that $n$ grows at most polynomially with $T$} we obtain the bound \[ \sup_{i=1,...,n} |\check \alpha_i - \alpha_{0i}| = O_P((\log n)^{1/2}/T^{1/2}). \] Now if $i,j \in I_k$ then $\alpha_{0i} = \alpha_{0j}$ and thus \[ 1/\check\Lambda_S = \Big\{\inf_{k}\inf_{i,j \in I_k} |\check\alpha_{i} - \check\alpha_{j}|^{-2} \lambda\Big\}^{-1} = \lambda^{-1} \sup_{k}\sup_{i,j \in I_k} |\check\alpha_{i} - \check\alpha_{j}|^2 = O_P\Big(\frac{\log n}{\lambda T}\Big). \] Moreover, under (C) we have \[ \inf_{k\neq k'} \inf_{i \in I_k, j \in I_{k'}} |\alpha_{0i} - \alpha_{0j}| \geq \varepsilon_0 > 0 \] and thus \[ \check \Lambda_D \leq \lambda \Big\{ \inf_{k\neq k'} \inf_{i \in I_k, j \in I_{k'}} |\check \alpha_{i} - \check\alpha_{j}|\Big\}^{-2} \leq \lambda/(\varepsilon_0 - o_P(1))^2 = O_P(\lambda). \]
Given this choice of weights, the conditions $\frac{\Lambda_D}{\Lambda_S} = o_P(1)$ is satisfied provided that $T/\log n \to \infty$. The other conditions in (ref) take the form \[ T^{1/2} \gg n\lambda \gg T^{-1/4} (\log n)^{7/4}. \] Assuming that $(\log n)^{7/3} = o(T)$, this provides a range of possible values for $\lambda$ which will ensure that (ref) holds.
In this section we provide theoretical guarantees for the performance of the information criterion based estimators $\hat \beta^{IC}, \hat \alpha_k^{IC}, \hat{I}_k^{IC}, \hat K^{IC}$ Our main result shows that, under fairly general conditions on the penalty parameter $p_{n,T}$, the IC procedure selects the correct number of groups with probability tending to one. Moreover, the estimators $(\hat\alpha_1^{IC},...,\hat\alpha_{\hat K^{IC}}^{IC},\hat\beta^{IC})$ are shown to enjoy the 'oracle property', i.e. they have the same asymptotic distribution as estimators which are based on the true (but unknown) grouping of individuals. Before making this statement more formal, we need some additional notation. Let
denote the infeasible 'oracle' which uses the true group membership. The asymptotic variance of the oracle estimator is conveniently expressed in terms of the following two limits which we assume to exist\footnote{Appropriate forms of asymptotic normality of the oracle and IC estimators continue to hold without assuming that the limits exist. This assumption is made for notational convenience.}
where $\tilde Z_{ik} := (e_k^\top, X_{i1}^\top)^\top$ and $e_k$ denotes the $k$'th unit vector in $\mathbb{R}^K$. Additionally, we need the following condition on the grid $\lambda_1,...,\lambda_L$
Assumption (G) is fairly mild. It only requires that among the candidate values for $\lambda$ there exists one value so that $\check\lambda_{i,j}$ satisfies the assumptions of Theorem (ref). In practice, we recommend choosing a grid of values that results in sufficiently many different numbers of groups.
To implement the proposed quantile panel data regression with group fixed effect, we need to solve the optimization problem stated in ((ref)). A natural normalization of the objective and the penalty function leads to
this is equivalent to the objective ((ref)) except that $\tilde \lambda$ is adjusted according to $n$ and $T$ so that we can use a generic grid for $\tilde \lambda$ rather than letting the grid support change with $(n, T)$. In practice, the grid support of $\tilde \lambda \in \{0, \tilde \lambda_1, \dots, \tilde \lambda_{\ell},\dots, \tilde \lambda_{L}\}$ is chosen such that the number of distinct values of the solution $\{\hat \alpha_{1, \ell}, \dots, \hat \alpha_{n, \ell}\}$ for $\ell = 1, \dots, L$ takes all possible integer values in the set $\{1, 2, \dots, n\}$. This is always achievable as long as the grid width of $\tilde \lambda$ is small enough. Since for each fixed $\tilde \lambda_{\ell}$, ((ref)) is a linear programming problem which can be efficiently solved in any reliable solvers, this is not computationally expensive.
To make this section self contained, we provide some details of the primal problem stated in ((ref)) and its corresponding dual problem. Define $\lambda_{ij} := |\check{\alpha_i} - \check{\alpha_j}|^{-2}$ and observe that we can re-write $|\alpha_i - \alpha_j| = 2\Big (\frac{1}{2} - 1\{0 \leq \alpha_i - \alpha_j\}\Big) \Big(0 - (\alpha_i - \alpha_j)\Big)$. With this notation (ref) can be equivalently expressed as follows
with $\boldsymbol{\theta}$ being a vector of length $\frac{n(n-1)}{2}$ that consists entries $(\alpha_i- \alpha_j) \lambda_{ij}$ for $i < j$. We can represent $\boldsymbol{\theta}$ as $A \boldsymbol{\alpha}$ where $A$ is a $\frac{n(n-1)}{2}\times n$ matrix taking the form \[ A =
\] The corresponding dual problem of ((ref)) can be stated as:
with $Z$ being the incidence matrix that identifies the $n$ individuals. The solution for $\boldsymbol{\alpha}$ and $\boldsymbol{\beta}$ in the primal problem is then the dual solutions of the dual problem. We implement the dual problem using the Mosek optimization software of \citeasnoun{mosek} through the R interface Rmosek of \citeasnoun{Rmosek}. We have also implemented the estimation procedure using the quantreg package in R and the code will be made available for public use.\footnote{Mosek is a commercial state-of-the-art convex optimization solver that provides a free academic license. We use its interior point algorithm to solve our linear programming problem. The estimation procedure implemented using the quantreg package calls the sfn method, which uses the Frisch-Newton algorithm and exploits the sparse algebra to compute iterates.}
To assess the finite sample performance of the proposed convex clustering panel quantile regression estimator, we apply the method to simulated data sets. In particular, we consider data generated from two models and two error distributions for a range of $n$ and $T$. The responses, $Y_{it}$, are generated by either a location shift model
or a location scale shift model
where the individual latent effects $\alpha_i$ are generated from three groups taking values $\{1,2,3\}$ with equal proportions. The covariate $X_{it}$ is generated such that it has a non-zero interclass correlation coefficient. In particular, \[ X_{it} =\rho \alpha_i + \gamma_i + v_{it} \] with $\gamma_i$ and $v_{it}$ independent and identically distributed over $i$ and $i,t$ respectively. We conduct simulation experiments with $\rho \in \{0, 0.5\}$ to investigate both cases where the fixed effect is independent or correlated with the covariate.\footnote{\citeasnoun{Koenker2004} used a similar data generating process with $\rho = 0$ for $X$ and pointed out that the interclass correlation induced by $\gamma_i$ is crucial for the penalized quantile regression fixed effect estimator to have superior performance than the unpenalized QRFE estimator. } The true parameters are $\beta = 1$ and $\gamma = 1/10$. The error terms $u_{it}$ are i.i.d. following either a standard normal distribution or a student t distribution with three degrees of freedom. Results reported are based on 2000 repetitions.
We first investigate the performance of using the information criterion for estimating the number of groups. Table (ref) and Table (ref) report the proportion of estimated number of groups under the two models for different combinations of $n$ and $T$ for $\tau = 0.5$ and $\tau = 0.75$, respectively. Throughout the simulations, we used an equally spaced grid of $\lambda$ values with width $1/200$ and support $[0, 0.35]$. This grid was chosen to ensure that the number of groups estimated for different $\lambda$ values covers the integers in range $[1, n]$. For the IC criteria, the sparsity function $\hat s(\tau)$ is estimated with bandwidth chosen based on the \citeasnoun{HallSheather} rule implemented in the quantreg package (see discussion in \citeasnoun{koenker2005}).
The results suggest that the probability of getting the correct number of groups for $\tau = 0.5$ is slightly better than for $\tau = 0.75$. When $T \geq 30$, the estimates for the number of groups for both error distributions and for different $n$ are mostly satisfactory. The performance for $t$ error deteriorates compared to those with normal error, especially for higher quantiles. For $T = 15$ and quantiles other than the median the proposed method should be used with caution. Including correlation between individual effects and predictors does not lead to dramatic changes in the accuracy for estimating the number of groups and group membership.
Table (ref) and Table (ref) summarize the finite sample properties of $\hat \beta^{IC}(\tau)$ for $\tau = 0.5$ and $\tau = 0.75$, respectively and compare with the QRFE estimator where no penalization on the individual fixed effect is used (i.e. $\lambda = 0$).
The standard errors used for constructing confidence intervals (nominal coverage $95\%$) {are based on the nid option with Hall-Sheather bandwidth in the package quantreg}. Results based on the Bofinger bandwidth selection are similar and not reported here. For the QRFE, the Bofinger bandwidth rule was used since the Hall-Sheather rule resulted in substantial under-coverage with $T=30$ for some of the models.
When covariates and fixed effects are independent, the RMSE of the PQR-FEgroup estimator, $\hat \beta(\tau)$, is smaller than that of the QRFE for all settings considered. This shows that our penalization gains efficiency for estimating $\beta$ when there is group structure in the fixed effects. The results do not change much from normal error to $t$ error and from median to higher quantiles.
Introducing correlation between predictors and group membership leads to a bias for the grouped effect estimator, while the fixed effects estimator does not suffer from additional bias. This bias can be quite noticeable for small values of $T$, especially at the $75\%$ quantile. The bias becomes negligible as $T$ increases, so there is no contradiction to our asymptotic theory. An intuitive explanation for this behaviour is that for smaller $T$ it is difficult to get a perfect grouping, and a wrong grouping leads to bias since there is dependence between predictors and group structure.
Last, we report in Table (ref) and Table (ref) the proportion of perfect classification of individual effects and the average value of the percentage of correct classification together with their standard errors. Since the comparison of the estimated membership and the true membership only makes sense when $\hat K = K_0$, the estimated membership are based on $\lambda$ for which $\hat K = K_0$ (see \citeasnoun{su2016} for a similar approach). Results suggest that for $T\geq 30$ and $\tau = 0.5$, the group membership estimation is quite satisfactory. While the proportion of perfect matches is low even for $T = 30$, the average proportion of correct classification shows that those effects are typically due to very few misclassified individuals. For $T= 15$ perfect classification is almost impossible while average correct classification rates remain reasonable. Adding correlation between individual effects and covariates leads to a deterioration of the probability for achieving a perfect grouping for location-scale models, especially at higher quantiles, but does not have a strong impact on other results. Overall the simulations suggest that for small $T$, there is just not enough information available for each individual to hope for perfect classification.
The discussion at the end of Section (ref) provides a motivation for the tuning parameters $\hat C$ in the IC criteria. Here we further investigate the impact of rescaling $p_{n,T}$ by different factors. For illustration we consider DGP1 in the previous section where data are generated based on the model ((ref)). We use a grid of constants $c \in [0.01, 0.3]$ with width 0.01 and plot the associated performance of the estimated number of groups, the RMSE of $\hat \beta^{IC}(\tau)$ and the coverage rate for $p_{n,T} = cnT^{1/4}$. Figure (ref) and Figure (ref) contain corresponding results for the location-scale shift model with $t$ errors and $\tau=0.5,0.75$, respectively. For $T$ as small as 15, the performance is quite sensitive to the chosen constant. As predicted by the theory this dependence becomes somewhat less prominent as $T$ increases. Overall the choice $p_{n,T} = nT^{1/4}/10$ shows good performance for settings that we tried in this simulation. The patterns are similar for those with the normal error and the location-shift models and likewise for DGP2, results are reported in Figure (ref)- Figure (ref) in the Appendix for the sake of completeness.
To further illustrate our proposed methodology, we revisit the empirical inquiry on the much-debated "More Guns Less Crime" hypothesis. \citeasnoun{LM97} provide the first empirical analysis which claims that the adoption of Right-to-Carry (RTC) laws, which allows local authorities to issue a concealed weapon permit to all applicants that are eligible, reduces crime. Ever since its publication, there has been much academic and political debate that challenges the findings. \citeasnoun{AD03} shows that the negative effect of RTC laws has no statistical significance under a more reasonable model specification and inference using both the state and county level data between 1977 - 1999 for 51 U.S. states. Their conclusion is echoed by the National Research \citeasnoun{NRC2004} (NRC) report which finds little reliable statistical support for the “More Guns Less Crime" hypothesis. Recently, \citeasnoun{ADZ14} revisited the hypothesis using the updated panel data from 1977 - 2010. They correct several mistakes in the dataset used in earlier analysis and discuss the shortcoming of using county level crime data. We refer the readers to more details provided in \citeasnoun{ADZ14}. For our empirical analysis, we use their updated state level data kindly provided by the authors.\footnote{The data and the detailed description of the data source can be downloaded from \url{https://works.bepress.com/john_donohue/107/}.}
The main parameter of interest is the effect of the indicator of the RTC laws (denoted as lawind) on crime rates. In our analysis, we focus on the violent crime rate. A similar analysis is possible for other categories of offense. In addition to the law indicator, there is also information on the incarceration rate (sentenced prisoners per 100,000 residents; denoted as prisoner) in the states in the previous year, real per capita personal income (denoted as rpcpi) and other demographic variables on population proportions in different age-gender groups for various ethnicities. As argued in \citeasnoun{AD03} and \citeasnoun{ADZ14}, to avoid multicollinearity and accounting for the fact that 90% of violent crimes in the U.S. are committed by male offenders, we follow their specification and control for only the proportion of African-American male in the age group 10-19, 20-29 and 30-39. Inevitably, there are many different model specifications that can be considered for this empirical inquiry and it is impossible to report all the results. Our goal is to emphasize the heterogeneous law adoption effect for states that have low violent crime rate versus those with high crime rate. We also provide evidence of clustering behavior of the state fixed effects which are incorporated to capture states' unobserved heterogeneity and show that by taking advantage of the dimension reduction of grouping the fixed effects, the parameters of interest on other control variables enjoy better statistical precision.
Our main model specification is
where the additional control variables $X_{it}$ include lagged incarceration rate (prisoners), real per capita personal income (rpcpi) and the three demographic variables (afam1019, afam2029, afam3039). We note that high incarceration rates in a given state may be a feedback towards rising violent crime, and therefore including those as control variables might lead to endogeneity issues. As a robustness check, we also report the corresponding results in the appendix for model ((ref)) without the incarcerating rate as a control covariate. The effects stay mostly unchanged.
Figure (ref) reports the panel data quantile regression estimates with state fixed effects for $\tau \in \{0.1, 0.25, 0.5, 0.75, 0.9\}$ and their associated point-wise confidence intervals. Similar to the findings in \citeasnoun{ADZ14}, we see that the RTC law-adoption has a positive effects on violent crime rate. In addition this effect is significant for lower quantiles while there is no statistical significance for such an effect for states at higher quantiles of the violent crime rate. All other control variables have expected signs, with a negative effect of the lagged incarceration rate indicating that states with stricter laws have lower violent crime rates, although this effect is not very precisely estimated. Higher real per capita personal income has a negative effect on violent crime at lower quantiles; this effect is diminishing for higher quantiles of violent crime rate. The proportion of African-American population at age group 30 - 39 has a positive effect on the violent crime and this effect becomes more prominent for higher crime rate states. Figure (ref) shows the corresponding state fixed effect estimates for different quantile levels. There is some evidence of clustered behavior of these fixed effect estimates, although these are estimates with statistical errors. Using our proposed methodology, the estimated number of groups for the five quantile levels are respectively $\{20, 20,16,17,19\}$.
Figure (ref) plots the corresponding panel data quantile regression estimates with the estimated optimal grouped state fixed effects for $\tau \in \{0.1, 0.25, 0.5, 0.75, 0.9\}$ and their associated point-wise confidence intervals. The pattern of the quantile effects for all the control variables stay roughly the same compared to the fixed effect quantile regression estimates, while the variance of these estimates under the grouped fixed effects are noticeably smaller. This is similar to what we observe in the simulation section where the common parameters in quantile panel data regression with grouped fixed effects have lower variances than those of the estimates based on “individual heterogeneity". However, we would also like to point out that the standard errors are based on refitted models with estimated group structure, which does not account for the uncertainty of model selection on the fixed effects and should be interpreted with caution. A proper inference method based on the proposed methodology accounting for such uncertainty is important and is part of our future research agenda.
The present paper suggests a simple and computationally efficient way to incorporate group fixed effects into a panel data quantile regression by means of a convex clustering penalty. We develop theoretical results on consistent group structure estimation and discuss the asymptotic properties of the resulting joint and group-specific estimators.
There are several directions that we plan to explore in the future. First, our theory focused on individual fixed effects while assuming common slope coefficients. It is equally interesting to allow for group structure in some of the slope coefficients while keeping other slope coefficients common across individuals, perhaps even allowing for individual fixed effects. This can be achieved by straightforward modifications of the penalization approach which we explored so far, but a more detailed theoretical analysis of this approach remains beyond the scope of the present paper.
Second, one can take the standpoint that in many applications there is no exact group structure. In such settings, an alternative interpretation of the penalty which we investigated is as a way of regularizing problems that have too many parameters. Such an interpretation is in the spirit of the proposals of \citeasnoun{Koenker2004} and \citeasnoun{Lamarche}, and a detailed investigation of the resulting bias-variance trade-off warrants further research.
Finally, a deeper analysis of issues that are related to uniformity of distributional approximation in the entire parameter space was not addressed here, but remains an important theoretical and practical question which we hope to address in the future.
We begin by collecting some useful facts and defining additional notation. We will repeatedly make use of Knight's identity (see koenker2005, p. 121)) which holds for $u \neq 0$:
Additionally, let $\gamma_{0i} := (\alpha_{0i},\beta_0^\top)^{\top}$. The symbols $a_n \lesssim b_n, a_n \gtrsim b_n$ will mean that there exists a non-random constant $C \in (0,\infty)$ which is independent of $n,T,\tau$ such that $P(a_n \leq Cb_n) = 1$ and $P(a_n \geq Cb_n) = 1$, respectively. Define $\varepsilon_{it}^\tau := Y_{it} - Z_{it}^\top \gamma_{i0}(\tau)$ and let $F_{\varepsilon_{it}^\tau|X_{it}}(u|X_{it}) = F_{Y_{it}|X_{it}}( Z_{it}^\top \gamma_{i0}(\tau) + u | X_{it})$ denote the conditional cdf of $\varepsilon_{it}^\tau$ given $X_{it}$. When there is no risk of confusion, we will also write $\varepsilon_{it}$ instead of $\varepsilon_{it}^\tau$. Define $\psi_\tau(x) := (\1\{x\leq 0\} - \tau)$.
We begin by stating some useful technical results which will be proved at the end of this section.
Proof of Theorem (ref)
Step 1: first bounds In this step we shall prove that
Combine the results in Lemma (ref) and Lemma (ref) to find that any minimizer of $\Theta(\alpha_1,...,\alpha_n,\beta)$ must satisfy
Let $N_\Delta := \# \{i: \|\beta - \beta_0\|^2 + |\alpha_i - \alpha_{0i}|^2 \geq \Delta\}$. Then for any $0 < \Delta < \varepsilon^2$ \[ T N_\Delta \Delta = O_P( n s_{n,1} + \Lambda_D n^2), \] i.e. by Lemma (ref) \[ N_\Delta = n O_P((T^{-1/2}(\log n)^{1/2}) + \Lambda_D n/T) \Delta^{-1} \] and in particular $N_\Delta = o_P(n)$ as long as $\Delta \gg T^{-1/2}(\log n)^{1/2} + \Lambda_D n/T$. Provided that $\Lambda_D n/T = o_P(1)$ we obtain
Define $D_{n,T} := T^{-1/2}(\log n)^{1/2} + \Lambda_D n/T$. Next we will prove that, provided {$n \Lambda_D = o_P(T^{1/2}), T^{3/4}(\log n)^{3/4}/(n\Lambda_S) = o_P(1)$} also \[ \sup_i |\hat \alpha_i - \alpha_{i0}|^2 = O_P(D_{n,T}). \] To this end, it suffices to prove that $\sup_i |\hat \alpha_i - \alpha_{i0}|^2 = O_P(d_{n,T})$ for any $n\Lambda_S \gg d_{n,T} \gg D_{n,T}$. Define \[ \tilde\alpha_i = \left\{
\right. \] Define the set $E := \{i: \tilde\alpha_i = \hat\alpha_i\}$. Observe that
Thus
Now since $N_{d_n} = o_P(n)$ and by definition of $\tilde\alpha_n$ it follows that under (C) \[ \max_k \Big| \frac{|I_k \cap E|}{n\mu_k} - 1 \Big| = o_P(1), \] and since by assumption $\Lambda_D/\Lambda_S = o_P(1)$ we obtain \[ \sum_{i,j} \lambda_{i,j}\Big\{|\hat \alpha_i - \hat \alpha_j| - |\tilde \alpha_i - \tilde \alpha_j| \Big\} \gtrsim n\Lambda_S \sum_{i \in E^C} \{|\hat\alpha_{i} - \alpha_{0i}| - d_{n,T}^{1/2}\}. \]
Next we note that for any $i$ with $|\hat \alpha_i - \alpha_{0i}|^2 \geq (2+c_1/c_0)d_{n,T}$ we have
with probability tending to one by (ref) and the definition of $d_{n,T}$. For $i$ with $|\hat \alpha_i - \alpha_{0i}|^2 < (2+c_1/c_0)d_{n,T}$ note that by Lemma (ref)
where the $O_P$ terms are uniform in $i$. Thus
Summarizing we have proved that
Under the conditions {$n \Lambda_D = o_P(T^{1/2}), T^{3/4}(\log n)^{3/4}/(n\Lambda_S ) = o_P(1)$} the last line is strictly positive with probability tending to one unless $E^C = \emptyset$ with probability tending to one. Thus the proof of (ref) is complete.\\
Step 2: recovery of clusters with probability to one
To simplify notation, assume that individual $1,...,N_1$ belongs to cluster 1, individual $N_1+1,...,N_1+N_2$ to cluster 2 and so on. Since all cluster can be handled by similar arguments we only consider the first cluster. Let $\hat \alpha_{(1)},...,\hat \alpha_{(L)}$ denote the distinct values of $\hat \alpha_1,...,\hat\alpha_{N_1}$, ordered in increasing order, and let $n_{1,k} := \#\{i: \hat \alpha_i = \hat \alpha_{(k)}\}$. Again, to simplify notation assume w.o.l.g. that $\hat \alpha_1 = ... = \hat\alpha_{n_{1,1}} = \hat \alpha_{(1)}$. To prove the result, we proceed in an iterative way. We will prove by contradiction that $L=1$, i.e. all estimators of individuals from cluster 1 take the same value. Assume that $L\geq 2$.
We will now prove by contradiction that $n_{1,1} > N_1/2$. Assume that $n_{1,1} < N_1/2$. Define $\tilde\alpha_i = \hat\alpha_{(2)}$ for $i=1,...,n_{1,1}$ and $\tilde\alpha_i = \hat \alpha_i$ for $i > n_{1,1}$. By (ref) Lemma (ref) we find that
Next, observe that by construction, under (C) and using the fact that $\sup_i |\hat \alpha_i - \alpha_{i0}| = o_P(1)$,
From this we obtain
where the last inequality holds for sufficiently large $n,T$ since by assumption {$\Lambda_D/ \Lambda_S = o_P(1), n\Lambda_S \gg T^{3/4}(\log n)^{1/4}$} and since we assumed $n_{1,1} < N_1/2$ so that $N_1-n_{1,1} \geq N_1/2 \gtrsim n$. However, this is a contradiction to the fact that $\hat\alpha_1,...,\hat\alpha_n,\hat\beta$ minimizes $\Theta$.
In a similar fashion, one can prove that $n_{1,L} > N_1/2$. Just define $\tilde\alpha_{N_1},...,\tilde\alpha_{N_1-n_{1,L}+1} = \hat\alpha_{(L-1)}$ and proceed as above. Since $n_{1,L} + n_{1,1} \leq N_1$ and we have already proved that $n_{1,1} > N_1/2$ this leads to a contradiction with $L\geq 2$, and hence $L=1$. All other clusters can be handled in a similar fashion and that completes the proof of the second step. $\Box$
Proof of Lemma (ref) Apply Knight's identity (ref) to find that
Hence it follows that
Now by a Taylor expansion
so the bound on $\tilde r_{n,i}^{(1)}(a_1,a_2)$ is established. Next define the classes of functions
Note that the class of functions $\mathcal{G}_2$ has envelope function $F \equiv 1$. Thus by Lemma 2.6.15 and Theorem 2.6.7 of vdVW the class of functions $\mathcal{G}_2$ satisfies, for any probability measure $Q$, $N(\varepsilon,\mathcal{G}_2,L_2(Q)) \leq K(1/\varepsilon)^V$ for some finite constants $K,V$(here, $N(\varepsilon,\mathcal{G}_2,L_2(Q))$ denotes the covering number, see Section 2.1 of vdVW). Moreover, $\mathcal{G}_1 \subseteq \{g_1-g_2|g_1,g_2\in \mathcal{G}_2\}$, and elementary computations with covering numbers show that $N(\varepsilon,\mathcal{G}_1,L_2(Q)) \leq \tilde K(1/\varepsilon)^{\tilde V}$ for some finite constants $\tilde V, \tilde K$. Hence we find that by Theorem 2.14.9 of vdVW), for any $h > 0$, \[ P^* \Big( \sup_{g \in \mathcal{G}_1} \frac{1}{\sqrt{T}} \Big|\sum_t g(Y_{it},X_{it}) - \mathbb{E}[g(Y_{it},X_{it})]\Big| \geq h \Big) \leq \Big(\frac{Dh}{\sqrt{\tilde V}}\Big)^{\tilde V} e^{-2h^2} \] for some constant $D$ that depends only on $\tilde K$ (here, $P^*$ denotes outer probability). Letting $h = \sqrt{\log n}$ and applying the union bound for probabilities we obtain
Hence
Thus the bound on $\tilde r_{n,i}^{(2)}(a_1,a_2)$ follows and the proof is complete. $\Box$
Proof of Lemma (ref) Observe that by Knight's identity (ref)
Now under assumption (A2) $|F_{\varepsilon_{it}|X_{it}}(s|X_{it}) - F_{\varepsilon_{it}|X_{it}}(0|X_{it})| \leq s \overline{f'}$ a.s., and thus given (A1)
This shows the upper bound in (ref). For the lower bound, note that $s \mapsto F_{\varepsilon_{it}|X_{it}}(s|X_{it})$ is non-decreasing almost surely. Moreover, $f_{\varepsilon_{it}|X_{it}}(0|X_{it}) \geq f_{min}$ a.s. by (A3) and thus by (A2) and (A3) we have almost surely \[ \inf_{|s| \leq f_{min}/2\overline{f'}} f_{\varepsilon_{it}|X_{it}}(s|X_{it}) \geq \frac{f_{min}}{2}. \] Define $\delta_i := (\gamma - \gamma_{0i})\min\{1,f_{min}/(2M\overline{f'}\|\gamma - \gamma_{0i}\|)\}$. Noting that $s \mapsto F_{\varepsilon_{it}|X_{it}}(s|X_{it})$ is non-decreasing almost surely, it follows that a.s.
where the last inequality follows since by definition $|\delta_i^\top Z_{it}| \leq f_{min}/(2\overline{f'})$ a.s. Finally, under assumption (A1), $\mathbb{E}[(\delta_i^\top Z_{it})^2] \geq \|\delta_i\|^2 c_\lambda$. Summarizing, we find \[ \mathbb{E}\Big[\rho_{\tau}(Y_{it} - Z_{it}^\top \gamma) - \rho_{\tau}(Y_{it} - Z_{it}^\top \gamma_{0i})\Big] \geq \frac{f_{min}c_\lambda}{4}\|\delta_i\|^2 = \frac{f_{min}c_\lambda}{4}\Big(\|\gamma - \gamma_{0i}\| \wedge \frac{f_{min}}{2M\overline{f'}} \Big)^2 \] which proves the lower bound in (ref). Thus the proof of the Lemma is complete. $\Box$
Proof of Lemma (ref) Consider the class of functions \[ \mathcal{G}_B := \Big\{ (y,z) \mapsto g_{\gamma}(y,z) := \frac{(\rho_\tau(y - z^\top \gamma) - \rho_\tau(y))\1\{|z| \leq M\} + MB}{2MB} ~\Big|~ \|\gamma\| \leq B \Big\}. \] Note that by construction $0 \leq g_{\gamma}(y,z) \leq 1$ for all $\|\gamma\| \leq B$ and moreover $\sup_{y,z} |g_{\gamma}(y,z) - g_{\gamma'}(y,z)| \leq \|\gamma-\gamma'\|/(2B)$. This shows the existence of constants $V,K_B < \infty$ such that for all $i=1,...,n$ $N_{[~]}(\varepsilon,\mathcal{G}_B,L_1(P_i)) \leq (K_B/\varepsilon)^V$ for $0<\varepsilon<K_B$ where $K_B$ depends on $B$ only and $P_i$ denotes the measure corresponding to $(Y_{i1},Z_{i1})$. Thus we have by Theorem 2.14.9 of vdVW, \[ P^* \Big( \sup_\gamma \frac{1}{\sqrt{T}} \Big|\sum_t g_\gamma(Y_{it},Z_{it}) - \mathbb{E}[g_\gamma(Y_{it},Z_{it})]\Big| \geq h \Big) \leq \Big(\frac{D_Bh}{\sqrt{V}}\Big)^V e^{-2h^2} \] where the constant $D_B$ depends only on $K_B$ and $P^*$ denotes outer probability. Set $h = \sqrt{\log n}$ to bound the right-hand side above by $o(n^{-1})$. Defining the events \[ E_{i,n} := \Big\{ \sup_\gamma \frac{1}{\sqrt{T}} \Big|\sum_t g_\gamma(Y_{it},Z_{it}) - \mathbb{E}[g_\gamma(Y_{it},Z_{it})]\Big| \geq \sqrt{\log n} \Big\} \] we obtain \[ P^*(\cup_i E_{i,n}) \leq n \sup_i P^*(E_{i,n}) \leq n o(n^{-1}) = o(1). \] Finally, note that under (A1) we have a.s. \[ \frac{\rho_\tau(Y_{it} - Z_{it}^\top \gamma) - \rho_\tau(Y_{it}) - \mathbb{E}[\rho_\tau(Y_{it} - Z_{it}^\top \gamma) - \rho_\tau(Y_{it})]}{2MB} = g_\gamma(Y_{it},Z_{it}) - \mathbb{E}[g_\gamma(Y_{it},Z_{it})] \quad \forall i,t. \] This completes the proof. $\Box$
We begin by stating a useful technical result that will be proved at the end of this section.
Proof of Theorem (ref) The proof proceeds in several steps. First, we note that the 'oracle' estimation problem (ref) corresponds to a classical, fixed-dimensional quantile regression with true parameter vector $(\alpha_{(01)},...,\alpha_{(0K)},\beta_0^\top)$ and $nT$ independent observations $(Y_{it},\tilde Z_{it})$ where $\tilde Z_{it}^\top = (e_k^\top,X_{it}^\top), i \in I_k, t=1,...,T$ where $e_k$ denotes the k'th unit vector in $\mathbb{R}^K$. A straightforward extension of classical proof techniques in parametric quantile regression shows that under assumptions (A1)-(A3) and (C) the oracle estimator is asymptotically normal as claimed.
Second, we observe that by definition of the optimization problem the estimated group structure $\hat I_{1,\ell},...,\hat I_{K_\ell,\ell}$ is the same for all values of $\ell$ with $\lambda_\ell$ that give rise to the same number of groups. Since the value of $IC(\ell)$ depends only on $\hat I_{1,\ell},...,\hat I_{K_\ell,\ell}$, it suffices to minimize $IC$ over those values of $\ell$ that correspond to different numbers of groups. Denote the distinct estimated numbers of groups by $\hat K_1,...,\hat K_R$, the corresponding estimated groupings by $\hat I_{(1 \hat K_r)},...,\hat I_{(\hat K_r \hat K_r)}$, and the corresponding values of $IC$ by $IC_{\hat K_1},...,IC_{\hat K_R}$. By assumption (G) and Theorem (ref), the probability of the event
Hence it suffices to prove that
Once this result is established, we directly obtain \[ P\Big((\hat\alpha_{1}^{IC},...,\hat\alpha_{\hat K^{IC}}^{IC},(\hat\beta^{IC})^\top) = (\hat\alpha_{(1)}^{(OR)},...,\hat\alpha_{(K)}^{(OR)},(\hat\beta^{(OR)})^\top)\Big) \to 1, \] and thus the asymptotic distribution of $(\hat\alpha_{1}^{IC},...,\hat\alpha_{\hat K^{IC}}^{IC},(\hat\beta^{IC})^\top)$ matches that of the oracle estimator.
We will now prove (ref). From Theorem 3.2 in \citeasnoun{Kato} we know that under (A1)-(A3) and the additional assumptions that $n \to \infty$ but $T$ grows at most polynomially in $n$ \[ \check\beta - \beta_0 = O_P((T/\log n)^{-3/4}\vee (nT)^{-1/2}). \] If $n \to \infty$ and $T$ grows at most polynomially in $n$ it follows that $\check\beta - \beta_0 = o_P(T^{-1/2})$. Moreover, standard quantile regression arguments show that \[ \check\alpha_i - \alpha_{0i} = - \frac{1}{\mathbb{E}[f_{\varepsilon_{it}^\tau|X_{it}}(0|X_{it})]} \frac{1}{T}\sum_t \psi_\tau(\varepsilon_{it}^\tau) + R_{n,i} \] where $\sup_i |R_{n,i}| = O_p\Big ( \Big (\frac{\log T}{T}\Big )^{3/4}\Big )$. Next apply Lemma (ref) to find that provided $\frac{(\log T)^3 (\log n)^2}{T} \to 0$,
Next, observe that by asymptotic normality of the oracle estimator \[ \sup_{k=1,...,K} \|\hat\gamma_{(k)}^{(OR)} - \gamma_{(0k)}\| = O_P((nT)^{-1/2}) \] where we defined $\hat\gamma_{(k)}^{(OR)} := (\hat\alpha_{(k)}^{(OR)}, \hat\beta^{(OR)})$. Again applying Lemma (ref) we obtain
Combining the results obtained so far we have
Next, let $V_n(L)$ denote the set of all disjoint partitions of $\{1,...,n\}$ into $L$ subsets. Observe that by (ref) we have under assumption (C)
Finally, note that by Lemma (ref)
Summarizing, we find that under (C)
The final result follows from a combination of (ref), (ref) and (ref). First, observe that for $K_r > K$ we have by (ref), (ref) and the assumptions on $p_{n,T}, \hat C$, with probability tending to one \[ IC_{K_r} - \inf_s IC_{K_s} \gtrsim p_{n,T} - n + o_P(n) \gg 0. \] It follows that, with probability tending to one, $\mathop{\rm arg\,min}_\ell IC(\ell) \leq K$. Moreover, for $K_r < K$ we have by (ref) and the assumptions on $p_{n,T}, \hat C$, with probability tending to one \[ IC_{K_r} - \inf_s IC_{K_s} \gtrsim - K p_{n,T} + nT - O_P(nT^{1/2}(\log n)^{1/2}) \gg 0. \] Hence, with probability tending to one, $K \geq \mathop{\rm arg\,min}_r IC_{K_r} \geq K$ and thus (ref) follows. $\Box$
Proof of Lemma (ref) By Knight's identity (ref) we have
Define
By a Taylor expansion we obtain
and thus the bound on $r_{n,2}^{(i)}$ is established. Next we note that
since the conditional expectation given $Z_{it}$ equals zero almost surely and moreover
Note that in particular for $\|\delta\|\leq 1$ we have \[ |I_{it}(\delta)| \leq M(M+1)\|\delta\|, \quad \mathbb{E}[I_{it}^{2}(\delta)] \leq 2(M^4 + 4M^3 \overline{f'})\|\delta\|^3. \] Define $c_{1,M} := M(M+1), c_{2,M} := 2(M^4 + 4M^3 \overline{f'})$ and apply the Bernstein inequality to show that for any $1 \geq \|\delta\| \geq T^{-1}\ell_{n,T}^2, 0 < a < \infty$
For $0 < a < 3 \ell_{n,T}^{1/2}c_{2,M}/c_{1,M}$ the last line above is bounded by $2 (n \vee T)^{-a^2/(4 c_{2,M})}$. Denote by $G_T$ a grid of values $\delta_1,...,\delta_{|G_T|}$ such that $T^{-1} \leq \|\delta_j\| \leq 1$ for all $j \in G_T$ and \[ \sup_{\ell_{n,T}^2T^{-1} \leq \|\delta\|\leq 1} \inf_{\tilde\delta\in G_T} \|\delta -\tilde\delta\| = o(T^{-2}). \] Note that it is possible to find such a $G_T$ with $|G_T| = O(T^{2(d+1)})$. It follows that
Finally, note that for $0 < a < 3 \ell_{n,T}^{1/2}c_{2,M}/c_{1,M}$
Since $\ell_{n,T} \to \infty$ we can pick $a$ such that the last line above is $o(1)$, and hence \[ \sup_i \sup_{\delta\in G_T}\frac{\Big|\sum_t I_{it}(\delta)\Big|}{\|\delta\|^{3/2}} = O_P(\ell_{n,T}^{1/2}T^{1/2}). \] Finally, observe that, denoting by $\|A\|_\infty$ the maximum norm of the entries of the matrix $A$,
where the last line follows by a straightforward application of the Hoeffding inequality. Thus the proof of Lemma (ref) is complete. $\Box$
\newgeometry{a4paper,left=1in,right=1in,top=1in,bottom=1in}
\restoregeometry