Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
31,498 characters · 8 sections · 45 citation commands
Nonparametric Quantile Regressions for Panel Data Models with Large $T$
\affil{School of Economics, Shanghai University of Finance and Economics}
This paper studies the estimation of nonparametric quantile panel data models. To facilitate the discussion, consider the following model:
where $Y_{it}\in\mathbb{R}$ is the observed dependent variable, $X_{it}\in \mathcal{X} \subset \mathbb{R}^d$ is the observed regressors, $\alpha_i\in\mathbb{R}$ is the unobserved individual effect representing individual heterogeneity, and $\epsilon_{it}| (X_{it},\alpha_i)\sim \mathcal{U}(0,1) $. Similar models have also been studied by altonji2005cross and chernozhukov2013average under different assumptions. Assuming that the mapping $\tau \mapsto Q(x, a, \tau)$ is strictly increasing for almost all $(x,a)$ in the support of $(X_{it},\alpha_i)$, then almost surely, \[ \mathsf{Q}_{Y_{it}}[\tau|X_{it}=x,\alpha_i=a] = Q(x, a, \tau) \overset{\text{def}}{=} Q_{\tau}(x,a),\] where $\mathsf{Q}_{Y_{it}}[\tau|\cdot]$ denotes the $\tau$-th conditional quantile of $Y_{it}$. Our main object of interest is the quantile partial effects (QPE, hereafter) of $X_{it}$ on $Y_{it}$ while controlling for the individual effects, i.e., $\partial Q_{\tau}(x,a)/\partial x$ for $\tau\in(0,1)$.
Recent development in the literature of quantile panel data models with large $T$, including koenker2004quantile, lamarche2010robust, galvao2010penalized, galvao2011quantile, canay2011simple, kato2012asymptotics and galvao2016smoothed, has mainly focused on the linear models where $Q_{\tau}(x,a) = \beta(\tau)'x + \lambda_\tau (a)$. This linearity specification for $Q_{\tau}(x,a)$ is convenient for constructing estimators of the QPE based on quantile regressions and analyzing their asymptotic properties, but it entails two possibly strong restrictions. First, in these models, $\partial Q_{\tau}(x,a)/\partial x =\beta(\tau)$, i.e., the QPE is homogeneous across $x$ and $a$. Second, the linearity assumption on $Q_{\tau}(x,a)$ usually impose strong restrictions on the regressors. For example, consider location-scale shifting models: $Y_{it} =\beta' X_{it} + \alpha_i + g(X_{it}) \cdot \epsilon_{it}$, where $\epsilon_{it}$ is independent of $(X_{it},\alpha_i)$. In order to have $Q_{\tau}(x,a)$ linear in $x$ for all $\tau$, we need $g(x) = \gamma' x >0 $ for some $\gamma \in \mathbb{R}^d$ and almost all $x$ in the support of $X_{it}$. Thus, for $d=1$, $X_{it}$ must be positive almost surely if $\gamma>0$.
To overcome the limitations of the linearity assumption, in this paper, we consider the following more general specification:
which is a separable nonparametric model in the sense that $q_{\tau}$ and $\lambda_\tau$ are both unknown functions. In this case, \[ \partial Q_{\tau}(x,a)/\partial x = \partial q_{\tau}(x)/\partial x \overset{\text{def}}{=} \beta_\tau(x).\] Thus, the QPE is allowed to be heterogeneous across $x$. Two estimators of $\beta_\tau(x)$ are proposed. The first one is based on local linear quantile regressions (LLQR, hereafter) and the second one is based on local linear smoothed quantile regressions (LLSQR, hereafter). The main advantage of the proposed estimators is that computationally, they are as efficient as the estimators of kato2012asymptotics and galvao2016smoothed for linear quantile panel models.
Despite being computationally simple, analyzing the asymptotic properties of the LLQR estimator and the LLSQR estimator in the large $T$ framework is a nontrivial task, mainly due to the well-known problem of “incidental parameters" --- see and lancaster2000incidental, hahn2004jackknife and fernandez2018fixed. Another major contribution of this paper is that it provides a set of regularity conditions under which the proposed estimators are shown to be asymptotically normally distributed. In particular, for the LLQR estimator, the incidental-parameter biases are hard to characterize (see the discussions of kato2012asymptotics) and we need $N\ll T^{\frac{2}{d+4}}$ to ignore the asymptotic biases. On the other hand, under the assumption that $N\asymp Th^d$ ($h$ is the bandwidth parameter in the local linear regression), we are able to derive the asymptotic bias of the LLSQR estimator for the boundary points of $\mathcal{X}$. Interestingly, the LLSQR estimator for the interior points of $\mathcal{X}$ are shown to be free of asymptotic bias. Moreover, our asymptotic analysis provides the theoretical basis of using split-panel jackknife (see dhaene2015split) for bias corrections.
Other Related Literature
As pointed out by arellano2011nonlinear, the identification of nonlinear panel data models with fixed $T$ is a nontrivial problem. Similarly, in the “small $T$” framework, the identification of $\partial Q_{\tau}(x,a)/\partial x$ is not straightforward. Invoking the result of hu2008instrumental, one can show that for $T=3$, if $\epsilon_{i1}, \epsilon_{i2}$ and $\epsilon_{i3}$ are mutually independent conditional on $X_i\overset{\text{def}}{=} (X_{i1},\ldots,X_{iT})'$ and some other high level conditions are satisfied, the general model (ref) is nonparametrically identified, i.e., all the conditional densities $\mathsf{f}_{Y_{it}|X_i,\alpha_i}$ for $t=1,2,3$ and $\mathsf{f}_{\alpha_i|X_i}$ are identified (see Proposition 2.1 of arellano2016nonlinear). Given this result, the identification of $\partial Q_{\tau}(x,a)/\partial x$ follows easily. evdokimov2010identification considers a separable model where $Q(X_{it}, \alpha_i, \epsilon_{it}) =m(X_{it}, \alpha_i) +U_{it}$ and $U_{it} \overset{\text{def}}{=}U(X_{it},\epsilon_{it} )$. In this model, $Q_{\tau}(x,a) = m(x,a)+ \mathsf{Q}_{U_{it}}[\tau|X_{it}=x] $. For $T=2$, evdokimov2010identification provides sufficient conditions for the identification of $m(x,a)$ and $\mathsf{f}_{U_{it} |X_{it}}$, which implies the identification of $\partial Q_{\tau}(x,a)/\partial x$. yan2018nonparametric considers a similar model with $Q(X_{it}, \alpha_i, \epsilon_{it}) =m(X_{it})+\alpha_i +\sigma(X_{it}) \epsilon_{it} $. They propose a multiple-step estimator of the conditional quantile function: $m(x)+\sigma(x) \mathsf{Q}_{\epsilon}(\tau) $, but no asymptotic theory was provided for this estimator. Moreover, varying-coefficients quantile panel models where $Q_{\tau}(x,a) = \beta_\tau(x_2)'x_1+a$ and $x=(x_1',x_2')'$ is studied by su2016sieve and cai2018semiparametric.
The identification of the quantile treatment effects (QTE) in nonseparable panels with fixed $T$ is considered by chernozhukov2013average and chernozhukov2015nonparametric. Note that the QTE considered in these papers is the derivative of the quantile structural function: $Q^{\ast}_{\tau}(x)$, which is defined by $P[Y_{it} \leq Q^{\ast}_{\tau}(x)|X_{it}=x] =\tau $. Therefore, it is obvious that $Q^{\ast}_{\tau}(x)\neq Q_{\tau}(x,a)$, and the QTE is different from the QPE. More recently, graham2018quantile considers the case where $ Q^{\ast}_{\tau}(x) = \beta_{\tau}(x)'x$ and focuses on the identification and estimation of the average conditional quantile effects (ACQEs) defined as $\mathbb{E}[ \beta_\tau(X_{it})]$.
Last but not least, this paper extends a large literature on nonparametric quantile regressions (see chaudhuri1991nonparametric, fan1994robust, yu1998local, honda2000nonparametric, su2009nonparametric, qu2015nonparametric, etc.) to panel data models with fixed effects.
Structure of the Paper
The rest of the paper is organized as follows: Section 2 introduces the models and provides some illustrative examples. Section 3 defines the estimators, whose asymptotic properties are established in Section 4. In Section 5, A Monte Carlo simulation is used to evaluate the performance of the proposed estimators and the bias-correction method. Finally, Section 6 concludes. All the proofs are collected in the appendix.
Specification (ref) implies the following panel data model:
where the error terms satisfy the following quantile restrictions: \[ P[ u_{it}(\tau)\leq 0|X_{it},\alpha_i] = \tau.\] It follows that the conditional quantile of the outcomes $Y_{it}$ given the observed covariates $X_{it}$ and the individual effect $\alpha_i$ can be written as \[ \mathsf{Q}_{Y_{it}}[\tau|X_{it}=x,\alpha_i=a] =Q_{\tau}(x,a)= q_{\tau}(x) + \lambda_{\tau}(a),\] and as discussed in the introduction, our main object of interest is the QPE: $\beta_{\tau}(x) = \dot{q}_{\tau}(x)$\footnote{To simply the notations we use $\dot{q}_{\tau}(x)$ and $\ddot{q}_{\tau}(x)$ to denote the first and second order derivatives of $q_\tau(\cdot)$ respectively.} for all $x\in\mathcal{X}\subset\mathbb{R}^d$ and all $\tau\in\mathcal{T}$, where $\mathcal{T}$ is a compact subset of $[0,1]$.
Consider the following 3 examples:
where $\epsilon_{it}$ is independent of $X_{it},\alpha_i $ with quantile function $\mathsf{Q}_{\epsilon}$. It follows that
Note that both Example 1 and Example 2 are nested by our model. In particular, the function $\lambda_{\tau}(\cdot)$ in Example 1 are invariant across $\tau \in (0,1)$. Example 3 is not nested by our model, since the conditional quantile function is not additive separable as functions of $x$ and $a$. Our model implies that the QPE is a function of $X_{it}$ only, while in Example 3 the QPE depends on both $X_{it}$ and $\alpha_i$.
Suppose that we have a random sample of $(Y_{it},X_{it})$ for $i=1,\ldots, N$ and $t=1,\ldots,T$, where the realized values of the individual effects are $(\alpha_{01},\ldots,\alpha_{0N})$. We follow a fixed effects approach, treating $(\lambda_{01,\tau},\ldots,\lambda_{0N,\tau}) \operatorname*{ \overset{\text{def}}{=}} (\lambda_\tau(\alpha_{01}),\ldots,\lambda_\tau(\alpha_{0N}))$ as fixed parameters, and consider the asymptotic framework where both dimensions of the panel data go to infinity, i.e., $N,T\rightarrow \infty$.
Focus on a single point $x \in \mathcal{X}$. Expanding $q_\tau(X_{it})$ around $x$, we have \[ q_{\tau}(X_{it}) = q_{\tau}(x) +\dot{q}_{\tau}(x)' (X_{it}-x) + 0.5 (X_{it}-x)' \ddot{q}_{\tau}(x)(X_{it}-x) + R_{\tau}(x,X_{it})\] where $R_{\tau}(x,X_{it})$ is the remainder term. It follows that \[ Y_{it} = \lambda_{0i,\tau} + q_{\tau}(x) +\dot{q}_{\tau}(x)' (X_{it}-x) + 0.5 (X_{it}-x)' \ddot{q}_{\tau}(x)(X_{it}-x) + R_{\tau}(x,X_{it}) +u_{it}(\tau), \] The above representation motivates that following LLQR estimator for $\eta_{0i,\tau}(x) = \lambda_{0i,\tau} + q_\tau(x)$ and $\beta_\tau(x)$: \[ (\hat{\eta}_{1,\tau}(x),\ldots,\hat{\eta}_{N,\tau}(x), \hat{\beta}_{\tau}(x))=\operatorname*{arg\min}_{\eta_1,\ldots,\eta_N,\beta} \sum_{i=1}^{N}\sum_{t=1}^{T}\rho_{\tau}(Y_{it} - \eta_i - (X_{it}-x)' \beta)\cdot K\bigg(\frac{X_{it}-x}{h}\bigg) \] where $\rho_\tau(u)=(\tau-\mathbf{1}(u\leq 0))u$ is the check function, $K(\cdot)$ is a multivariate kernel function, and $h$ is a bandwidth parameter. Note that
where $K_{it} =K(( X_{it}-x)/h)$, $\tilde{Y}_{it} = Y_{it} K_{it}$, and $\tilde{X}_{it} =(X_{it}-x)K_{it} $. Thus, the LLQR estimator can be easily calculated by running a standard quantile regression of $\tilde{Y}_{it}$ on $\tilde{X}_{it}$ and $N$ additional regressors: $\mathbf{1}(i=1)K_{it}, \ldots, \mathbf{1}(i=N)K_{it}$, therefore it is very computationally efficient.
Inspired by galvao2016smoothed, we also consider the following LLSQR estimator:
where $G(z) = 1-\int_{-\infty}^{z}g(u)du$, $g(\cdot)$ is a continuously differentiable function with support $[-1,1]$, and $b$ is a bandwidth parameter. The idea of smoothed quantile regression (see amemiya1982two and horowitz1998bootstrap) is to approximate the non-smooth indicator function with a smooth cumulative distribution function.
Before presenting the asymptotic results, it is useful to define some new notations. Let $x_{\partial}$ be on the boundary of $\mathcal{X}$. The boundary points are defined as \[ x=x_{\partial}+c h \text{ for some } c\in \text{supp}(K),\] and the domain for integration is defined as \[ \mathcal{B} = \{ v\in\mathbb{R}^d: (c+v)h\in \mathcal{X} \} \cap \text{supp}(K).\] Define \[c_0= \int_{\mathcal{B}} K(u)du, \text{ }\mathcal{C}_1= \int_{\mathcal{B}} uK(u)du, \text{ } \mathcal{C}_2= \int_{\mathcal{B}} uu'K(u)du ,\text{ } \mathcal{C}= \mathcal{C}_2-\mathcal{C}_1 \mathcal{C}_1'/c_0,\] \[ \bar{\mathcal{C}}_2 =
, d_0= \int_{\mathcal{B}} K^2(u)du, \mathcal{D}_1= \int_{\mathcal{B}} uK^2(u)du, \mathcal{K}_1=\int uu'K(u)du , \mathcal{K}_2=\int uu'K^2(u)du. \]
Write $u_{it}$ instead of $u_{it}(\tau)$ to simply the notations. Let $B_{\epsilon}$ be a neighbourhood of $0$. We first impose the following assumptions:
{Remark 1.1:} The above assumptions, except (A2) and (A7), are standard in the literature of local linear quantile regressions and quantile regressions. Note that we only need the existence and smoothness of the conditional density of $u_{it}$ given $X_{it}$, thus the estimator is robust to heavy tails and outliers in $u_{it}$.
{Remark 1.2:} The independence assumption (A2) is also adopted by kato2012asymptotics and it excludes time-invariant regressors. This independence assumption can be relaxed to allow for $\beta$-mixing on the time dimension along the line of galvao2016smoothed at the cost of much lengthier proofs. Thus, to keep the proofs tractable, (A2) is maintained throughout the paper.
{Remark 1.3:} Assumption (A7) ensures that $\log N \ll Th^{d+4}$, $N\ll Th^{d+2}$, $NTh^{d+6}\rightarrow 0$, and $N^2 \ll Th^d$. These conditions are needed to prove Theorem 1 below. For example, for $d=2$ and $c_N=1/4$, we can choose $c_h=1/5$. Note that due to the nonparametric nature of our estimator, the condition $N^2 \ll Th^d$ imposed here is stronger than the condition $N^2 \ll T$ required by kato2012asymptotics, since the order of the incidental-parameter bias is approximately $(Th^d)^{-3/4}$, while in kato2012asymptotics the bias is approximately of order $T^{-3/4}$. Such conditions are hard to justify in practice, this is why we also consider the LLSQR estimator, whose asymptotic distribution can be established under more realistic assumptions about the relative sizes of $N$ and $T$.
The following theorem gives the asymptotic distribution of the LLQR estimator.
{Remark 1.4:} The LLQR estimator suffers from two types of biases: a bias due to the estimation of incidental parameters, and another one due to local linear approximations. As discussed in Remark 1.3, the first bias can be ignored at the expense of a very strong condition: $N\ll T^{\frac{2}{d+4}} $. The term $h B^{(1)}$ is the leading bias term in the local linear approximations (see fan1994robust for example). This bias can be further reduced by using local polynomial regressions. Note that $B^{(1)}=0$ for the interior points, thus the leading bias term for the estimators of the interior points is $O(h^2)$.
{Remark 1.5:} In general, it is straightforward to construct consistent estimators of the asymptotic variances, since $\mathcal{K}_1,\mathcal{K}_2$ and $\Omega$ only depend on the kernel function $K(\cdot)$, and $\sigma(x)$ can be consistently estimated using standard nonparametric methods. In particular, if the distribution of $(u_{it},X_{it})$ are identical across $i$, i.e., $ f_{u,i}(0|x) = f_{u}(0|x)$ and $f_{X,i}(x)=f_{X}(x)$ for all $i$, then $ \sigma(x) = \left( f_{X}(x) \right)^{-1}\cdot \left( f_{u}(0|x) \right)^{-2}$.
We impose the following assumptions:
{Remark 2.1:} Assumptions (B2) and (B3) are also imposed in galvao2016smoothed. In particular, we need $g(\cdot)$ to be a fourth (or higher) order kernel function. Assumption (B4) is new. Condition (ref) implies that $Th^{d+1}\rightarrow\infty$, $Th^{d+3}\rightarrow0$, $Th^db^3\rightarrow\infty$, $Th^db^m\rightarrow0$ and $ b^m \ll h^2 \ll b$. These conditions will be used in the proof of Theorem 2. Moreover, $m\geq4$ and (ref) ensure that $c_h$ lies in a non-empty set. For example, for $m=4$ and $d=2$, one can choose $c_b=1/6$ and $c_h\in(1/5,1/4)$.
The following theorem gives the asymptotic distribution of the LLSQR estimator.
{Remark 2.2:} It can be seen that the asymptotic distributions of the LLSQR estimators and the LLQR estimators are very similar, with one noticeable difference: the LLSQR estimator for the boundary points suffers from an asymptotic bias: $\kappa B^{(2)}$, which is the consequence of estimating incidental parameters. In the proof of Theorem 2, it is found that the incidental-parameter bias of the LLSQR estimator for the boundary points is of order $(Th^d)^{-1}$ rather than $T^{-1}$ --- this is why we need $N\asymp Th^d$ to derive the analytical expression of the asymptotic bias. Interestingly, the LLSQR estimators for the boundary points at $\tau=0.5$ and the LLSQR estimators for the interior points at all $\tau$s are all free of asymptotic biases. These findings are further confirmed by a Monte Carlo simulation in Section 5.
{Remark 2.3:} Theorem 2 provides the theoretical basis for bias corrections using the split-panel jackknife method proposed by dhaene2015split. In particular, divide the whole sample into two subsamples: $(Y_{it},X_{it})$ for $i=1,\ldots,N; t=1,\ldots, T/2$, and $(Y_{it},X_{it})$ for $i=1,\ldots,N; t=T/2+1,\ldots, T$, and let $\check{\beta}_{\tau,1}(x), \check{\beta}_{\tau,2}(x)$ denote the LLSQR estimators using the two subsamples respectively\footnote{The bandwidth parameter $h$ should be the same for $\check{\beta}_{\tau}(x)$, $\check{\beta}_{\tau,1}(x)$ and $\check{\beta}_{\tau,2}(x)$.}. The bias-corrected estimator is simply given by \[ \check{\beta}_{\tau}^{bc}(x) = 2\check{\beta}_{\tau}(x) - 0.5 [\check{\beta}_{\tau,1}(x)+\check{\beta}_{\tau,2}(x)]. \] Under Assumptions (B1) to (B4) we can show that for the boundary points, \[ \sqrt{NTh^{d+2}} \big[ \check{\beta}_{\tau}^{bc}(x) - \beta_{\tau}(x) - h B^{(1)}\big] \overset{d}{\rightarrow} \mathcal{N}\Big(0, \tau(1-\tau) \sigma(0) \Omega \Big). \]
{Remark 2.4:} Assumption (B4) requires that $N\asymp Th^d$, which is much less restrictive than Assumption (A6) which imposes $N \ll \sqrt{Th^d}$. As discussed in Remark 2.1, for $d=2$, Assumption (B4) admits the choice: $c_h=1/4.5$ and therefore $N\asymp T^{5/9}$. Thus, given the nonparametric nature of the problem, Assumption (B4) is still more stringent than the usual assumption $N\asymp T$ imposed for nonlinear fixed-effects estimators (see hahn2004jackknife and fernandez2018fixed).
In this section, we evaluate the performance of the proposed estimators in finite samples using the following data generating process (DGP): \[ Y_{it} = \beta X_{it} + \alpha_i + \sqrt{ 1+X_{it}^2} \cdot \epsilon_{it}, \] where $X_{it}\sim i.i.d \text{ } \mathcal{N}(0,1)\cdot \mathbf{1}\{|X_{it}|\leq 2\}$, $\alpha_i\sim i.i.d \text{ } \mathcal{N}(0,1)$. It is easy to see that $\beta_{\tau}(x) = 1+ \mathsf{Q}_{\epsilon}(\tau)\cdot x/\sqrt{1+x^2}$, where $\epsilon_{it}$ are i.i.d with quantile function $\mathsf{Q}_{\epsilon}$. We consider two different distributions of $\epsilon_{it} $: (i) $\mathcal{N}(0,1)$, and (ii) t distribution with 3 degrees of freedom, and compare the biases and mean-square errors (MSEs) of four different estimators: the LLQR estimator $\hat{\beta}_{\tau}$, the LLSQR estimator $\check{\beta}_{\tau}$, and the bias-corrected versions of these two estimators, denoted as $\hat{\beta}_{\tau}^{bc}$ and $\check{\beta}_{\tau}^{bc}$ respectively.
To same space, we only report the results for $N=T=100$, $\tau=0.25,0.5,0.75$ and $x=-2,-1.6,\ldots,1.6,2$. For all estimators, we choose $h=0.8$, and for the LLSQR estimators, we choose $b=0.5$ and consider the following fourth-order kernel function: \[ k(u)=\frac{105}{64}\left(1-5 u^{2}+7 u^{4}-3 u^{6}\right) 1(|u| \leq 1). \] We have also tried other bandwidth values and find that the results is more sensitive to the choice of $h$ than the choice of $b$.
Table 1 reports the results for $\tau=0.25$ and $\epsilon_{it}\sim\mathcal{N}(0,1)$ while Table 2 reports the results for $\tau=0.25$ and $\epsilon_{it}\sim T(3)$. The results for $\tau=0.5$ and $\tau=0.75$ are reported in Table 3 to Table 6.
For $\tau=0.25$, four conclusions can be drawn from the results in Tables 1 and 2. (i) The performance of $\hat{\beta}_{\tau}$ and $\check{\beta}_{\tau}$, in terms of biases and MSEs, are very close. (ii) As predicted by our Theorem 2, the bias-correction method significantly reduces the biases of the LLSQR estimators, especially at the boundary points (e.g., $|x|=2,1.6$). Interestingly, the bias-correction method can also effectively reduce the biases of the LLQR estimators. (iii) The bias-corrected estimators for the boundary points have much lower MSEs due to the large decrease in biases. (iv) The performance of the estimators are robust to heavy tails of $\epsilon_{it}$. Similar conclusions are supported by the results for $\tau=0.75$. However, for $\tau=0.5$, the bias correction at the boundary points is not very effective --- this is predicted by Theorem 2, which shows that the LLSQR estimator for the boundary points is free of asymptotic biases at $\tau=0.5$ (see Remark 2.2).
To the best of our knowledge, this is the first paper that considers nonparametric quantile regressions in the context of large $T$ panels. Our model is additively separable as unknown functions of the regressors and the individual effects, and it allows the QPE to be heterogeneous across individuals. We propose two estimators of the QPE based on local linear approximations, and establish their asymptotic distributions under a set of regularity assumptions. Our theoretical results highlight the importance of incidental-parameter biases and justify the use of convenient jackknife method to correct the asymptotic biases. The good performance of the bias-correction method in finite samples is confirmed using a Monte Carlo simulation.
Like any other nonparametric estimators, the choice of bandwidth is crucial in practice. In this paper we have focused on the theoretical conditions that the bandwidth parameters have to satisfy, but how to choose those bandwidths in practice is an important question that is left for further investigation.