Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.
82,064 characters · 14 sections · 40 citation commands
{
}
\onehalfspacing
We consider a longitudinal data model of conditional quantiles with individual intercepts. Variations of this model have been extensively studied in the literature since at least NS48. Recent contributions to the literature using this model for quantile regression have emphasized the drawbacks of estimating a large number of individual intercepts ($N$) when the number of time periods ($T$) is small (see Galvao and Kato, 2018,\nocite{Galvao2018} for an excellent survey). Koenker (2004)\nocite{rK04} proposed an estimator where $N$ individual parameters are regularized by a Lasso-type penalty, shrinking them towards a common value. As in the case of the Gaussian random effect estimator, shrinkage can reduce the variability of the estimator of the slope parameter in the quantile regression model (Koenker, 2004). In models with short $T$, shrinkage can reduce the bias of the fixed effects estimator of the slope parameter as well (Harding and Lamarche, 2019\nocite{harding2019}).
Although the regularization procedure has advantages, the asymptotic distribution of the estimator is difficult to approximate. It is known that Lasso-type estimators have non-standard limiting distributions (Knight and Fu, 2000\nocite{kK00}), but in the case of quantile regression, there are new challenges. Because individual intercepts are treated as parameters, the increasing dimension of the parameter vector as the number of units increases can be an issue. In the case of estimators without regularization, Kato, Galvao, and Montes-Rojas (2012) and Galvao, Gu, and Volgushev (2020)\nocite{AGalvao2020} found that $T$ must grow faster than $N$ for consistency and asymptotic normality at rates that are, at best, similar to standard non-linear panel data models Newey2004. Second, the covariance matrix of quantile regression estimators typically depends on conditional densities and the penalized estimator of Koenker (2004) is no exception. Inference based on the asymptotic distribution requires non-parametric estimation of nuisance parameters, which can lead to important size distortions (He, 2018\nocite{He2018}).
Motivated by these limitations, cross-sectional pairs (or block) bootstrap, which samples sets of covariate and response vectors over individuals with replacement, appears to be a natural alternative method for inference. However, we demonstrate that the cross-sectional pairs bootstrap does not approximate well the limiting distribution of the penalized estimator. We consider instead a wild residual bootstrap procedure, which was previously employed by Feng, He, and Hu (2011)\nocite{xumingHe2011}, and Wang, Van Keilegom, and Maidman (2018)\nocite{lanWang2018} in cross-sectional settings. We investigate the application of the procedure to longitudinal data and show that the proposed wild bootstrap procedure is a consistent estimator of the distribution of the penalized estimator.
We begin by deriving consistency and asymptotic normality results for $\ell_1$ penalized estimators of a longitudinal model in which individual effects can be correlated with the regressors. Although our model might be considered to be high-dimensional, the number of parameters is smaller than the number of observations, as in the pioneering work by Koenker (2004), and thus our results are obtained without assuming sparsity in terms of the individual intercepts. Consistency and asymptotic normality with $T$ growing faster than $N$ are achieved by letting the penalty parameter that controls shrinkage diminish in importance asymptotically. Thus, relative to rK04, the asymptotic bias of the estimator is zero in our case. The consistency and asymptotic normality results are new — they extend the heuristic results in Koenker (2004) obtained for a model with individual effects as location shifts and are not included in kK12 and AGalvao2020 because they did not consider penalized estimation.
The main theoretical contribution is to show that the distribution of the wild bootstrap estimator consistently estimates the asymptotic distribution and covariance of the penalized estimator. The results include the special case of no penalization, and thus, these results also show the consistency of the wild bootstrap for the quantile regression estimator with fixed effects. The consistency of the wild bootstrap is established using developments that are critically different to those used in Wang, Van Keilegom, and Maidman (2018)\nocite{lanWang2018}. We also consider bootstrap estimation of the asymptotic covariance matrix of the slope parameter estimator, which is novel in the panel quantile literature. As emphasized in GoncalvesWhite5, Andreas2017, and HahnLiao21, the weak convergence of the bootstrap estimator does not necessarily imply convergence of the bootstrap second moment estimator. Therefore, we provide conditions and establish a result that supports using the second moment of the bootstrap distribution to estimate the asymptotic variance of the estimator.
Several penalized estimators for quantile regression models have been proposed in the literature since rK04. belloni2011 propose quantile regression estimators for high-dimensional sparse models using cross-sectional data. WANG2013 considers a penalized least absolute deviation estimator, and Wang2019 derives error bounds for the penalized estimator under weak conditions. cL06 investigates the selection of a regularization parameter, and SLee2018 study estimation of a high-dimensional quantile regression model with a change point, or threshold. harding2016,harding2019 investigate estimation of models with attrition and correlated random effects. GU2019 propose a method for estimation of models with unknown group membership. CHEN2009 establish the validity of a related weighted bootstrap procedure for the limiting distribution of a penalized sieve estimator and consider applications using quantile regression CHEN2015. The literature on penalized estimation methods for linear panel data models has also grown in the last decade kock2013,KOCK2016,aBelloni2016, lSu2016, lSu2018, CANER2018143, kock_tang_2019
This paper is organized as follows. The next section provides background and discusses the motivation of our study. It also introduces the proposed wild residual bootstrap approach. Section 3 presents theoretical results. Section 4 investigates the small sample performance of the method, showing that the estimator has satisfactory performance under different specifications and it performs better than the cross-sectional pairs bootstrap procedure. Section 5 presents extensions to the basic model. Section 6 illustrates the theory and provides practical guidelines from an application of the method. Considering data from the U.S. Census, we estimate a quantile function with more than eighty thousand parameters to study how wages of U.S. workers have been affected by the North American Free Trade Agreement. Finally, Section 7 concludes. One appendix contains proof of the main results, while a supplementary appendix contains additional technical results and proofs.
In this section, we first introduce the model and the estimator, and then we discuss the validity of a cross-sectional pairs bootstrap method. Motivated by the limitations of existing procedures, we propose a new approach to estimate the asymptotic distribution of the estimator.
We observe repeated measures $\{ ( y_{it},\bm{x}_{it}') \}_{t=1}^{T}$ for each subject $1 \leq i \leq N$. The variable $y_{it} \in \mathbb{R}$ denotes the response for $i$ at time $t$ and $\bm{x}_{it}$ denotes a $p$-dimensional vector of covariates. Although the number of repeated observations does not vary with $i$, the analysis can be trivially extended to consider $T_i$ as long as $\max T_i / \min T_i$ is bounded for $1 \leq i \leq N$ (Gu and Volgushev, 2019). The model considered in this paper is
where $\tau \in (0,1)$ and $ Q_y(\tau | \bm{x}_{it}) $ is the $\tau$-th quantile of the conditional distribution of $y_{it}$ given $\bm{x}_{it}$. It is assumed that the vector $\bm{x}_{it}$ does not contain an intercept. The parameter of interest is $\bm{\beta}_0(\tau) \in \mathbb{R}^p$ and $\alpha_{i0}(\tau)$ is treated as a nuisance parameter. Because we consider just one value of $\tau$, we supress the dependence of the parameters on $\tau$ in the sequel.
Let $\bm{\theta} = (\bm{\beta}',\bm{\alpha}')' \in \bm{\Theta} \subseteq \mathbb{R}^{p+N}$, where $\bm{\alpha} = (\alpha_{1},...,\alpha_{N})'$, and let $\bm{\theta}_0 = (\bm{\beta}_0',\bm{\alpha}_0')'$. To estimate $\bm{\theta}_0$, we consider the following estimator:
where $\rho _{\tau }(u)=$ $u(\tau -I(u<0))$ is the quantile regression loss function. The tuning parameter $\lambda_T \geq 0$ depends on $T$ and it can also depend on data, as discussed below.
The penalty term in (2.2) helps improve the finite sample performance of the fixed effects estimator, which is defined for $\lambda_T=0$. Shrinkage of the individual effects can lead to reductions of the variance of the estimator. In models with incidental parameters, the penalty term reduces the noise in the estimation of individual intercepts, and consequently, it can also reduce the bias of the fixed effects estimator of $\bm{\beta}_0$. The online appendix presents simulation evidence to illustrate finite sample improvements when the time dimension is short, complementing the evidence presented in Koenker (2004) and Harding and Lamarche (2019). See Bester and Hansen (2009)\nocite{cH2009} for a related penalty approach to bias reduction in nonlinear models with fixed effects.
We establish conditions that result in a tractable asymptotic distribution for the estimator defined in (ref). However, we expect that resampling methods offer a more accurate description of the distribution of the estimator in finite samples. In practice, the cross-sectional pairs bootstrap, which samples over $i$ with replacement keeping the entire block of time series observations for each $i$, has been used as a method for inference, primarily in the fixed effects case when $\lambda_T = 0$. However, the cross-sectional pairs bootstrap does not provide a good approximation to the sampling distribution of the penalized estimator (ref), as in the case of the pairs bootstrap procedure for the Lasso estimator camponovo15.
We now offer a heuristic illustration of some problems with using a cross-sectional pairs bootstrap and the penalized quantile regression estimator. The cross-sectional pairs bootstrap can be used successfully to estimate the distribution of the quantile regression model with unpenalized fixed effects, but it will be shown below that the penalty causes problems for this approach to resampling. We fix $N$ in this section to avoid the effect of a diverging number of parameters as the sample size increases (later, asymptotic approximations will be found assuming that $T$ grows faster than $N$). This allows us to see problems with the cross-sectional pairs without the additional incidental parameters problem. Define $\bm{\gamma} = (\bm{\delta}', \bm{\eta}')' \in \mathbb{R}^{p + N}$, where $\bm{\delta} = \sqrt{NT} (\bm{\beta} - \bm{\beta}_0)$ and for $i = 1, \ldots N$, $\eta_i = \sqrt{T} (\alpha_i - \alpha_{i0})$. Then let
where $u_{it} = y_{it} - \bm{x}_{it}' \bm{\beta}_0 - \alpha_{i0}$. This objective function is equivalent to (ref). kK00 developed a method for dealing with the asymptotic behavior of this objective function, stated here as a lemma.
To examine the validity of the cross-sectional pairs bootstrap, consider an analog loss function for resampled data. Letting $\bm{y}_i$ and $\bm{X}_i$ denote the vector and matrix of response and covariate observations corresponding to unit $i$, a cross-sectional pairs bootstrap procedure resamples $N$ pairs $(\bm{y}_i, \bm{X}_i)$ for $1 \leq i \leq N$ with replacement. Let $n_i^*$ denote the number of times unit $i$ is redrawn from the original sample. Thus, the asymptotic distribution of $\hat{\bm{\gamma}}$ is approximated with $\tilde{\bm{\gamma}} = ( \sqrt{NT} (\tilde{\bm{\beta}} - \hat{\bm{\beta}})', \sqrt{T} ( \tilde{\bm{\alpha}} - \hat{\bm{\alpha}})' )'$ where
Since $n_i^*$ is a multinomial weight with probability $1/N$, it is straightforward to calculate that the expected value of the objective function with respect to the bootstrap weights (i.e., conditional on the observations) is minimized at $\hat{\bm{\theta}} = (\hat{\bm{\beta}}, \hat{\bm{\alpha}})$. However, a finite sample problem is associated with the presence of the penalty in the objective function. To see this, let $\alpha_i^\ast = n_i^* |\alpha_i|$ and $\mathcal{A}=\{i : \alpha_i^* \neq 0\}$ denote the “active" set corresponding to the penalty term in (ref). In each bootstrap repetition, the cardinality of $\mathcal{A} < N$, leading to solutions $\tilde{\bm{\theta}}$ that can be potentially very different than the minimizer $\hat{\bm{\theta}}$. This may be especially so when $\alpha_i$ is correlated with $\bm{x}_{it}$.
To see other problems with the cross-sectional bootstrap, we can find the weak limit of the bootstrap objective function (ref) similarly to Lemma (ref). When we recenter (ref) employing $\hat{\bm{\theta}}$, using the $i$ chosen by resampling, we find a naive bootstrap analog of the original objective function (ref), denoting $\hat{u}_{it} = y_{it} - \hat{\bm{\beta}}' \bm{x}_{it} - \hat{\alpha}_i$:
As $T \rightarrow \infty$, assuming $\hat{\eta}_i = \sqrt{T}(\hat{\alpha}_i - \alpha_{i0}) \overset{d}{\longrightarrow} A_i$ for $i = 1, \ldots N$ as $T \rightarrow \infty$, $\tilde{\mathbb{V}}_T$ converges weakly to
However, there are two key differences with the resulting expression. The first problem with this limiting objective function is that $\tilde{\bm{B}} \neq \bm{B}$ and $\tilde{\bm{D}_1} \neq \bm{D}_1$ from Lemma (ref), due to the fact that recentering uses $\hat{\bm{\theta}}$, which is asymptotically biased if $\lambda_0>0$. Second, there is additional randomness arising from variable selection and resampling. (In the online appendix, we illustrate these issues with fixed $N$ and $T$). In the next section, we propose a wild residual bootstrap that does not suffer from these shortcomings. Then we expect that the distribution of the wild bootstrap estimator $\bm{\gamma}^\ast$ provides a better approximation to the distribution of $\hat{\bm{\gamma}}$ in Lemma (ref).
Let $\hat{u}_{it} = y_{it} - \bm{x}_{it}' \hat{\bm{\beta}} - \hat{\alpha}_i$ be the $\tau$-th quantile residual. Let $u_{it}^\ast = w_{it} | \hat{u}_{it} |$ denote bootstrap residuals, where $w_{it}$ is drawn randomly from a pre-determined distribution $G_W$ that satisfies the following conditions:
Several weight distributions have been proposed in the quantile regression literature that satisfy these conditions. xumingHe2011 propose, for $1/8 \leq \tau \leq 7/8$, the continuous weight density $g_W(w) = - w I(-2 \tau - 1/4 \leq w \leq -2 \tau + 1/4) + w I(2 (1-\tau) - 1/4 \leq w \leq 2 (1-\tau) + 1/4)$. Another distribution that satisfies (ref)-(ref) is the two-point distribution at $w = 2 (1-\tau)$ with probability $\tau$ and at $w = - 2 \tau$ with probability $(1-\tau)$. We adopt this distribution in the numerical examples. See Appendix 3 in lanWang2018 for additional examples of the weight distribution.
Using the bootstrap sample of residuals and the penalized quantile estimator as defined in equation (ref), we can form $y_{it}^\ast = \bm{x}_{it}' \hat{\bm{\beta}} + \hat{\alpha}_i + u_{it}^\ast$ to obtain the bootstrap estimator:
Given a bootstrap sample $\{\bm{\beta}_b^\ast\}_{b=1}^B$, we can obtain confidence intervals that are asymptotically valid, as demonstrated in Theorem (ref) below. Let $G_{j}^*(\alpha/2)$ and $G_{j}^*(1-\alpha/2)$ be the $(\alpha/2)$-th quantile and $(1-\alpha/2)$-th quantile of the bootstrap distribution of $\sqrt{NT} ( \beta_{j}^\ast - \hat{\beta}_{j})$ for $j = 1,2,\hdots,p$. We obtain asymptotically valid $100 (1-\alpha)\%$ confidence intervals for $\beta_{j}$ by $[ \hat{\beta}_{j} - (NT)^{-1/2} G_{j}^*(1 - \alpha/2), \hat{\beta}_{j} - (NT)^{-1/2} G_{j}^*(\alpha/2)]$. Alternatively, Theorem (ref) shows that we may also estimate the covariance matrix of $\sqrt{NT}(\hat{\bm{\beta}} - \bm{\beta}_0)$ using the estimated covariance matrix from the bootstrap sample, which can be used to estimate the variance without requiring density estimation and to construct bootstrap-$t$ statistics for inference.
We may also consider a threshold estimator for $1 \leq i \leq N$, $\alpha_i^{**} = \hat{\alpha}_i I( | \hat{\alpha}_i | \geq a_{T})$, where $a_T$ is a constant that satisfies $a_T \to 0$ as $T \to \infty$. Define $v_{it}^\ast = w_{it} | \hat{v}_{it} |$, where $\hat{v}_{it} = y_{it} - \bm{x}_{it}' \hat{\bm{\beta}} - \alpha_i^{**}$. The response variable is generated as $y_{it}^{**} = \bm{x}_{it}' \hat{\bm{\beta}} + \alpha_i^{**} + v_{it}^\ast$, and the threshold estimator is defined as
As in the case of the estimator defined in (ref), we estimate the distribution of $\hat{\bm{\theta}}$ based on the estimator $\bm{\theta}^{**}$. Given the similarities between estimators (ref) and (ref), we derive below consistency and asymptotic normality results for (ref) only. The performance of the bootstrap with this estimator is examined in the online appendix.
The tuning parameter $\lambda_T$ controls the degree of shrinkage of the individual effect $\alpha_i$ towards zero and the penalty helps to control the bias and variance of $\hat{\bm{\beta}}$. We restrict the tuning parameter to $\lambda_T \in \mathcal{L} \subset [0, \lambda_U]$, where $\lambda_U$ is an upper bound. As shown in Lemma (ref) in the supplementary appendix, $\lambda_U = \max\{\tau, 1 - \tau\} T$ is a natural choice because if $\lambda_T$ is set larger than this value, all the individual effects will be set equal to zero. If the number of observed time periods $T_i$ vary over $i$, then one would need to replace the $T$ in these bounds with $\max_i T_i$. This estimator accommodates the choice of $\lambda_T = 0$, which means that the results below continue to hold for the corresponding unpenalized estimator.
The selection $\lambda_T$ in related settings has been investigated in several papers cL06,ERLee2014,lanWang2018. We follow lanWang2018 and employ cross-validation for tuning parameter selection. To the best of our knowledge, theory has not yet been developed for the stochastic order of $\lambda_T$ when chosen using cross-validation, but in extensive simulations we have found that it tends to grow much more slowly than $T$, as required in Theorems (ref) and (ref) below.
This section investigates the large sample properties of the proposed estimator. We consider the following assumptions:
These conditions are standard in the literature on quantile regression with individual effects. Conditions (ref) and (ref) are the same as Assumptions (A1) and (A3) in Kato, Galvao, and Montes-Rojas (2012). Condition (ref) is relaxed in Kato et al. (2012) and in Section 5 below to allow for time dependence. Condition (ref) is an identification condition and it is sufficient for consistency. Slightly weaker than the assumption that $F_i$ has a continuous density given $\bm{x}_{it}$, it allows an expansion that guarantees the convexity of the limiting objective function, and therefore, the uniqueness of $(\bm{\beta}_0',\alpha_{i0})$ for all $1 \leq i \leq N$. Assumption (ref) is a simple way to assume appropriate moment conditions on the covariates and it is similar to (B1) in Kato, Galvao, and Montes-Rojas (2012) and (A1) in GU2019. The condition can be relaxed as in Kato, Galvao and Montes-Rojas (2012). Condition B3 can be replaced with the moment condition $\sup_{i \geq 1} \textnormal{E} \left[ \| \bm{x}_{i1} \|^{2s} \right] < \infty$ for some $s \geq 1$. The implication of this weaker condition is that $N / T^s \to 0$ instead of $\log(N)/T \to \infty$ to achieve consistency, as demonstrated in Theorem (ref).
The consistency of the estimator $\hat{\bm{\theta}}$ is needed to establish the main result stated in Theorem (ref).
We now focus our attention on weak convergence and we present a series of results to facilitate the estimation of standard errors and confidence intervals. To show asymptotic normality of the estimator, it is necessary to strengthen the conditions required for consistency slightly with the following conditions routinely adopted in the panel quantile regression literature (see, e.g., assumptions (B2) and (B3) in Kato, Galvao, and Montes-Rojas, 2012, and assumption (A2) in Gu and Volgushev, 2019).
Then we have the following result:
The wild residual bootstrap procedure is consistent as an estimator of the asymptotic distribution of $\hat{\bm{\beta}}$, as the next theorem shows.
Theorem (ref) only shows consistency of the bootstrap distribution estimator. Theorem (ref) ahead shows that the bootstrap covariance matrix, defined as
may be used to estimate the covariance of $\sqrt{NT}(\hat{\bm{\beta}} - \bm{\beta}_0)$. In practice, one simply uses the sample covariance of all the bootstrap repetitions, increasing the number of repetitions to bring the sample average as close as desired to the bootstrap expectation. Variance estimation using the bootstrap was formally investigated for quantile regression with clustered data in Andreas2017, but the model in this paper is complicated by the diverging number of individual effects as $N \rightarrow \infty$ and the penalty term in (ref).
The assumptions that are required for Theorem (ref) are slightly stronger than those used in Theorem (ref). The requirement on $\lambda_T$ is due to its presence in asymptotic expansions leading to the Bahadur representation of $\bm{\beta}^\ast$ and is similar to the moment requirement made on the covariates in Andreas2017. The compactness assumption must be made to ensure that expansions used in the asymptotic approximation are uniformly bounded.
In this section, we report the results of several simulation experiments designed to evaluate the performance of the method in finite samples. We consider a data generating process similar to the ones considered in Koenker (2004) and Kato, Galvao and Montes-Rojas (2012). The dependent variable is $y_{it} = \alpha_i + x_{it} + (1 + \zeta x_{it}) u_{it}$, where $x_{it} = 0.5 \alpha_i + z_i + \epsilon_{it}$, and $z_i$ and $\epsilon_{it}$ are i.i.d. random variables distributed as $\chi^2$ with 3 degrees of freedom ($\chi_3^2$). The corresponding quantile regression function is $Q_y (\tau | x_{it}) = \alpha_{0i} + \beta_0 x_{it}$, where $\alpha_{0i} = \alpha_i + F_u(\tau)^{-1}$, $\beta_0 = 1 + \zeta F_u(\tau)^{-1}$, and $F_u(\cdot)$ denotes the distribution of the error term, $u_{it}$.
We generate data from several variations of the basic model. In one variant of the model, $\alpha_i$ is an i.i.d. Gaussian random variable. In another, we generate $\alpha_i = i/N$ for $1 \leq i \leq N$ as in Galvao, Gu, and Volgushev (2020). We use $\zeta \in \{0,0.5\}$, and thus, $\beta_0 = 1$ in the location shift version of the model and $\beta_0 = 1 + 0.5 F_u(\tau)^{-1}$ in the location-scale shift case. Lastly, we consider three different distributions for the error term. We assume that $u_{it}$ is distributed as $\mathcal{N}(0,1)$, a $t$ distribution with 3 degrees of freedom ($t_3$), or $\chi_3^2$.
Tables (ref), (ref), (ref), and (ref) present coverage probabilities for a nominal 90% confidence interval for the slope parameter $\beta_{0}$. We present coverage probabilities using the empirical distribution of the bootstrap estimator (Tables (ref) and (ref)), as well as coverage probabilities of the asymptotic Gaussian confidence interval (Tables (ref) and (ref)). In the latter case, the coverage is constructed using the standard error of the corresponding bootstrap procedure. Tables (ref) and (ref) present results for the location shift model ($\zeta=0$), while Tables (ref) and (ref) present results for the location-scale shift model ($\zeta=0.5$). The tables present results for $\tau \in \{0.50,0.75\}$, based on different combinations of $N\in \{25,50,100,200\}$ and $T\in \{5,10,50,100\}$. The number of bootstrap repetitions is set to 400, and the results are obtained by using 1000 random samples.
The tables show results for two bootstrap methods. The cross-sectional pairs bootstrap (CS) samples over $i$ with replacement, keeping the entire block of time series observations. The wild bootstrap (WB) is implemented as discussed in Section (ref). We first obtain residuals $\hat{u}_{it}$ using the penalized quantile regression (ref), which is labeled `PQR' in the tables. The tuning parameter is obtained as $\hat{\lambda}_T = b_T \tilde{\lambda}$ where $\tilde{\lambda}$ is obtained by cross-validation and $b_T=0.5 T^{-\nu}$ controls the bias. The selection of $\nu = 1$ performed well in the simulations and it is consistent with Theorem (ref). As in the case of the wild bootstrap estimator proposed by Feng, He, and Hu (2011), a finite sample correction is recommended. We adopt an adjustment following closely the {\tt R} package {\tt quantreg} by Koenker (2021)\nocite{rK21}. In our case, we adjust the residuals with the influence function and sign function following the Bahadur representation of the estimator derived in Theorem (ref). Then, we generate $u^\ast_{it} = w_{it} | \hat{u}_{it} |$, where $w_{it}$ is an i.i.d. random variable distributed as a two-point distribution with probabilities $\tau$ and $1-\tau$ at $w_{it} = -2 \tau$ and $w_{it} = 2 (1-\tau)$. Lastly, we generate the dependent variable as $y_{it}^\ast = \hat{\alpha}_i + \hat{\beta} x_{it} + u^\ast_{it}$. The performance of the estimator (ref) was similar and the results are not presented here to save space. Finally, we include the estimator (ref) defined for $\lambda_T = 0$ and it is labeled `FE'.
Following the result presented in Theorem (ref), the coverage probabilities in Table (ref) are obtained considering the quantiles of the empirical distribution of $\sqrt{NT} (\beta^\ast - \hat{\beta})$. As can be seen in the upper block of Table (ref), the performance of the WB bootstrap estimators are excellent, and they are in general around the specified coverage probability. Furthermore, performance improves with $T$, and tends to be similar for both $0.5$ and $0.75$ quantiles. On the other hand, the performance of the CS estimator is poor, with estimates not approaching to specified nominal values. In the lower parts of the table, we present the performance of the estimators for different distributions $F_u$. The WB method continues to perform better than CS, and, as expected, the estimation of the higher quantile is more challenging in the $\chi_3^2$ case. In all the variations of the model considered in the table, the WB estimator performs much better than the CS estimator.
The results for the location-scale shift model presented in Table (ref) are similar. We continue to see that the WB bootstrap performs better than the CS method. This conclusion holds when we consider asymptotic Gaussian confidence intervals obtained using bootstrap standard errors $\mbox{se}(\beta^\ast)$ (see Tables (ref) and (ref)). Moreover, the tables confirm two results that were expected. First, as $T$ increases relative to $N$, the coverage of the WB improves. Second, the performance of WB in the case of $\lambda_T=0$ reveals that, in general, the procedure proposed in this paper is valid for approximating the distribution of the fixed effects estimator.
We finish the section by briefly documenting the relative performance of the estimators of the standard errors. We generate data from a location-scale shift model ($\zeta=0.5$) when the error term $u_{it} \sim \mathcal{N}(0,1)$ and $\alpha_i \sim \mathcal{N}(0,1)$, by setting $N=100$, $T=10$, and $\tau = 0.5$. The left panel of Figure (ref) shows CS and WB bootstrap estimates of the standard error, $\mbox{se}(\beta^\ast)$, and the standard deviation of the penalized estimator, $\mbox{sd}(\hat{\beta})$. The figure shows the advantage of the penalized estimator relative to the fixed effects estimator, as the standard deviation of the estimator is decreasing as $\lambda_T$ increases. We also see that the WB procedure performs better than CS when $\lambda_T$ is relatively small, and the performance of the WB estimator does not seem to change over the degree of shrinkage of the individual effects, as the bias appears to be roughly constant over $\lambda_T$. Using the right panel in Figure (ref), we explore further the difference in performance between approaches. The empirical distribution obtained by the CS procedure is not centered at the true value, and the distribution of the standard error of the WB is centered at $\mbox{se}(\hat{\beta}) = 0.081$ (with $\lambda_T = 0.05$).
In this section, we investigate the consistency of the wild bootstrap under different conditions. First, we extend the results of Theorems (ref) and (ref) to allow for dependent data, and then we focus on the consistency of the wild bootstrap. In such case, we use the following assumptions:
Theorem (ref) presents both consistency and asymptotic normality results for the estimator with dependent error terms.
Theorem (ref) shows consistency of the bootstrap distribution estimator in the case of dependent errors. This more complex situation requires another assumption:
Assumption (ref) is a high-level assumption on the distribution of bootstrap weights. The assumption guarantees that the variance of the bootstrap estimator is bounded and sufficiently close to the true variance, because the weights mimic the within-unit dependence structure of the errors. A feasible version could use a plug-in estimate of the average of the joint conditional CDFs of $(u_{it}, u_{it+j})$ to generate weights that satisfy the average probability.
Finally, we investigate if the conditions on the size of $T$ relative to $N$ needed for the asymptotic normality in Theorem (ref) can be improved, especially in the light of recent work by AGalvao2020. If instead of focusing on the stochastic order of the terms of the Bahadur representation of the penalized estimator, we focus on the expected values of the remainder terms, it is possible to show that the rates can be improved substantially. In order to show asymptotic normality, we employ the following assumption about the behavior of the penalty parameter.
Assumption (ref) dictates the rate at which the probability of observing large a $\lambda_T$ becomes small asymptotically. As illustrated in remark (ref), it is needed to provide a tail bound for the distribution of individual effects, which figure in the remainder terms of the Bahadur representation used to find the asymptotic distribution of $\bm{\hat{\beta}}$ (such a bound holds naturally for terms related to minimizing the quantile regression objective function with bounded regressors, a fact used extensively in AGalvao2020). In the theorem below, we require $\lambda_T = O_p(\log T) = o_p(T^{1/2} (\log T)^{1/2})$, so this assumption only mildly strengthens the other regularity conditions.
The proof in Theorem (ref) uses an infeasible estimator $\tilde{\alpha}_i$ that is obtained considering $T$ observations $y_{it} - \bm{x}_{it}'\bm{\beta}_0$. The difference between $\tilde{\alpha}_i$ and $\hat{\alpha}_i$ converges to zero as the slope coefficient $\hat{\bm{\beta}}$ converges in probability towards $\bm{\beta}_0$, under the condition on $\lambda_T$. Therefore, the remainder terms of the corresponding Bahadur representations are sufficiently close, leading to the improvements in the rates first obtained in AGalvao2020 for the fixed effects estimator. We now show the consistency of the bootstrap distribution estimator under these relatively closer orders of $N$ and $T$.
In recent years, policy makers and the general public have been debating and re-evaluating several aspects of trade, including the benefits of trade agreements mB2001, sH2016. An important question is whether workers have been negatively affected by the North American Free Trade Agreement (NAFTA), which was signed by the governments of the United States of America, Canada, and Mexico in 1993. Hakobyan and McLaren (2016) find that the effect of NAFTA on average wage growth in the period 1990-2000 was negative. In this section, we use similar data and apply our approach to study the distributional impact of NAFTA. Our findings suggest that the agreement increased wage inequality. Low-wage workers experienced significant negative wage growth, while high-wage workers experienced, in general, significant positive wage growth. Our results are similar to evidence on the effect of Chinese imports on low-wage American workers denisChet2016.
Following Hakobyan and McLaren (2016), we use a 5% sample from the U.S. Census. We employ two cross-sectional samples in the year 1990 and 2000, and therefore, workers in the sample are observed once. The longitudinal nature of the analysis comes from exploiting the fact that we observe multiple individuals in a given industry and location. The sample includes workers between 25 and 64 years of age who reported positive income. We have demographic information including age, gender, marital status, race, and educational attainment of the worker classified in four categories: high school dropout, high school graduate, some college, and college graduate.
The data on U.S. tariffs and Mexico's revealed comparative advantage (RCA) are obtained from Hakobyan and McLaren (2016). Using their data, we have access to average U.S. tariffs by industry of employment of the worker and location (or Consistent Public-Use Microdata Area, abbreviated conspuma) of residence of the worker. In 1990, the average tariff by industry in 1990 was 2.1% percent (with a standard deviation of 3.9%), while the average local tariff by conspuma level was 1.03% (with a standard deviation of 0.67%). In the period 1990-2000, the tariffs decreased 1.7% at the industry level and 0.9% at the conspuma level. These descriptive statistics are used in the next section to estimate the percentage change in wages associated with the reduction in tariffs. We consider all industries with the exception of agriculture.
To investigate the effect of NAFTA on the wages of American workers, we consider a specification that allows for the impact of the trade agreement to vary by industry, location, and educational attainment of the worker. To that end, we consider the following model as in Hakobyan and McLaren (2016):
where the response variable $y_{ijc}$ is the logarithm of wages for worker $i$, who is employed in industry $j$ and resides in conspuma $c$, $\bm{L}_{ic}$ and $\Delta \bm{L}_{ic}$ are location variables to be described below, $\bm{I}_{ij}$ and $\Delta \bm{I}_{ij}$ are industry variables, $\bm{X}_{ijc}$ is the vector of control variables considered in Hakobyan and McLaren (2016), and $\alpha_{jc}$ is a industry-conspuma effect. The error term is denoted by $u_{ijc}$.
The location variables are defined as $\bm{L}_{ic} = (L_{ic,1},L_{ic,2},L_{ic,3},L_{ic,4})'$, where $L_{ic,k}$ is the product of an indicator for educational category $k$ of worker $i$, an indicator variable for whether $i$ is in the 2000 sample, and the average tariff in the conspuma of residence of worker $i$. Similarly, we can define $\Delta \bm{L}_{ic} = (\Delta L_{ic,1},\Delta L_{ic,2},\Delta L_{ic,3},\Delta L_{ic,4})'$, as the change in $\bm{L}_{ic}$ due to the change in tariffs between 1990 and 2000 in the conspuma of residence of worker $i$. In terms of the industry variables, $\bm{I}_{ij} = (I_{ij,1},I_{ij,2},I_{ij,3},I_{ij,4})'$, where $I_{ij,k}$ is the product of an indicator for educational category $k$, the RCA in industry $j$, an indicator variable for whether $i$ is in the 2000 sample, and the tariff of the industry that employs worker $i$. Similarly, we define $\Delta \bm{I}_{ij} = (\Delta I_{ij,1},\Delta I_{ij,2},\Delta I_{ij,3},\Delta I_{ij,4})'$, as the change in $\bm{I}_{ij,k}$ due to the tariff change between 1990 and 2000 in the industry that employs worker $i$.
Because industry latent factors and trends in some areas can affect wages and also the changes in tariffs, we employ the penalized estimator (ref) to estimate a high-dimensional model with more than 84,000 parameters $\alpha_{jc}$. The parameters of interest in equation (ref) are $\bm{\beta}_{1L}$, $\bm{\beta}_{2L}$, $\bm{\beta}_{1I}$, and $\bm{\beta}_{2I}$, which measure the initial effect of tariffs by location and industry ($\bm{\beta}_{1L}$ and $\bm{\beta}_{1I}$), and the impact effect of a reduction of tariffs by location and industry ($\bm{\beta}_{2L}$ and $\bm{\beta}_{2I}$). Using these parameters, it is possible to obtain the effect of the trade agreement on wages. For instance, for locations that lost all of their protection after the introduction of NAFTA, the effect of the local average tariff is measured by $\bm{\beta}_{1L} - \bm{\beta}_{2L}$. Similarly, for industries that lost all of their protection, the effect of the industry tariff is $\bm{\beta}_{1I} - \bm{\beta}_{2I}$.
Table (ref) reports results for the coefficients $\bm{\beta}_{1I}$, and $\bm{\beta}_{2I}$ for the four educational categories. The table also shows results for $\beta_{1I,k} - \beta_{2I,k}$ for each educational category $k$ and p-values (in brackets) of Wald-type tests for the null hypothesis $\mbox{H}_0: \beta_{1I,k} = \beta_{2I,k}$. The variance of the test is obtained using the proposed wild residual bootstrap procedure. The first column presents mean fixed effects regression results, that is, estimation of model (ref) by least squares methods. The last five columns show penalized quantile regression (PQR) results with $\lambda_T$ selected by cross-validation. The standard errors are obtained by the proposed wild residual bootstrap procedure. To save space, we do not present results on the control variables included in the vector $\bm{L}_{ic}$, $\Delta \bm{L}_{ic}$, and $\bm{X}_{ijc}$, but the fixed effects results shown in the first column are similar to the results in Table 4 (column (2)) in Hakobyan and McLaren (2016).
Looking at the first set of estimates in the first rows, we see that an initial tariff estimate equal to 2.02 and an impact effect of 3.57. Based on the standard deviation of tariffs at the industry level, a 1% standard deviation increase in the initial industry tariff has an effect of reducing wages by $3.9\% \times -1.55$, or $-6.05\%$ in the period 1990-2000. This implies that, among industries with tariff declining after the introduction of NAFTA, average wage growth is negative for high school dropouts. The results, however, show that the average response does not summarize well the distributional impact of NAFTA. While the industry effect, which is measured as the difference between the initial effect and the impact effect, is negative ($-1.93$, or $-7.50\%$) and significant for high school dropouts at the 0.1 quantile, it is small ($-0.17$, or $-0.65\%$) and insignificant at the 0.9 quantile. Moreover, we find that the largest differences between the 0.1 and 0.9 effects are among college graduates in industries that lost all of their protection, suggesting that wage growth has been also unequal by educational attainment.
Lastly, using Figure (ref), we report point estimates and confidence intervals for the location and industry effects for high school dropouts and college graduates. The evidence reveals that inequality increased in the period after the implementation of the trade agreement.
In this article, we address the problem of estimating the distribution of the penalized quantile regression estimator for longitudinal data using a wild residual bootstrap procedure. Originally introduced by Koenker (2004) as a convenient alternative to the quantile regression estimator with fixed effects, the practical use of the penalized estimator has been limited by challenges involving inference. We show that the wild bootstrap procedure is asymptotically valid for approximating the distribution of the penalized estimator. We derive a series of new asymptotic results and carry out a simulation study that indicates that the wild residual bootstrap performs better than an alternative bootstrap approach commonly used in practice for similar estimators that do not include a penalty term.
Although the paper makes an important contribution by providing a valid method for statistical inference, there are several questions that remain to be answered. We believe that the procedure leads to valid inference in the case of $J$ quantiles estimated simultaneously, but we leave this to future research. Moreover, under an assumption of sparsity as in other high-dimensional models, we expect changes in the consistency and asymptotic normality results. In terms of theoretical developments, we did not consider the case where $\alpha_i$ is a random effect. Lastly, the practical implementation of the wild bootstrap in the case of dependent data involves a few challenges. We hope to investigate these directions in future work.