EconBase
← Back to paper

Panel Quantile Regression with Common Shocks

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

65,099 characters · 10 sections · 77 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Panel Quantile Regression with Common Shocks

abstractThis paper develops an asymptotic and inferential theory for fixed-effects panel quantile regression (FEQR) that delivers inference robust to pervasive common shocks. Such shocks induce cross-sectional dependence that is central in many economic and financial panels but largely ignored in existing FEQR theory, which typically assumes cross-sectional independence and requires $T \gg N$. We show that the standard FEQR estimator remains asymptotically normal under the mild condition $(\log N)^2/T \to 0$, thereby accommodating empirically relevant regimes, including those with $T \ll N$. We further show that common shocks fundamentally alter the asymptotic covariance structure, rendering conventional covariance estimators inconsistent, and we propose a simple covariance estimator that remains consistent both in the presence and absence of common shocks. The proposed procedure therefore provides valid robust inference without requiring prior knowledge of the dependence structure, substantially expanding the applicability of FEQR methods in realistic panel data settings.

JEL Classification: C15, C23, C31, C80

Keywords: quantile regression, panel data, common shocks, asymptotic theory.

Introduction

Quantile regression koenker1978regression offers a flexible framework for analysing heterogeneity across the conditional distribution of an outcome and is particularly valuable for studying tail behaviour and outcomes with heavy tails or outliers. In many economic and financial applications, data are naturally organised as panels, making panel quantile regression with individual fixed effects (FEQR) an important tool for modelling distributional heterogeneity while controlling for unobserved unit-specific heterogeneity (Koenker04). A salient empirical characteristic of such panels, however, is the prevalence of strong common shocks.

Common shocks—aggregate disturbances or latent factors that affect many cross-sectional units within a given period—are a central and recurring feature of economic and financial data. Much empirical work is built around this structure; for example, asset-pricing frameworks such as the two-pass procedure emphasise time variation in risk premia driven by economy-wide conditions rather than purely idiosyncratic firm shocks fama1973risk. More broadly, macroeconomic announcements, monetary policy changes, business-cycle fluctuations, regulatory reforms, and shifts in global risk sentiment generate substantial comovement across units, so regression disturbances typically display cross-sectional dependence even after conditioning on observables. As emphasised in andrews2003cross , this pervasiveness reflects the fact that key determinants of behaviour—interest rates, inflation, financial conditions, policy actions, institutional and legal changes, political events, environmental and health shocks, and technological change—operate at an aggregate level, simultaneously influencing wages, consumption, investment, costs, wealth, and credit conditions for large sets of agents. Although exposure varies with sector, location, or balance-sheet characteristics, these forces link units through shared systemic influences, and with deepening economic integration their transmission has become more widespread. Hence, cross-sectional independence is rarely a credible assumption in practice; common shocks are a structural feature of economic environments rather than an exceptional complication.

These features have first-order econometric consequences. In cross-sectional settings, andrews2005cross shows that least squares can be inconsistent when an unobserved common shock affects both regressors and errors, inducing systematic dependence between them. Intuitively, the key requirement for consistency--that regressors are uncorrelated with the error term--is violated once this dependence varies with the realisation of the common shock, so that averaging across individuals no longer eliminates the bias. An analogous issue arises in quantile regression, where the score object replaces the product of regressor and error; if this object also depends on the common shock in a systematic way, similar inconsistency emerges. In panel data with common shocks, the error process violates cross-sectional independence and often lies outside the weak-dependence frameworks underlying classical panel asymptotics. driscoll1998consistent show that conventional covariance estimators are generally inconsistent under such dependence and propose standard errors that are robust to broad forms of cross-sectional and temporal correlation. In empirical finance settings, petersen2008estimating demonstrate that ignoring common time effects leads to substantial size distortions, with severely overstated t-statistics. This concern is especially acute in panels with fairly large $N$ but moderately large $T$, where limited time variation restricts information about aggregate components.

Against this background, the theoretical literature on FEQR has progressed considerably. Koenker04 first studies the estimation problem of a linear quantile regression model with individual fixed effects, proposing an $\ell_1$-regularised estimator to mitigate the incidental parameter problem and analysing its asymptotic properties. canay2011simple propose a two-step estimator that eliminates fixed effects under a location-shift restriction, delivering standard large-$N$, large-$T$ asymptotics. Without imposing a location shift, kato2012asymptotics develop a general asymptotic theory for Koenker-type FEQR with individual specific effects, establishing consistency and asymptotic normality under joint growth of $N$ and $T$ subject to a long-panel rate condition such as $N^2(\log N)^3/T \to 0$, and highlighting the challenges posed by incidental parameters and the non-smooth objective. galvao2016smoothed introduce a kernel-smoothed version of the quantile objective to enable higher-order expansions, characterise the incidental-parameter bias, and propose analytic bias correction. More recently, galvao2020unbiased obtain asymptotically unbiased normal approximations under a weaker long-panel requirement, $N(\log T)^2/T \to 0$, which is standard in the nonlinear panel data models literature (see, e.g., ArellanoHahn07, ArellanoBonhomme11, and FernandezValWeidner18). Nevertheless, these results still rely on regimes in which the time dimension dominates the cross section, effectively requiring $T \gg N$, which limits applicability in many economic and financial panels. Other recent approaches include MachadoSantosSilva19 considering a location-scale shift effect for the individual effects, while adopting a location-scale shift specification also for the slope parameter, and GuVolgushev19 and ZhangWangLianLi2023 investigating the estimation of panel quantile regression models with group effects.\footnote{The panel quantile literature offers different identification and estimation strategies including penalized fixed effects, random effects, correlated random effects, group effects, instrumental variables, and factor models (see, e.g., Koenker04; AbrevayaDahl08; Lamarche10; KimYang11; ChetverikovLarsenPalmer16; ArellanoBonhomme16; GrahamHahnPoirierPowell2015; demetrescu2023tests).}

Despite its increasing empirical relevance, the theory of fixed-effects panel quantile regression still exhibits two notable gaps. First, most existing results abstract from common shocks and therefore do not account for the cross-sectional dependence that is inherent in many economic and financial panels. In the linear regression context, theoretical analysis of common time effects in cross-sectional regressions was developed by andrews2005cross. Comparable results, however, appear to be unavailable for panel quantile regression models. Second, even the most advanced analyses rely on stringent long-panel conditions such as $N(\log T)^2/T \to 0$ galvao2020unbiased, effectively requiring the time dimension to dominate, which is hardly satisfied in empirical applications. Indeed, besstremyannaya2019reconsideration review $81$ empirical papers with panel quantile models and document that only $6$ involve data with $N<T$, whereas $55$ of them have $N\ge 10T$.

In this paper, we develop a novel framework for FEQR models that accommodates common time effects while preserving asymptotic normality as both \(N\) and \(T\) diverge, thereby covering a wide range of asymptotic regimes, including \(N \asymp T\) and \(T \ll N\). The inclusion of common shocks alters both the convergence rate and the form of the asymptotic covariance, rendering existing inference procedures for conventional FEQR inconsistent. This relaxation of conventional rate restrictions is driven by a phenomenon we term data-generating process (DGP)–induced smoothing: conditional averaging over latent shocks inherently smooths the empirical criterion and concentrates its stochastic variation around a smooth component, thereby enabling refined stochastic expansions. Exploiting this structure, we establish uniform convergence and asymptotic normality of the standard unregularised FEQR estimator. In addition, we propose a new simple robust covariance estimator that is consistent under both common-shock and classical FEQR settings. Consequently, valid inference does not require prior knowledge of whether common shocks are present. These results contribute to the broader literature on nonlinear panel models with large panels \((N,T \to \infty)\).

In cross-sectional settings with common shocks, the logic of andrews2005cross implies that quantile regression may be inconsistent. The panel setting, however, fundamentally changes this conclusion. When units are observed repeatedly over time and the common shocks vary independently across periods, temporal averaging attenuates their influence and restores consistency of FEQR. This does not mean that the shocks are asymptotically negligible. Rather, their effect remains first-order, entering the limiting distribution at a rate determined by the time dimension and thereby shaping the estimator's leading asymptotic behaviour.

Another strand of the literature attempts to accommodate common shocks through factor-based approaches to panel quantile regression. chen2021quantile develop identification and estimation for quantile factor models with interactive individual and time effects, requiring large \(N\) and \(T\). Related contributions include AndoBai2020, who use principal components for heterogeneous parameter models; MaLintonGao2021, who study semiparametric quantile factor models; HardingLamarchePesaran2020, who analyse dynamic quantile models with factors; BelloniChenPadillaWang2023, who consider high-dimensional latent heterogeneity; and Chen2024, who proposes a two-step estimator for models with interactive fixed effects.

Our paper differs from this line of work in that we do not impose an explicit factor structure. Instead, within a classical linear quantile regression specification with individual fixed effects first proposed by Koenker04, we allow time-specific common shocks to enter the regressors and the error term in a highly flexible and unspecified manner. This modelling choice preserves the familiar and interpretable fixed-effects quantile regression structure while accommodating rich forms of cross-sectional dependence. The trade-off is that we do not attempt to identify or recover an underlying factor or interactive fixed-effects structure. In our view, this approach strikes a reasonable balance between interpretability and flexibility. At the same time, it remains closely aligned with specifications that are widely utilised by empirical researchers and is simple and straightforward to implement computationally.

The reminder of the paper is organised as follows. Section (ref) presents the model and estimation procedure. In Section (ref) we formalise the main statistical properties of the estimator. Section (ref) provides practical inference procedures and establishes their validity. Section (ref) provides Monte Carlo simulations. Finally, Section (ref) concludes.

Notations

Let $\mathbb{N}$ denote the set of positive integers and $\mathbb{R}$ the real line. For a finite-dimensional vector $a$, let $\|a\|$ denote its Euclidean norm, and for a matrix $A$, let $A'$ denote its transpose. The probability spaces we consider are assumed to be standard Borel. Throughout the remainder of the paper, all asymptotic statements refer to regimes in which both $N$ and $T$ diverge to infinity.

Model and Estimation

We begin by introducing the modelling framework; further motivation is provided in Remark (ref) below. Consider the data-generating process

align[align omitted — 72 chars of source]

where the regressor vector $X_{it}$ excludes an intercept. The latent variables $\{A_i\}_{i\in\mathbb{N}}$, $\{B_t\}_{t\in\mathbb{N}}$, and $\{U_{it}\}_{(i,t)\in\mathbb{N}^2}$ are mutually independent and unobserved by the econometrician, and $g:\mathcal A\times \mathcal B\times\mathcal U\to \mathbb{R}^{p+1}$ is an unknown Borel-measurable function. For simplicity, the common time shocks $B_t \in \mathcal{B}$ are assumed i.i.d.\ across $t$, and the idiosyncratic shocks $U_{it} \in \mathcal{U}$ are i.i.d.\ across $(i,t)$. No identical sampling or particular dependence structure is imposed on the unit-specific effects $A_i \in \mathcal{A}$. We impose no smoothness-type restrictions on the function $g$ with respect to these latent shocks. Moreover, the spaces $\mathcal A$, $\mathcal B$, and $\mathcal U$ are allowed to be high-dimensional and structurally complex.

In the spirit of Koenker04, we specify a quantile regression model with individual fixed effects:

align[align omitted — 111 chars of source]

Here, similar to kato2012asymptotics, the fixed effects are allowed to vary across quantile levels. For the purpose of asymptotic analysis, we condition on a realisation of the individual effects, $\{A_i\}_{i=1}^N = \{a_i\}_{i=1}^N$. Under this conditioning, the model can be written as

align*[align* omitted — 139 chars of source]

where $\alpha_{i0}(\tau) = \alpha_0(\tau; a_i)$. As a consequence, the regressor--error pair then admits a representation,

align[align omitted — 83 chars of source]

where $g_i(b,u) = g(a_i, b, u)$. We impose no parametric restrictions on the relationship between the individual fixed effects $\alpha_{i0}(\tau)$ and the regressors $X_{it}$; in particular, the fixed effects are allowed to vary flexibly with $\tau$. Moreover, under this framework, common shocks are permitted to exert heterogeneous effects across population units, with the magnitude and direction of these effects depending on unit-specific characteristics, which may be either observed or unobserved.

remark[The rationale behind the structures of Equations (ref) and (ref)] The nonparametric representation in (ref) is closely related to andrews2005cross, with similar modelling frameworks considered in chernozhukov2013average and hausman2025linear in panel data contexts. Fundamentally, it is highly general and can be viewed as an application of the de Finetti’s theorem under exchangeability. In particular, for each fixed $i$, de Finetti’s theorem implies that exchangeable observations are conditionally independent given a latent variable, yielding a representation of the form in (ref)--see Lemma 2 in Section 7 in andrews2005cross. Exchangeability arises naturally in a wide range of economic models. For example, in dynamic oligopoly models with investment, athey2001investment show that exchangeable profit functions are consistent with standard environments such as Cournot and differentiated-product competition, reflecting a symmetry restriction whereby firms’ payoffs depend on rivals’ actions and states but not their identities. This symmetry is closely related to the notion of anonymity in cooperative game theory and social choice theory (e.g.\ moulin1991axioms). Similarly, in differentiated product markets, berry1995automobile argue that demand and cost functions are exchangeable in competitors’ characteristics and show that equilibrium uniqueness implies exchangeability in both observed and unobserved components. Likewise, the factor-type representation in (ref) is motivated by the literature on exchangeable arrays and network-type dependence bickel2011method, graham2024sparse, as well as by work on two-way clustering and dependence structures in panel data menzel2021bootstrap,davezies2021empirical, chiang2024standard,athey2025identification. Under a suitable exchangeability condition, the existence of such data-generating processes is guaranteed by the Aldous--Hoover representation theorem (see Chapter 7 of kallenberg2005probabilistic). $\lozenge$

Given data $\{(Y_{it}, X_{it}') : i=1,\ldots,N,\; t=1,\ldots,T\}$, we jointly estimate the common parameter $\beta_0(\tau)$ and the fixed effects $\{\alpha_{i0}(\tau)\}_{i=1}^N$ using the standard Koenker-type FEQR estimator Koenker04 without regularisation:

align[align omitted — 228 chars of source]

where $\rho_\tau(u) = u\big(\tau - \mathbbm{1}\{u < 0\}\big)$ is the check loss function. Unless stated otherwise, we fix $\tau \in (0,1)$ throughout and suppress its dependence in the notation.

From a computational perspective, the FEQR estimator is straightforward to implement using existing optimisation routines. The objective function is convex in $(\bm{\alpha}, \beta)$, so estimation reduces to a standard linear programming problem. Consequently, the estimator can be readily computed using widely available quantile regression packages that accommodate fixed effects or high-dimensional controls, such as quantreg in R, statsmodels and cvxpy in Python, as well as built-in procedures in standard econometric software. Modern algorithms for large-scale convex optimisation make it feasible to handle panels with a large number of cross-sectional units, ensuring practical scalability in empirically relevant applications.

Asymptotic Theory

In this section, we study the theoretical properties of the FEQR estimator under the common shock framework.

We denote by $F_i$ and $f_i$ the distribution function and density of $\epsilon_{it}=g_i^\epsilon(B_t,U_{it})$. Further, let $F_{i,X}(\cdot \mid x)$ and $f_{i,X}(\cdot \mid x)$ denote the conditional distribution function and density of $\epsilon_{i1}$ given $X_{i1}=x$, and let $F_{i,B}(\cdot \mid b)$ and $f_{i,B}(\cdot \mid b)$ denote the corresponding objects given $B_1=b$. Likewise, let $F_{i,XB}(\cdot \mid x,b)$ and $f_{i,XB}(\cdot \mid x,b)$ denote the conditional distribution function and density of $\epsilon_{i1}$ given $X_{i1}=x$ and $B_1=b$. When no confusion is likely to arise, we suppress the subscripts $X$ and/or $B$ and write, with slight abuse of notation, expressions such as $f_i(\cdot \mid x)$ and $F_i(\cdot \mid x,b)$. Finally, let $\mathcal{X}_i$ and $\mathcal{E}_i$ denote the support of $X_{i1}$ and $\epsilon_{i1}$, and write $\mathcal{X} = \bigcup_{i \geq 1} \mathcal{X}_i \subset \mathbb{R}^p$ and $\mathcal{E} = \bigcup_{i \geq i} \mathcal{E}_i \subset \mathbb{R}$.

Uniform Consistency

We first present the assumptions needed for uniform consistency.

assumption[Bounded regressors] $\mathcal{X}$ is a bounded set in $\mathbb{R}^p$ for a fixed $p$.
assumption[Identification] For each $\delta>0$, \begin{align*} \epsilon_\delta:=\inf_{i\ge 1} \inf_{|\alpha|+\|\beta\|_1}\mathbbm{E}\left[\int_{0}^{\alpha+X_{i1}'\beta}\{F_i(s|X_{i1})-\tau\}ds\right]>0, \end{align*} where $\|\cdot\|_1$ is the $\ell_1$ norm.

Assumption (ref) is imposed for technical convenience and could in principle be relaxed to sub-gaussian tail bounds or suitable polynomial moment conditions, albeit at the expense of substantially more involved arguments. This assumption is recurrent in the quantile regression literature, see, for example, Koenker04 and ChernozhukovHansen06. Assumption (ref) is a standard identification condition that can be found in, e.g. fernandez2005bias and kato2012asymptotics.

Under Assumptions (ref) and (ref), the following lemma shows that both the individual effects estimators $\hat{\alpha}_i$ and the common parameter estimator $\hat{\beta}$ are uniformly consistent for their true values.

proposition[Uniform consistency] Under Assumptions (ref) and (ref) and suppose that $(\log N)^2 / T \to 0$, we have $\max_{1 \leq i \leq N}|\hat{\alpha}_i - \alpha_{i0}| \vee \|\hat{\beta} - \beta_0\| \overset{p}{\rightarrow} 0$.

A proof is provided in Section (ref) of the Appendix. The argument closely parallels that of Theorem 3.1 in kato2012asymptotics, with only minor modifications. In particular, the presence of common shocks does not materially complicate the consistency proof. By contrast, establishing asymptotic normality requires a substantially different line of argument.

Asymptotic Distribution

To establish asymptotic distributional theory, we impose the following additional assumptions. Write $P_i$ as the joint law of $(\epsilon_i, X_{i1}, B_1)$ on $\mathcal{E}_i \times \mathcal{X}_i \times \mathcal{B}$.

assumption[Joint densities] For each $i$, $(\epsilon_{i1}, X_{i1}, B_1)$ admits a joint density on $\mathcal{E}_i \times \mathcal{X}_i \times \mathcal{B}$ with respect to $\lambda \otimes P_{i, XB}$, where $\lambda$ denotes Lebesgue measure on $\mathbb{R}$ and $P_{i,XB}$ is the marginal distribution of $(X_{i1}, B_1)$.

Henceforth, “for all” is understood to mean $P_i$-almost everywhere.

assumption[Smoothness of densities] For each $i \in \mathbb{N}$, \begin{enumerate} • $f_i(\epsilon \mid x, b)$ is continuously differentiable with respect to $\epsilon$ for all $(x, b)$ in $\mathcal{X}_i \times \mathcal{B}$. • There exists finite constant $C_f$ and $L_f$ such that $$f_i(\epsilon \mid x, b) \leq C_f, \quad\text{and } \quad|\partial f_i(\epsilon \mid x,b) / \partial \epsilon| \leq L_f$$ uniformly over $(\epsilon, x, b) \in \mathcal{E}_i \times \mathcal{X}_i \times \mathcal{B}$ and $i \in \mathbb{N}$; • $f_i(0 \mid b)$ is bounded from below by some positive constant independent of $i$ and $b$. \end{enumerate}
remarkAssumption (ref) is an if and only if condition for the almost-sure existence of a conditional density of $\epsilon_{i1}$ given $(X_{i1}, B_1)$. Note that the distribution of $X_{i1}$ may have point masses, so the regressors are allowed to be discrete-valued. Moreover, we do not require $\mathcal{B}$ to be finite-dimensional. Since $\mathcal{E}_i \subset \mathbb{R}$ is a standard Borel space, a regular conditional distribution of $\epsilon_{i1}$ given $(X_{i1}, B_1)$ always exists; see, for example, Theorem 8.5 of kallenberg2021foundation. Assumption (ref) ensures that this conditional distribution is absolutely continuous with respect to Lebesgue measure for $P_{i, XB}$-almost every $(x, b)$.

Assumptions (ref) and (ref) are mild, and primarily ensure the existence and sufficient smoothness of the conditional density of $\epsilon_{i1}$ given $(X_{i1}, B_1)$.

Note that the existence of the conditional density of $\epsilon_{i1}$ given $(X_{i1}, B_1)$ implies the existence of the conditional density of $\epsilon_{i1}$ given $X_{i1}$ and the marginal density of $\epsilon_{i1}$. Let us define $\gamma_i = \mathbbm{E}[f_i(0 \mid X_{i1})X_{i1}] / f_i(0)$, which plays a key role in the expression of the asymptotic covariance introduced below.

assumption[Variance with common shocks] Suppose that $\sup_{i\in \mathbb N} \|\gamma_i\|<\infty$. Further assume the following matrix exists and is nonzero \begin{align} \Sigma= \lim_{N \to \infty} {Var}\left( \frac{1}{N} \sum_{i = 1}^N \mathbbm{E}[(\tau - \mathbbm{1}\{\epsilon_{i1} \leq 0\})(X_{i1} - \gamma_i)\mid B_1]\right) . \end{align}

Assumption (ref) implicitly requires that the common time effect has a non-negligible impact on the distribution of the data. This condition rules out the classical fixed-effects quantile regression framework analysed in kato2012asymptotics and subsequent studies, where—after controlling for individual fixed effects—the disturbances are independent across $(i,t)$ pairs, corresponding to a degenerate special case of our setting. The assumption is conceptually related to the non-degeneracy conditions imposed in the multiway clustering literature; see, for example, chiang2024standard.

In the cross-sectional setting without panel data, andrews2005cross shows that the OLS estimator is inconsistent when the common shock generates a nonzero conditional covariance between the regressor and the error term. This is because the consistency of OLS relies on $\mathbbm{E}[X_i \epsilon_i] = 0$, and hence if $X_i \epsilon_i$ systematically depends on the common shock, OLS is no longer consistent. In our notation, this can be written as $\text{\upshape{Var}}\left( \mathbbm{E}[X_i \epsilon_i \,|\,B]\right) \neq 0$, where $B$ is the common shock. The anagolous object in quantile regression is the score $(\tau - \mathbbm{1}\{\epsilon_i \leq 0\})X_i$, so the corresponding condition would be $\text{\upshape{Var}}\left( \mathbbm{E}[(\tau - \mathbbm{1}\{\epsilon_i \leq 0\}) X_i\,|\,B]\right) \neq 0$.

In our setting, we allow the joint distribution of $(X_{it}, \epsilon_{it})$ to vary across $i$, and so the analogous assumption in panel data quantile regression is captured by (ref). Moreover, in panel data, the availability of repeated observations over time preserves the consistency of $\hat{\beta}$, as shown in Proposition (ref). However, while consistency is retained, common shocks continue to affect the estimator through its asymptotic distribution.

remark[Structure of asymptotic covariance] This term $\Sigma$ constitutes the middle component of the sandwich-form asymptotic covariance of $\hat\beta$. In contrast to the corresponding middle component in kato2012asymptotics and galvao2020unbiased, which, in our notation, is \begin{align} \Omega=\lim_{N \to \infty}\frac{1}{N}\sum_{i=1}^N {Var} \bigl((\tau - \mathbbm{1}\{\epsilon_{i1}\le 0\})(X_{i1}-\gamma_i)\bigr), \end{align} it admits a fundamentally different structural representation. To see this, consider the variance term appearing inside the limit in the definition of $\Sigma$: \begin{align*} &{Var}\!\left( \frac{1}{N} \sum_{i = 1}^N \mathbbm{E}[(\tau - \mathbbm{1}\{\epsilon_{i1} \le 0\})(X_{i1} - \gamma_i)\mid B_1]\right)\\ &= \frac{1}{N^2} \sum_{i = 1}^N {Var}\!\left( \mathbbm{E}[(\tau - \mathbbm{1}\{\epsilon_{i1} \le 0\})(X_{i1} - \gamma_i)\mid B_1]\right) \\ &\quad + \frac{1}{N^2} \sum_{i = 1}^N\sum_{j\ne i} {Cov}\!\Big( \mathbbm{E}[(\tau - \mathbbm{1}\{\epsilon_{i1} \le 0\})(X_{i1} - \gamma_i)\mid B_1], \mathbbm{E}[(\tau - \mathbbm{1}\{\epsilon_{j1} \le 0\})(X_{j1} - \gamma_j)\mid B_1] \Big). \end{align*} The distinction is not merely the presence of the cross-sectional covariance terms. Even the first sum differs in structure. By the law of total variance, \begin{align*} {Var}\!\left( \mathbbm{E}[(\tau - \mathbbm{1}\{\epsilon_{i1} \le 0\})(X_{i1} - \gamma_i)\mid B_1]\right) &= {Var}\!\left((\tau - \mathbbm{1}\{\epsilon_{i1} \le 0\})(X_{i1} - \gamma_i)\right) \\ &\quad - \mathbbm{E}\!\left[\text{{Var}}\!\left((\tau - \mathbbm{1}\{\epsilon_{i1} \le 0\})(X_{i1} - \gamma_i)\mid B_1\right)\right]. \end{align*} Unless the conditional variance in the second term on the right-hand side vanishes, the aggregate variance contribution differs from the i.i.d.-type structure in kato2012asymptotics. The common time shocks therefore alter the structure of the covariance term even before accounting for cross-unit covariance. $\lozenge$
assumption[Jacobian matrix] The following matrix exists and is invertible \begin{align*} \Gamma= \lim_{N \to \infty} \frac{1}{N}\sum_{i = 1}^N \mathbbm{E}[f_i(0 \mid X_{i1})X_{i1}(X_{i1} - \gamma_i)'] . \end{align*}

Assumption (ref) requires existence of the Jacobian matrix and the usual full-rank condition on the bread matrix in the sandwich-form asymptotic covariance. In contrast to the middle component $\Sigma$ in Assumption (ref), $\Gamma$ retains the same structure as its counterpart in classicial FEQR theory as in kato2012asymptotics. Also that $\Gamma$ is a symmetric matrix--observe that $\mathbbm{E}[f_i(0 \mid X_{i1})\gamma_{i}(X_{i1} - \gamma_i)'] = 0$ for all $i$, which implies $\mathbbm{E}[f_i(0 \mid X_{i1})X_{i1}(X_{i1} - \gamma_i)'] = \mathbbm{E}[f_i(0 \mid X_{i1})(X_{i1} - \gamma_i)(X_{i1} - \gamma_i)'].$

comment\begin{assumption}[Nondegeneracy] For any $(\alpha, \beta) \in \mathcal{A} \times \mathcal{B}$: \begin{align*} Var(\mathbbm{E}[\mathbbm{1}\{Y_{it} \leq \alpha_i + X_{it}'\beta\} \mid B_t])>0,\quad Var(\mathbbm{E}[\mathbbm{1}\{Y_{it} \leq \alpha_i + X_{it}'\beta\}X_{it} \mid B_t])>0. \end{align*} \end{assumption} {\color{red}[**$\mathcal{A},\mathcal{B} $ above not defined. And do we need it to hold for any any $(\alpha, \beta) \in \mathcal{A} \times \mathcal{B}$ or just at the true value?? ]} {\color{blue}[HDC: It seems $\mathcal{A}, \mathcal{B}$ are defined in Section 2 for different purposes--here they are supposely subsets of parameter space and in Section 2 they are the support for the latent shocks $A_i,B_t$. One of them has to change. Also, I think we only need Assumption 8 to hold at the true parameter (or at most a small neighbourhood of it).]} {\color{teal}[CW: Yes, you are right, the $\mathcal{A}$ and $\mathcal{B}$ and section 2 were different. I think you are also right that the Assumption only has to hold at or around the true parameter. We only need those assumptions for the final Lemma 9. Indeed, we only need the average of them across $i$ to be nondegenerate. But it seems that it is already covered by Assumption 5.]} {\color{teal}[CW: I guess this assumption becomes relevant when we do not condition on a fixed sequence of $\{A_i\}$. Because then the expression in Assumption 5 becomes something like: \begin{align*} &{Var}\left( \mathbbm{E}[(\tau - \mathbbm{1}\{\epsilon_{it} \leq 0\})(X_{it} - \gamma(A_i)) \,|\,B_t]\right). \end{align*} where averaging over $A_i$'s is replaced by expectation. Note that with Assumption 8, even if $X_{it}$ is independent of $B_t$ (there is no omitted variable bias caused by not taking into account of $B_t$), the variance above would still be nonzero. This means $\hat{\beta}$ would still be $O_p(1/\sqrt{T})$ instead of $O_p(1 / \sqrt{NT})$ rate.]} Assumption (ref) requires that the common time shock exerts a non-trivial influence on the distribution of the data. This condition excludes the classical FEQR framework studied in kato2012asymptotics and subsequent work, where--after conditioning on individual fixed effects—the disturbances are independent across \((i,t)\) pairs. The assumption is conceptually analogous to conditions imposed in the multiway clustering literature, for example davezies2021empirical.

Denote the subgradients of the check function evaluated at $(\bm{\alpha}', \beta')$ by

align*[align* omitted — 310 chars of source]

In classical FEQR environments, such as in kato2012asymptotics, the non-differentiability of the check loss precludes standard Taylor expansions. Combined with the faster $\sqrt{NT}$-rate implied by cross-sectional independence, the empirical process terms arising from differences between subgradients evaluated at the estimator and at the true parameter become difficult to control, thereby typically requiring asymptotic regimes in which $T$ grows substantially faster than $N$. Under the common time-effect structure in (ref), however, these non-smooth subgradients admit approximations via suitably defined differentiable projections. This smoothness restores local differentiability, thereby permitting Taylor expansions and substantially simplifying the arguments and weaken the conditions.

Explicitly, let $\bm{B}_T = \{B_t\}_{t=1}^T$ and define the conditional (on common shocks) projections of the subgradients as

align*[align* omitted — 667 chars of source]

Conditioning on $\bm{B}_T$ removes the non-smoothness stemming from the indicator function at the level of the stochastic fluctuations. The projected objects $\tilde{\mathbb{H}}_{Ni}^{(1)}$ and $\tilde{\mathbb{H}}_{N}^{(2)}$ are expectations of indicator functions and therefore can be expressed in terms of conditional distribution functions of $Y_{it}$ given $B_t$. Under mild regularity conditions—such as conditional smoothness of these distributions and bounded densities—the projections approximate the original subgradients uniformly as $N$ diverges while, at the same time, being smooth functions of $(\alpha_i,\beta)$. This structure permits standard Taylor-type expansions around the true parameters despite the non-smoothness of the original objective function, while creating a limiting distribution driven by the common shocks, it yields simple and general asymptotic theory. We refer to this phenomenon as DGP-induced smoothing: conditional averaging over the common time factors smooths the empirical criterion through the randomness of the idiosyncratic components, resulting in stochastic fluctuations of order $\sqrt{T}$ around its mean. This mechanism is central to obtaining refined stochastic expansions and, ultimately, to establishing asymptotic normality under panel asymptotic regimes that allow $N$ to be large relative to $T$.

We now present the main asymptotic distributional result.

theorem[Asymptotic distribution] Suppose Assumptions (ref)--(ref) hold and suppose that $(\log N)^2 / T \to 0$, then we have \begin{align*} \sqrt{T}(\hat{\beta} - \beta_0) \overset{d}{\longrightarrow} \mathcal{N}(0, V),\quadwhere V= \Gamma^{-1}\Sigma \Gamma^{-1}. \end{align*}

A proof can be found in Section (ref) in the appendix. Some remarks are in order.

remark[Comparison with classical FEQR results] Compared with the asymptotic normality results for FEQR obtained in the absence of time effects, such as kato2012asymptotics and galvao2020unbiased, several important differences arise. First, the presence of pervasive common shocks changes the effective convergence rate from $\sqrt{NT}$ to $\sqrt{T}$. Despite of the slower rate, this is not necessarily detrimental, as the bias induced by the incidental parameter problem is attenuated in our setting and has a less pronounced impact on inference. Furthermore, standard inference procedure for FEQR, therefore do not deliver a consistent estimate of the asymptotic covariance. Second, because the projected objective is smooth, Taylor expansion arguments become available. As a consequence, the resulting linear representation—and hence the asymptotic covariance—explicitly incorporates the effect of estimating the fixed effects. Finally, under the classical setup, even the refined analysis in galvao2020unbiased imposes the restriction \( N(\log T)^2 / T \to 0 \). By comparison, our requirement \( (\log N)^2 / T \to 0 \) is markedly weaker, thereby allowing for a wide class of asymptotic sequences with \( T \ll N \). $\lozenge$
remark[Relationship with andrews2005cross] andrews2005cross shows that, in a cross-sectional setting, the OLS estimator is inconsistent when a common shock induces a nonzero conditional covariance between the regressor and the error term. Specifically, under a common shock, \begin{align*} \hat{\beta} \overset{p}{\longrightarrow} \beta_0 + \gamma(B), \end{align*} $\gamma(B) \neq 0$ whenever $\text{\upshape{Var}}\left( \mathbbm{E}[X_i \epsilon_i \,|\,B] \right) \neq 0$. In contrast, Proposition (ref) shows that in the panel setting, aggregating observations over time with i.i.d. $\{B_t\}$ averages out the effect of common shocks, so that consistency of FEQR estimator is retained. Theorem (ref) further shows that this effect remains first-order asymptotically relevant: it appears with $\sqrt{T}$ rate, and dominates the first order asymptotics of the FEQR estimator.
remark[Intuition behind Theorem (ref)] Here we provide intuition for the proof. The key step is to work with the projected score, obtained by conditioning on the common time effects: \begin{align*} \tilde{\mathbb{H}}_{Ni}^{(1)}(\alpha_i, \beta) & = \frac{1}{T} \sum_{t = 1}^T \mathbbm{E}\!\left[\tau - \mathbbm{1}\{Y_{it} \le \alpha_i + X_{it}'\beta\} \mid B_t\right]. \end{align*} We show that this projected score is uniformly close to the original score $\mathbb{H}^{(1)}_{Ni}(\alpha_i,\beta)$. Furthermore, under mild conditions, $\tilde{\mathbb{H}}_{Ni}^{(1)}(\alpha_i,\beta)$ is a differentiable function of $(\alpha_i,\beta)$, even though the original score is non-differentiable. This DGP-induced smoothing permits a Taylor expansion around the true parameter, in contrast to Equation (A.6) of kato2012asymptotics, where the absence of smoothness precludes such a linearisation: \begin{align*} \hat{\alpha}_i - \alpha_{i0} =& \left( \frac{1}{T} \sum_{t = 1}^T f_i(0 \mid B_t) \right)^{-1} \tilde{\mathbb{H}}_{Ni}^{(1)}(\alpha_{i0}, \beta_0) \\ &- \left( \frac{1}{T} \sum_{t = 1}^T f_i(0 \mid B_t) \right)^{-1} \left( \frac{1}{T} \sum_{t = 1}^T \mathbbm{E}[f_i(0 \mid X_{it}, B_t)X_{it}' \mid B_t] \right) (\hat{\beta} - \beta_0) \\ &+ higher-order terms. \end{align*} Such an expansion is not available in earlier FEQR analyses, where the lack of smoothness of the objective necessitated a coarser rate based on concentration inequalities. The projected-score approach, therefore replaces non-smooth empirical process arguments with a smooth stochastic expansion. Together with the $\sqrt{T}$-asymptotic normality induced by the common shocks, this permits substantially weaker restrictions on the relative growth rates of $N$ and $T$ $\lozenge$
remark[Common shocks with dependence] For simplicity, we assume that the common shocks are i.i.d. The asymptotic theory can be extended in a fairly direct way to the case in which the common shocks form a martingale difference sequence. More generally, stationary weakly dependent common shocks, such as $\beta$-mixing processes, should also be admissible, though at the cost of additional technical work. The required modifications are mainly technical and would rely on standard decoupling and blocking arguments for temporal dependence. We leave a full analysis to future work $\lozenge$

Statistical Inference

This section suggests simple practical methods for inference for the quantile model with common shocks. To utilise Theorem (ref) for statistical inference, one needs a consistent estimator of the covariance matrix $V$. We now introduce the new covariance estimator and establish its asymptotic guarantees.

For a kernel function (a probability density function) $K:\mathbb{R}\to \mathbb{R}$, denote $K_h(u)=h^{-1}K(u/h)$ and $\hat \epsilon_{it}=Y_{it}-\hat\alpha_i-X_{it}'\hat\beta$. We define

align[align omitted — 191 chars of source]

We estimate $\Sigma$ and $\Gamma$ by

align[align omitted — 267 chars of source]

where

align*[align* omitted — 164 chars of source]

Hence, the asymptotic covariance in Theorem (ref) can be estimated by the robust covariance estimator:

align[align omitted — 107 chars of source]

We emphasise that this estimator in (ref) differs from the standard covariance estimators commonly used in the literature; see, for example, kato2012asymptotics. As discussed in Remark (ref), the presence of common shocks fundamentally alters the structure of the asymptotic covariance relative to the classical setting with cross-sectional independence. Consequently, conventional covariance estimators are generally inconsistent in this environment and lead to invalid inference. We illustrate this discrepancy in the simulation study reported in Section (ref). The following result establishes the asymptotic properties of the proposed covariance estimator.

theorem[Robust covariance estimation] Suppose the kernel $K$ is continuous, bounded, and of bounded variation on the real line, and the sequence of $h$ satisfies $h\to 0$ and $(\log N)/(Th)\to 0$ as $N,T\to \infty$, then \begin{enumerate} • Under Assumptions of Theorem (ref) and $\log T/N\to 0$, it holds that \( \hat V \xrightarrow{p} V . \) • In the absence of common shocks $\{B_t\}_t$, under the assumptions of Theorem 1 in galvao2020unbiased, it holds that \( N\hat V \xrightarrow{p} \tau(1-\tau)\Gamma^{-1}\Omega\Gamma^{-1}, \) where $\Omega$ is as defined in (ref). \end{enumerate}

A proof can be found in Section (ref). A couple of remarks are in order.

remark[Estimating $\Sigma$] Recall that $\Sigma$ defined in (ref) contains the conditional expectation \( \mathbbm{E}\!\left[ (\tau - \mathbbm{1}\{\epsilon_{i1} \leq 0\})(X_{i1} - \gamma_i) \mid B_1 \right] \). To estimate it, in addition to the need to replace $\epsilon_{it}$ and $\gamma_i$ by their estimators, we also need to approximate the conditional expectation given the common shock $B_1$. Direct estimation of this object is challenging because $B_t$ is unobserved and may be high-dimensional, rendering standard nonparametric estimation challenging. Fortunately, for each fixed $t$, all cross-sectional units share the same realisation of $B_t$. This allows the conditional expectation given $B_t$ to be consistently approximated by its cross-sectional average over $i=1,\ldots,N$, thereby avoiding the need to estimate it nonparametrically. $\lozenge$
remark[Robustness of the covariance estimator] Theorem (ref) implies that, even in the absence of common shocks, inference based on the proposed variance estimator $\hat V$ remains valid under the classical fixed effects quantile regression (FEQR) framework of galvao2020unbiased. In particular, valid inference does not require the practitioner to determine ex ante whether common shocks are present in the data-generating process. For this reason, we recommend the use of our robust covariance estimator of (ref) in empirical applications in place of existing alternatives. $\lozenge$

Monte Carlo Simulations

We conduct Monte Carlo simulations to examine the finite-sample performance of the fixed effects quantile regression estimator and to evaluate the accuracy of the proposed variance estimator in the presence of common time shocks. The simulation design reflects empirically relevant panel dimensions with a large cross-sectional dimension and a moderate time dimension.

We consider the following location-scale shift model similar to kato2012asymptotics

align[align omitted — 136 chars of source]

where $\{\alpha_i\}_{i=1}^N$ are individual fixed effects and $X_{it}$ is a scalar regressor. The regressors are generated according to $ X_{it} = \chi^2_{it}(3) + 0.3\,\alpha_i, $ where $\chi^2_{it}(3)$ are independent chi-square random variables with three degrees of freedom, and $\alpha_i \sim \mathrm{Uniform}(0,1)$. The error term is generated as

align*[align* omitted — 64 chars of source]

where the idiosyncratic component $\varepsilon_{it} \sim N(0,1)$ and common shock $\eta_t \sim N(0,1)$ are mutually independent and i.i.d. across $i$ and $t$. The common shock component $\eta_t$ induces cross-sectional dependence while preserving the conditional quantile structure. We focus on parameters $ \beta = 1, $ $ \gamma = 0.2 $, and estimate the model for three quantile levels $\tau = \{0.25, 0.50, 0.75\}$. Under this specification, the true quantile slope coefficient equals $\beta(\tau) = \beta + \gamma q_\tau,$ where $q_\tau$ denotes the $\tau$-quantile of the standard normal distribution.

The panel dimensions are set to $ N \in\{ 250, 500, 1,000\}$, $ T \in\{ 25,50\} $, and each experiment is repeated over $2,000$ Monte Carlo replications. For each simulated sample, we estimate the FEQR model using the standard Koenker-type estimator as defined in (ref). We consider two covariance estimators: (i) The conventional sandwich variance estimator from kato2012asymptotics. (ii) The proposed robust covariance estimator of (ref). In both covariance estimators, a Gaussian kernel is used with the bandwidth chosen according to Silverman's rule-of-thumb, $h=1.06 \cdot \mathrm{sd}(\hat{\epsilon})\cdot T^{-1/5}$, while imposing a lower bound of $0.05$ to prevent undersmoothing and ensure numerical stability in finite samples. For both variance estimators, we construct nominal $95\%$ confidence intervals using asymptotic normal approximations. We evaluate FEQR estimator performance using bias, root mean squared error (RMSE), and the two covariance estimators by their coverage probabilities. The results are displayed in Tables (ref) and (ref).

table[table omitted — 2,015 chars of source]

The performance of the FEQR estimator under the common shock model is evaluated first by its bias and RMSE. Table (ref) reports the results for the three quantiles and different sample sizes. The results show numerical evidence that the FEQR estimator has small bias for every sample under consideration. Nevertheless, we highlight the important numerical evidence that for a fixed cross-section dimension, the bias decreases as the time-series dimension increases. On the other hand, for a given time-series, the bias is relatively stable as the cross-section dimension increases. Here we recall that from Lemma (ref) above, the main condition on the sample size growth for establishing consistency of the estimator is that $(\log N)^2 / T \to 0$. Hence, the numerical results corroborate the theoretical prediction. In addition, the RMSE decreases monotonically as either $N$ or $T$ increases. These results suggest that the FEQR estimator performs well in small samples for the common shock model, and even for the case of large $N$ relative to $T$ the model.

comment\begin{table}[t] \begin{threeparttable} \caption{Monte Carlo inference accuracy: coverage of robust and standard covariance estimators} \begin{tabular}{lcccccc} \toprule &\multicolumn{2}{c}{$\tau=0.25$}&\multicolumn{2}{c}{$\tau=0.50$}&\multicolumn{2}{c}{$\tau=0.75$}\\ \cmidrule(lr){2-7} $(N,T)$ & Robust & Standard & Robust & Standard & Robust & Standard \\ \midrule \addlinespace[0.25em] $(250,10)$ & 0.865 & 0.625 & 0.908 & 0.648 & 0.862 & 0.644 \\ $(250,25)$ & 0.903 & 0.646 & 0.922 & 0.633 & 0.908 &0.623 \\ $(250,50)$ & 0.911 & 0.617 & 0.927& 0.612 & 0.930 & 0.627\\ $(500,10)$ & 0.868 & 0.498 & 0.924 & 0.490 & 0.859 & 0.496 \\ $(500,25)$ & 0.900 & 0.518 & 0.916 & 0.489 & 0.909 & 0.488 \\ $(500,50)$ & 0.905 & 0.521 & 0.911 & 0.525 & 0.923 & 0.566 \\ $(1000,10)$ & 0.872 & 0.393 & 0.921 & 0.395 & 0.851 & 0.395 \\ $(1000,25)$ & 0.910 & 0.379 & 0.911 & 0.335 & 0.898 & 0.338 \\ $(1000,50)$ & 0.920 & 0.365 & 0.932 & 0.357 & 0.924 & 0.375 \\ \bottomrule \end{tabular} \begin{tablenotes}[flushleft] • Notes: Entries report empirical coverage probabilities of nominal 95% confidence intervals for $\beta(\tau)$ constructed using (i) the proposed robust covariance estimator and (ii) the conventional sandwich covariance estimator that ignores cross-sectional dependence. Results are based on $2,000$ Monte Carlo replications under the DGP in Section (ref). \textcolor{red}{Bandwidth: $h=1.06 \cdot \mathrm{sd}(\hat{\epsilon})\cdot N^{-1/5}$.} \end{tablenotes} \end{threeparttable} \end{table}
table[table omitted — 2,004 chars of source]

Results evaluating the finite sample performance of the proposed inference procedure are collected in Table (ref). If the asymptotic inference procedure correctly approximates the finite sample distribution of $\hat\beta(\tau) - \beta_{0}(\tau)$, the coverage rate should be close to the nominal level of significance ($95$%). Results show that, for a given $N$, empirical coverage improves as $T$ increases. Moreover, as one investigates the diagonal of the table, for example, the pairs $(100,10)$, $(250,25)$, $(500,50)$, and $(1000,100)$, one sees that coverage rates for the Robust case improve, while empirical coverages worsen for the Standard case. Overall, the confidence intervals display accurate finite-sample coverage, with empirical coverage under the proposed variance estimator closely aligned with the nominal $95$% level.

Next, to investigate the robustness of the numerical results we consider a DGP model without common shocks. In particular, we use the same process as described in equation (ref), but use a simpler process for the innovations such that $U_{it}=\varepsilon_{it}$, where $\varepsilon_{it} \sim N(0,1)$. The corresponding results for bias and RMSE are provided in Table (ref), and those for coverage rates in Table (ref).

table[table omitted — 2,061 chars of source]

Results in Table (ref) are similar to those in Table (ref) with the difference that RMSE are smaller. This is an expected result since consistency of the estimator holds under a mild condition that $(\log N)^2 / T \to 0$ described in Proposition (ref).

Coverage rates are collected in Table (ref). Under no common shocks, results for the Standard method are substantially improved relative to the previous case. The table also shows that empirical coverages for the Robust case are close to the nominal 95% coverage. It is also interesting to see that both the Robust method seems to improve more as $T$ increases, for a fixed $N$. These results confirm numerically the robustness of the Robust method.

table[table omitted — 2,048 chars of source]

Conclusion

This paper develops an asymptotic theory for fixed-effects panel quantile regression with pervasive common shocks, a feature central to many economic and financial panels but largely absent from existing FEQR theory. By modelling outcomes through unit, time, and idiosyncratic latent components, we allow for general cross-sectional dependence while retaining the standard FEQR estimator. The key insight is that conditioning on common time shocks induces DGP-induced smoothing, which restores local differentiability of projected scores and enables Taylor expansions despite the non-smooth objective. This yields consistency and asymptotic normality under joint asymptotics with \(N,T \to \infty\), including regimes with \(T \ll N\). Common shocks fundamentally alter the stochastic structure of FEQR: the estimator concentrates around a $\sqrt{T}$-order asymptotic component, with an asymptotic covariance matrix that differs from that in classical settings, rendering conventional FEQR variance estimator inconsistent. We therefore develop inference procedures, including a novel robust variance estimator, that deliver valid confidence intervals and hypothesis tests in the presence of systemic time shocks while remaining applicable under the classical FEQR framework.

Several directions for future research appear promising. First, the framework could be extended to other panel models with non-smooth objective functions -- such as censored or threshold-type estimators -- where analogous forms of inherent smoothing may emerge in the presence of common shocks. Second, an important extension is to permit serially correlated time factors and to develop inference procedures that simultaneously accommodate both cross-sectional and temporal dependence in non-smooth panel settings. Finally, incorporating two-way fixed effects into panel quantile regression under full generality remains theoretically demanding and represents a particularly interesting open problem.

Appendix