EconBase
← Back to paper

Estimation and Inference for the $τ$-Quantile of Individual Heterogeneous Coefficient

Extracted main text — title through conclusion, appendix excluded. This is what our citation measures are computed over, published so the extraction can be checked by eye.

104,341 characters · 15 sections · 57 citation commands

Rendered from LaTeX for readability, not typeset faithfully. Citation keys are highlighted; maths is left as source; figures, tables and equation environments are summarised rather than reproduced; unrecognised commands are greyed out so nothing is silently dropped. Email addresses are removed.

Estimation and Inference for the $$-Quantile of Individual Heterogeneous Coefficient

abstractThis paper proposes estimation and inference procedures for the quantiles of individual heterogeneous slope coefficients within panel data. We develop a two-step quantile estimation framework for analyzing heterogeneity in individual coefficients. Unlike conventional panel quantile regression, which focuses on outcome heterogeneity, our approach targets the $\tau$-quantile of the cross-sectional distribution of individual-specific slopes. We establish asymptotic theory under both stochastic and deterministic designs, with convergence rates $\sqrt{N}$ and $\sqrt{N\sqrt{T}}$, respectively. We also develop two corresponding bootstrap procedures for practical inference, and formally establish their validity. The suggested methods are of practical interest since they require weaker sample size growth conditions than standard fixed-effect quantile regression, and accommodate large $N$ settings. Numerical simulations and an application to mutual fund performance illustrate the proposed methods and the heterogeneity patterns they reveal across quantiles. JEL Classification: C22, C23 Keywords: quantile regression, heterogeneous coefficients, panel data, asymptotic theory.

{18pt}{10pt} \belowdisplayskip \abovedisplayskip {5pt} \abovedisplayshortskip \belowdisplayshortskip {8pt} \belowdisplayskip \abovedisplayskip{4pt}

\thispagestyle{empty}

Introduction

Quantile regression (QR) has emerged as a central framework in econometric analysis for modeling heterogeneous relationships across different points of the conditional distribution of an outcome variable (koenker1978regression; koenker2005quantile). Unlike mean regression approaches, it provides a more comprehensive view of the data-generating process by allowing the effects of covariates to vary across quantiles. In panel data settings, quantile regression offers flexible means to capture individual-specific unobserved heterogeneity and distributional dynamics over time (see, e.g., Koenker04; canay2011simple; galvao2011quantile). These methods are particularly useful for uncovering how explanatory variables influence the outcome differently across its conditional distribution, thereby providing richer insights than conventional mean-based methods.

Standard panel quantile regression models are, however, designed to primarily address heterogeneity in the outcome variable conditional on covariates through only allowing individual fixed effects that shift the intercept, while slope parameters are typically assumed homogeneous across individuals. As a result, existing panel quantile regression methods focus on heterogeneity in the conditional distribution of outcomes. In contrast, in this paper we study heterogeneity in the structural parameters themselves by estimating quantiles of individual-specific coefficients.

This paper examines a different dimension of heterogeneity---across individuals $i$---in panel data quantile models by focusing on the distribution of individual-specific slope coefficients, while account for individual specific fixed effects. We are interested in how the structural effects themselves vary across units, rather than how effects differ across outcomes. This distinction is crucial for understanding underlying differences in behavioral responses, treatment effects, or policy sensitivities. For instance, in education or labor economics, researchers may wish to assess how the effect of class size or job training varies across individuals, rather than how the conditional distribution of outcomes changes with these factors. Despite its empirical importance, this type of heterogeneity has received limited formal attention within a unified econometric framework.

We propose a two-step estimation framework for the \(\tau\)-quantile of the cross-sectional distribution of heterogeneous individual coefficients in panel data. In the first step, we obtain unit-specific estimators \(\{\hat{\theta}_{Ti}\}_{i=1}^N\). In the second step, we do not pool these estimates to recover a common coefficient (e.g. Galvao and Wang, galvao2015efficient), nor do we use \(\tau\) to index the conditional quantile of the outcome variable as in the standard panel quantile regression (e.g. GalvaoGuVolgushev20). Instead, we apply the quantile operator $Q_{\tau}$ across \(i\) to estimate the \(\tau\)-quantile of the latent cross-sectional distribution of \(\theta_{i0}\):

equation*[equation* omitted — 83 chars of source]

Thus, $\tau$ indexes location in the distribution of heterogeneous coefficients across individuals, not location in the conditional distribution of $y_{it}$ given regressors. By considering several values of $\tau$, we summarize different parts of the distribution of structural effects, such as lower-tail, median, and upper-tail coefficients across units.

We develop an asymptotic theory for the proposed estimator under two distinct designs: stochastic and deterministic. These correspond to two common empirical scenarios. The stochastic-design case applies when the heterogeneous parameters, $\{\theta_{i0}\}_{i=1}^N$, represent random draws from a larger population, and inference targets the population distribution of individual effects. In this case, the estimator achieves $\sqrt{N}$-consistency and asymptotic normality under a mild condition $\sqrt{N}/T=O(1)$, which is weaker than the standard requirement $N/T=O(1)$ in fixed-effect quantile regression; see, e.g., Galvao and Kato galvao2016smoothed. The deterministic-design case is relevant when the researcher observes the full large cross-sectional population of interest and therefore treats the heterogeneous parameters as fixed. It is also relevant when the object of interest is the empirical distribution of the heterogeneous parameters. In this case, the estimator converges at the novel rate $\sqrt{N\sqrt{T}}$, reflecting the absence of cross-sectional randomness, but requiring a more restrictive growth condition $T^{1/2}\ll N\ll \frac{T^{3/2}}{(\log T)^2}$ to ensure bias control and limit validity.

To conduct practical inference, we introduce two bootstrap procedures tailored to the two designs considered: the stochastic-design quantile bootstrap (SQB) and the deterministic-design quantile bootstrap (DQB). Both methods are formally shown to provide consistent approximations to the corresponding asymptotic distributions of the estimators. The SQB accounts for both cross-sectional and time-series randomness, while the DQB conditions on the realized heterogeneity and resamples over the time dimension.

We provide numerical simulations to evaluate the proposed methods in finite samples. Simulation results demonstrate that the two bootstrap procedures complement each other: SQB performs well in stochastic designs, while DQB yields stable and accurate coverage in deterministic settings across a wide range of $(N,T)$ configurations.

An empirical application to mutual fund performance illustrates the usefulness of the proposed method. By estimating fund-specific coefficient quantiles over a range of $\tau$ values, we document substantial cross-sectional variation in return- and liquidity-timing abilities, whereas volatility-timing and abnormal-return heterogeneity are limited. The results indicate that slope heterogeneity among fund managers is asymmetric.

Our work is related to the literature on heterogeneous effects and their distributions, including the sorted effects framework of Chernozhukov, Fernández-Val, and Luo (chernozhukov2018sorted), Fernández-Val et al. (fernandez2022dynamic), as well as the broader literature on random functions, structural estimation, and quantile effects (e.g., Matzkin, matzkin2003nonparametric; Heckman and Vytlacil, heckman2005structural; Graham and Powell, graham2012identification; Arellano and Bonhomme, arellano2012identifying; Chernozhukov, Fernández-Val, and Melly, chernozhukov2013inference). Unlike these approaches, which study heterogeneity in outcome distributions, treatment effects, identification, or functionals of coefficients, we focus on quantiles of heterogeneous structural parameters that are themselves estimated in a first step using panel data. This creates a distinct econometric problem characterized by two-step estimation and joint asymptotics in both the cross-sectional and time-series dimensions.

Our work is also related to the classical literature on heterogeneous panel-data models with unit-specific coefficients, including random-coefficient and mean-group approaches; see, among others, Swamy (swamy1970efficient), Pesaran, Shin, and Smith (pesaran1999pooled), Pesaran (pesaran2006estimation), Pesaran and Yamagata (pesaran2008testing), Liao and Yang (liao2017uniform), and Li, Cui, and Lu (li2020efficient). These contributions allow for slope heterogeneity across units and provide important tools for estimation, aggregation, and testing in heterogeneous panels. Their focus, however, is different from ours. In that literature, the main object is typically a common parameter, an average effect, a long-run mean effect, or a test of whether slope heterogeneity is present. By contrast, our object of interest is the \(\tau\)-quantile of the cross-sectional distribution of the heterogeneous coefficients themselves. Thus, even though our first step also estimates unit-specific coefficients, the second step is not designed to recover a pooled or average effect. Instead, it is used to study the cross-sectional pattern of $\{\theta_{i0}\}_{i=1}^N$ through quantiles evaluated at specific values of $\tau$.

The remainder of the paper is structured as follows. Section (ref) outlines the model and estimation procedures. Section (ref) establishes the asymptotic results under a set of high-level conditions. Section (ref) presents the bootstrap-based inference procedures. Section (ref) presents an application to the least squares estimator as the first-step estimator. Section (ref) examines the finite-sample performance of the proposed method. Section (ref) reports an empirical application. Section (ref) concludes. Mathematical derivations are provided in the Appendix.

Model and Estimation

This section presents the model of interest and the proposed two-step estimator.

Model

We consider a linear panel-data model with heterogeneous coefficients:

equation[equation omitted — 193 chars of source]

where $y_{it}$ denotes the outcome variable, $\alpha_{i0}$ is the $i$-th individual specific fixed effect, $\bm{\beta}_{i0}$ is the $i$-th individual specific slope coefficient, $\bm{Z}_{it}=(1,\bm{z}_{it}^{\top})^{\top}$ is a $K$-dimensional regressor vector, and $\varepsilon_{it}$ is an unobserved innovation independent of $\bm{Z}_{it}$. The restriction imposed on $\varepsilon_{it}$ depends on the estimator under consideration. For example, for mean regression, we require $E(\varepsilon_{it}|\bm{Z}_{it})=0$; for quantile regression, we instead assume $Q_{\tau}(\varepsilon_{it}|\bm{Z}_{it})=0$, where $Q_{\tau}(\varepsilon_{it}|\bm{Z}_{it})$ denotes the conditional $\tau$-quantile of $\varepsilon_{it}$ given $\bm{Z}_{it}$. In the model, $(\bm{z}_{it}^{\top},\varepsilon_{it})$ are i.i.d. across both $i$ and $t$, and all cross-sectional heterogeneity is captured by $\{\bm{\theta}_{i0}\}_{i}$. Moreover, for each $i$, the sequence $\{(\bm{z}_{it}^{\top},\varepsilon_{it})\}_{t}$ is independent of $\bm{\theta}_{i0}$.

The parameter vector $\bm{\theta}_{i0}\in\Theta\subseteq\mathbb{R}^{K}$ collects the unit-specific intercept and slope coefficients. Throughout, the parameter space $\Theta$ is assumed to be compact. We denote the $p$-th coordinate of $\bm{\theta}_{i0}$, for $p\in\{1,\ldots,K\}$, by $\theta_{i0,p}$. Thus, $\theta_{i0,p}$ may represent the intercept $\alpha_{i0}$ or any component of the slope vector $\bm{\beta}_{i0}$. The empirical distribution function of $\{\theta_{i0,p}\}_{i=1}^{N}$ is defined as:

equation*[equation* omitted — 87 chars of source]

For $\tau\in(0,1)$, the main parameter of interest is the $\tau$-th quantile of the cross-sectional distribution $\{\theta_{i0,p}\}_{i=1}^{N}$. The sequence $\{\bm{\theta}_{i0}\}_{i=1}^{N}$ may be treated as either deterministic or stochastic, depending on the specification adopted. In this context, two scenarios arise depending on the researcher's objective and data structure:

(i) Stochastic $\{\theta_{i0,p}\}$: This setting applies when a random sample of $N$ units is drawn from a large population, and the target is the population distribution of $\theta_{i0,p}$. If $\{\theta_{i0,p}\}_{i=1}^{N}$ are i.i.d. draws, the empirical distribution function $F_{N}(x,p)$ converges to the true population CDF as $N\to\infty$:

equation*[equation* omitted — 68 chars of source]

where $P(\theta_{i0,p}\leq x)$ is the true population CDF. The $\tau$-th quantile is then:

equation[equation omitted — 102 chars of source]

where the superscript ${\rm {S}}$ indicates the stochastic nature of the quantile.

(ii) Deterministic $\{\theta_{i0,p}\}$: This case occurs when the full population is observed (e.g., all countries' oil reserves) or when the researcher is interested in the limiting distribution of an observed sample of size $N$ (e.g., students' IQs in a college of size $N$). When $\{\theta_{i0,p}\}_{i=1}^{N}$ are deterministic, the empirical distribution function converges to a limiting distribution $F(x,p)$ as $N\to\infty$, which may differ from the population distribution:

equation*[equation* omitted — 52 chars of source]

The quantile is then given by:

equation[equation omitted — 86 chars of source]

where the superscript ${\rm {D}}$ denotes the deterministic nature of the quantile.\footnote{{Both stochastic and deterministic treatments of heterogeneous coefficients appear in the literature, although the specification is sometimes left implicit. On the one hand, a stochastic interpretation is adopted in the random-coefficient and heterogeneous-panel literature; see, e.g., Pesaran, Shin, and Smith (pesaran1999pooled), Hsiao and Pesaran (hsiao2008random), and Li, Cui, and Lu (li2020efficient). On the other hand, a deterministic interpretation is common in panel-data settings where the heterogeneous parameters are treated as fixed unknown quantities attached to the observed units. This perspective underlies the individual fixed effects in standard panel linear regression or panel quantile regression, and also appears in models with individual-specific slope effects; see, e.g., Polachek and Kim (polachek1994panel). See also Su, Shi, and Phillips (su2016identifying) for a useful discussion of the contrast between homogeneous, random-coefficient, and grouped heterogeneous coefficient models.}}

remarkIn the deterministic setting, the limiting distribution $F(x,p)$ is determined by the observed data, and the quantile $\theta_{\tau,p}^{{\rm {D}}}$ reflects information within this fixed set of observations. It is important to note that $F(x,p)$ may differ from the population distribution $P(\theta_{i0,p}\leq x)$, as it depends on the sample. For example, if students' IQs are randomly distributed nationwide, but in a high-quality college of size $N$, the IQ distribution is fixed and influenced by specific criteria (e.g., high admissions standards), the empirical distribution within the college will converge to a limiting distribution $F(x,p)$ when $N$ is large, which differs from the national distribution $P(\theta_{i0,p}\leq x)$.

To clarify and illustrate the empirical scope of our framework, next we provide examples illustrating its application and contrasting it with conventional methods, highlighting the interpretational gains from targeting quantiles of structural effects rather than conditional outcome quantiles.

\paragraph{Example 1: Quantile of heterogeneous slopes.}

Suppose that student's GPA, $y_{it}$, depend on class size ${Z}_{it}$ via

equation*[equation* omitted — 72 chars of source]

In this specification, $\beta_{i0}$ reflects the extent of student $i$'s sensitivity to class size. We then rank the coefficients $\{\beta_{i0}\}_{i}$ according to their magnitudes, from the least to the most sensitive. Some students may be less affected ($\tau=0.1$), while others more dependent ($\tau=0.9$). Our target is $\beta_{\tau}$, the $\tau$-quantile of these sensitivities.

We estimate $\beta_{\tau}$ in two steps:

equation*[equation* omitted — 280 chars of source]

where $\rho(\cdot)$ is the standard check function, $\rho_{\tau}(u)=u(\tau-\mathbf{1}\{u\leq0\})$.

By contrast, the standard fixed-effect QR imposes

equation*[equation* omitted — 84 chars of source]

so that $\beta(\tau)$ captures the effect for the $\tau$-quantile conditional GPA student (i.e., $\tau$-quantile $y_{it}$ conditional on $Z_{it}$), not the $\tau$-quantile class-size sensitive student (i.e., $\tau$-quantile $\beta_{i}$). Thus, conventional QR does not target the desired distribution of sensitivities. Equivalently, in conventional panel quantile regression, \(\tau\) indexes heterogeneity in the conditional distribution of the outcome \(y_{it}\), while the slope parameter is typically common across individuals at each fixed \(\tau\). In our framework, \(\tau\) instead indexes heterogeneity in the cross-sectional distribution of the coefficients \(\beta_{i0}\) themselves. These are conceptually different objects: one concerns which outcome quantile of $y_{it}$ is being modeled, while the other concerns which quantile of the cross-sectional distribution of structural effects is being summarized.

\paragraph{Example 2: Quantile of average wages.}

Suppose wages follow

equation*[equation* omitted — 53 chars of source]

with $\theta_{i0}$ the long-run average wage of individual $i$. Apply the estimator $\widehat{\theta}_{Ti}=\frac{1}{T}\sum_{t=1}^{T}Y_{it}$. Our object of interest is the $\tau$-quantile of wage levels across individuals, based on the collection of individual-specific long-run wage estimates:

equation*[equation* omitted — 127 chars of source]

By contrast, pooled quantile regression estimates

equation*[equation* omitted — 129 chars of source]

which corresponds to the overall $\tau$-quantile computed from all observations \(\{Y_{it}\}_{i,t}\), and therefore ignores cross-sectional heterogeneity in long-run wage levels.

Likewise, individual-specific or time-specific quantile regression captures quantiles within a given unit or within a given period, but does not recover the cross-sectional quantile of long-run averages. For example, individual-specific quantile regression yields the $\tau$-quantile wage over time for each individual \(i\),

equation*[equation* omitted — 177 chars of source]

while time-specific quantile regression yields the $\tau$-quantile wage across individuals at each time period \(t\),

equation*[equation* omitted — 177 chars of source]

These objects characterize within-individual or within-period outcome heterogeneity, rather than the cross-sectional distribution of individual-specific long-run wage levels.

comment\paragraph{Example 3: Quantile of quantiles (electricity demand).} Consider household electricity usage model \begin{equation*} \qquad y_{it}=\alpha_{i0}(\tau_{1})+{Z}_{it}\beta_{i0}(\tau_{1})+\varepsilon_{it}, \end{equation*} where $y_{it}$ is log electricity consumption and ${Z}_{it}$ is log price. Here $\beta_{i0}(\tau_{1})$ represents household $i$'s price elasticity at usage state $\tau_{1}$: $\tau_{1}=0.10$: low-usage states (e.g., mild days). $\tau_{1}=0.90$: peak-usage states (e.g., cold winter days with heating). We are not interested in each household's slope individually, but in the distribution across households. For example, the lower-tail ($\tau=0.25$) of $\{\beta_{i0}(0.90)\}_{i}$ identifies the least price-responsive households during peak usage periods households effectively “locked in” with little scope to reduce demand in peak days. In contrast, the median or upper quantiles ($\tau=0.50,0.75$) describe typical or highly responsive households during peaks.

The current formulation defines $\theta_{\tau,p}$ as an unconditional quantile across individual specific slope parameters. A natural extension considers conditional quantiles that depend on observed cross-sectional characteristics $\bm{\omega}_{i}$:

equation[equation omitted — 134 chars of source]

where $\eta(\tau)$ and $\bm{\lambda}(\tau)$ denote the $\tau$-th specific intercept and slope coefficients, respectively. It is important to note that this specification represents a cross-sectional relationship rather than a panel data model, since the time series observations are used only to estimate $\theta_{i0}$ for each unit $i$. The unconditional quantile model discussed earlier is a special case obtained by setting $\bm{\omega}_{i}=\bm{0}$, in which case $\eta(\tau)=\theta_{\tau,p}$.

Estimation

We now describe how to estimate the parameters of interest, $\theta_{\tau,p}^{{\rm {S}}}$ and $\theta_{\tau,p}^{{\rm {D}}}$, defined in equations (ref) and (ref), respectively. Define $\bm{X}_{it}\equiv\bigl(y_{it},\,\bm{Z}_{it}^{\top}\bigr)^{\top}$. Since $y_{it}$ depends on the unknown parameter $\bm{\theta}_{i0}$, the vector $\bm{X}_{it}$ also inherits this dependence. For clarity, we may explicitly write $y_{it}(\bm{\theta}_{i0})$ and $\bm{X}_{it}(\bm{\theta}_{i0})$ when needed.

We then propose the following two-step estimation procedure for the distributional quantile $\theta_{\tau,p}$:

description• Two-step Estimation Procedure

\paragraph{Step 1 (Individual estimation).}

For each individual $i$, obtain an estimator of $\theta_{i0}$: $\widehat{\bm{\theta}}_{Ti}=\widehat{\bm{\theta}}_{Ti}\left(\left\{ \bm{X}_{it}\left(\bm{\theta}_{i0}\right)\right\} _{t}\right).$

\paragraph{Step 2 (Quantile aggregation).}

Given the collection $\{\widehat{\bm{\theta}}_{Ti}\}_{i=1}^{N}$, estimate the $\tau$-quantile of them by

equation[equation omitted — 175 chars of source]

This two-step procedure first recovers unit-level parameters, and then aggregates them to form quantile estimates across heterogeneous units. Note that we do not restrict the first-step estimator $\widehat{\bm{\theta}}_{Ti}$, but we will impose some standard high-level conditions on the properties of the estimator. We further present an example that satisfies these assumptions in Section (ref).

It is useful to distinguish our estimator from standard panel quantile regression. In conventional panel QR, the index \(\tau\) refers to the conditional \(\tau\)-quantile of the outcome \(y_{it}\) given regressors, and the slope coefficient is typically common across individuals. In our framework, the first-step estimator may itself be based on mean regression, quantile regression, or another method, but the second-step index \(\tau\) always refers to the cross-sectional \(\tau\)-quantile of the heterogeneous coefficients \(\theta_{i0}\). Thus, the same symbol \(\tau\) plays a different conceptual role in the two literatures.

The role of Step 2 is fundamentally different from that in the usual panel-data literature. In the existing panel quantile regression, one first estimates individual-specific objects \(\hat{\theta}_{Ti}\), and then aggregates them by the sample mean or a minimum-distance criterion in order to recover a common parameter or an average effect; see, e.g., Galvao and Wang (galvao2015efficient). Such procedures are appropriate when the underlying model considers \(\theta_{i0}=\theta_0\) for all \(i\). In contrast, we allow \(\theta_{i0}\) to differ across individuals and take this heterogeneity as the primary object of interest. Accordingly, Step 2 uses the sample quantile, rather than the sample mean, because our goal is to summarize the cross-sectional heterogeneity in $\theta_{i0}$ across $i$, rather than to collapse it into a single pooled coefficient. In particular, different values of $\tau$ correspond to different parts of the cross-sectional distribution of the heterogeneous coefficients.\footnote{Our framework is also distinct from the unconditional quantile regression of Firpo, Fortin, and Lemieux (firpo2009unconditional). Their method studies how covariates affect quantiles of the unconditional distribution of the outcome via recentered influence function regressions. By contrast, our object of interest is the $\tau$-quantile of the cross-sectional distribution of latent heterogeneous coefficients $\{\theta_{i0}\}_{i=1}^N$. Thus, here $\tau$ indexes heterogeneity in structural parameters across individuals, rather than heterogeneity in the unconditional distribution of observed outcomes.}

If one were to replace the quantile operator in Step 2 by the sample mean, the resulting estimator would target a different object, namely the cross-sectional average \(E(\theta_{i0})\) in the stochastic case or its empirical analogue in the deterministic case, rather than a distributional feature such as the median or tail behavior. It is important to note that this estimator is still different from the minimum-distance estimator in Galvao and Wang (galvao2015efficient). The key distinction lies in the maintained assumption on \(\theta_{i0}\). In Galvao and Wang (galvao2015efficient), the parameter is homogeneous across individuals, so that \(\theta_{i0}=\theta_0\) for all \(i\), and the sample mean is used to recover the common true value \(\theta_0\). In our framework, by contrast, \(\theta_{i0}\) is allowed to vary across individuals. As a result, the sample mean no longer targets a common structural parameter, but instead the average of the heterogeneous coefficients, $E(\theta_{i0})$. We briefly discuss the asymptotic behavior of this alternative estimator, and how it differs from our quantile-based estimator, in Section (ref).

Note that the same estimator $\widehat{\theta}_{\tau,p}$ is used to estimate both $\theta_{\tau,p}^{{\rm S}}$ and $\theta_{\tau,p}^{{\rm D}}$. At first sight, this may seem paradoxical: if one estimator converges to two parameters, one might expect the parameters to be trivially identical. The answer is that the two parameters are related but not the same, and the distinction emerges because they are learned at different scales of the data. In particular, $\widehat{\theta}_{\tau,p}$ approximates $\theta_{\tau,p}^{{\rm S}}$ at rate $\sqrt{N}$, whereas it approximates $\theta_{\tau,p}^{{\rm D}}$ at the faster rate $\sqrt{N\sqrt{T}}$.

To illustrate, consider the mutual fund application. If the observed funds are viewed as a random sample drawn from a broader population of fund managers, then the parameter of interest is the $\tau$-quantile of the population distribution of managerial skill, denoted $\theta_{\tau,p}^{\rm S}$. In this case, even if each fund-specific coefficient were estimated without error, cross-sectional sampling uncertainty would still remain, since the observed funds represent only a subset of the underlying population.

By contrast, suppose that the observed funds are treated as the fixed population of interest, for example, the universe of major mutual funds actively managed by Wall Street institutions. Then the parameter of interest is the $\tau$-quantile of the limiting empirical distribution of the fund-specific coefficients, denoted $\theta_{\tau,p}^{\rm D}$. Under this interpretation, the cross-sectional units are not viewed as random draws from a larger population, but rather as the population whose heterogeneity is itself the object of study. Consequently, once the fund-specific coefficients are estimated more precisely, the uncertainty in the second-step quantile estimator is reduced accordingly. Thus, although the estimator is identical in form in both settings, it may converge at different rates to different target parameters, because the underlying source of uncertainty differs across the two frameworks.

For notational simplicity, we first consider the scalar case $K=1$ and suppress the subscript $p$ throughout. When $K>1$, the same argument applies elementwise to each coordinate $p=1,\ldots,K$.

Asymptotic Theory

We now establish the asymptotic properties of the two step estimator considering two settings separately: (i) stochastic $\{{\theta}_{i0}\}_{i}$ and (ii) deterministic $\{{\theta}_{i0}\}_{i}$.

Stochastic $\left\{ \mathbf{\theta}_{i0}\right\}_{i}$

In the stochastic setting, $\bm{X}_{it}$, and consequently $\widehat{{\theta}}_{\tau}$, involve two layers of randomness: the first due to the randomness of ${\theta}_{i0}$ itself, and the second due to sampling $\bm{X}_{it}({\theta}_{i0})$ conditional on ${\theta}_{i0}$. We introduce conditions that ensure the consistency and asymptotic normality of $\widehat{{\theta}}_{\tau}$.

assumption(i) ${\theta}_{i0}$ is i.i.d. over $i$ with a distribution $F$. (ii) ${\theta}_{\tau}^{{\rm {S}}}$ is the unique minimizer of $E\left[\rho_{\tau}\left({\theta}_{i0}-{\theta}\right)\right]$ over the compact set $\Theta$. (iii) $\theta_\tau^\mathrm{S}$ is an interior point of $\Theta$.
assumption(i) $\theta_{i0}$ has a continuous density $f$ over $\Theta$; (ii) $f\left(\theta_{\tau}^\mathrm{S}\right)\in(0,\infty)$ and $f\left(\theta\right)$ is continuously differentiable and bounded on a neighborhood of $\theta_{\tau}^\mathrm{S}$.

Assumption (ref) requires $\theta_{i0}$ to be stochastic, and imposes an identification condition for the parameter of interest $\theta_{\tau}$. Denote the cdf and pdf of the standard Gaussian distribution by $\Phi$ and $\phi$, respectively. Define the asymptotic variance function ${\sigma}\left({\theta}\right)^{2}=\lim_{T\to\infty}Var\left(\sqrt{T}\widehat{{\theta}}\vert{\theta}\right)$. We now consider the high-level conditions for the first step estimation. Assumption $\ref{as: continuously differentiable}$ places continuity and smoothness conditions on the density of $\theta_{i0}$ that appear in the unconditional quantile regression literature.

assumption[High-level conditions for the first-step estimation, stochastic] Let $W_{T}\left({\theta}\right)={\sigma}\left({\theta}\right)^{-1}\sqrt{T}\left(\widehat{{\theta}}-{\theta}\right)$. There exist ${\kappa}_{3, \theta}$ and ${\kappa}_{4, \theta}$ such that \begin{equation*} R_{T,1}\left(x,{\theta}\right)=P\left(W_{T}\left({\theta}\right)\le x\vert{\theta}\right)-\Phi\left(x\right) \end{equation*} and \begin{equation*} R_{T,2}\left(x,{\theta}\right)=P\left(W_{T}\left({\theta}\right)\le x\vert{\theta}\right)-\left[\Phi\left(x\right)+T^{-1/2}p_{1, \theta}\left(x\right)\phi\left(x\right)+T^{-1}p_{2, \theta}\left(x\right)\phi\left(x\right)\right], \end{equation*} where ${p}_{1, \theta}\left(x\right)=\frac{{\kappa}_{3, \theta}}{6}\left(1-x^{2}\right)$ and ${p}_{2, \theta}\left(x\right)=\frac{{\kappa}_{4, \theta}}{24}\left(-x^3+3x\right)+\frac{{\kappa}_{3,\theta}^2}{72}\left(-x^5+10x^3-15x\right)$ satisfying: \begin{enumerate}[(i)] • (Berry-Esseen Bound) $E_{\theta}\left[\sup_{x\in\mathbb{R}}\left|R_{T,1}\left(x,{\theta}\right)\right|\right]=o\left(1\right).$ • (Bounded variance) $0<\inf_{{\theta}\in\Theta}{\sigma}\left({\theta}\right)^{2}\le\sup_{{\theta}\in\Theta}{\sigma}\left({\theta}\right)^{2}<\infty.$ • (Differentiable variance) ${\sigma}\left({\theta}\right)^{2}$ is continuously differentiable over $\Theta$. • (Edgeworth Expansion) $E\left|\kappa_{3, \theta}\right|<\infty$, $E\left|\kappa_{4, \theta}\right|<\infty$, and $E\left(\sup_{x\in\mathbb{R}}\left|R_{T}\left(x,{\theta}\right)\right|\right)=o\left(T^{-1}\right).$ \end{enumerate}

Assumption (ref)(i) ensures a uniform first-order Gaussian approximation for $W_{Ti}$. Assumption $\ref{as: high level random-1}$(ii) imposes eigenvalue bounds on ${\sigma}_{i}$, ruling out degenerate scaling. Assumption $\ref{as: high level random-1}$(iii) requires that the asymptotic variance is differentiable. Assumption $\ref{as: high level random-1}$(iv) strengthens (ii) by requiring a two-term Edgeworth expansion with a $o\left(T^{-1}\right)$ remainder; hence it is a strictly stronger, second-order refinement of the normal approximation.

theoremUnder Assumptions (ref), (ref), and (ref)(i)-(ii), as $N,T\rightarrow\infty$, \begin{equation*} \widehat{\theta}_{\tau}\xrightarrow{P}\theta_{\tau}^{{\rm {S}}}. \end{equation*}

Theorem (ref) establishes the consistency of $\widehat{\theta}_{\tau}$. Note that no restriction is imposed on the ratio of $N$ and $T$. Let $f'$ and $\sigma'$ denote the derivative of the density function and the asymptotic variance function, respectively.

theoremUnder Assumptions (ref), (ref), and (ref), as $N,T\rightarrow\infty$ with $\sqrt{N}/T=O(1)$, \begin{equation*} \sqrt{N}\left(\widehat{\theta}_{\tau}-\theta_{\tau}^{{\rm {S}}}\right)\xrightarrow{d}\mathcal{N}\left(B_{R},f\left(\theta_{\tau}\right)^{-2}\tau\left(1-\tau\right)\right). \end{equation*} where \begin{align*} B_{R} & =\lim_{N,T\to\infty}-\frac{\sqrt{N}}{T}\left(\frac{f\left(\theta_{\tau}\right)^{-1}f'(\theta_{\tau})}{2}\sigma\left(\theta_{\tau}\right)^{2}+\sigma\left(\theta_{\tau}\right)\sigma'\left(\theta_{\tau}\right)\right). \end{align*}

Although the asymptotic distribution in Theorem (ref) is concise, its proof is non-trivial. In particular, the argument requires a careful treatment of higher-order terms that arise throughout the analysis.

The asymptotic variance is $f(\theta_{\tau})^{-2}\tau(1-\tau),$ which is the same as in the infeasible benchmark, where the latent parameters $\{\theta_{i0}\}_{i}$ are directly observed. The key intuition is that the first-step estimation error in $\widehat{\theta}_{Ti}$ does not contribute to the asymptotic variance at the first order. Because each $\theta_{i0}$ is estimated separately, this error is idiosyncratic across $i$, and its contribution to the variance averages out sufficiently fast (at the $T^{-1/2}$ scale) relative to the second-step cross-sectional quantile estimation. Instead, the effect of first-step estimation appears through the bias term $B_{R}$, which captures the discrepancy

equation*[equation* omitted — 109 chars of source]

Thus, estimating $\theta_{i0}$ in the first step changes the centering of the second-step quantile estimator, but not its leading stochastic fluctuation. As a result, the asymptotic variance coincides with that in the oracle case based on directly observed $\{\theta_{i0}\}_{i}$.

Regarding the required growth condition on the sample size, in the stochastic case, the estimation error of $\widehat{\theta}_{Ti}$ around $\theta_{i0}$ averages out across the $N$ units because of cross-sectional independence, which permits a relatively milder restriction than in conventional FE-QR settings ($N/T=O(1)$). The condition $\sqrt{N}/T=O(1)$ ensures that the bias term $B_{R}$ remains bounded. If, in addition, $\sqrt{N}/T=o(1)$, then the bias vanishes asymptotically, and

equation*[equation* omitted — 161 chars of source]

By contrast, if $\sqrt{N}/T\to c\in(0,\infty),$ then $B_{R}$ is asymptotically non-negligible and induces a centering shift, which may invalidate standard inference if left unaccounted for. Nonetheless, as shown below, the bootstrap is able to replicate both this bias term and asymptotic variance. We will discuss the requirements on the sample size and convergence rates further in Section (ref) below.

Deterministic $\left\{ \theta_{i0}\right\} _{i}$

In the deterministic scenario, the only source of randomness comes from sampling $\bm{X}_{it}$ conditional on $\{\theta_{i0}\}_{i}$. We introduce conditions that ensure the consistency of $\widehat{\theta}_{\tau}$.

assumption(i) $\left\{ \theta_{i0}\right\} _{i}$ are deterministic over the compact set $\Theta$. (ii) $\theta_{\tau}^{{\rm {D}}}$ is the unique minimizer of \( S\left(\theta\right)=\lim_{N\rightarrow\infty}\frac{1}{N}\sum_{i=1}^{N}\rho_{\tau}\left(\theta_{i0}-\theta\right). \) (iii) $\theta_\tau^\mathrm{D}$ is an interior point of $\Theta$.

Assumption (ref) is an identification condition of $\theta_{\tau}$.

assumption[High-level conditions for the first step estimation, deterministic] Conditional on $\{{\theta}_{i0}\}_{i}$, \begin{enumerate}[(i)] • (Uniform consistency) $\sup_{i\le N}\left\vert \widehat{{\theta}}_{Ti}-{\theta}_{i0}\right\vert =o_{P}\left(1\right).$ • (Bounded variance) $0<\inf_{i\le N}{\sigma}_{i}\le\sup_{i\le N}{\sigma}_{i}<\infty$, where ${\sigma}_{i}^{2}={\sigma}\left({\theta}_{i0}\right)^{2}$. • (Differentiable variance) ${\sigma}\left({\theta}\right)^{2}$ is continuously differentiable over $\Theta$. • (Edgeworth Expansion) $\sup_{i\le N}\left|\kappa_{3, \theta_{i0}}\right|<\infty$, $\sup_{i\le N}\left|\kappa_{4, \theta_{i0}}\right|<\infty$, $\sup_{i\le N}\sup_{x\in\mathbb{R}}\left|R_{T,2}\left(x,{\theta}_{i0}\right)\right|=o\left(T^{-1}\right).$ \end{enumerate}

Assumption (ref) is the deterministic-design counterpart of Assumption (ref), with two main differences. First, for consistency of \(\widehat{\theta}_{\tau}\), we impose uniform consistency of the individual estimators rather than a Berry-Esseen type condition. Second, the required conditions are assumed to hold uniformly over \(i\), rather than in expectation.

theoremUnder Assumptions (ref) and (ref)(i), as $N,T\rightarrow\infty$, it holds that \begin{equation*} \widehat{\theta}_{\tau}\xrightarrow{P}\theta_{\tau}^{{\rm {D}}}. \end{equation*}

Theorem (ref) demonstrates consistency of $\widehat{\theta}_{\tau}$ under uniform consistency of heterogeneous estimators. A condition of the form \( \frac{\log N}{T}=o(1) \) is standard for establishing uniform consistency over $i$, i.e., Assumption (ref)(i); see, for instance, Kato, Galvao, and Montes-Rojas (KATO201276). Such a ratio restriction typically appears when the parameter sequence is deterministic.

As in the stochastic case, the asymptotic normality result requires an assumption on the distributional behavior of $\{\theta_{i0}\}$. In the present setting, however, $\{\theta_{i0}\}$ are deterministic rather than random, so the assumption must be formulated directly in terms of their limiting empirical distribution.

assumptionThere exists $\varepsilon>0$ such that (i) The limiting distribution \(F=\lim_{N\to\infty}F_N\) is twice continuously differentiable on a neighborhood \(\mathcal{N}_{\varepsilon}(\theta_\tau)\) of \(\theta_\tau\), with density \(f=F'\). Moreover, \begin{equation*} f(\theta_\tau)\in(0,\infty) \qquadand\qquad \sup_{\theta\in\mathcal{N}_{\varepsilon}(\theta_\tau)} |f(\theta)|<\infty. \end{equation*} (ii) Let $\theta_{(1)}\leq\cdots\leq\theta_{(N)}$ denote the order statistics of the fixed array $\{\theta_{i0}\}_{i=1}^{N}$ and $\Delta_i\equiv \theta_{(i+1)}-\theta_{(i)}$, then \begin{equation*} \max_{\theta_{(i)}\in \mathcal{N}_{\varepsilon}(\theta_{\tau})}\left|\theta_{(i)}-F^{-1}\!\left(\tfrac{i}{N}\right)\right|=O\!\left(N^{-1}\right), \end{equation*} and \begin{equation*}\max_{\theta_{(i)},\theta_{(j)}\in \mathcal{N}_{\varepsilon}(\theta_{\tau})}|\Delta_{i}-\Delta_j|=O(N^{-2}).\end{equation*}

Assumption (ref) is closely related to Assumptions (ref), but here the focus is on the distribution of the deterministic array $\{\theta_{i0}\}$. Part (i) ensures that the limiting distribution has a strictly positive and smooth density around $\theta_{\tau}$, which is essential for quantile identification.

Part (ii) further guarantees that the empirical distribution of $\{\theta_{i0}\}$, $F_{N}$, is sufficiently “smooth” around $\theta_{\tau}$, thereby ensuring that the empirical cdf approximates $F$ closely in the relevant neighborhood. Intuitively, if the $\theta_{i0}$ were too sparse or irregular near $\theta_{\tau}$, the $\tau$-th quantile may be determined by finite numbers of order statistics at the $\tau$-th quantile, and its distribution would not stabilize as $N$ grows.

Together, Assumption (ref) ensures that the empirical quantile around $\theta_{\tau}$ behaves as if drawn from a smooth underlying distribution with density $f(\theta_{\tau})>0$, making subsequent asymptotic expansions and limit arguments valid.

theoremUnder Assumptions (ref), (ref) and (ref), as $N,T\rightarrow\infty$, $T^{1/2}\ll N\ll \tfrac{T^{3/2}}{(\log T)^2}$, \begin{equation*} \sqrt{N\sqrt{T}}\left(\widehat{\theta}_{\tau}-\theta_{\tau}^{{\rm {D}}}\right)\xrightarrow{d}\mathcal{N}\left(0,\frac{\sigma(\theta_{\tau})}{\sqrt{\pi}f(\theta_{\tau})}\right). \end{equation*}

It is worth emphasizing that the validity of Theorem (ref) differs substantially from that of Theorem (ref). When \(\theta_{i0}\) is deterministic, the proof becomes considerably more delicate for at least three reasons. First, one must handle the nonsmooth empirical distribution function directly. Second, the objective function, which involves an indicator function, is itself nonsmooth, and in the present setting this lack of smoothness cannot be circumvented by passing to its expectation; accordingly, a separate argument is required. Third, the usual stochastic equicontinuity arguments are no longer directly applicable under the unconventional rate \(\sqrt{N\sqrt{T}}\), so the proof must proceed by a different route.

The growth-rate restrictions on sample size $N$ and $T$ serve two distinct purposes: (i) The condition $T^{1/2}\ll N$ is required for the CLT of the non-centered empirical component. Intuitively, this ensures that a sufficiently large number of individuals lie within a $\sqrt{T}$-neighborhood of $\theta_{\tau}$, which is necessary for stable estimation of the quantile functional. In the extreme case where $T=\infty$, each $\theta_{i0}$ can be estimated perfectly, i.e., $\widehat{\theta}_{Ti}=\theta_{i0}$. Consequently, $\widehat{\theta}_{\tau}$ exactly equals the population quantile $\theta_{\tau}^{\mathrm{D}}$, leaving no stochastic fluctuation and hence no CLT to derive. (ii) The condition $N\ll \frac{T^{3/2}}{(\log T)^2}$ controls the asymptotic bias, which arises from two sources: the estimation error in $\widehat{\theta}_{Ti}$ and the smoothing approximation error incurred when replacing the indicator function $\mathbf{1}\{u\ge0\}$ by its continuous counterpart $\Phi(u)$.

Comparison with the Standard FE-QR Case

To gain further intuition, we now compare the standard fixed-effects panel quantile regression (FE-QR) model with the heterogeneous slope panel model studied in this paper. In particular, we focus on the different convergence rates and required sample size ratios for asymptotic unbiasedness, as presented in Table (ref).

Before comparing convergence rates, it is useful to emphasize that the two literatures target different parameters. In standard FE-QR, the parameter of interest is a common slope \(\beta_0(\tau)\) for the conditional \(\tau\)-quantile of the outcome distribution. In our framework, the parameter of interest is instead \(\theta_\tau\), the \(\tau\)-quantile of the cross-sectional distribution of unit-specific coefficients \(\{\theta_{i0}\}\). Therefore, although both approaches involve panel data, fixed effects, and quantile methods, they address different notions of heterogeneity: outcome heterogeneity in the former and coefficient heterogeneity across individuals in the latter.

\paragraph{Sample size rate restrictions.}

For the standard panel quantile regression model with individual fixed effects, GalvaoGuVolgushev20 established the asymptotic properties of the conventional fixed-effects QR estimator. Because the slope parameter $\bm{\beta}_{0}(\tau)$ is homogeneous across $i$ and $t$, the estimator achieves the $\sqrt{NT}$ convergence rate. However, they impose the restriction $N/T=o(1)$, which arises from the difficulty of controlling higher-order terms when smoothing the objective function.

Galvao and Wang galvao2015efficient proposed a minimum distance estimator based on individual time-series QR estimates $\widehat{\bm{\beta}}_{i0}(\tau)$; although the minimum distance estimator shares similarity with our two-step approach, they assume homogeneous $\bm{\beta}_{0}(\tau)$ and their estimator requires the same strong condition on time series growing faster than cross-section. Galvao and Kato galvao2016smoothed later showed that smoothing the objective function can reduce the restriction to $N/T=O(1)$, as the bias term is of order $\sqrt{N/T}$ and must be controlled.

table[table omitted — 877 chars of source]

Compared with the standard FE-QR model, the proposed heterogeneous-slope models relax the ratio restriction, but this flexibility comes at the cost of a slower convergence rate. The reason is straightforward: for all models, \( E(\widehat{\bm{\beta}}_{\tau}-\bm{\beta}_{\tau}) \) is of order (around) $T^{-1}$. Hence, after multiplying by the relevant rates $\sqrt{NT}$, $\sqrt{N\sqrt{T}}$, and $\sqrt{N}$, the bias terms become (around) $\ O\!\left(\tfrac{\sqrt{N}}{\sqrt{T}}\right),O\!\left(\tfrac{\sqrt{N}}{T^{3/4}}\right),O\!\left(\tfrac{\sqrt{N}}{T}\right),$ respectively. These orders therefore determine the corresponding rate restrictions required for asymptotic unbiasedness.

\paragraph{Intuition for the stochastic- and deterministic-design rates.}

To see this distinction, we first clarify why increasing $T$ helps recover $\theta_{\tau}$ in the deterministic case but not in the stochastic case. Suppose, hypothetically, that $T=\infty$, so that the first-step estimates yield the exact values $\{{\theta}_{i0}\}_{i=1}^{N}$. When $\{{\theta}_{i0}\}$ are treated as fixed, observing them exactly allows us to compute their empirical $\tau$-quantile $\theta_{\tau}^{\mathrm{D}}$ without any sampling error. Thus, increasing $T$ directly sharpens first-step estimation and leads to more accurate recovery of $\theta_{\tau}^{\mathrm{D}}$. When $\{{\theta}_{i0}\}$ are themselves random draws from an underlying population distribution, even observing them exactly ($T=\infty$) does not reveal the population $\tau$--quantile $\theta_{\tau}^{\mathrm{S}}$. Consequently, increasing $T$ improves the estimation of each ${\theta}_{i0}$ but does not reduce the sampling variability across $i$, so it does not help in identifying $\theta_{\tau}^{\mathrm{S}}$ beyond the usual $\sqrt{N}$ rate.

We have explained why $T$ affects the convergence rate when $\{{\theta}_{i0}\}$ are treated as deterministic. The appearance of the $T^{1/4}$ factor, however, is somewhat unusual in the quantile regression literature. A useful heuristic is to view each first-step estimator as a noisy observation of the latent heterogeneous coefficient,

equation[equation omitted — 89 chars of source]

where $T^{-1/2}e_i$ represents the noise of order $T^{-1/2}$. Then $\widehat{\theta}_\tau$ behaves as a quantile estimation of noisy observations, with typical variance function

equation*[equation* omitted — 236 chars of source]

where $P_i=P(\widehat{\theta}_{Ti}\le \theta_\tau)$. If $\theta_{i0}$ is stochastic, the first-step noise is asymptotically negligible, so $P_i\approx P(\theta_{i0}\le \theta_\tau)=\tau>0$, and $\frac{1}{N^2}\sum_{i=1}^N P_i(1-P_i) \asymp \frac{1}{N}$, yielding a convergence rate of $\sqrt{N}$.

By contrast, when $\theta_{i0}$ is deterministic, the randomness comes solely from the noise. In that case, given the magnitude of the noise, only units with $\theta_{i0}$ lying within a $T^{-1/2}$ neighborhood of $\theta_\tau$ make a non-negligible contribution. For units outside this neighborhood, we have $P_i\in\{0,1\}$, and hence $P_i(1-P_i)=0$. The number of informative units is of order $NT^{-1/2}$. Therefore, \( \frac{1}{N^2}\sum_{i=1}^N P_i(1-P_i) \asymp \frac{1}{N^2}\cdot \frac{N}{\sqrt{T}} = \frac{1}{N\sqrt{T}}, \) which implies the convergence rate $\sqrt{N\sqrt{T}}$.

\paragraph{Applying the sample mean in Step 2 of Algorithm 1.} If, in Step 2 of Algorithm (ref), we replace the quantile operator by the sample mean, namely

equation*[equation* omitted — 92 chars of source]

then the estimator no longer targets a quantile of the cross-sectional distribution of \(\{\theta_{i0}\}\). Instead, in the stochastic case it targets the expectation \(E(\theta_{i0})\), and in the deterministic case it targets the corresponding empirical average. In this case, the convergence rate may differ from that of our original quantile-based estimator.

To illustrate this point, we use the noisy-observation representation in (ref). Then \( Var(\widehat{\theta}_{\tau}) =Var\!\left(\frac{1}{N}\sum_{i=1}^N \widehat{\theta}_{Ti}\right)=\frac{1}{N}Var\!\left( \widehat{\theta}_{Ti}\right). \)

In the stochastic case, the cross-sectional randomness in \(\theta_{i0}\) is the leading source of variation, while the first-step estimation noise is asymptotically negligible. Hence, \( Var(\widehat{\theta}_{\tau}) \asymp \frac{1}{N}Var(\theta_{i0}), \) which yields the convergence rate \(\sqrt{N}\), the same as in the quantile case.

In the deterministic case, by contrast, the only source of randomness comes from the first-step estimation noise. Using (ref), we obtain

equation*[equation* omitted — 115 chars of source]

and therefore \( Var(\widehat{\theta}_{\tau}) =\frac{1}{NT}Var(e_i). \) This implies the convergence rate \(\sqrt{NT}\), which is faster than the rate \( \sqrt{N\sqrt{T}} \) obtained in the quantile case. This difference arises because, unlike quantile regression, which is driven mainly by units in a neighborhood of the target quantile, mean regression depends equally on all units. A related discussion for the mean regression case is provided in Section 3.2.1 of Fernández-Val et al. (fernandez2022dynamic).

Bootstrap Inference

In this section, we describe bootstrap procedures for constructing confidence intervals under two scenarios: stochastic design and deterministic design. Below we will provide conditions to establish the consistency of both procedures.

Consider the following procedure for the stochastic-design.

description• Stochastic-Design Quantile Bootstrap (SQB)
enumerate[Step 1] • Compute the original estimate $\widehat{\theta}_{\tau}$. • For each $i$, generate the first-step bootstrap sample $\left\{ \bm{X}_{it}^{*b}:t\geq1\right\} $ by sampling with replacement from the original sample $\left\{ \bm{X}_{it}:t\geq1\right\} $. This resampling is performed independently across $i$. • Generate the second-step bootstrap sample $\left\{ \bm{X}_{it}^{**b}:t\geq1,\,i\geq1\right\} $ by drawing units $i$ with replacement from the index set $\{1,\dots,N\}$, and each selected $i$ includes the entire time series from the first-step sample $\left\{ \bm{X}_{it}^{*b}:t\geq1\right\} $. • Compute the bootstrap estimate $\widehat{\theta}_{\tau}^{**b}$ following a similar procedure as $\widehat{\theta}_{\tau}$, with $\bm{X}_{it}$ being replaced by $\bm{X}_{it}^{**b}$. • Repeat Step 2 to Step 4 for $B$ times. The bootstrap confidence interval of $\sqrt{N}\left(\widehat{\theta}_{\tau}-\theta_{\tau}^{{\rm {S}}}\right)$ is then constructed based on $\left\{ \sqrt{N}\left(\widehat{\theta}_{\tau}^{**b}-\widehat{\theta}_{\tau}\right)\right\} _{b}$.

Next, we consider the following procedure for the deterministic-design.

description• Deterministic-Design Quantile Bootstrap (DQB)
enumerate[Step 1] • Compute the original estimate $\widehat{\theta}_{\tau}$. • For each $i$, generate the bootstrap sample $\left\{ \bm{X}_{it}^{*b}:t\geq1\right\} $ by sampling with replacement from the original time series $\left\{ \bm{X}_{it}:t\geq1\right\} $. This resampling is performed independently across $i$. • Compute the bootstrap estimate $\widehat{\theta}_{\tau}^{*b}$ following a similar procedure as $\widehat{\theta}_{\tau}$, with $\bm{X}_{it}$ being replaced by $\bm{X}_{it}^{*b}$. • Repeat Step 2 to Step 3 for $B$ times. The bootstrap confidence interval of $\sqrt{N\sqrt{T}}\left(\widehat{\theta}_{\tau}-\theta_{\tau}^{{\rm {D}}}\right)$ is then constructed based on $\left\{ \sqrt{N\sqrt{T}}\left(\widehat{\theta}_{\tau}^{*b}-\widehat{\theta}_{\tau}\right)\right\} _{b}$.

Unless otherwise stated, we compute the critical values using the equal-tail bootstrap-based confidence intervals. Let $P^{*}$, $E^{*}$, and $Var^{*}$ denote the bootstrap probability, expectation, and variance, respectively, conditional on the original sample, $\left\{ \bm{X}_{it}\right\} _{it}$. Similarly, let $P^{**}$, $E^{**}$, and $Var^{**}$ denote the bootstrap probability, expectation, and variance conditional on the first step bootstrap data, $\left\{ \bm{X}_{it}^{*}\right\} _{it}$. Let ${\sigma}_{i}^{*2}=\lim_{T\to\infty}Var^{*}\left(\sqrt{T}\widehat{{\theta}}_{Ti}^{*}\right)$ denote the SQB first-step and DQB (two are the same procedure) bootstrap asymptotic variance.

assumption[Bootstrap high-level conditions] Conditional on $\{{\theta}_{i0}\}_{i}$, let $W_{Ti}^{*}={\sigma}_{i}^{*-1}\sqrt{T}\left(\widehat{{\theta}}_{Ti}^{*}-\widehat{{\theta}}_{Ti}\right)$ be the first-step bootstrap statistic. \begin{enumerate}[(i)] • (Identical asymptotic variance function) $ \sigma_i^2= \sigma_i^{*2}.$ • (Edgeworth Expansion) Let $\widehat{p}_{1,i}\left(x\right)=\frac{\widehat{\kappa}_{3,i}}{6}\left(1-x^{2}\right)$ and $\widehat{p}_{2,i}\left(x\right)=\frac{\widehat{\kappa}_{4,i}}{24}\left(-x^3+3x\right)+\frac{\widehat{\kappa}_{3,i}^2}{72}\left(-x^5+10x^3-15x\right)$, with $\sup_{i\le N}\widehat{\kappa}_{3,i}=O_{P}\left(1\right)$ and $\sup_{i\le N}\widehat{\kappa}_{4,i}=O_{P}\left(1\right)$, then \begin{equation*} \sup_{i\le N}\sup_{x\in\mathbb{R}}\left|P^{*}\left(W_{Ti}^{*}\le x\right)-\left[\Phi\left(x\right)+T^{-1/2}\widehat{p}_{1,i}\left(x\right)\phi\left(x\right)+T^{-1}\widehat{p}_{2,i}\left(x\right)\phi\left(x\right)\right]\right|=o_{P}\left(T^{-1}\right). \end{equation*} \end{enumerate}

Assumption (ref) is the bootstrap counterpart of the Assumptions (ref) and (ref). We do not impose any explicit relationship between $\widehat{p}_{1,i}$ and $p_{1,i}$, or between $\widehat{p}_{2,i}$ and $p_{2,i}$. In fact, for our argument, what matters is only that these terms retain the polynomial form arising in the Edgeworth expansion. This structure is sufficient, after appropriate treatment, to show that their contributions are asymptotically negligible.

theorem(i) When $\{\theta_{i0}\}_{i}$ are stochastic, under Assumptions (ref), (ref), (ref), and (ref), as $N,T\rightarrow\infty$ with $\sqrt{N}/{T}=O(1)$, \begin{equation} \sup_{x\in\mathbb{R}}\left\vert P^{*}\left(\sqrt{N}(\widehat{\theta}_{\tau}^{**}-\widehat{\theta}_{\tau})\leq x\right)-P\left(\sqrt{N}(\widehat{\theta}_{\tau}-\theta_{\tau}^{{\rm {S}}})\leq x\right)\right\vert \xrightarrow{P}0. \end{equation} (ii) Conditional on $\{\theta_{i0}\}_{i}$, under Assumptions (ref), (ref), (ref), and (ref), as $N,T\rightarrow\infty$ with $T^{1/2}\ll N\ll \tfrac{T^{3/2}}{(\log T)^2}$, \begin{equation} \sup_{x\in\mathbb{R}}\left\vert P^{*}\left(\sqrt{N\sqrt{T}}(\widehat{\theta}_{\tau}^{*}-\widehat{\theta}_{\tau})\leq x\right)-P\left(\sqrt{N\sqrt{T}}(\widehat{\theta}_{\tau}-\theta_{\tau}^{{\rm {D}}})\leq x\right)\right\vert \xrightarrow{P}0. \end{equation}

Theorem (ref) establishes the validity of the two bootstrap procedures in both settings under the same ratio restriction. It is worth noting that, when $\theta_{i0}$ is stochastic and $\sqrt{N}/T \to c \in (0,\infty)$, the bias term is asymptotically non-negligible. Nevertheless, the bootstrap remains valid in this case and correctly captures the bias as well.

An Application to the Least Squares Case

This section provides an example considering the ordinary least squares estimator as the first-step estimator, given its important role in empirical work. In this example we check the high-level assumptions provided previous to illustrate the methods. Consider the following condition:

assumption[First-step least squares estimation] For model (ref), \begin{enumerate}[(i)] • (Moment conditions) $E\left(\bm{Z}_{it}\varepsilon_{it}\right)=\bm{0}$, $E\left(\left\Vert \bm{Z}_{it}\right\Vert ^{8}\right)<\infty$, and $E\left(\left|\varepsilon_{it}\right|^{8}\right)<\infty$, $E\left(\bm{Z}_{it}\bm{Z}_{it}^{\top}\right)$ and $Var\left(\bm{Z}_{it}\varepsilon_{it}\right)$ exist and are non-singular. • (Cramér's condition) Let $\bm{\mathcal{Z}}_{it}=\left(\bm{Z}_{it}^{\top}\varepsilon_{it},vech\left(\bm{Z}_{it}\bm{Z}_{it}^{\top}\right)^{\top}\right)^{\top}$. For every nonzero vector $\bm{t}\in\mathbb{R}^{K+K\left(K+1\right)/2}$, \begin{equation*} \limsup_{\left\Vert \bm{t}\right\Vert \to\infty}\left|E\left(\exp\left(i\bm{t}^{\top}\bm{\mathcal{Z}}_{it}\right)\right)\right|<1. \end{equation*} \end{enumerate}

Assumption (ref) imposes standard moment conditions and a Cramér-type nonlattice condition for the OLS estimator. The component involving $vech\left(\bm{Z}_{it}\bm{Z}_{it}^{\top}\right)$ is included mainly for technical convenience, so that the Edgeworth expansion can be derived by treating OLS as a smooth function of the sample moments. It is not essential to the underlying result and could be relaxed with a more involved proof.

commentFor the first step-estimator, we discuss an example of the M-estimator, which encompasses several different estimators, as e.g., least squares, maximum likelihood, quantile regression, and many other widely used estimators used in practice; see Section 5 of Van der Vaart van2000asymptotic. Based on the data-generating process (ref), and the imposed assumptions, we construct a criterion function: \begin{equation*} g\!\left(\bm{X}_{it}({\theta}_{i0}),\,{\theta}\right):\mathbb{R}^{K+1}\times\mathbb{R}^{K}\to\mathbb{R}, \end{equation*} such that the true parameter $\bm{\theta}_{i0}$ uniquely minimizes the population objective \begin{equation} \bm{\theta}_{i0}\;=\;\arg\min_{\bm{\theta}\in\Theta}\;\mathbb{E}\!\left[g\!\left(\bm{X}_{it}(\bm{\theta}_{i0}),\,\bm{\theta}\right)\right]. \end{equation} For notational simplicity, let $\nabla_{\theta}^{j}$ denote a vector of all distinct partial derivatives with respect to $\bm{\theta}$ of order $j$. Define the score $\psi(\bm{X}_{it},\bm{\theta})\equiv\nabla_{\theta}g(\bm{X}_{it},\bm{\theta})$ and the objective $\mathbb{Q}_{Ti}(\bm{\theta})=T^{-1}\sum_{t=1}^{T}g(\bm{X}_{it},\bm{\theta})-\frac{1}{T}\sum_{t}g\left(\bm{X}_{it},\bm{\theta}_{i0}\right)$. \begin{assumption}[First-step M-estimation, uniform in $i$] Conditional on $\{\bm{\theta}_{i0}\}_{i}$, \begin{enumerate} • (Uniform identification) For every $\delta>0$ there exists $\varepsilon_{\delta}>0$ such that \begin{equation*} \inf_{1\le i\le N}\ \inf_{\|\bm{\theta}-\bm{\theta}_{i0}\|=\delta}\ \mathcal{Q}_{i}(\bm{\theta})\ \ge\ \varepsilon_{\delta},\qquad\mathcal{Q}_{i}(\bm{\theta})=E\!\left[\mathbb{Q}_{Ti}(\bm{\theta})\mid\bm{\theta}_{i0}\right]. \end{equation*} • (Interior) $\bm{\theta}_{i0}\in\mathrm{int}(\Theta)$ for all $i$. $\Theta$ is compact. • (Continuity) With probability one, $\psi(\bm{X}_{it},\bm{\theta})$ is continuous at each $\bm{\theta}\in\Theta$. • (Envelope) There exists an envelope $G(\bm{X}_{it})$ with $E\!\left[G(\bm{X}_{it})\right]<\infty$ such that \begin{equation*} \sup_{\bm{\theta}\in\Theta}|g(\bm{X}_{it},\bm{\theta})|\le G(\bm{X}_{it}). \end{equation*} • (Uniform nondegeneracy) $0<\inf_{\bm{\theta}\in\Theta}\lambda_{\min}\!\left(Var\!\left(\psi(\bm{X}_{it},\bm{\theta})\mid\bm{\theta}\right)\right)$ and $0<\inf_{\bm{\theta}\in\Theta}\lambda_{\min}\!\left(E\!\left[\nabla_{\theta}\psi(\bm{X}_{it},\bm{\theta})\mid\bm{\theta}\right]\right).$ • (Smoothness and moments) There exist $\Psi\left(\bm{X}_{it}\right)$ with $\sup_{\bm{\theta}\in\Theta}E\!\left[\Psi(\bm{X}_{it})^{6}\mid\bm{\theta}_{i0}\right]<\infty$ and $\varepsilon>0$ such that for $0\le j\le4$ and all $\bm{X}_{it}$, $\nabla_{\theta}^{j}\psi(\bm{X}_{it},\bm{\theta})$ exists on a neighborhood of $\bm{\theta}_{i0}$, $\mathcal{N}_{\varepsilon}\left(\bm{\theta}_{i0}\right)$. $\sup_{\bm{\theta}\in\mathcal{N}_{\varepsilon}\left(\bm{\theta}_{i0}\right)}\left\Vert \nabla_{\theta}^{j}\psi(\bm{X}_{it},\bm{\theta})\right\Vert \le\Psi\left(\bm{X}_{it}\right)$, and for each $\bm{\theta}\in\mathcal{N}_{\varepsilon}\left(\bm{\theta}_{i0}\right)$, \begin{equation*} \left\Vert \nabla_{\theta}^{4}\psi(\bm{X}_{it},\bm{\theta})-\nabla_{\theta}^{4}\psi(\bm{X}_{it},\bm{\theta}_{i0})\right\Vert \le\Psi(\bm{X}_{it})\left\Vert \bm{\theta}-\bm{\theta}_{i0}\right\Vert . \end{equation*} • (\emph{Integrable characteristic function}) Let $\phi_{X,i}\left(\bm{t}\right)=E\left(\exp\left(i\bm{t}^{\top}\psi(\bm{X}_{it},\bm{\theta}_{i0})\right)\right)$. For every nonzero vector $\bm{t}\in\mathbb{R}^{K}$ and for some $v\ge1$, \begin{equation*} \sup_{\bm{\theta}\in\Theta}\int_{\mathbb{R}^{K}}\left|\phi_{X,i}\left(\bm{t}\right)\right|^{v}d\bm{t}\le\infty. \end{equation*} \end{enumerate} \end{assumption} Assumption (ref)(i)--(v) are standard conditions required for consistency and asymptotic normality of the first-step M-estimator, uniformly over $i$. Assumptions (ref)(vi)-(vii) are required for a uniform higher-order expansion. The former imposes additional smoothness and moment conditions on the relevant derivatives, whereas the latter is a regularity condition used to justify the existence and continuity of density; see, for example, Section XVI.2 of Feller (feller1966introduction). It also implies the Cramér condition for Edgeworth expansion for the distribution function.
theorem(i) Under Assumptions (ref), (ref), and (ref), as $N,T\rightarrow\infty$ with $\sqrt{N}/{T}=O(1)$, results in Theorems (ref), (ref), and (ref)(i) continue to hold. (ii) Under Assumptions (ref), (ref), and (ref), as $N,T\rightarrow\infty$ with $T^{1/2}\ll N\ll \frac{T^{3/2}}{(\log T)^2}$, results in Theorem (ref), (ref), and (ref)(ii) continue to hold.

Theorem (ref) shows that, under identification and cdf conditions for $\theta_{\tau}$, together with standard moment and Cram\'er's conditions for the OLS estimator and an appropriate ratio restriction, consistency, asymptotic normality, and bootstrap validity all hold in both the stochastic and deterministic cases. Overall, the assumptions imposed are relatively mild.

Simulation Experiments

Quantile of Sample Mean

We now present simulation results to demonstrate the performance of the bootstrap procedure for inference regarding the sample mean. We consider the following DGP. For each $i$,

equation*[equation* omitted — 201 chars of source]

such that $E(X_{it})=\theta_{i0}$ and $\mathrm{Var}(X_{it})=\sigma_{i0}^{2}$. Hence, we estimate $\widehat{\theta}_{Ti}=\frac{1}{T}\sum_{t=1}^{T}X_{it}.$ In this case, we may represent the data-generating process as \( X_{it}=\theta_{i0}+\sigma_{i0}\varepsilon_{it}, \) where $(\theta_{i0},\sigma_{i0})$ are heterogeneous across $i$, while $\varepsilon_{it}$ is i.i.d. over both $i$ and $t$. This specification is not covered exactly by model (ref), our simulation results indicate that the proposed method still performs well in this setting, suggesting potential scope for future extensions to more general models and estimators.

In the stochastic design, we generate $\theta_{i0}$ i.i.d. following the uniform distribution $U(0,2)$. In contrast, under the deterministic design, we set

equation*[equation* omitted — 53 chars of source]

The default quantile index is $\tau=0.70$, and the baseline sample sizes are $N=T=40$, with one dimension varied at a time. Table (ref) reports the coverage probabilities for the two proposed bootstrap methods: SQB and DQB.

Panels A and B of Table (ref) report the bias and coverage probabilities of two methods under the stochastic design. In this setting, SQB performs well, whereas DQB performs noticeably less well. In Panel A, we set $\sigma_{i0}=1$ for all $i$. For a given $T$, both bias and coverage accuracy improve as $N$ increases. By contrast, for a given $N$, increasing $T$ tends to worsen performance. For example, when $T=20$, increasing $N$ from 20 to 80 reduces the bias from $-0.0156$ to $0.0023$ and increases the SQB coverage probability from $90.75\%$ to $93.10\%$. In contrast, when $N=20$, increasing $T$ from 20 to 80 changes the bias from $-0.0156$ to $-0.0185$ and lowers the SQB coverage probability from $90.75\%$ to $83.74\%$. The DQB coverage probabilities exhibit a similar pattern, although their overall performance is poor, with coverage rates typically ranging from \(0.40\) to \(0.60\). This is expected, since DQB ignores the additional randomness arising from sampling the heterogeneous coefficients \(\theta_{i0}\). The pattern is nevertheless consistent with the \(\sqrt{N}\) convergence rate underlying both the CLT and the bootstrap CLT. By contrast, the time-series dimension \(T\) introduces additional sampling and bootstrap variability through the first-step estimators \(\widehat{\theta}_{Ti}\).

Panel B introduces additional heterogeneity by setting $\sigma_{i0}\overset{i.i.d.}{\sim}U(0,2)$. The results exhibit a qualitatively similar pattern to Panel A, and the coverage probabilities remain comparable, suggesting that moderate heterogeneity in $\sigma_{i0}$ has only a limited effect on finite-sample performance.

table[table omitted — 2,259 chars of source]

Panels C and D of Table (ref) report the results under the deterministic design. Panel C assumes homoskedasticity ($\sigma_{i0}=1$ for all $i$), whereas Panel D introduces heteroskedasticity. In this setting, the SQB method becomes overly conservative, with coverage probabilities exceeding $99\%$, highlighting its invalidity when $\theta_{i0}$ is deterministic rather than stochastic. By contrast, DQB remains valid, with most coverage probabilities around $80\%$, although they are somewhat lower than the SQB results under the stochastic design. Another consistent pattern is that the coverage probabilities of SQB are uniformly higher than those of DQB across all scenarios, as expected, since SQB incorporates one additional layer of randomness relative to DQB.

Interestingly, Panel C reveals the opposite pattern from the stochastic case: increasing $T$ improves bias and coverage (e.g., coverage increases from $83.69\%$ to $86.88\%$ when $N=20$ as $T$ increases from 20 to 80), whereas increasing $N$ worsens bias and coverage (e.g., coverage reduces from $83.69\%$ to $81.14\%$ when $T=20$ as $N$ rises from 20 to 80). When the individual parameters $\theta_{i0}$ are deterministic, the cross-sectional averaging effect disappears. Consequently, inference precision primarily depends on the time dimension $T$, and increasing $N$ no longer improves finite-sample inference. Panel D exhibits a similar trend, with slightly weaker DQB performance, indicating that heterogeneity in $\sigma_{i0}$ further reduces coverage accuracy.

Table (ref) examines sensitivity to the quantile index \(\tau\), holding \(N=T=40\) fixed and otherwise maintaining the same design as in the previous panels. Across all panels, coverage remains fairly stable for \(\tau\) between \(0.30\) and \(0.70\), but becomes more erratic as \(\tau\) moves toward the tails, such as \(\tau=0.10\) or \(\tau=0.90\). The bias exhibits a similar pattern, being small in magnitude for central quantiles but larger in the tails. This finding is consistent with the requirement that \(\tau\in(0,1)\), since more extreme quantiles are estimated from less informative data about \(\theta_{\tau}\).

table[table omitted — 2,206 chars of source]

To estimate $\theta_{\tau}$, the most informative observations are those for individuals near the $\tau$th quantile. For example, when $\tau=0.50$, individuals between the 30th and 70th percentiles contribute most to identifying the parameter. However, when $\tau$ is large (e.g., $\tau=0.99$), only a few upper-tail observations inform the upper bound of $\theta_{\tau}$, leading to potential large $\widehat{\theta}_{\tau}$ and overrejection. Conversely, when $\tau$ is very small, limited lower-tail information tends to produce small $\widehat{\theta}_{\tau}$ and underrejection. Thus, although results for large $\tau$ may appear slightly better (e.g. SQB yields a very good probability of 95.94% at $\tau=0.90$ in Panel A), this is likely due to sampling variability.

Overall, SQB performs reasonably well in the stochastic design, with frequencies above $85\%$ in most cases, while DQB remains stable with most frequencies above $80\%$ in the deterministic design.

Quantile of Least Squares Estimator

As an important application, we next consider inference in a regression model with a scalar dependent variable $y_{it}$ and $K$ regressors $\bm{z}_{it}\in\mathbb{R}^{K}$, where $\bm{z}_{it}=(z_{it,k})_{k=1}^{K}$ and $z_{it,1}=1$ denotes the intercept. The model is specified as

align[align omitted — 134 chars of source]

The parameter of interest is the $\tau$th quantile of the heterogeneous slope coefficient $\theta_{i0}$ on the last regressor $z_{it,K}$, whose distribution across individuals $i$ characterizes the extent of heterogeneity in the sensitivity of $y_{it}$ to this regressor.

To enhance robustness and avoid overreliance on specific distributional assumptions, we adopt a different distribution for $\theta_{i0}$ than that used in Section (ref). In the stochastic design, we generate $\theta_{i0}\overset{i.i.d.}{\sim}\Phi$, where $\Phi$ denotes the standard normal cumulative distribution function. In the deterministic design, we fix the heterogeneity pattern by setting $\theta_{i0}=\Phi^{-1}(i/(N+1))$ for $i=1,\ldots,N$, which evenly spans the support of $\Phi^{-1}$ across individuals. For each $i$, we compute the corresponding least squares estimator $\widehat{\theta}_{Ti}$ from the regression of $y_{it}$ on $\bm{x}_{it}$ over $t=1,\ldots,T$.

We set $K=10$ to ensure a moderate-dimensional regression design that avoids the degenerate case of a constant term and a single regressor. This choice follows the recommendation of MacKinnon, Nielsen, and Webb (mackinnon2023fast), who emphasize that very small values of $K$ (e.g., $K=2$) can lead to spuriously optimistic finite-sample performance. A relatively larger $K$ introduces realistic estimation noise and more variability in the estimated coefficients, thereby providing a more stringent test for the proposed quantile inference procedures. Correspondingly, we select larger values of $(N,T=40,80,160)$ to balance the estimation precision given the higher-dimensional regressors.

The simulation results are reported in Table (ref). Overall, the patterns closely resemble those observed for the mean parameter in Section (ref). Under the stochastic design, SQB exhibits satisfactory performance, with bias and coverage probability improving as $N$ increases and $T$ decreases. Its behavior is nearly identical across the homoskedastic and heteroskedastic cases.

Under the deterministic design, where the heterogeneity pattern of $\theta_{i0}$ is fixed across individuals, DQB performs reasonably well. Inference accuracy improves as $N$ decreases, since smaller $N$ mitigates the variation induced by deterministic heterogeneity. When $N=40$, increasing $T$ leads to better coverage, whereas for $N=160$, increasing $T$ from 40 to 160 does not further improve performance, which seems to contradict the result in the sample mean case. However, in unreported results, we find that increasing $T$ can enhance performance, as observed in the sample mean case, though this requires a much larger $T$ relative to $N$ (e.g., $T\geq500$).

Overall, the results confirm that both SQB and DQB deliver robust inference across a wide range of $(N,T)$ configurations, though their relative advantages differ between the stochastic and deterministic designs.

table[table omitted — 2,349 chars of source]

Empirical Studies

We examine mutual fund performance. Although a large body of research has recognized substantial heterogeneity in skill across funds; see, for example, Kaplan and Schoar (kaplan2005private), Kosowski et al. (kosowski2006can), Fama and French (fama2010luck), and Hounyo and Lin (hounyo2025can), these studies do not focus directly on the cross-sectional pattern of managerial skill. Our goal is to examine how managerial performance and style exposures vary across funds through the cross-sectional quantiles of fund-specific coefficients. This perspective is informative along at least three dimensions. First, even when average skill is approximately zero, the quantiles can reveal the fraction of funds that fail to cover costs, roughly break even, or deliver positive abnormal performance. Second, it allows us to assess the economic magnitude of cross-sectional differences, such as whether top-percentile fund managers are substantially better than others or only slightly so, and how poor the performance of bottom-tail funds is. Third, it can reveal asymmetry and nonlinearities in the cross section, helping distinguish broad dispersion from genuine concentration of skill in particular parts of the distribution.

Due to potential model misspecification in the classical models of Fama and French fama1993common and Carhart carhart1997persistence, a large body of research has proposed alternative approaches to measuring fund performance (see, e.g., Kothari and Warner kothari2001evaluating, Cremers and Petajisto cremers2009active, Cremers, Petajisto, and Zitzewitz cremers2012should). Mutual fund outperformance may arise from multiple dimensions of managerial skill, including stock selection, market timing, and liquidity timing. Following the literature, we consider the following timing-augmented performance model:

align[align omitted — 324 chars of source]

where $r_{it}$ denotes the excess return of fund $i$ over the risk-free rate at time $t$. The factors include the market excess return ($RMRF_{t}$), the size and value-growth factors ($SMB_{t}$ and $HML_{t}$) from Fama and French fama1993common, and the momentum factor ($MOM_{t}$) from Carhart carhart1997persistence. $V_{t}$ and $L_{t}$ denote monthly market volatility and liquidity, respectively. Volatility is computed as the standard deviation of daily demeaned market excess returns within the month, while liquidity follows the P{\'a}stor and Stambaugh pastor2003liquidity measure.\footnote{We also confirm qualitatively similar results using the Amihud amihud2002illiquidity measure.} The long-run averages $\overline{V}$ and $\overline{L}$ are calculated from the past 60 months.

In this specification, $\gamma_{i,1}$, $\gamma_{i,2}$, and $\gamma_{i,3}$ capture the fund manager's ability to time market return, volatility, and liquidity, respectively. A positive $\gamma_{i,1}$ indicates successful return timing, a negative $\gamma_{i,2}$ reflects effective volatility timing (reducing exposure in turbulent markets), and a positive $\gamma_{i,3}$ suggests successful liquidity timing.

We let $\bm{\theta}_{i0}=(\alpha_i,\beta_{i,1},\beta_{i,2},\beta_{i,3},\beta_{i,4},\gamma_{i,1},\gamma_{i,2},\gamma_{i,3})^\top$. Our objective is to study cross-sectional heterogeneity in these coefficients through their quantiles, both within our sample of 187 mutual funds and, under the stochastic interpretation, in the broader mutual fund population.

Define

equation*[equation* omitted — 167 chars of source]

whose elements represent the $\tau$th quantile of the corresponding component in $\{\bm{\theta}_{i0}\}_{i}$ across funds. To this end, we estimate each parameter $\widehat{\bm{\theta}}_{Ti}$ individually and then compute coefficient quantiles over $\tau\in[0.01,0.99]$, which provide a detailed descriptive picture of fund performance and timing heterogeneity across funds. Stochastic-design (SQB) and deterministic-design (DQB) 95% confidence intervals are constructed following the bias-corrected percentile formulas defined earlier.

Our sample consists of a balanced monthly panel covering $T=228$ periods from January 1984 to December 2002, drawn from the Center for Research in Security Prices (CRSP) Survivor-Bias-Free U.S. Mutual Fund Database. This period is less affected by post-publication bias, as much of the subsequent debate in finance suggests that factors lose their excess returns after publication; see, e.g., McLean and Potiff mclean2016does. Following Ferson and Lin ferson2014alpha and Busse and Tong busse2012mutual, we exclude index funds to focus on actively managed funds. To mitigate incubation bias, we further exclude funds with total net assets below \$15 million, as recommended by Elton et al. elton2001first. The resulting balanced panel includes $N=187$ funds. Factor return data are obtained from Ken French's data library.\footnote{\url{https://mba.tuck.dartmouth.edu/pages/faculty/ken.french/data_library.html}}

\paragraph{Intercept and Timing Ability Parameters.}

Figures (ref) and (ref) report the quantile-specific estimates and confidence intervals. The black line shows the point estimates, while the red and blue dashed lines depict the corresponding quantile-specific stochastic- and deterministic-design intervals. Consistent with our simulation evidence, the stochastic-design intervals are generally wider, reflecting the additional sampling variability for random heterogeneity.

Panel (a) of Figure (ref) presents the cross-sectional quantiles of $\widehat{\alpha}_{\tau}$. The estimated quantile curve is centered near zero, indicating that the typical (median) fund does not generate significant abnormal performance. The tail estimates indicate that a small subset of funds achieves economically meaningful positive $\alpha_{i}$, although the associated confidence intervals widen substantially, reflecting greater uncertainty in the extremes. At the 90$^{\text{th}}$ percentile, the lower bound of the 95% confidence interval under the stochastic design is $0.05\%$, indicating that, in the broader mutual fund market, top-performing funds earn at least $0.05\%$ abnormal return per month. By comparison, the corresponding lower bound under the deterministic design is $0.09\%$, suggesting that, within our sample of 187 mutual funds, the top-performing funds achieve statistically significant abnormal returns of at least $0.09\%$.

figure[figure omitted — 764 chars of source]

Interpreting $\widehat{\alpha}_{\tau}$ as the estimated $\tau$-quantile of the cross-sectional distribution of managerial skill, the results suggest that fund performance is largely homogeneous in the median but exhibits considerable dispersion in the tails. The relatively symmetric shape around $\tau=0.5$ and a Kolmogorov--Smirnov test $p$-value of 0.15 (for variance-rescaled $\widehat{\alpha}_{\tau}$) suggest that the cross-sectional pattern of abnormal returns does not deviate significantly from normality with zero mean, supporting the view that persistent skill is rare.

Panels (b)-(d) of Figure (ref) display the quantile estimates of the timing parameters $\widehat{\gamma}_{1,\tau}$, $\widehat{\gamma}_{2,\tau}$, and $\widehat{\gamma}_{3,\tau}$.Panel (b) reports the quantile estimates for $\widehat{\gamma}_{1,\tau}$, which captures return-timing behavior. The estimates are close to zero for the lower quantiles and become positive beyond the midrange, indicating that a substantial share of funds exhibit positive return-timing coefficients. For the lower bound of the 95% confidence interval, both designs cross zero at approximately the 40th percentile. This implies that in the broader mutual fund market (stochastic design) and within our sample of 187 mutual funds (deterministic design), around 60% of funds display statistically positive return-timing ability. These estimates suggest that return-timing behavior is widespread, though the strength of the timing response varies considerably across funds.

Panel (c) presents the results for $\widehat{\gamma}_{2,\tau}$, which measures volatility-timing behavior. In this setting, effective volatility timing is associated with negative coefficients, as skilled managers are expected to reduce market exposure during periods of heightened volatility (Busse busse1999volatility). Only a small fraction of funds in the lower tail exhibit negative and statistically significant coefficients, suggesting that few managers actively decrease exposure when volatility rises. For the majority of funds, the positive $\widehat{\gamma}_{2,\tau}$ values imply procyclical behavior, that is, maintaining or even increasing market exposure when volatility increases. Overall, the results point to weak or inconsistent volatility-timing ability across the fund universe.

Panel (d) displays the quantile estimates for $\widehat{\gamma}_{3,\tau}$, which measures liquidity-timing behavior. The estimates are positive across most quantiles, and the lower bound of both confidence intervals crosses zero near the 50th percentile. This indicates that roughly 50% of funds exhibit statistically significant positive liquidity-timing coefficients in both the mutual fund market and our sample, suggesting that many managers adjust exposure in response to liquidity conditions. Taken together, Panels (b)-(d) show that timing activity, especially in return and liquidity dimensions, is prevalent, though its strength varies across funds.

\paragraph{Standard Factor Loadings.}

Figure (ref) presents the quantile estimates for the four traditional Fama--French--Carhart factor loadings, $\widehat{\beta}_{j,\tau}$ ($j=1,\ldots,4$). Overall, these loadings are stable but asymmetric across quantiles, reflecting heterogeneous exposures among funds.

Panel (a) shows $\widehat{\beta}_{1,\tau}$ (market factor), which remains positive and precisely estimated throughout, confirming that $RMRF$ is the dominant driver of mutual fund excess returns. The estimates lie mostly between 0.6 and 1.2, but a marked drop at the bottom quantiles indicates that a subset of conservative or benchmark-constrained funds exhibit weaker market sensitivity, whereas others maintain full market exposure.

figure[figure omitted — 749 chars of source]

Panel (b) displays $\widehat{\beta}_{2,\tau}$ (SMB), which shows a pronounced J-shaped pattern. Lower quantiles, typically associated with large-cap funds, have near-zero loadings and narrow confidence intervals, implying homogeneous exposure among large-cap portfolios. Higher quantiles, linked to small-cap strategies, display more dispersion and wider confidence intervals, reflecting heterogeneous and often aggressive small-firm exposures.

Panels (c) and (d) show $\widehat{\beta}_{3,\tau}$ (HML) and $\widehat{\beta}_{4,\tau}$ (MOM). For $\widehat{\beta}_{3,\tau}$, estimates are negative below the 40th percentile and positive above, implying that about 40% of funds tilt toward growth stocks while most favor value. Dispersion widens in the tails, indicating heterogeneous style preferences. The momentum loading $\widehat{\beta}_{4,\tau}$ stays negative for the lower 60% quantiles but turns mildly positive in the upper tail, suggesting limited yet diverse use of momentum strategies. Both factors display weaker but still asymmetric cross-sectional patterns compared with the more pronounced pattern in SMB.

In summary, the factor loadings exhibit systematic yet asymmetric cross-sectional heterogeneity. Timing activity, particularly in the return and liquidity dimensions, is prevalent among funds, whereas abnormal return and volatility-timing abilities appear relatively weak. Meanwhile, the size, value, and momentum exposures display nonlinear and uneven patterns, most notably the J-shaped structure in SMB, underscoring the diverse style positioning and heterogeneous risk profiles across the mutual fund universe.

Conclusion

Quantile regression is a useful tool for analyzing heterogeneous effects in panel data models, particularly when slope heterogeneity is present. However, conventional panel quantile regression focuses on heterogeneity in the outcome variable $y_{it}$ conditional on regressors. This study considers a different dimension of heterogeneity-across individuals \(i\)-by examining the distribution of individual-specific slope coefficients. Although this form of heterogeneity is empirically relevant, its analysis remains relatively limited in the existing literature.

We propose a two-step quantile estimation framework for the $\tau$-quantile of individual heterogeneous coefficients in panel data. The procedure combines first-step unit-level estimation with second-step cross-sectional quantile aggregation, allowing inference on quantile-based summaries of structural effects across units. This approach provides a direct way to study individual-level heterogeneity beyond conventional fixed-effect or pooled quantile regression methods.

Large-sample properties are established under two distinct designs: stochastic and deterministic. The stochastic-design case applies when the sample represents random draws from a population and inference targets the population distribution. In this case, the estimator achieves $\sqrt{N}$-consistency and asymptotic normality under a mild condition $\sqrt{N}/T=O(1)$, which is weaker than the conventional requirement $N/T=O(1)$ in fixed-effect quantile regression (Galvao and Kato galvao2016smoothed). The deterministic-design case applies when the full population is observed or when inference concerns the empirical distribution of the observed sample of size $N$. In this setting, the convergence rate becomes $\sqrt{N\sqrt{T}}$, a new rate in the literature that reflects the absence of cross-sectional randomness. Although faster than the standard $\sqrt{N}$-rate, this regime requires a more restrictive growth condition, $T^{1/2}\ll N\ll \frac{T^{3/2}}{(\log T)^2}$, to control estimation bias and sampling variability. We present an application in which the commonly used OLS estimator serves as the first-step estimator.

We further establish the validity of two bootstrap procedures corresponding to the two designs: the stochastic-design quantile bootstrap (SQB) and the deterministic-design quantile bootstrap (DQB). Both methods provide consistent inference for the asymptotic distribution of the proposed estimator. Simulation results show that the SQB performs well under stochastic designs, while the DQB yields reasonable and stable coverage under deterministic designs across a range of $(N,T)$ configurations.

The empirical application to mutual fund performance illustrates the practical relevance of the proposed method. The estimated quantile distributions of fund alphas and timing coefficients reveal pronounced heterogeneity in return- and liquidity-timing abilities, while volatility-timing and abnormal-return heterogeneity are limited. The results indicate an asymmetric pattern of slope heterogeneity across funds.

Overall, the proposed framework offers a practical tool for analyzing heterogeneity in structural effects within panel data. It complements existing quantile regression methods by focusing on the quantiles of individual effects rather than conditional outcome heterogeneity. Future work may extend these results to dependent time series or cross-sectional structures, as well as dynamic panel settings.