EconBase
← Back to paper

Fixed-k Inference for Conditional Extremal Quantiles

The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.

80,812 characters

Fixed-$k$ Inference for Conditional Extremal Quantiles



\title{Fixed-$k$ Inference for Conditional Extremal Quantiles\thanks{
We thank Federico Bugni, Xiaohong Chen, Tim Christensen, Yanqin Fan, Yoonseok
Lee, Zhijie Xiao, Yichong Zhang, and participants at the seminar/conference
at Boston College, PSU, SMU, EC$^{2}$ Conference 2019, Econometric Society North American Summer Meeting 2019,
and Greater New York Metropolitan Area Econometrics Colloquium 2019, for
very helpful comments and advice. Wang gratefully acknowledges the financial
support by the Applyby-Mosher fund.}}
\author{Yuya Sasaki\thanks{
Associate professor of economics, Vanderbilt University. Email:
[email removed]} \ and Yulong Wang\thanks{
Assistant professor of economics, Syracuse University. Email:
[email removed].}}
\date{First arXiv version: August 2019\\
This version: July 2020}
\maketitle

\begin{abstract}
\setlength{\baselineskip}{5.6mm} {\small We develop a new extreme value
theory for repeated cross-sectional and panel data to construct asymptotically
valid confidence intervals (CIs) for conditional extremal quantiles from a
fixed number $k$ of nearest-neighbor tail observations. As a by-product, we
also construct CIs for extremal quantiles of coefficients in linear random
coefficient models. For any fixed $k$, the CIs are uniformly valid without
parametric assumptions over a set of nonparametric data generating processes
associated with various tail indices. Simulation studies show that our CIs
exhibit superior small-sample coverage and length properties than
alternative nonparametric methods based on asymptotic normality. Applying
the proposed method to Natality Vital Statistics, we study factors of
extremely low birth weights. We find that signs of major effects are the
same as those found in preceding studies based on parametric models, but
with different magnitudes.  }

{\small { \ \ \newline
\textbf{Keywords: } conditional extremal quantile, confidence interval,
extreme value theory, fixed $k$, random coefficient} }
\end{abstract}

\setcounter{page}{1}  \small \normalsize
\newpage

\section{Introduction}

Tail risks and extreme events are important research topics in economics. In
many applications with multivariate analysis, features of interest are
conditional tail properties such as conditional extremal quantiles. This
article provides a new method to construct confidence intervals for
conditional extremal quantiles from a fixed number $k$ of nearest-neighbor tail observations.
Advantages of the proposed method are three-fold: first, it is robust
against flexible distributional assumptions unlike parametric methods;
second, the procedure yields
asymptotically valid confidence intervals for any fixed tuning parameter $k$ unlike existing kernel methods that rely on sequences of moving tuning parameters for asymptotically valid inference; and third, our confidence intervals enjoy a uniform coverage property over a set of data generating processes involving a set of values of the tail index.
In the existing literature, methods of inference about conditional quantiles concern about middle quantiles, e.g., \citeasnoun{Qu2015} -- also see \citeasnoun{Qu2019} -- based on local quantile estimators of \citeasnoun{Fan1994} and \citeasnoun{Yu1998}.
We aim to complement this existing literature by proposing a method of inference about conditional extremal quantiles.

Compared with unconditional tail features, the conditional tail counterparts
are more difficult to study. This is because conditional tails depend on
both marginal distributions and their joint behavior. Although marginal
distributions can be generally assumed to be approximately Pareto near the
tails,\footnote{
This statement follows from the Pickands-Balkema-de Haan Theorem (\citeasnoun
{BalkemadeHaan1974} and \citeasnoun{Pickands1975}). See \citeasnoun{deHaan07} for an
overview.} joint distributions cannot be generally assumed to be
approximated by a fully parametric joint distribution and thus are harder to
study given very limited tail observations. To model a covariate-dependent
yet tractable tails, the seminal paper by \citeasnoun{Chernozhukov05} extends the
quantile regression (QR) estimator of \citeasnoun{Koenker78} to tails, and proposes a
method called the extremal quantile regression (EQR). \citeasnoun{Chernozhukov11}
further investigate the EQR to construct confidence intervals (CIs) based on
subsampling.

The EQR approach is based on the assumption that the conditional extremal quantile can be well approximated by a parametric location-scale shift model:
\begin{equation}
Q_{Y|X=x}\left( \tau \right) \sim \mu \left( x\right) +\sigma \left(x\right)
(1-\tau )^{-\xi }  \label{EQR cond}
\end{equation}
for $\tau \rightarrow 1$, where $\mu \left( x\right) $ and $\sigma \left(x\right) $ are parametric functions that capture the location and scale, respectively.
The element $(1-\tau )^{-\xi }$ can be treated as the quantile function of a standard Pareto distribution, that is, $\mathbb{P}\left( Y>y\right) \sim y^{-1/\xi }$ where $1/\xi $ is the Pareto exponent and $\xi $ is the tail index.
This single parameter captures the tail shape in the way that a larger $\xi $ implies a heavier tail.
The assumption of model (\ref{EQR cond}) simplifies the conditional tail distribution so that the covariate $X$ only affects the location and scale, but not the shape.\footnote{\citeasnoun{WangLi13} formally establish that the location-shift model
assumption is equivalent to assuming $\xi $ remains constant across $x$.}
This is satisfied if $X$ and $Y$ are jointly normal but violated by many
other joint distributions. Unlike mid-sample features, misspecification bias
could be substantial in studying tail ones.\footnote{With this said, we remark that the existing literature suggests a couple of ways in which one can rationalize a possibly misspecified quantile regression. \citeasnoun{Angrist06} show that the parametric linear quantile regression function minimizes a weighted distance to the true nonparametric quantile regression function. \citeasnoun{Kato17} show that the linear quantile regression parameter is a weighted average of the slopes of the true
nonparametric quantile regression function.}
In this paper, we consider a wider class of flexible joint distribution models using a repeated cross-sectional or panel data structure.

There are a number of reasons for which we want to study conditional tail features, such as conditional extremal quantiles, under flexible joint distribution models. First, conditional value-at-risk (VaR) is a risk measure commonly used in financial management, insurance, and actuarial science.
Estimation and inference are studied by \citeasnoun{Chernozhukov01} and \citeasnoun{Engle04}, among others.
\citeasnoun{Adrian16} propose a new measure for systemic risk, $\Delta $-CoVar, defined as the difference between two conditional VaRs.
The tail shape governs the third-and higher-order moments of the portfolio return, which typically depend on other economic factors, e.g., business cycles.
As this is excluded by the location-scale model (\ref{EQR cond}), it is preferred to accommodate a larger class of joint distributions.
Second, \citeasnoun{Kelly14} find that extreme event risk affects asset pricing in the U.S. stock market.
The shape parameter measures tail risk and varies with other stock characteristics such as stock size.
Third, macroeconomists are interested in analyzing lower tails of the conditional distributions of GDP growth rate given financial conditions in the recent growth-at-risk literature -- see \citeasnoun{Adrian19} for example.
Fourth, top wealth inequality is an active research question in macro-finance literature (see, for example, \citeasnoun{Piketty03}, \citeasnoun{Gabaix16}, and \citeasnoun{Jones18}).
The tail of the wealth distribution is well documented to follow Pareto, and the exponent is in general a function of fundamentals in general equilibrium models.
For example, \citeasnoun{Beare17} derive a formula for the Pareto exponent and comparative statics results, and \citeasnoun{Toda19} applies that formula in a general equilibrium context.
Finally, investigating factors of infants' birth weights, such as mother's demographic characteristics and maternal behaviors, is an important question
in health economics (e.g., \citeasnoun{Abrevaya01} and \citeasnoun{Koenker01}).
The lower tails of the conditional distribution are especially of interest for
their critical health consequences -- see \citeasnoun{Chernozhukov11}.
Other economic issues about conditional tail features can be found in the
comprehensive review by \citeasnoun{Chernozhukov17}.

The existing literature suggests alternative approaches besides those based
on the parametric location-scale specification (\ref{EQR cond}). To our best
knowledge, they all focus on estimation, as opposed to inference, and can be
roughly categorized into two classes. The first class maintains some
parametric form but relaxes the location-shift model to allow for some
nonlinearity. \citeasnoun{WangTsai09} assume that $\xi (x)$ equals to $
\exp(x^{\intercal }\theta_{0}) $ for some unknown parameter $\theta _{0}$.
\citeasnoun{WangLi13} assume that the Box-Cox transformed $Y$ has linear
conditional quantiles in $X$. The second class is fully nonparametric and
constructs some local smooth estimators, including, for example, \citeasnoun
{Beirlant04}, \citeasnoun{Gardes10}, \citeasnoun{Gardes12}, \citeasnoun{Daouia13}, and \citeasnoun
{Martins-Filho18}.

In this article, we focus on statistical inference rather than estimation,
and provide confidence intervals (CIs) of a conditional extremal quantile
that have preferred coverage and length properties. Our proposed method
applies to both repeated cross-sectional data and panel data. The main idea is
very intuitive. Consider the case of using panel data of $(Y,X)$ to fix
ideas, and suppose that one is interested in the conditional extremal
quantile of $Y $ given $X=x_{0}$, denoted by $Q_{Y|X=x_{0}}(\tau)$. If, for
every individual, there exists some time period in which $X$ takes the value $x_{0}$,
then we can simply collect the associated $Y$'s and form a cross-sectional
sample from $F_{Y|X=x_{0}}$. Since this is infeasible especially when $X$ is
continuous, we instead collect from each individual's time series the
induced $Y$ associated with $X$ that is the nearest neighbor (NN) of $x_{0}$
. These induced $Y$'s are now \textit{approximately} stemming from $
F_{Y|X=x_{0}}$, and the large (respectively, small) order statistics from
them can be used for inference about the upper (respectively, lower)
conditional extremal quantile $Q_{Y|X=x_{0}}(\tau)$. For multi-dimensional
covariates, this is done by defining the NN measured by a certain choice of
metric, such as the one induced by the Euclidean norm. If a linear
regression model is appropriate, then the NN can also be defined using the
linear index.

The above approximation approach is formalized by establishing a new extreme
value (EV) theory. The theory is based on the large-$n$ and large-$T$
asymptotics, where $n$ and $T$ denote the sample sizes in cross-sectional
and time-series dimensions, respectively. A large $T$ guarantees that the NN
is close enough to the query point $x_{0}$, while a large $n$ provides
enough observations from a more accurate tail sample. Given the new EV
theory, we apply it to construct new confidence intervals for the conditional extremal
quantiles.

Our proposed approach only requires some smoothness condition on the joint
distribution and hence enjoys more robustness against functional form
specification than existing methods. A natural question is how much
efficiency we lose by using only one out of $T$ observations in each time
series. It turns out that if the tail shape depends on the covariate highly
nonlinearly,\footnote{
See Section \ref{sec MC} for concrete numerical settings which this
qualitative phrase stands for.} then our proposed NN method dominates
existing methods in both coverage and length when $T$ is only moderately
large, say 50. When $T$ is very large, say 500, the new CIs also deliver
comparable lengths to the kernel regression method with the optimal
bandwidth -- see the Monte Carlo results ahead in Section \ref{sec MC} for
more details.

As a by-product of our main result, we also develop CIs for extremal quantiles of the coefficients in a random coefficient regression model.
In particular, suppose that $Y_{it}$ and $X_{it}$ are generated from the model $Y_{it}=\alpha _{i}+X_{it}^{\intercal }\beta _{i}+u_{it}$, where $(\alpha_{i},\beta _{i}^{\intercal })^{\intercal}$ is a random vector drawn from some unknown distribution. We first construct the least squares estimators of $\alpha_{i} $ and $\beta _{i}$ using the $i$-th time series for all $i$ and collect the largest (smallest) order statistics from these estimates.
We then show that the estimation error is negligible under the large $n$ and large $T$
framework, and hence the largest (smallest) order statistics among these estimates again satisfy the desired EV theory, which further supports the application of the fixed-$k$ CIs for extremal quantiles of $\alpha _{i}$ and $\beta _{i}$.
This complements the existing literature focusing on the mid-sample properties of heterogeneous effects (e.g., \citeasnoun{Hsiao04} and  \citeasnoun{Wooldridge05}).

Applying the proposed methods, we study the tail risk of extremely low birth weight conditional on mothers' behavioral and demographics characteristics.
We find that signs of major effects are the same as those found in preceding studies based on parametric models.
On the other hand, we find that some effects exhibit different magnitudes from those reported in the previous studies based on parametric models.

The rest of the paper is organized as follows.
Section \ref{sec main} presents the main results of this paper.
Section \ref{sec MC} presents Monte Carlo simulation studies.
Section \ref{sec:application} presents an empirical application.
Section \ref{sec conclusion} concludes the paper.
All mathematical proofs and additional details are found in the appendix.

\textbf{Notation} Let $\overset{p}{\rightarrow }$ denote convergence in
probability and $\overset{d}{\rightarrow }$ denote convergence in
distribution as $n,T\rightarrow \infty $. Let $\mathbf{1}[A]$ denote the
indicator function of a generic event $A$. Let $\left\vert \left\vert
B\right\vert \right\vert $ denote the Euclidean norm of a vector or matrix $
B $, and let $C$ denote a generic constant whose value may change across
lines. Let $B_{\delta }(x)$ denote a generic open ball centered at $x$ with
radius $\delta $. When $X$ denotes a column vector and $c$ a scalar, the
notation $X-c$ is understood as the vector $X-(c,\ldots ,c)^{\intercal }$.

\section{Main result\label{sec main}}

We present the main result of this paper in this section. Let $X$ denote a $
\dim (X)\times 1$ vector of continuous random variables with uniformly
positive joint PDF.\footnote{
While we focus on continuous random variables in our presentation, the
method can also accommodate discrete random variables. Suppose that the
covariate vector is written as $(X^{\prime },W^{\prime })^{\prime }$ where
the subvector $X$ consists of continuous random variables and the subvector $
W$ consists of discrete random variables. Suppose that one is interested in
the conditional extremal quantiles given $(X^{\prime },W^{\prime })^{\prime
}= (x_0^{\prime },w_0^{\prime })^{\prime }$. We can then extract the
subsample with $W=w_0$, and then apply our proposed method for the subsample.
} The main object of interest is the conditional extremal quantile $
Q_{Y|X=x_{0}}\left( \tau \right) $ of $Y$ given $X = x_{0}$ for a
pre-specified $x_{0}\in \mathbb{R} ^{\dim \left( X\right) }$ and $\tau
\rightarrow 1$. For ease of exposition, we consider a balanced\footnote{
This is only for notational ease. The new approach is valid as long as $T$
is large for all $i$.} repeated cross-sectional or panel data set $\{Y_{it},X_{it}\}_{i=1:n,t=1:T}$ that
is i.i.d. across $i$ and strictly stationary and weakly dependent across $t$
. Section \ref{sec general} presents an informal overview of our proposed
method, and Section \ref{sec condition and result} gives a formal
theoretical justification.
Finally, Section \ref{sec ext} presents an extension of the main theoretical results to linear random coefficient models.

\subsection{Overview\label{sec general}}

Our method consists of the following three steps. In the first step, we
make use of the repeated cross-sectional or panel data structure by selecting a subsample induced by the distances of the covariates $\{X_{it}\}$ to the query point $x_{0}$.
The subsample will be a $k\times 1$ random vector denoted as $\mathbf{Y}$. In
the second step, by appealing to the extreme value theory, we show that
after some normalization, $\mathbf{Y}$ converges in distribution to a
well-defined limiting random variable, $\mathbf{V}$, whose distribution $f_{
\mathbf{V}}$ is parametric and uniquely determined by the tail features of $
F_{Y|X=x_{0}}$. In particular, $f_{\mathbf{V}}$ will be uniquely
characterized by a scalar parameter $\xi $ that fully captures the tail
heaviness of $F_{Y|X=x_{0}}$. Note that $\xi $ depends on the query point $
x_{0}$, which will be suppressed in our notations for simplicity when there
is no confusion. Since $Q_{Y|X=x_{0}}\left( \tau \right) $ can also be
uniquely expressed as a function of $\xi $ after suitably normalizing $\tau $
, the asymptotic problem becomes conceptually straightforward: constructing
inference for a function of $\xi (x_{0})$ given a random draw $\mathbf{V}$.
This type of problems is studied by \citeasnoun{EMW15}, who provide a generic
argument to construct optimal inference when there exists a nuisance
parameter under a null hypothesis. In the third and final step, we tailor
their arguments to inference about $Q_{Y|X=x_{0}}\left( \tau \right) $ with $
\xi $ being the nuisance parameter.

The next three subsubsections introduce details of these three steps in order. Section \ref{sec condition and result} then follows up
by presenting regularity conditions and the main theoretical result of the
paper that guarantees that our confidence interval constructed in the
three-step procedure controls coverage asymptotically and uniformly over a set of data generating processes.

\subsubsection{Step 1: subsample selection based on NN}

First, we select our subsample $\mathbf{Y}$ as follows.

\begin{itemize}
\item Collect, for each $i$, the induced $Y$ associated with the NN of $
\{X_{it}\}_{t=1}^{T}$ to $x_{0}$, where the NN is measured by the Euclidean
distance $\left\vert \left\vert X_{it}-x_{0}\right\vert \right\vert$. Denote
them by $\{Y_{i,[x_{0}]}\}_{i=1}^{n}$.\footnote{
Details of this step with more notations are as follows. For each $i \in
\{1,...,n\}$ and $t \in \{1,...,T\}$, compute $d_{it} = \|X_{it}-x_0\|$
where $\|\cdot\|$ denotes the Euclidean distance.  Then, for each $i \in
\{1,...,n\}$, let $t_i^\ast$ denote the argument $t$ that minimizes $d_{it}$
. We denote $Y_{i,[x_0]} = Y_{it_i^\ast}$.}

\item Take the largest $k$ order statistics from $\{Y_{i,[x_0]}\}_{i=1}^{n}$
and denote the vector of them by
\begin{equation}
\mathbf{Y}=\mathbf{(}Y_{(1),[x_{0}]},Y_{(2),[x_{0}]},...,Y_{(k),[x_{0}]})^{
\intercal },  \label{Y_k}
\end{equation}
where $Y_{(1),[x_{0}]}\geq Y_{(2),[x_{0}]}\geq \ldots \geq Y_{(n),[x_{0}]}$
are the order statistics of $\{Y_{i,[x_{0}]}\}_{i=1}^{n}$.
\end{itemize}

The key idea for such a selection is heuristically illustrated by the
following derivation - a formal argument is presented as a proof of Theorem \ref{thm
main} in Appendix \ref{sec:proof}. For each $i$, denote the NN among $
\{X_{it}\}_{t=1}^{T} $ to $x_{0}$ as $X_{i,(x_{0})}$. Then for any $y\in
\mathbb{R} $,
\begin{align*}
&\mathbb{P}\left( Y_{i,[x_{0}]}\leq y\right) \\
=&\mathbb{E}_{X_{i,\left( x_{0}\right) }}\left[ \mathbb{P}\left(
Y_{i,[x_{0}]}\leq y|X_{i,\left( x_{0}\right) }\right) \right] \\
=&\mathbb{E}_{X_{i,\left( x_{0}\right) }}\left[ F_{Y|X=X_{i,\left(
x_{0}\right) }}\left( y\right) \right] &&\text{ (by strict stationarity)} \\
=&F_{Y|X=x_{0}}\left( y\right) +\mathbb{E}_{X_{i,\left( x_{0}\right) }}
\left[ \left. \frac{\partial F_{Y|X=x}\left( y\right) }{\partial
x^{\intercal }}\right\vert _{x=\dot{x}_{i}}(X_{i,\left( x_{0}\right) }-x_{0})
\right] &&\text{ (by mean value expansion)} \\
\rightarrow &F_{Y|X=x_{0}}\left( y\right) &&\text{ as }T\rightarrow \infty ,
\end{align*}
where $\dot{x}_{i}$ lies between $X_{i,\left( x_{0}\right)}$ and $x_{0}$.
The first equality is by the definition of conditional expectation. The
second one follows under the strict stationarity. The third equality is
valid if the conditional CDF is smooth. The last convergence holds if the NN
converges to its query point $x_{0}$ and if the CDF is smooth with
bounded derivatives.

The above derivation states that the collection of the induced order
statistics $Y$ associated with the NN of $x_{0}$ can be treated as
approximately stemming from the true conditional CDF $F_{Y|X=x_{0}}$
asymptotically. Thus the largest (cross-sectional) order statistics $\mathbf{
Y}$ can be treated as draws from the tail of $F_{Y|X=x_{0}}$.

\subsubsection{Step 2: asymptotic distribution of the subsample}

To proceed with the second step, we need some regularity conditions about $
F_{Y|X=x_{0}}$. For readability, we introduce only one of the conditions
here with the remaining of them discussed in Section \ref{sec condition and
result}. Specifically, we assume that $F_{Y|X=x_{0}}$ is within the domain
of attraction (DOA) of the extreme value distribution (denoted by $
F_{Y|X=x_{0}}\in \mathcal{D}(G_{\xi })$), in the sense that there exist
sequences of constants $a_{n}$ and $b_{n}$ such that for every $v$,
\begin{equation*}
\lim_{n \rightarrow \infty}F_{Y|X=x_{0}}(a_{n}v+b_{n})= G_{\xi }(v),
\end{equation*}
where
\begin{equation}
G_{\xi }(v)=\left\{
\begin{array}{ll}
\exp (-(1+\xi v)^{-1/\xi })\text{, } & 1+\xi v>0\text{, for }\xi \neq 0 \\
\exp (-e^{-v})\text{, } & v\in \mathbb{R} \text{, }\xi =0.
\end{array}
\right.  \label{def_G}
\end{equation}
This DOA condition is extensively studied in the statistics literature and
is satisfied by many commonly used distributions, including, for example,
Pareto, Student-t, F, Gaussian, and even uniform distributions. See Chapter
1 in \citeasnoun{deHaan07} for a complete review.

Under the DOA assumption and the cross-sectional i.i.d. assumption, we show
that for any fixed $k$,
\begin{equation}
\frac{\mathbf{Y-}b_{n}}{a_{n}}\overset{d}{\rightarrow }\mathbf{V}=\left(
\begin{array}{c}
V_{1} \\
\vdots \\
V_{k}
\end{array}
\right) \text{, as }n,T\rightarrow \infty ,  \label{evt_k}
\end{equation}
where the joint probability density function (PDF)~of $\mathbf{V}$ is given
by
\begin{equation}
f_{\mathbf{V}}(v_{1},\ldots ,v_{k};\xi )=G_{\xi
}(v_{k})\prod_{i=1}^{k}g_{\xi }(v_{i})/G_{\xi }(v_{i})  \label{evt_pdf}
\end{equation}
for $v_{k}\leq v_{k-1}\leq \ldots \leq v_{1}$ with $g_{\xi }(v)=\partial
G_{\xi }(v)/\partial v$, and zero otherwise.

Note that the constants $a_{n}$ and $b_{n}$ depend on $\xi $, and their
estimates thus tend to exhibit large magnitudes of sensitivity to tail
observations. For example, $a_{n}$ is $n^{\xi }$ if $F_{Y}$ is standard
Pareto. Since a small estimation error in $\xi $ is amplified by the $n$
-power, inference relying on a good estimate of $\xi $ and the scale usually
requires a large $k$ and a even larger sample size $n$. Besides, the $G_{\xi
}(v_{k})$ term in (\ref{evt_pdf}) suggests that the largest $k$ order
statistics are not asymptotically independent, given any fixed $k$.
\footnote{We also derived the estimation and inference method based on the increasing-$k$ asymptotics in a previous version of this article. Their performance is dominated by the fixed-$k$ approach, especially in case with only moderate sample sizes. Therefore, we present the fixed-$k$ result exclusively in this version for illustrational simplicity.}

\subsubsection{Step 3: construction of the asymptotic inference}

We aim for a $(1-\alpha)$ CI for the conditional extremal quantile $
Q_{Y|X=x_{0}}(\tau )$ for $\tau $ close to 1. Specifically, we rewrite $\tau$
as $1-h/n$ for some $h>0$ following \citeasnoun{Chernozhukov05} and \citeasnoun
{Chernozhukov11}. This setup means that the extremal quantile is of the same
order of the sample maximum from $n$ random draws from the true conditional
CDF $F_{Y|X=x_{0}}$.

Our objective is to construct a confidence set $S(\mathbf{Y})\subset \mathbb{
R} $ such that $\mathbb{P}(Q_{Y|X=x_{0}}(\tau )\in S(\mathbf{Y}))\geq
1-\alpha + o(1)$, as $n\rightarrow \infty $ and $T\rightarrow \infty $.
Under the DOA assumption, calculations show that
\begin{equation*}
\frac{Q_{Y|X=x_{0}}(1-h/n)-b_{n}}{a_{n}}\rightarrow q(\xi ,h)\equiv \left\{
\begin{array}{ll}
\frac{h^{-\xi }-1}{\xi } & \text{if }\xi \neq 0 \\
-\log (h) & \text{if }\xi =0.
\end{array}
\right.
\end{equation*}
Note that $q(\xi ,h)$ is the $\exp (h)$ quantile of $V_{1}$. Since it is
shared by both $\mathbf{Y}$ and $Q_{Y|X=x_{0}}(1-h/n)$, we can impose
location and scale equivariance on the CI to cancel them out. Specifically,
we impose that for any constants $a>0$ and $b$, $S(a\mathbf{Y}+b)=aS(\mathbf{
Y})+b$, where $aS(\mathbf{Y})+b=\{y:(y-b)/a\in S(\mathbf{Y})\}$. Under this
equivariance constraint, we can write
\begin{eqnarray*}
&&\mathbb{P}(Q_{Y|X=x_{0}}(1-h/n)\left. \in \right. S(\mathbf{Y})) \\
&=&\mathbb{P}\left( \frac{Q_{Y|X=x_{0}}(1-h/n)-Y_{(k),[x_{0}]}}{
Y_{(1),[x_{0}]}-Y_{(k),[x_{0}]}}\in S\left( \frac{\mathbf{Y}-Y_{(k),[x_{0}]}
}{Y_{(1),[x_{0}]}-Y_{(k),[x_{0}]}}\right) \right) \\
&\rightarrow &\mathbb{P}_{\xi }\left( V^{q}\in S(\mathbf{V}^{\ast })\right),
\end{eqnarray*}
where we introduce the self-normalized statistics
\begin{eqnarray*}
V^{q} &=&\frac{q(\xi ,h)-V_{k}}{V_{1}-V_{k}} \\
\mathbf{V}^{\ast } &\mathbf{=}&\left( \frac{V_{1}-V_{k}}{V_{1}-V_{k}},\frac{
V_{2}-V_{k}}{V_{1}-V_{k}},...,\frac{V_{k}-V_{k}}{V_{1}-V_{k}}\right) ,
\end{eqnarray*}
and highlight with the subscript $\xi $ that the densities of $V^{q}$ and $
\mathbf{V}^{\ast }$ now depend solely on $\xi $. These can be computed by
using (\ref{evt_k}), (\ref{evt_pdf}), and a change of variables.

Since $\xi $ is unknown, we impose the size constraint uniformly for all the
values of $\xi $ that are empirically relevant. In this sense the fixed-$k$
approach is more robust against misspecification, especially when the sample
size is not large enough to support a precise estimation of $\xi$. Let $\Xi
\subset\mathbb{R}$ be the set of tail indices for which we impose the
asymptotically correct coverage.\footnote{
We use\ $\Xi =[-1/2,1/2]$ for inference about conditional extremal quantiles
in later applications, which covers all the distributions with finite
variance. This range can be easily extended.} The asymptotic problem then is
to construct a location and scale equivariant $S$ that satisfies
\begin{equation}
\mathbb{P}_{\xi }\left( V^{q}\in S(\mathbf{V}^{\ast })\right) \geq 1-\alpha
\text{ for all }\xi \in \Xi ,  \label{asy_cov}
\end{equation}
since any $S$ that satisfies (\ref{asy_cov}) also satisfies $\lim
\inf_{n\rightarrow \infty ,T\rightarrow \infty }\mathbb{P}
(Q_{Y|X=x_{0}}(1-h/n)\in S(\mathbf{Y}))\left. \geq \right. 1\left.
-\right.\alpha $ by the continuous mapping theorem. Among all the solutions
to this problem, we choose the optimal one that minimizes the weighted
average expected length criterion
\begin{equation}
\int \mathbb{E}_{\xi }[\text{lgth}(S(\mathbf{V}))]dW(\xi )\text{,}
\label{asy_length}
\end{equation}
where $W$ is a positive measure with support on $\Xi $,\footnote{
We use the uniform weight in later sections.} and $\text{lgth}(A)=\int
\mathbf{1}[y\in A]dy$ for any Borel set $A\subset \mathbb{R}$. The
equivariance of $S$ further implies $\mathbb{E}_{\xi }[\text{lgth}(S(\mathbf{
V}))]\left. =\right. \mathbb{E}_{\xi }[(V_{1}-V_{k})\text{lgth}(S(\mathbf{V}
^{\ast }))]$. Thus the program of minimizing (\ref{asy_length}) subject to (
\ref{asy_cov}) among all equivariant sets $S$ asymptotically becomes
\begin{equation}
\begin{tabular}{l}
$\min_{S(\cdot )}\int_{\Xi }\mathbb{E}_{\xi }[(V_{1}-V_{k})\text{lgth}(S(
\mathbf{V}^{\ast }))]dW(\xi )$ \\
\multicolumn{1}{r}{$\text{s.t. }\mathbb{P}_{\xi }(V^{q}\in S(\mathbf{V}
^{\ast }))\geq 1-\alpha \text{ for all }\xi \in \Xi ,$}
\end{tabular}
\label{program}
\end{equation}
where we abuse the notation of $\mathbb{E}_{\xi }$ and $\mathbb{P}_{\xi }$
to emphasize that the distributions of $V^{q}$ and $\mathbf{V}^{\ast }$
depend on $\xi $ (and further on $x_{0}$). Note that any solution to (\ref
{program}) also provides the form of $S$, that is, $S(\mathbf{V}
)=(V_{1}-V_{k})S(\mathbf{V}^{\ast })+V_{k}$. Once $S(\cdot )$ is determined,
therefore, the confidence interval can be constructed in practice by
plugging in
\begin{equation*}
(Y_{(1),[x_{0}]}-Y_{(k),[x_{0}]})S\left( \frac{\mathbf{Y}-Y_{(k),[x_{0}]}
}{Y_{(1),[x_{0}]}-Y_{(k),[x_{0}]}}\right) +Y_{(k),[x_{0}]}.
\end{equation*}

In solving (\ref{program}), we write the problem in the following Lagrangian
form:
\begin{equation*}
\min_{S(\cdot )}\int_{\Xi }\mathbb{E}_{\xi }[(V_{1}-V_{k})\text{lgth}(S(
\mathbf{V}^{\ast }))]dW(\xi )+\int_{\Xi }\mathbb{P}_{\xi }\left( V^{q}\in S(
\mathbf{V}^{\ast })\right) d\Lambda (\xi ),
\end{equation*}
where the non-negative measure $\Lambda $ denotes the Lagrangian weights
that guarantee the asymptotic coverage constraint. By defining $\kappa (\mathbf{V
}^{\ast };\xi )=\mathbb{E}_{\xi }[V_{1}-V_{k}|\mathbf{V}^{\ast }]$ and
writing the expectations above as integrals over the densities $f_{\mathbf{V}
^{\ast }}$ and $f_{V^{q},\mathbf{V}^{\ast }}$ of $\mathbf{V}^{\ast }$ and $
(V^{q},\mathbf{V}^{\ast })$, respectively, the solution of the above problem is given by
\begin{equation}
S(\mathbf{v}^{\ast })=\left\{ y:\int_{\Xi }\kappa (\mathbf{v}^{\ast };\xi
)f_{\mathbf{V}^{\ast }}(\mathbf{v}^{\ast };\xi )dW(\xi )<\int_{\Xi }f_{V^{q},
\mathbf{V}^{\ast }}(y,\mathbf{v}^{\ast };\xi )d\Lambda (\xi )\right\} .
\label{S_Lambda}
\end{equation}
The integrals can be numerically computed by Gaussian quadrature. To find
suitable Lagrangian weights $\Lambda $, we appeal to the generic algorithm
in \citeasnoun{EMW15}, who provide a numerical method to construct $\Lambda $. We
tailor their arguments to our conditional extreme tail inference problem and
provide the corresponding MATLAB program on the author's website. The
computation cost is only several seconds using a modern PC. Note that $
\Lambda $ only needs to constructed once by the author but not the empirical
users. See Section \ref{sec computation} for more details.

Other tail-related quantities, such as the conditional tail expectations,
are also covered by our proposed method as long as they can be expressed as
functions of the conditional tail index. We discuss such an extension in
Section \ref{sec ext}. In the following subsection, we formally introduce all the
regularity conditions and formalize the uniform coverage property of the
confidence interval (\ref{S_Lambda}).

\subsection{Conditions and main theoretical results\label{sec condition and
result}}

Our asymptotic theory requires the following four conditions.

\begin{description}
\item[Condition 1.1] $(Y_{i1},X_{i1}^{\intercal })^{\intercal
},\ldots,(Y_{iT},X_{iT}^{\intercal })^{\intercal }$ are i.i.d.\ across $i$. $
(Y_{it},X_{it}^{\intercal })^{\intercal }$ for each $t=1,\ldots ,T$ is
strictly stationary and $\beta $-mixing with the mixing coefficient
satisfying $\beta \left( t\right) =O(t^{-2-\varepsilon })$ for some $
\varepsilon >0$. In addition, $f_{X}(x)$ is uniformly continuously
differentiable and bounded away from $0$ in an open ball centered at $x_{0}$.
\end{description}

Condition 1.1 requires the data to be independent across $i$ and weakly
dependent across $t$, which is plausibly satisfied by the Natality Vital
Statistics that we use for our empirical application. In addition, this condition also
requires the density of $X$ to be positive in an open neighborhood around
the query point $x_{0}$. This condition is sufficient to establish that the
NN converges to the query point $x_{0}$ almost surely at some power rate. To
the best of our knowledge, this is the first result about the (almost sure
and L$^{2}$) convergence rate of the NN under weak dependence, whose proof
is non-trivial. We formalize this result as Lemma \ref{lemma NN} in Appendix
A.1, which might be of independent research interest.\footnote{
The $\beta $-mixing condition allows the application of Berbee's lemma (\citeasnoun
{Berbee87}) in establishing Lemma \ref{lemma NN}. This is assumed to avoid
technical complexity and can be relaxed to other forms of weak dependence.}
Note that we intentionally choose only one NN to allow for weak dependence
across $t$. If data are independent across both $i$ and $t$, then more than
one NNs can be chosen to enlarge the effective sample. We leave this for future research.

\begin{description}
\item[Condition 1.2] $F_{Y|X=x_{0}}\in \mathcal{D}\left( G_{\xi \left(
x_{0}\right) }\right) $ with $\xi (x_{0})\in \Xi \,$, a compact subset of $
\mathbb{R}
$.
\end{description}

This condition requires that the underlying conditional distribution is in
the domain of attraction of the generalized EV distribution. This is a mild
condition as it is satisfied by many commonly used joint distributions. In
particular, it generalizes the conditional location-scale shift model (\ref
{EQR cond}) by allowing $\mu (x),\sigma (x),$ and $\xi (x)$ to be all
unknown (but smooth) functions of $x$. The case of negative $\xi (x_{0})$ is
included only for comprehensiveness, since the $Y$ in most applications involving
tail features has an unbounded support that entails a non-negative $\xi(x_{0})$. To illustrate the mildness of this
condition, we discuss the following three examples. Our Condition 1.2 is
satisfied in all three of them, but the location-scale model assumption (\ref
{EQR cond}) is not.

\begin{description}
\item[Example 1 (Joint Normal)] Suppose that $(Y,X)$ is jointly normal with
zero means, unit variances, and correlation $\rho $. Then the conditional
distribution of $Y$ given $X=x$ is normal with mean $\rho x$, and variance $
1-\rho ^{2}$. The conditional tail index is $\xi (x)=0$ for all $x\in
\mathbb{R}$. The conditional quantile is $Q_{Y|X=x}\left( \tau \right) =\rho
x+\sqrt{1-\rho ^{2}}\Phi ^{-1}\left( \tau \right) $, where $\Phi ^{-1}\left(
\cdot \right) $ is the quantile function of the standard normal
distribution. Thus, the location-scale model assumption (\ref{EQR cond}) is
satisfied.

\item[Example 2 (Joint Student-t)] Suppose that $(Y,X)$ is jointly Student-t
distributed with d.f.$\ v$, zero means, unit variances, and correlation $
\rho \neq 0$. Then the conditional distribution of $Y$ given $X=x$ is
Student-t distributed with d.f.\ $v+1$, mean $\rho x$, and variance $(1-\rho
^{2})(v+x^{2})/(v+1)$. The conditional tail index is $\xi (x)=1/(v+1)$ for
all $x\in \mathbb{R}$.\footnote{
See \citeasnoun{Ding16} for the exact expression for the PDF.} The conditional
quantile is $Q_{Y|X=x}\left( \tau \right) =\rho x+\sqrt{(1-\rho
^{2})(v+x^{2})/(v+1)}Q_{t(v)}(\tau )$, where $Q_{t(v)}(\cdot )$ is the
quantile function of the standard Student-t distribution with d.f. $v$. This
specification satisfies the location-scale shift model (\ref{EQR cond}) but
the scale function is highly nonlinear in $x$.

\item[Example 3 (Conditional Pareto)] Suppose that $X$ is half-normal with
positive support and $Y$ given $X=x$ is the Pareto distribution such that $
\mathbb{P}(Y\leq y|X=x)=1-(y+1)^{-1/x}$ for $y\geq 0$ and any $x>0$. Then
the conditional tail index is $\xi (x)=x$ and the conditional quantile is $
Q_{Y|X=x}\left( \tau \right) =-1+(1-\tau )^{-x}$, which violates the
location-scale shift model (\ref{EQR cond}).
\end{description}

Let $y_{0}$ denote the end-point of the conditional CDF, that is, $
y_{0}=Q_{Y|X=x_{0}}\left( 1\right) \leq \infty $. The next condition is a
high level regularity assumption on the smoothness of the conditional tail.

\begin{description}
\item[Condition 1.3] $f_{Y|X=x}(y)$ is uniformly bounded and continuously
differentiable in $x$ and $y$. In addition, for any fixed $y\left. >\right.
0 $ with $u_{n}\left. =\right. a_{n}y\left. +\right. b_{n}\left. \rightarrow
\right. y_{0}$, and any open ball $B_{\eta _{T}}\left( x_{0}\right) $
centered at $x_{0}$ with radius $\eta _{T}\left. \equiv \right. O(T^{-\eta
}) $ for some $\eta \left. >\right. 0$, $\lim_{u_{n}\rightarrow y_{0}}\left.
\sup_{x\in B_{\eta _{T}}\left( x_{0}\right) }\right. \left. T^{-\eta
}\left\vert \left\vert \frac{\partial F_{Y|X=x}\left( u_{n}\right) /\partial
x}{1-F_{Y|X=x_{0}}\left( u_{n}\right) }\right\vert \right\vert \right.
\left. =\right. 0$ and $\lim_{u_{n}\rightarrow y_{0}}\left. \sup_{x\in
B_{\eta _{T}}\left( x_{0}\right) }\right. \left. T^{-\eta }\left\vert
\left\vert \frac{\partial f_{Y|X=x}\left( u_{n}\right) /\partial x}{
f_{Y|X=x_{0}}\left( u_{n}\right) }\right\vert \right\vert \right. \left.
=\right. 0$ as $n\left. \rightarrow \right. \infty $ and $T\left.
\rightarrow \right. \infty .$
\end{description}

Condition 1.3 requires that the derivatives of the conditional CDF and PDF
are smooth and decay quickly. This is a mild condition again, which is
satisfied by the above examples by straightforward calculation. For
readability, we provide low-level primitive assumptions as sufficient
conditions for Condition 1.3 and discuss them in Appendix \ref{sec:primitive_conditions}.

\begin{description}
\item[Condition 1.4] $n\rightarrow \infty $, $T\rightarrow \infty $, and $
T/n\rightarrow \lambda $ for some $\lambda \in (0,\infty )$.
\end{description}

Condition 1.4 requires both $n$ and $T$ to be large. A large $n$ guarantees
that the error due to the EV approximation is negligible, and a large $T$
controls the distance between the NN and the query point. The parameter $
\lambda $ can be any constant in the open unit interval, and hence $T$ can
be much smaller than $n$.

Under the above conditions, we establish the asymptotically correct uniform coverage
by confidence interval (\ref{S_Lambda}) in the following theorem, which is
the main result of this paper.

\begin{theorem}
\label{thm main}Suppose that Conditions 1.1-1.4 hold. For any fixed $k$ and
any $F_{Y|X=x_{0}}$ that satisfies these conditions,
\begin{equation*}
\left. \underset{n,T\rightarrow \infty }{\lim \inf }\right. \mathbb{P}
\left(Q_{Y|X=x_{0}}(1-h/n)\in S(\mathbf{Y})\right) \geq 1-\alpha ,
\end{equation*}
where $S\left( \cdot \right) $ is determined in (\ref{S_Lambda}).
\end{theorem}

We conclude this subsection with a discussion of main properties, advantages
and disadvantages of the our proposed fixed-$k$ confidence interval. First,
since the confidence interval is based on a fixed number $k$ of tail
observations, its length does not decrease in $n$. On the other hand, the
length decreases in $k$. Second, unlike kernel regression approaches that
require a sequence of moving tuning parameter (i.e., the bandwidth parameter tending
to zero), our method only relies on a `fixed' tuning parameter which is $k$.
While common data-driven choice rules for bandwidths are not theoretically
compatible with inference for their failure to undersmooth estimates, our
method based on any fixed $k$ of a researcher's choice guarantees asymptotically valid inference. Third, our fixed-$k$ approach allows for the confidence
interval to have a uniform size control property over a set of data generating processes involving a set of values of the tail index, while the existing methods have not been shown to share this uniformity property.
This property of our fixed-$k$ approach is useful because $\xi$ is practically unknown to researchers
and thus the size control should be uniform for all the values of $\xi$ that
are empirically relevant.

\subsection{Extension to linear random coefficients models
\label{sec ext}}

The proof strategy for our main result, namely Theorem \ref{thm main}, and
thus our proposed method of constructing confidence intervals apply to
other contexts. Among others, inference for extremal quantiles of
random coefficients in linear regression models is also possible with
our proposed strategy. In this section, we study this class of models which
have been widely used in empirical studies in economics.

Consider the model
\begin{equation}
Y_{it}=\alpha _{i}+X_{it}^{\intercal }\beta _{i}+u_{it},  \label{reg}
\end{equation}
where $(\alpha _{i},\beta _{i}^{\intercal })^{\intercal }$ denotes random
coefficients and $u_{it}$ denotes an error term. This setup has been studied
by numerous papers in the literature, and covers the classic panel linear
regression model with fixed effects in which $\beta _{i}=\beta_{0}$ for all $
i$.
As long as Conditions 1.1-1.4 are satisfied, the previously introduced
methods naturally apply here for inference on the conditional extremal
quantiles of $Y$. In addition, the model (\ref{reg}) allows
us to conduct inference on the unconditional tail features of the random coefficients, $\alpha _{i}$
and $\beta _{i}$. The remainder of this subsection illustrates a procedure
to this end.

Let $(\hat{\alpha}_{i},\hat{\beta}_{i}^{\intercal })^{\intercal }$ be the
OLS estimator by regressing $Y_{it}$ on $(1,X_{it}^{\intercal })^{\intercal}$
using the time series associated with the $i$-th individual. Collect $\{(
\hat{\alpha}_{i},\hat{\beta}_{i}^{\intercal })^{\intercal }\}_{i=1}^{n}$ and
sort each series of estimates in the descending order. We then define
\begin{equation*}
\mathbf{A}=(\hat{\alpha}_{\left( 1\right) },...,\hat{\alpha}_{\left(k\right)
})^{\intercal },
\end{equation*}
that is, the largest $k$ order statistics of $\{\hat{\alpha}_{i}\}$, and
\begin{equation*}
\mathbf{B}_{j}=(\hat{\beta}_{j,\left( 1\right) },...,\hat{\beta}
_{j,\left(k\right) })^{\intercal },
\end{equation*}
that is, the largest $k$ order statistics of the $j$-th coordinate of $\{
\hat{\beta}_{i}\}_{i=1}^{n}$, for each $j$. Without loss of generality, we focus on the
first coordinate of $\beta _{i}$, and suppress the subscript $j$ from our
notations for simplicity.

Now, we substitute $\mathbf{A}$ or $\mathbf{B}$ in $S(\cdot )$ as in (\ref
{S_Lambda}) to construct the confidence interval for extremal quantiles of $
\alpha _{i}$ or $\beta _{i}$. The following conditions are imposed for a
theoretical guarantee of correct asymptotic coverage.

\begin{description}
\item[Condition 2.1] $(\alpha _{i},\beta _{i}^{\intercal
},u_{it},X_{it}^{\intercal })^{\intercal }$ are i.i.d.\ across $i$ and
strictly stationary and weakly dependent across $t$;

\item[Condition 2.2] $F_{\alpha }\in \mathcal{D}\left( G_{\xi _{\alpha
}}\right) $ and $F_{\beta }\in \mathcal{D}\left( G_{\xi _{\beta }}\right) $
with $\xi _{\alpha } \in \Xi$ and $\xi _{\beta } \in \Xi$;

\item[Condition 2.3] $\sup_{i}\left\vert
\left\vert (\hat{\alpha}_i,\hat{\beta}_{i}^\intercal)^\intercal-(\alpha_i,\beta _{i}^\intercal)^\intercal\right\vert \right\vert=o_{p}(1)$, $
\sup_{i}\left\vert \bar{u}_{i}\right\vert =o_{p}(1)$, and $
\sup_{i}\left\vert \left\vert \bar{X}_{i}\right\vert \right\vert =O_{p}(1)$,
where $\bar{X}_{i}=T^{-1}\sum_{t=1}^{T}X_{it}$ and $\bar{u}
_{i}=T^{-1}\sum_{t=1}^{T}u_{it}$. In addition, if $\xi _{\omega }=0$, $
\sup_{i}\left\vert
\left\vert (\hat{\alpha}_i,\hat{\beta}_{i}^\intercal)^\intercal-(\alpha_i,\beta _{i}^\intercal)^\intercal\right\vert \right\vert/f_{\omega }\left( Q_{\omega
}(1-1/n)\right) \left. =\right. o_{p}(1)$ and $\sup_{i}\left\vert \bar{u}
_{i}\right\vert /f_{\omega }\left( Q_{\omega }(1-1/n)\right) \left. =\right.
o_{p}(1)$ for $\omega =\alpha $ or $\beta $, where $Q_{\omega }(\cdot )$ and
$f_{\omega }(\cdot )$ denote the quantile function and the PDF of $\omega $,
respectively. Alternatively, if $\xi _{\omega }<0$, $n^{-\xi _{\omega
}}T^{-1/2}\rightarrow 0$, $\sup_{i}\left\vert \left\vert \bar{X}
_{i}\right\vert \right\vert =O_{p}\left( 1\right) $, $\sup_{i}\left\vert
\left\vert (\hat{\alpha}_i,\hat{\beta}_{i}^\intercal)^\intercal-(\alpha_i,\beta _{i}^\intercal)^\intercal\right\vert \right\vert =O_{p}\left(
T^{-1/2}\right) $ and $\sup_{i}\left\vert \bar{u}_{i}\right\vert
=O_{p}\left( T^{-1/2}\right) $.
\end{description}

Condition 2.1 is similar to Condition 1.1. Since the objects of interest are
the unconditional extremal quantiles of $\alpha _{i}$ and $\beta _{i}$, we
do not need the NN condition on covariates. The dependence structure is left
unspecified as long as it is sufficient for Condition 2.3. Condition 2.2
assumes that the distributions of $\alpha _{i}$ and $\beta _{i}$ are in the
domains of attraction of $G_{\xi _{\alpha }}$ and $G_{\xi _{\beta }}$,
respectively.
Condition 2.3 requires that the estimator $
(\hat{\alpha}_{i},\hat{\beta}_{i}^\intercal)^\intercal$ is consistent for all $i$ and the moments of sample
averages of $u_{it}$ and $X_{it}$ across $t$ are bounded. If the tail index
is non-positive, then these bounds need to be stronger to accommodate the fact that $
a_{n}\rightarrow 0$.\footnote{
A straightforward calculation yields that normal distribution satisfies
Condition 2.3, if $\sup_{i}\left\vert
\left\vert (\hat{\alpha}_i,\hat{\beta}_{i}^\intercal)^\intercal-(\alpha_i,\beta _{i}^\intercal)^\intercal\right\vert \right\vert\left. =\right.
O_{p}(T^{-\varepsilon })$ and $\sup_{i}\left\vert \bar{u}_{i}\right\vert
\left. =\right. O_{p}(T^{-\varepsilon })$ for some $\varepsilon >0$ and if $
n/T\left. \rightarrow \right. \lambda $ for some $\lambda \left. \in
\right.(0,\infty )$. This can be seen by $1/f_{\alpha }\left( Q_{\alpha
}\left( 1\left.-\right. 1/n\right) \right) \left. \leq \right. O(\log (n))$
when $f_{\alpha}$ and $Q_{\alpha }$ are standard normal density and quantile
functions, respectively (cf. Example 1.1.7 in \citeasnoun{deHaan07}).}

Under these conditions, the following corollary establishes the asymptotic
coverage.

\begin{corollary}
\label{col ab}Suppose that Conditions 1.4 and 2.1-2.3 hold. For any fixed $k$
and any $F_{\alpha }$ and $F_{\beta }$ that satisfy these conditions,
\begin{eqnarray*}
\left. \underset{n,T\rightarrow \infty }{\lim \inf }\right. \mathbb{P}\left(
Q_{\alpha }(1-h/n)\in S(\mathbf{A})\right)  &\geq &1-\alpha  \\
\left. \underset{n,T\rightarrow \infty }{\lim \inf }\right. \mathbb{P}\left(
Q_{\beta }(1-h/n)\in S(\mathbf{B})\right)  &\geq &1-\alpha
\end{eqnarray*}
where $S\left( \cdot \right) $ is defined in (\ref{S_Lambda}).
\end{corollary}

\section{Monte Carlo simulation studies\label{sec MC}}

We conduct Monte Carlo experiments to examine the small sample performance
of the new approach. In Section \ref{sec MC no FE}, we first consider the
simple panel data $\{Y_{it},X_{it}\}$ without any fixed effect. In Section
\ref{sec MC lgth}, we compare the efficiency of the new approach with the
kernel estimator, which essentially uses more than one NNs. In Section \ref
{sec MC with RE}, we consider the linear random coefficient regression setup
(\ref{reg}).

\subsection{Conditional extremal quantiles\label{sec MC no FE}}

We continue to consider the three examples in Section \ref{sec general} as
the data generating processes (DGPs). In all experiments, generated data are
i.i.d. across $i$, but are dependent across $t$. The dependence structure across $t$ is specified as follows.

\begin{description}
\item[1. Joint Normal] $X_{it}=\rho X_{it-1}+u_{it}$ with $u_{it}\sim ^{iid}
\mathcal{N}\left( 0,1-\rho ^{2}\right) $ and $X_{i1}\sim \mathcal{N}\left(
0,1\right) $. $Y_{it}=r_{xy}X_{it}+\sqrt{1-r_{xy}^{2}}v_{it}$ where $
v_{it}\sim ^{iid}\mathcal{N}\left( 0,1\right) $ and independent of $u_{it}$.
Set $\rho =0.5$ and $r_{xy}=0.5.$

\item[2. Joint Student-t] $(X_{it},Y_{it})$ is i.i.d.\ across $t$ and
distributed as $t_{v}(\mu ,\Sigma )$ with $v=3$, $\mu =[0,0]^{\intercal }$,
and $\Sigma =[1,0.5;0.5,1]$.

\item[3. Conditional Pareto] $X_{it}=\rho X_{it-1}+u_{it}$ with $u_{it}\sim
^{iid}\mathcal{N}\left( 0,1-\rho ^{2}\right) $ and $X_{i1}\sim \mathcal{N}
\left( 0,1\right) $. $Y_{it}|X_{it}=x\sim \mathrm{Pa}(\xi (x))$, that is, $
\mathbb{P}(Y_{it}\leq y|X_{it}=x)=1-y^{-1/\xi (x)}$ for $y\geq 1$ where $\xi
\left( x\right) =x+0.5$.
\end{description}

We construct CIs for $Q_{Y|X=x_{0}}\left( 1-h/n\right) $ with $x_{0}=0$ and $
1.65$ (the 50\% and 95\% quantiles of $X$, respectively) and $h=1$ and $5$.
The sample sizes $n$ and $T$ are either 200 or 500, with smaller
combinations exercised in later experiments.

We compare results across three approaches: (i) the fixed-$k$ approach
(fixed-$k$) introduced in this paper, (ii) quantile regression (QR), and
(iii) bootstrapping the empirical quantile (Boot). We produce the fixed-$k$
CI using $k=20$ in most cases if not otherwise noted. The space of $\xi $ is
restricted to be $[-1/2,1/2]$. For the QR approach, we run a quantile regression of $
Y_{it}$ on $X_{it}$ and a constant at the $\tau $ quantile for each $i\,$.
The conditional quantile is estimated at $\hat{\beta}_{0i}+x_{0}\hat{\beta}
_{1i}$ where $\hat{\beta}_{0i}$ and $\hat{\beta}_{1i}$ are the coefficient
estimates using the $i$-th individual's observations. The CI is defined the
2.5\% and 97.5\% quantiles of these $n$ estimates. The bootstrap CI is based
on bootstrapping the empirical $\tau $ quantile in $\{Y_{i,[x_{0}]}
\}_{i=1}^{n}$. The bootstrap size is 200.

Tables \ref{tbl h5 tx05}-\ref{tbl h1 tx05} depict the coverage probabilities
(Cov) and the average lengths (Lgth) of the above three methods based on 500
simulation draws. The fixed-$k$ approach performs well in terms of both the
coverage and length across all the specifications. Regarding the QR method,
recall that the conditional quantile is a linear function of $X$ in the
first DGP but not in the other two. Therefore, not surprisingly, the CIs
based on QR perform well in the first DGP but deliver substantial
undercoverage and longer length in the other two due to misspecification.
The bootstrap approach is robust to misspecification but requires the
asymptotic normal approximation, which performs well only in the mid-sample
and does not near the tails. As such, the bootstrap intervals exhibit more
undercoverage for $h=1$ than for $h=5$.

We conclude the current subsection with a remark about the choice of $k$. A
larger $k$ leads to more tail observations and hence shorter confidence
intervals, but is subject to a larger approximation bias due to including
too many mid-sample data. This indicates that the choice of $k$ is
difficult, especially when $n$ is only moderate. It is actually impossible
to choose a uniformly best $k$ allowing the underlying CDF to be flexible
(see Theorem 1 of \citeasnoun{MuellerWang17}). The CDFs in our Monte Carlo designs
are all well behaved so that such a value of $k$ as large as 40\% of the
sample size performs well. This is seen in Table \ref{tbl h1 tx05}, which
reports the numbers for $k=20$ and $50$.

\begin{footnotesize}
\begin{table}[tbp]
\vspace{-1cm}
\begin{center}
\begin{footnotesize}
\caption{Finite sample performance of inference about conditional extremal quantile,
 no model specification}\label{tbl h5 tx05}

\begin{tabular}{lccccccccccc}
\hline\hline
$n$ &  & \multicolumn{4}{c}{200 (97.5\% quantile)} &  &  &
\multicolumn{4}{c}{500 (99\% quantile)} \\ \cline{3-12}
$T$ &  & \multicolumn{2}{c}{200} & \multicolumn{2}{c}{500} &  &  &
\multicolumn{2}{c}{200} & \multicolumn{2}{c}{500} \\
&  & Cov & Lgth & Cov & Lgth &  &  & Cov & Lgth & Cov & Lgth \\ \hline
&  & \multicolumn{10}{c}{Joint Normal} \\
fixed-$k$ &  & 0.97 & 0.63 & 0.96 & 0.66 &  &  & 0.95 & 0.56 & 0.96 & 0.56
\\
QR &  & 1.00 & 0.63 & 1.00 & 0.41 &  &  & 1.00 & 0.89 & 1.00 & 0.56 \\
Boot &  & 0.97 & 0.64 & 0.91 & 0.61 &  &  & 0.88 & 0.58 & 0.95 & 0.55 \\
\hline
&  & \multicolumn{10}{c}{Joint Student-t} \\
fixed-$k$ &  & 0.96 & 1.35 & 0.96 & 1.47 &  &  & 0.95 & 1.62 & 0.94 & 1.63
\\
QR &  & 0.95 & 2.20 & 0.00 & 1.31 &  &  & 1.00 & 4.76 & 0.01 & 2.87 \\
Boot &  & 0.91 & 1.36 & 0.95 & 1.34 &  &  & 0.89 & 1.51 & 0.94 & 1.68 \\
\hline
&  & \multicolumn{10}{c}{Conditional Pareto} \\
fixed-$k$ &  & 0.96 & 7.65 & 0.97 & 7.14 &  &  & 0.98 & 15.8 & 0.97 & 11.6
\\
QR &  & 0.00 & ${>}$10$^{3}$ & 0.00 & ${>}$10$^{3}$ &  &  & 0.00 & ${>}$10$
^{3}$ & 0.00 & ${>}$10$^{3}$ \\
Boot &  & 0.93 & 8.30 & 0.93 & 7.80 &  &  & 0.94 & 15.3 & 0.90 & 12.7 \\
\hline
\end{tabular}

\end{footnotesize}
\end{center}
 \normalsize \footnotesize
Note: Entries are coverages and lengths of the CIs for $Q_{Y|X=0}(1-5/n)$.
See the main text for the description of the three approaches and the data
generating processes. Confidence level is 5\%. Based on 500 simulation draws.
\end{table}
\end{footnotesize}

\begin{footnotesize}
\begin{table}[tbp]
\vspace{-1cm}
\begin{center}
\begin{footnotesize}
\caption{Finite sample performance of inference about conditional extremal quantile,
 no model specification}\label{tbl h5 tx095}

\begin{tabular}{lccccccccccc}
\hline\hline
$n$ &  & \multicolumn{4}{c}{200 (97.5\% quantile)} &  &  &
\multicolumn{4}{c}{500 (99\% quantile)} \\ \cline{3-12}
$T$ &  & \multicolumn{2}{c}{200} & \multicolumn{2}{c}{500} &  &  &
\multicolumn{2}{c}{200} & \multicolumn{2}{c}{500} \\
&  & Cov & Lgth & Cov & Lgth &  &  & Cov & Lgth & Cov & Lgth \\ \hline
&  & \multicolumn{10}{c}{Joint Normal} \\
fixed-$k$ &  & 0.96 & 0.65 & 0.96 & 0.65 &  &  & 0.95 & 0.57 & 0.94 & 0.57
\\
QR &  & 1.00 & 1.28 & 1.00 & 0.80 &  &  & 1.00 & 1.81 & 1.00 & 1.13 \\
Boot &  & 0.93 & 0.63 & 0.92 & 0.64 &  &  & 0.92 & 0.55 & 0.91 & 0.56 \\
\hline
&  & \multicolumn{10}{c}{Joint Student-t} \\
fixed-$k$ &  & 0.97 & 2.42 & 0.95 & 2.36 &  &  & 0.97 & 2.88 & 0.97 & 2.77
\\
QR &  & 1.00 & 3.53 & 1.00 & 2.26 &  &  & 1.00 & 6.23 & 1.00 & 3.94 \\
Boot &  & 0.94 & 2.31 & 0.93 & 2.25 &  &  & 0.93 & 2.83 & 0.95 & 2.72 \\
\hline
&  & \multicolumn{10}{c}{Conditional Pareto} \\
fixed-$k$ &  & 0.95 & 9.30 & 0.97 & 7.45 &  &  & 0.84 & 16.3 & 0.94 & 12.5
\\
QR &  & 0.00 & ${>}$10$^{3}$ & 0.00 & ${>}$10$^{3}$ &  &  & 0.00 & ${>}$10$
^{3}$ & 0.00 & ${>}$10$^{3}$ \\
Boot &  & 0.95 & 12.5 & 0.96 & 8.91 &  &  & 0.80 & 26.8 & 0.93 & 15.2 \\
\hline
\end{tabular}

\end{footnotesize}
\end{center}
 \normalsize \footnotesize
{\footnotesize Note: Entries are coverages and lengths of the CIs for }$
Q_{Y|X=1.65}(1-5/n)${\footnotesize . See the main text for the description
of the three approaches and the data generating processes. Confidence level
is 5\%. Based on 500 simulation draws.}
\end{table}
\end{footnotesize}

\begin{footnotesize}
\begin{table}[tbp]
\vspace{-1cm}
\begin{center}
\begin{footnotesize}
\caption{Finite sample performance of inference about conditional extremal quantile,
 no model specification}\label{tbl h1 tx05}

\begin{tabular}{lccccccccccc}
\hline\hline
$n$ &  & \multicolumn{4}{c}{200 (99.5\% quantile)} &  &  &
\multicolumn{4}{c}{500 (99.8\% quantile)} \\ \cline{3-12}
$T$ &  & \multicolumn{2}{c}{200} & \multicolumn{2}{c}{500} &  &  &
\multicolumn{2}{c}{200} & \multicolumn{2}{c}{500} \\
&  & Cov & Lgth & Cov & Lgth &  &  & Cov & Lgth & Cov & Lgth \\ \hline
&  & \multicolumn{10}{c}{Joint Normal} \\
fixed-$k$(k=20) &  & 0.95 & 1.82 & 0.96 & 1.83 &  &  & 0.97 & 1.69 & 0.96 &
1.70 \\
QR &  & 1.00 & 1.19 & 1.00 & 0.75 &  &  & 1.00 & 1.18 & 1.00 & 1.07 \\
Boot &  & 0.63 & 0.62 & 0.64 & 0.59 &  &  & 0.64 & 0.57 & 0.65 & 0.59 \\
\hline
&  & \multicolumn{10}{c}{Joint Student-t} \\
fixed-$k$(k=20) &  & 0.96 & 4.71 & 0.96 & 4.69 &  &  & 0.96 & 5.62 & 0.97 &
5.61 \\
fixed-$k$(k=50) &  & 0.94 & 3.91 & 0.92 & 3.90 &  &  & 0.95 & 4.85 & 0.92 &
4.73 \\
QR &  & 1.00 & 8.51 & 0.68 & 5.51 &  &  & 1.00 & 8.47 & 1.00 & 11.5 \\
Boot &  & 0.62 & 2.01 & 0.60 & 2.02 &  &  & 0.63 & 2.57 & 0.61 & 2.56 \\
\hline
&  & \multicolumn{10}{c}{Conditional Pareto} \\
fixed-$k$(k=20) &  & 0.98 & 27.6 & 0.98 & 26.1 &  &  & 0.94 & 48.1 & 0.97 &
40.5 \\
QR &  & 0.00 & ${>}$10$^{3}$ & 0.00 & ${>}$10$^{3}$ &  &  & 0.00 & ${>}$10$
^{3}$ & 0.00 & ${>}$10$^{3}$ \\
Boot &  & 0.71 & 25.9 & 0.63 & 30.4 &  &  & 0.78 & 76.2 & 0.77 & 43.9 \\
\hline
\end{tabular}

\end{footnotesize}
\end{center}
 \normalsize \footnotesize
Note: Entries are coverages and lengths of the CIs for $Q_{Y|X=0}(1-1/n)$.
See the main text for the description of the three approaches and the data
generating processes. Confidence level is 5\%. Based on 500 simulation draws.
\end{table}
\end{footnotesize}

\subsection{Comparison with kernel smoothing\label{sec MC lgth}}

Our new approach takes only one NN in each time series, which raises the
question of efficiency loss. We answer this by comparing our fixed-$k$
approach with the kernel smoothing method proposed by \citeasnoun{Gardes10}. In
particular, we first pool the panel data into a cross-sectional sample.
Suppose the object of interest is still $Q_{Y|X=x_{0}}(\tau )$. We follow
\citeasnoun{Gardes10} to pick the bin $B_{b_{nT}}(x_{0})$ centered at $x_{0}$ with
a bandwidth $b_{nT}$. Since there is no theoretical justification for the
optimal choice of $b_{nT}$, we take the rule-of-thumb choice $c(nT)^{-1/5}$
with different values of the constant $c$. Now a certain choice of $b_{nT}$
leads to a certain collection of $Y^{\prime }s$ whose paired $X^{\prime }s$
are in the bin $B_{b_{nT}}(x_{0})$. Sort these induced $Y^{\prime }s$ in the
descending order into $\{Y_{(1)}\geq Y_{(2)}\geq ...\geq Y_{(m)}\}$ where $m$
denotes the local sample size determined by the bandwidth. Such local sample
size is approximately $nTb_{n}$ in the kernel smoothing (as opposed to $n$
in our new approach).

Given the induced $Y$'s, the conditional quantile is estimated as $\hat{Q}
_{Y|X=x_{0}}(\tau ) {\footnotesize =} Y_{(\left\lfloor (1-\tau
)m\right\rfloor )}$, that is, the $\left\lfloor (1-\tau )m\right\rfloor $-th
largest order statistics in the induced $Y$'s where $\left\lfloor (1-\tau
)m\right\rfloor $ denotes the integer part of $\left( 1-\tau \right) m$.
\citeasnoun{Gardes10} show that, under $m(1-\tau )\rightarrow \infty $ and some
other regularity conditions,
\begin{equation*}
\sqrt{m(1-\tau )}\left( \frac{{\footnotesize \hat{Q}}_{{\footnotesize Y|X=}
x_{0}}{\footnotesize (\tau )}}{{\footnotesize Q}_{{\footnotesize Y|X=}x_{0}}
{\footnotesize (\tau )}}-1\right) \overset{d}{\rightarrow }\mathcal{N}
(0,1/\xi _{0}^{2}(x_{0})).
\end{equation*}
Then, the CI of $Q_{{\footnotesize Y|X=}x_{0}}(\tau )$ is constructed by the
delta method and plugging in some consistent estimator of $\xi _{0}$. One
choice which they propose is the Hill-type estimator
\begin{equation}
1/\hat{\xi}=\frac{1}{k-1}\sum_{i=1}^{k-1}i\log (Y_{(i)}/Y_{(i+1)})
\label{Smooth Hill}
\end{equation}
for some choice of $k<m$.

For comparisons, we implement our fixed-$k$ approach by using the panel data
and the above kernel approach by pooling the data. In particular, we
implement the conditional Pareto DGP in the previous experiment with $n=200$
and $T$ ranging from 50 to 500. For the fixed-$k$ CI, we set $k=50$. For the
kernel method, we implement $c\in \{0.1,0.25,0.5,1,2\}$ and set $k$ (in the
Hill-type index estimator (\ref{Smooth Hill})) as the largest integer less
than or equal to $m/4$.

Table \ref{tbl kernel one x} presents the coverages and the lengths of the
fixed-$k$ and the kernel CIs. Several interesting observations can be made.
First, the kernel approach is sensitive to the choice of the bandwidth. In
particular, a correct coverage relies on a narrow window of the bandwidth
choice. A larger choice can lead to a substantial undercoverage since the
smoothing bias dominates quickly in the tail. Second, when $T$ is only
moderately large (say 25 and 50), the fixed-$k$ CIs are much shorter than
the kernel one and both of them have good coverage properties. This is
because the fully nonparametric method ignores the domain-of-attraction
information, which is utilized by our fixed-$k$ method. Third, when $T$ is
very large, say 500, choosing only one NN does incur an efficiency loss as
we compare the lengths between our fixed-$k$ and the kernel CIs. But such a
loss is approximately in a factor of two or three instead of $T^{1/2}$. This
means that a general covariate-dependent tail is very difficult to estimate
in a fully nonparametric way.

\begin{footnotesize}
\begin{table}[tbp]

\begin{center}
\begin{footnotesize}
\caption{Finite sample performance of inference about conditional extremal quantile,
 comparison with kernel method}\label{tbl kernel one x}

\begin{tabular}{lccccccccc}
\hline
$T$ &  & \multicolumn{2}{c}{50} & \multicolumn{2}{c}{100} &
\multicolumn{2}{c}{200} & \multicolumn{2}{c}{500} \\
&  & Cov & Lgth & Cov & Lgth & Cov & Lgth & Cov & Lgth \\ \hline
fixed-k &  & 0.97 & 21.1 & 0.97 & 19.8 & 0.97 & 16.8 & 0.98 & 15.3 \\
NP(c=0.1) &  & 0.91 & 50.1 & 0.89 & 30.0 & 0.94 & 24.4 & 0.93 & 14.6 \\
NP(c=0.25) &  & 0.94 & 33.1 & 0.96 & 19.0 & 0.93 & 13.3 & 0.96 & 9.15 \\
NP(c=0.5) &  & 0.93 & 17.1 & 0.94 & 12.7 & 0.93 & 9.28 & 0.95 & 6.36 \\
NP(c=1) &  & 0.93 & 13.9 & 0.90 & 9.84 & 0.89 & 7.14 & 0.88 & 4.77 \\
NP(c=2) &  & 0.37 & 15.6 & 0.24 & 10.1 & 0.14 & 6.86 & 0.11 & 4.18 \\ \hline
\end{tabular}

\end{footnotesize}
\end{center}
 \normalsize \footnotesize
Note: Entries are coverages and lengths of the CIs for $Q_{Y|X=0}(1-1/n)$
under the conditional Pareto DGP. See the main text for the description of
the two approaches and details of the DGP. Confidence level is 5\%. Based on
500 simulation draws.
\end{table}
\end{footnotesize}

In Table \ref{tbl kernel two x}, we consider a two-dimensional standard
normal $X$ and generate $Y_{it}$ by $Y_{it}|X_{it}=x\sim \pm $Pa$(\xi (x))$
with $\xi (x_{1},x_{2})=x_{1}+x_{2}+0.5$. The kernel method is illustrated
with $c\in \{0.5,1,2,4\}$. All the other parameter choices for both methods
remain unchanged as those used for Table \ref{tbl kernel one x}. The results
clearly suggest that our fixed-$k$ method together with the NN choice
dominates the kernel method in terms of both the coverage probabilities and
lengths. In particular, the kernel method suffers from the curse of
dimensionality as the dimension of $X$ increases.

\begin{footnotesize}
\begin{table}[tbp]
\begin{center}
\begin{footnotesize}
\caption{Finite sample performance of inference about conditional extremal quantile,
 comparison with kernel method, two-dimensional X}\label{tbl kernel two x}

\begin{tabular}{lccccccccc}
\hline
$T$ &  & \multicolumn{2}{c}{50} & \multicolumn{2}{c}{100} &
\multicolumn{2}{c}{200} & \multicolumn{2}{c}{500} \\
&  & Cov & Lgth & Cov & Lgth & Cov & Lgth & Cov & Lgth \\ \hline
fixed-k &  & 0.96 & 21.0 & 0.97 & 17.2 & 0.96 & 16.5 & 0.96 & 16.3 \\
NP(c=0.5) &  & 0.56 & 65.8 & 0.73 & 60.6 & 0.73 & 54.3 & 0.81 & 53.8 \\
NP(c=1) &  & 0.83 & 75.0 & 0.80 & 60.8 & 0.93 & 60.9 & 0.91 & 32.0 \\
NP(c=2) &  & 0.96 & 66.1 & 1.00 & 36.0 & 0.97 & 25.2 & 0.97 & 16.5 \\
NP(c=4) &  & 0.74 & 54.5 & 0.47 & 35.4 & 0.23 & 22.8 & 0.16 & 13.2 \\ \hline
\end{tabular}

\end{footnotesize}
\end{center}
 \normalsize \footnotesize
Note: Entries are coverages and lengths of the CIs for $Q_{Y|X=0}(1-1/n)$
under the conditional Pareto DGP. See the main text for the description of
the two approaches and details of the DGP. Confidence level is 5\%. Based on
500 simulation draws.
\end{table}
\end{footnotesize}

As a final remark of this subsection, we also implement the standard kernel
weighted quantile regression method designed for the mid-sample quantiles
(cf.\ Chapter 10 of \citeasnoun{LiRacine07}). Given a large $T$, the target $1-1/n$
conditional quantile is relatively in the mid-sample after pooling the panel
data into a cross-sectional one, and hence the confidence interval based on
asymptotic normality might work. However, unreported Monte Carlo simulations
show that this method works only if $T$ is substantially larger than (e.g,
five times as much as) $n$. In our experiments, it is strictly dominated by
the method proposed by \citeasnoun{Gardes10}.

\subsection{Extremal quantiles in a linear random coefficient model\label
{sec MC with RE}}

In this section, we consider the linear random coefficient model $
Y_{it}=\alpha _{i}+X_{it}\beta _{0}+u_{it}$, where the generated
observations including the random coefficients are i.i.d. across $i$. For
the time series dependence, we set $\alpha_{i}=T^{-1}\sum_{t=1}^{T}X_{it}$
and $X_{it}=\rho X_{it-1}+e_{it}$ with $e_{it}\sim ^{iid}\mathcal{N}\left(
0,(1-\rho ^{2})\right) $ and $X_{i0}\sim \mathcal{N}\left( 0,1\right) $. The
conditional distributions of $u_{it}$ given $X_{it}=x$ are specified as
follows.

\begin{description}
\item[1. Conditional Normal] $u_{it}|X_{it}=x\sim \mathcal{N}\left(
0,1+x^{2}\right) $.

\item[2. Conditional Student-t] $u_{it}|X_{it}=x\sim t\left( 2+\left\vert
x\right\vert \right) $.

\item[3. Conditional Pareto] $u_{it}|X_{it}=x\sim \pm \mathrm{Pa}(\xi (x))$,
that is, $\mathbb{P}(u_{it}\leq y|X_{it}=x)=1/2+(1-(1+y)^{-1/\xi (x)})/2$
for $y\geq 0$, and $\mathbb{P}(u_{it}\leq y|X_{it}=x)=\left( -y+1\right)
^{-1/\xi (x)}/2$ for $y\leq 0$ where $\xi \left( x\right) =x+0.5$.
\end{description}

We use the same set of the three approaches as in the Section \ref{sec MC no
FE} to construct CIs for the conditional extremal quantile $
Q_{Y_{it}|X_{it}=x_{0}}\left( \tau \right) =Q_{\varepsilon
_{it}|X_{it}=x_{0}}\left( \tau \right) +x_{0}\beta _{0}$, where $\varepsilon
_{it}$ denotes $\alpha _{i}+u_{it}$. Specifically, our fixed-$k$ approach is
conducted in two ways: with or without using the standard within least
squares estimator of $\beta _{0}$. For the former (fixed-$k$ w.\thinspace
LS), we first estimate $\beta _{0}$ using the standard within estimator $
\hat{\beta}$ and back out $\hat{\varepsilon}_{it}=Y_{it}-X_{it}\hat{\beta}$.
We then implement the steps in Section \ref{sec general} to construct the
CIs for the conditional quantiles of $\varepsilon _{it}$. The CIs for $
Q_{Y_{it}|X_{it}=x_{0}}\left( \tau \right) $ are obtained by adding back $
x_{0}\hat{\beta}$. For the one ignoring the linear regression structure
(fixed-$k$ w/o LS), we directly use $(Y_{it},X_{it})^\intercal$, and apply Steps 1-3
in Section \ref{sec general}.

Table \ref{tbl h1 tx05 RE} presents the results for $n\in \{100,200\}$ and $
T\in \{25,50,200,500\}$. Several interesting observations can be made.
First, the errors in the conditional t and conditional Pareto models do not
have finite variances when $x_{0}$ is 0, and hence the LS estimator of $
\beta _{0}$ behaves poorly. This leads to a poor performance of the fixed-$k$
approach if the linear regression model is utilized. This problem can be
solved by using the least absolute deviation (LAD) estimator as shown in
unreported results. In comparison, the fixed-$k$ CIs without using the
linear regression model always perform well given a large enough sample
size. Second, the QR approach still suffers from undercoverage in all three
specifications since the normal and the Student-t DGPs have nonlinear
heteroskedasticity and the conditional Pareto DGP violates the constant tail
shape condition. Finally, the bootstrap method performs poorly if the
extremal quantiles under investigation are too far in the tail.

\begin{footnotesize}
\begin{table}[tbp]
\vspace{-1cm}
\begin{center}
\begin{footnotesize}
\caption{Finite sample performance of inference about conditional extremal quantile,
 non-dynamic model with random effects}\label{tbl h1 tx05 RE}

\begin{tabular}{lccccccccccc}
\hline\hline
$n$ &  & \multicolumn{4}{c}{200 (99.5\% quantile)} &  &  &
\multicolumn{4}{c}{100 (99\% quantile)} \\ \cline{3-12}
$T$ &  & \multicolumn{2}{c}{200} & \multicolumn{2}{c}{500} &  &  &
\multicolumn{2}{c}{25} & \multicolumn{2}{c}{50} \\
&  & Cov & Lgth & Cov & Lgth &  &  & Cov & Lgth & Cov & Lgth \\ \hline
&  & \multicolumn{10}{c}{Conditional Normal} \\
fixed-$k$ w. LS &  & 0.94 & 2.23 & 0.93 & 2.13 &  &  & 0.80 & 2.85 & 0.88 &
2.34 \\
fixed-$k$ w/o LS &  & 0.93 & 2.22 & 0.92 & 2.14 &  &  & 0.53 & 3.17 & 0.76 &
2.62 \\
QR &  & 0.00 & 3.04 & 0.00 & 2.19 &  &  & 1.00 & 3.10 & 1.00 & 2.96 \\
Boot &  & 0.73 & 0.77 & 0.67 & 0.71 &  &  & 0.73 & 1.11 & 0.81 & 0.92 \\
\hline
& \multicolumn{11}{c}{Conditional Student-t} \\
fixed-$k$ w. LS &  & 0.95 & 15.3 & 0.94 & 15.5 &  &  & 0.93 & 10.0 & 0.96 &
10.4 \\
fixed-$k$ w/o LS &  & 0.95 & 15.3 & 0.94 & 15.5 &  &  & 0.91 & 10.0 & 0.93 &
10.3 \\
QR &  & 1.00 & 26.5 & 1.00 & 7.98 &  &  & 0.99 & 11.0 & 1.00 & 15.0 \\
Boot &  & 0.56 & 11.5 & 0.60 & 11.0 &  &  & 0.47 & 6.32 & 0.51 & 6.01 \\
\hline
& \multicolumn{11}{c}{Conditional Pareto} \\
fixed-$k$ w. LS &  & 0.00 & 16.2 & 0.00 & 5.90 &  &  & 0.02 & 59.6 & 0.01 &
58.1 \\
fixed-$k$ w/o LS &  & 0.97 & 18.8 & 0.97 & 16.9 &  &  & 0.95 & 16.0 & 0.96 &
15.4 \\
QR &  & 0.00 & ${>}$10$^{3}$ & 0.00 & ${>}$10$^{3}$ &  &  & 1.00 & ${>}$10$
^{3}$ & 1.00 & ${>}$10$^{3}$ \\
Boot &  & 0.71 & 16.2 & 0.67 & 16.8 &  &  & 0.76 & 313 & 0.78 & 31.4 \\
\hline
\end{tabular}

\end{footnotesize}
\end{center}
 \normalsize \footnotesize
Note: Entries are coverages and lengths of the CIs for $Q_{Y|X=0}(1-1/n)$.
See the main text for the description of different approaches and the data
generating processes. Confidence level is 5\%. Based on 500 simulation draws.
\end{table}
\end{footnotesize}

In Table \ref{tbl a and b}, we study the CIs for high quantiles of $
\alpha_{i}$ and $\beta _{i}$ with data generated from $Y_{it}=
\alpha_{i}+X_{it}\beta _{i}+u_{it}$, where $\left( \alpha _{i},\beta
_{i},X_{it},u_{it}\right) ^{\intercal }\sim ^{iid}\mathcal{N}
\left(0,I_{4}\right) $. The i.i.d. condition is across both $i$ and $t$ in
this setting. We first estimate $\alpha _{i}$ and $\beta _{i}$ by regressing
$Y_{it}$ on $(1,X_{it})^{\intercal }$ with $T$ observations from individual $
i$. We then collect the estimates, $\hat{\alpha}_{i}$ and $\hat{\beta}_{i}$,
for all $i$ and sort them in the descending order to apply each of the fixed-
$k$, QR, and bootstrap methods. The QR estimator is simply the empirical
quantile among the estimators for all $i$, whose asymptotic variance is
estimated by the standard kernel density estimator with the rule-of-thumb
bandwidth. The results suggest that the fixed-$k$ approach with NN dominates
the other two in both coverage and length, especially when the sample size
is only moderate.

\begin{footnotesize}
\begin{table}[tbp]
\vspace{-1cm}
\begin{center}
\begin{footnotesize}
\caption{Finite sample performance of inference about large quantiles of the
random coefficients}\label{tbl a and b}

\begin{tabular}{lccccccccccc}
\hline\hline
$n$ &  & \multicolumn{4}{c}{$200$} &  &  & \multicolumn{4}{c}{$500$} \\ \cline{3-12}
$T$ &  & \multicolumn{2}{c}{$10$} & \multicolumn{2}{c}{$20$} &  &  &
\multicolumn{2}{c}{$10$} & \multicolumn{2}{c}{$20$} \\
&  & Cov & Lgth & Cov & Lgth &  &  & Cov & Lgth & Cov & Lgth \\ \hline
&  & \multicolumn{10}{c}{CIs for $Q_{\alpha }(1-5/n)$} \\
fixed-$k$ &  & 0.92 & 0.77 & 0.95 & 0.76 &  &  & 0.92 & 0.69 & 0.93 & 0.67
\\
QR &  & 0.84 & 1.13 & 0.90 & 1.07 &  &  & 0.88 & 2.96 & 0.94 & 2.94 \\
Boot &  & 0.89 & 0.81 & 0.93 & 0.76 &  &  & 0.81 & 0.70 & 0.91 & 0.69 \\
\hline
&  & \multicolumn{10}{c}{CIs for $Q_{\beta }(1-5/n)$} \\
fixed-$k$ &  & 0.91 & 0.81 & 0.96 & 0.76 &  &  & 0.86 & 0.69 & 0.96 & 0.67
\\
QR &  & 0.81 & 1.15 & 0.88 & 1.08 &  &  & 0.88 & 3.21 & 0.94 & 2.83 \\
Boot &  & 0.85 & 0.82 & 0.92 & 0.76 &  &  & 0.78 & 0.73 & 0.91 & 0.68 \\
\hline
&  & \multicolumn{10}{c}{CIs for $Q_{\alpha }(1-1/n)$} \\
fixed-$k$ &  & 0.91 & 2.32 & 0.93 & 2.13 &  &  & 0.87 & 2.10 & 0.94 & 1.96
\\
QR &  & 0.89 & 1.45 & 0.91 & 1.42 &  &  & 0.89 & 1.33 & 0.88 & 1.29 \\
Boot &  & 0.57 & 0.52 & 0.58 & 0.51 &  &  & 0.57 & 0.47 & 0.54 & 0.47 \\
\hline
&  & \multicolumn{10}{c}{CIs for $Q_{\beta }(1-1/n)$} \\
fixed-$k$ &  & 0.88 & 2.32 & 0.94 & 2.28 &  &  & 0.85 & 2.16 & 0.93 & 1.93
\\
QR &  & 0.90 & 1.52 & 0.91 & 1.52 &  &  & 0.86 & 1.41 & 0.88 & 1.29 \\
Boot &  & 0.57 & 0.55 & 0.58 & 0.54 &  &  & 0.55 & 0.51 & 0.58 & 0.45 \\
\hline
\end{tabular}

\end{footnotesize}
\end{center}
 \normalsize \footnotesize
Note: The entries are coverage and length of the confidence intervals based
on (i) the fixed-$k$ approach using the largest k=20$\ $estimated
coefficients, (ii) empirical quantile of the estimated coefficients with
asymptotic normal approximation, and (iii) empirical quantile function of
the estimated coefficients and bootstrap. Data are generated from $
Y_{it}=\alpha _{i}+X_{it}\beta _{i}+u_{it}$ where $\left( \alpha _{i},\beta
_{i},X_{it},u_{it}\right) ^{\intercal }\sim ^{iid}\mathcal{N}\left(
0,I_{4}\right) $. The target is the 1-h/n quantile of $\alpha _{i}$ and $
\beta _{i}$ with $h=1$ and $5$, corresponding to 97.5\%, 98\%, 99\%, and
99.8\% quantiles\ given $n=200$\ and $500$, respectively. Confidence level
is 5\%. Based on 500 simulation draws.
\end{table}
\end{footnotesize}

\section{Empirical application to extremal birth weights\label
{sec:application}}

In this section, we reconsider the extremely low birth weights and their
relationships with mother's demographic characteristics and maternal
behaviors, which addresses an important question in health economics. We use
the detailed natality data published by the National Center for Health
Statistics, which has been used by \citeasnoun{Abrevaya01}, \citeasnoun{Koenker01}, and
\citeasnoun{Chernozhukov11} among many others. We follow these preceding studies,
but our analysis is different from theirs in two aspects. First, these
preceding studies use the cross-sectional data in one time period, while we
collect the repeated cross-sectional samples from January 1989 to December
2002.\footnote{
We chose this specific period for two reasons. First, this period contains
the time period of the cross-sectional data used by \citeasnoun{Abrevaya01}, \citeasnoun
{Koenker01}, and \citeasnoun{Chernozhukov11}. Second, these periods maintain the
identical variable definitions.} Second, the previous studies all made some
parametric model assumptions, including either the linear projection model
or the (extremal) quantile regression model. In contrast, our fixed-$k$
method is nonparametric, allowing for nonparametric joint distributions.
Accordingly, some of our findings are different from those in the previous
studies.

Details of our implementation are as follows. First, we follow the
previously mentioned literature -- \citeasnoun{Abrevaya01} in particular -- to choose included covariates. Our dependent
variable is the infant birth weight measured in kilograms, and the
continuous covariates include mother's age and net weight gain (wtgain)
during pregnancy. All the remaining covariates are discrete, and hence we
consider the subsamples constructed from various combinations of the
categorical variables. For comparison, we set a benchmark subsample in which
the infant is a boy, the mother is white and married, has levels of
education less than a high school degree, had her first prenatal visit in
the first trimester (natal1), and did not smoke during pregnancy. Second,
since the samples are repeated cross-sectional, it is more natural to switch
the labeling of the indices $i$ and $t$, and first take the NN within each month. The
query point is set at age equal to 27 and wtgain equal to 30, corresponding
to their respective median values. The NN is then measured by the Euclidean
norm after standardizing each of the two variables with mean zero and unit variance. Using the same notation as in Section \ref{sec main}, we have $n=168$
and $T$ is at least 100 in every subsample. Thus, our fixed-$k$ asymptotic
framework with a large $n$ and a large $T$ is suitable with this data.
Third, we set $k=30$ based on our simulation results in the previous section
and construct the 95\% fixed-$k$ confidence intervals for the conditional $p$
-quantiles with $p$ ranging from 1\% to 10\%. Figure \ref{fig natality}
depicts these confidence intervals in the benchmark subsample and six
alternative subsamples corresponding to one and only one of following
scenarios: the mother has at least high school diploma; the infant is a
girl; the mother is unmarried; the mother is black; the mother does not have
prenatal visit during pregnancy; and the mother smokes 10 cigarettes per day
on average.\footnote{
For the majority of the subsamples, the number of cigarettes as recorded in
data takes only a few discrete values, including 0, 5, 10, and 20.
Therefore, we treat it as a discrete random variable in our study.}

\begin{figure}[tbp]
\caption{Plot of confidence intervals for the conditional extremal quantile
of infant birth weight.}\label{fig natality} \vspace{-4ex}
\par
\begin{center}
\includegraphics[width=0.85
\textwidth]{figure_natality.jpg}
\end{center}
\par
\vspace{-4ex}
\par
{\footnotesize Note: This figure plots the 95\% fixed-$k$ confidence
intervals for the conditional $p$-quantile of infant birth weight with $p\in
\lbrack 0.01,0.1]$, conditional on mother's age being 27, net weight gain
during pregnancy being 30 pounds, and the other six discrete covarietes. See
the main text for more detailed descriptions of these six covariates. The
vertical axis is the birth weight in kilograms, and the horizontal axis is $p
$. Data are available at the National Center for Health Statistics:
https://www.cdc.gov/nchs/nvss/births.htm. }
\end{figure}

We can make the following observations in Figure \ref{fig natality}. First,
the effects of changing the covariates are found to have a similar pattern
as in the previous studies. In particular, compared with the benchmark
subsample, the conditional quantile of infant birth weight decreases
substantially if the mother is black, did not have a prenatal visit, and/or
smoked during pregnancy. These results reconfirm the signs of the effects reported by the previous
studies. Second, on the other hand, the magnitudes of these effects are larger than
those documented in the previous studies. Specifically, \citeasnoun{Abrevaya01}
finds that smoking ten cigarettes leads to approximately 200 fewer grams at
the 10th percentile of the infant birth weight, compared with smoking no
cigarettes. \citeasnoun{Chernozhukov11} finds the quantile regression coefficient
associated with the number of cigarettes is nearly zero at the 1st
percentile (their Figure 8). On the other hand, the last sub-figure in
Figure \ref{fig natality} suggests that the difference can be over 1000
grams at the 1st percentile if we compare the mid-value between the upper
and lower bounds of the confidence intervals between these two subsamples.
Finally, the effects on extremal birth weight quantiles induced by the
demographic characteristics vary across levels of quantiles, instead of
remaining fixed.

\section{Concluding remarks\label{sec conclusion}}

This paper develops a new nonparametric method of inference for conditional
extremal quantiles using a fixed number $k$ of nearest-neighbor tail
observations in repeated cross-sectional or panel data.
There are three advantages of our proposed method.
First, it is robust against flexible distributional assumptions unlike parametric methods.
Second, the procedure yields asymptotically valid confidence intervals for any fixed tuning parameter $k$, unlike existing kernel methods that rely on a sequence of moving tuning parameters for asymptotically valid inference.
Third, our confidence intervals enjoy the uniform coverage property over a set of data generating processes involving a set of values of the tail index.

The key insight is that the induced order statistics in each time series can be treated as
approximately stemming from the true conditional distribution, and the large
order statistics among these induced values can then be used to make
inference on extremal quantiles. By focusing on the induced order
statistics, we effectively reduce the conditional tail problem into an
unconditional one. Monte Carlo simulations show that the new method delivers
preferred small sample performance in terms of coverage probability and
length.

The new method is more flexible than the extremal quantile regression
because the latter assumes that the conditional extremal quantile is a
parametric location-shift model. If a linear regression model is imposed, then our proposed method can be easily combined with any existing consistent estimator of
structural parameters and applies to inference on extremal quantiles of the
random coefficients.

Applying the proposed method to Natality Vital Statistics, we reexamine
factors of extremely low birth weights that have been analyzed by preceding
studies. We find that signs of major effects are the same as those found in
preceding studies based on parametric models. On the other hand, we find
that some effects exhibit different magnitudes from those reported in the
previous studies based on parametric models.