The exact contents of citations.db main_text.text for this paper — one flattened LaTeX string, title through conclusion, appendix excluded, unmodified except for removing email addresses. This is what our citation measures are computed over.
70,707 characters
Inference for Extremal Conditional Quantile Models, with an Application to Market and Birthweight Risks
\pagestyle{empty}
\title[Inference for Extremal Quantile Regression]{
Inference for Extremal Conditional Quantile Models, with an Application
to Market and Birthweight Risks \\
}
\author[ ] {Victor
Chernozhukov$^\dag$ \ \ Iv\'an Fern\'andez-Val$^\S$}
\thanks{First version: May, 2002. This version: \today. This is a
revision of the paper with the original title ``Extreme Value
Inference for Quantile Regression.'' We would like to thank very
much Takeshi Amemiya, Brigham Fradsen, Hide Ichimura, Jerry Hausman, Peter Hinrichs,
Keith Knight, Roger Koenker, Joseph Romano, the editor Bernard
Salanie, and two referees.}
\thanks{$\dag$ Massachusetts Institute of Technology, Department of
Economics and Operations Research Center, University College London,
CEMMAP. E-mail: [email removed].}
\thanks{$\S$ Boston University, Department of Economics. E-mail:
[email removed].}
\begin{abstract}
Quantile regression is an increasingly important empirical tool in
economics and other sciences for analyzing the impact of a set of
regressors on the conditional distribution of an outcome. Extremal
quantile regression, or quantile regression applied to the tails, is
of interest in many economic and financial applications, such as
conditional value-at-risk, production efficiency, and adjustment bands
in (S,s) models. In this paper we provide feasible inference tools for
extremal conditional quantile models that rely upon extreme value
approximations to the distribution of self-normalized quantile
regression statistics. The methods are simple to implement and can
be of independent interest even in the non-regression case. We
illustrate the results with two empirical examples analyzing extreme fluctuations of a stock return
and extremely low percentiles of live infants'
birthweights in the range between 250 and 1500 grams.\\
\noindent {\sc Key Words}: \textsc{Quantile Regression, Feasible Inference, Extreme Value Theory} \\
\noindent {\sc JEL}: \textsc{ C13, C14, C21, C41, C51, C53} \\
\noindent {\sc Monte-Carlo programs and software are available at www.mit.edu/vchern} \\
\end{abstract}
\maketitle
\pagestyle{plain}\thispagestyle{empty}
\newpage\pagestyle{headings}\setcounter{page}{1}
\section{Introduction and Motivation}
Quantile regression (QR) is an increasingly important empirical tool
in economics and other sciences for analyzing the impact of a set of
regressors $X$ on features of the conditional distribution of an
outcome $Y$ (see Koenker, 2005). In many applications the features
of interest are the extremal or tail quantiles of the conditional
distribution. This paper provides practical tools for performing
inference on these features using extremal QR and
extreme value theory. The key problem we address is that
conventional inference methods for QR, based on the
normal distribution, are not valid for extremal QR. By using extreme
value theory, which specifically accounts for the extreme nature of
the tail data, we are able to provide inference methods that are
valid for extremal QR.
Before describing the contributions of this paper in more detail, we
first motivate the use of extremal quantile regression
in specific economic applications. Extremal
quantile regression provides a useful description of important
features of the data in these applications, generating both
reduced-form facts as well as inputs into estimation of structural
models. In what follows, $Q_{Y}(\tau|X)$ denotes the conditional
$\tau$-quantile of $Y$ given $X$; extremal conditional quantile
refers to the conditional quantile function $Q_{Y}(\tau|X)$ with the
quantile index $\tau = \epsilon$ or $1-\epsilon$, where $\epsilon$
is close to zero; and extremal quantile regression refers to the
quantile regression estimator of an extremal conditional quantile.
A principal area of economic applications of extremal quantile
regression is risk management. One example in this area is
conditional value-at-risk analysis from financial economics
\cite{vl,caviar}. Here, we are interested in the extremal quantile
$Q_{Y}(\epsilon|X)$ of a return $Y$ to a bank's portfolio,
conditional on various predictive variables $X$, such as the return
to the market portfolio and the returns to portfolios of other
related banks and mortgage providers. Unlike unconditional extremal
quantiles, conditional extremal quantiles are useful for
stress testing and analyzing the impact of adverse systemic events
on the bank's performance. For example, we can analyze the impact of
a large drop in the value of the market portfolio or of an associated
company on the performance of the bank's portfolio. The results of
this analysis are useful for determining the level of capital that
the bank needs to hold to prevent bankruptcy in unfavorable states
of the world. Another example comes from health economics, where we
are interested in the analysis of socio-economic determinants $X$ of
extreme quantiles of a child's birthweight $Y$ or other health
outcomes. In this example, very low birthweights are connected with
substantial health problems for the child, and thus extremal quantile regression is useful
to identifying which factors can improve these negative health
outcomes. We shall return to these examples later in the empirical
part of the paper.
Another primary area of economic applications of extremal quantile
regression deals with describing approximate or probabilistic
boundaries of economic outcomes conditional on pertinent factors. A
first example in this area comes from efficiency analysis in the
economics of regulation, where we are interested in the
probabilistic production frontier $Q_{Y}(1-\epsilon|X)$. This
frontier describes the level of production $Y$ attained by the most
productive $(1-\epsilon)\times 100$ percent of firms,
conditional on input factors $X$ \cite{timmer}. A second
example comes from the analysis of job search in labor economics,
where we are interested in the approximate reservation wage
$Q_{Y}(\epsilon|X)$. This function describes the wage level,
below which the worker accepts a job only with a small probability
$\epsilon$, conditional on worker characteristics and other
factors $X$ \cite{flinn}. A third example deals with estimating
$(S,s)$ rules in industrial organization and macroeconomics
\cite{caballero:engel3}. Recall that the $(S,s)$ rule is an optimal
policy for capital adjustment, in which a firm allows its capital
stock to gradually depreciate to a lower barrier, and once the
barrier is reached, the firm adjusts its capital stock sharply to an
upper barrier. Therefore, in a given cross-section of firms, the
extremal conditional quantile functions $Q_{Y}(\epsilon|X)$ and
$Q_{Y}(1-\epsilon|X)$ characterize the approximate
adjustment barriers for observed capital stock $Y$, conditional on a
set of observed factors $X$.\footnote{\citeasnoun{caballero:engel3}
study approximate adjustment barriers using distribution
models; obviously quantile models can also be used.}$^,$\footnote{In
the previous examples, we can set $\epsilon$ to 0 to recover the
exact, non-probabilistic, boundaries in the case with no unobserved
heterogeneity and no (even small) outliers in the data. Our
inference methods cover this exact extreme case, but we recommend
avoiding it because it requires very stringent assumptions.}
The two areas of applications described above are either
non-structural or semi-structural. A third principal area of
economic applications of extremal quantile regressions is structural
estimation of economic models. For instance, in procurement auction
models, the key information about structural parameters is contained in the extreme or
near-extreme conditional quantiles of bids given bidder and auction
characteristics (see e.g. \citeasnoun{chernozhukov:hong:denjump} and
\citeasnoun{hirano:porter}). We then can estimate and test a structural
model based on its ability to accurately reproduce a collection of
extremal conditional quantiles observed in the data. This indirect
inference approach is called the method-of-quantiles
\cite{koenker:book}. We refer the reader to
\citeasnoun{dp:superconsistent} for a detailed example of this
approach in the context of using $k$-sample extreme
quantiles.\footnote{The method-of-quantiles allows us to estimate
structural models both with and without parametric unobserved
heterogeneity. Moreover, the use of near-extreme quantiles instead
of exact-extreme quantiles makes the method more robust to a small
fraction of outliers or neglected unobserved heterogeneity.}
We now describe the contributions of this paper more specifically.
This paper develops feasible and practical inferential methods based
on extreme value (EV) theory for QR, namely, on the limit law theory
for QR developed in Chernozhukov (2005) for cases where the quantile
index $\tau\in (0,1)$ is either low, close to zero, or high, close
to 1. Without loss of generality we assume the former. By close to
0, we mean that the order of the $\tau$-quantile, $\tau T$, defined
as the product of quantile index $\tau$ with the sample size $T$,
obeys $\tau T \to k < \infty$ as $T \to \infty$. Under this
condition, the conventional normal laws, which are based on the
assumption that $\tau T$ diverges to infinity, fail to hold, and
different EV laws apply instead. These laws approximate
the exact finite sample law of extremal QR better than
the conventional normal laws. In particular, we
find that when the dimension-adjusted order of the $\tau$-quantile,
$\tau T/d$, defined as the ratio of the order of the $\tau$-quantile
to the number of regressors $d$, is not large, less than about $20$ or
$40$, the EV laws may be preferable to the normal law,
whereas the normal laws may become preferable otherwise.
We suggest this simple rule of thumb for choosing between the EV
laws and normal laws, and refer the reader to Section 5 for more
refined suggestions and recommendations.
Figure 1 illustrates the difference between the EV and normal
approximations to the finite sample distribution of the extremal QR estimators.
We plot the quantiles of these approximations against the quantiles of the exact finite
sample distribution of the QR estimator. We consider different
dimension-adjusted orders in a simple model with only one regressor,
$d=1$, and $T=200$. If either the EV law or the normal law were to
coincide with the true law, then their quantiles would fall
exactly on the 45 degree line shown by the solid line. We see from
the plot that when the dimension-adjusted order $\tau T/d$ is $20$
or $40$, the quantiles of the EV law are indeed very close to the 45
degree line, and in fact are much closer to this line than the
quantiles of the normal law. Only for the case when the effective
order $\tau T/d$ becomes $60$, do the quantiles of the EV law and
normal laws become comparably close to the 45 degree line.
\begin{figure}[!http!]\label{shock}
\psfigdriver{dvips}
\epsfig{figure=fig.noX.eps, width=6.2in, height=3in}
\caption{\textbf{Quantiles of the true law of QR vs. quantiles of EV and normal
laws}. The figure is based on a simple design with $Y = X + U$, where $U$
follows a Cauchy distribution and $X=1$. The solid line ``------" shows the actual quantiles of the true
distribution of QR with quantile index $\tau \in \{ .025, .2, .3\}$. The dashed line ``- \ - \ -" shows the quantiles
of the conventional normal law for QR, and the dotted line ``......" shows the quantiles
of EV law for QR. The figure is based on 10,000 Monte Carlo replications and plots quantiles over the $99\%$ range.
}
\end{figure}
A major problem with implementing the EV approach, at least
in its pure form, is its infeasibility for inference purposes. Indeed, EV approximations rely on canonical
normalizing constants to achieve non-degenerate asymptotic laws.
Consistent estimation of these constants is generally not possible,
at least without making additional strong assumptions. This
difficulty is also encountered in the classical non-regression case;
see, for instance, \citeasnoun{bertailetal} for discussion.
Furthermore, universal inference methods such as the bootstrap
fail due to the nonstandard behavior of extremal QR statistics; see
\citeasnoun{bickel} for a proof in the classical non-regression
case. Conventional subsampling methods with and without replacement
are also inconsistent because the QR statistic diverges in the
unbounded support case. Moreover, they require consistent estimation
of normalizing constants, which is not feasible in general.
In this paper we develop two types of inference approaches that
overcome all of the difficulties mentioned above: a resampling approach
and an analytical approach. We favor the first approach due to its ease of implementation in practice. At the heart of both approaches
is the use of self-normalized QR (SN-QR) statistics that employ random
normalization factors, instead of generally infeasible normalization
by canonical constants. The use of SN-QR statistics allows us to
derive feasible limit distributions, which underlie either of our
inference approaches. Moreover, our resampling approach is a
suitably modified subsampling method applied to SN-QR statistics.
This approach entirely avoids estimating not only the canonical
normalizing constants, but also all other nuisance tail parameters,
which in practice may be difficult to estimate reliably. Our
construction fruitfully exploits the special relationship between the
rates of convergence/divergence of extremal and intermediate QR statistics,
which allows for a valid estimation of the centering constants
in subsamples. For completeness we also provide inferential methods
for canonically-normalized QR (CN-QR) statistics, but we also show
that their feasibility requires much stronger assumptions.
The remainder of the paper is organized as follows. Section 2
describes the model and regularity conditions, and gives an
intuitive overview of the main results. Section 3 establishes the
results that underlie the inferential procedures. Section 4
describes methods for estimating critical values. Section 5
compares inference methods based on EV and normal approximations
through a Monte Carlo experiment. Section 6 presents empirical
examples, and the Appendix collects proofs and all other figures.
\section{The Set Up and Overview of Results}
\subsection{Some Basics} Let a real random variable $Y$ have a continuous
distribution function $F_Y(y) = \Pr[Y \leq y]$. A $\tau$-quantile of
$Y$ is $Q_Y(\tau)= \inf\{y: F_{Y}(y) > \tau\}$ for some $\tau \in
(0,1).$ Let $X$ be a vector of covariates related to $Y$, and
$F_{Y}(y|x)= \Pr[Y \leq y |X=x]$ denote the conditional distribution
function of $Y$ given $X=x$. The conditional $\tau$-quantile of $Y$
given $X=x$ is $Q_Y(\tau | x) = \inf\{y: F_{Y}(y | x)
> \tau\}$ for some $\tau \in (0,1).$ We refer to $Q_{Y}(\tau|x)$, viewed as a function of $x$, as the
$\tau$-quantile regression function. This function measures the
effect of covariates on outcomes, both at the center and at the
upper and lower tails of the outcome distribution. A conditional
$\tau$-quantile is \textit{extremal} whenever the probability index
$\tau$ is either low or high in a sense that we will make more
precise below. Without loss of generality, we focus the discussion
on low quantiles.
Consider the classical linear functional form for the conditional
quantile function of $Y$ given $X=x$:
\begin{eqnarray}\label{star}
Q_Y(\tau|x) = x'\beta(\tau), \quad \text{ for all } \tau \in
\mathcal{I}=(0, \eta], \text{ for some } \eta \in (0,1],
\end{eqnarray}
and for every $x \in \mathbf{X},$ the support of $X$. This linear
functional form is flexible in the sense that it has good
approximation properties. Indeed, given an original regressor
$X^*$, the final set of regressors $X$ can be formed as a vector of
approximating functions. For example, $X$ may include power
functions, splines, and other transformations of $X^*$.
Given a sample of $T$ observations $\{ Y_t, X_t, t=1,...,T\}$, the
$\tau$-quantile QR estimator $\widehat \beta(\tau)$ solves:
\begin{align}\begin{split}\label{canonical1} \widehat \beta(\tau) \in \arg \min_{\beta
\in \Bbb{R}^d } \quad \sum_{t=1}^T \rho_{\tau} \left( Y_t - X_t' \beta
\right),
\end{split}\end{align}
where $\rho_\tau(u) = (\tau - 1(u<0))u$ is the asymmetric absolute deviation
function. The median $\tau=1/2$ case
of (\ref{canonical1}) was introduced by
\citeasnoun{laplace:1818} and the general quantile formulation (\ref{canonical1}) by Koenker and Bassett (1978).
QR coefficients $\widehat \beta(\tau)$ can be seen as order statistics
in the regression setting. Accordingly, we will refer to $\tau T$
as the order of the $\tau$-quantile. A sequence of quantile index-sample size pairs
$\{\tau_T, T\}_{T=1}^\infty$ is said to be an {\it extreme order}
sequence if $\tau_T
\searrow 0$ and $ \tau_T T \rightarrow k \in (0, \infty)$ as $T \to \infty$;
an {\it intermediate order} sequence if $\tau_T
\searrow 0$ and $ \tau_T T \rightarrow \infty$ as $T \to \infty$; and a {\it central order} sequence if $\tau_T$
is fixed as $T \rightarrow \infty$. Each type of sequence
leads to different asymptotic approximations to the finite-sample
distribution of the QR estimator. The extreme order
sequence leads to an extreme value (EV) law in large samples, whereas the intermediate and central
sequences lead to normal laws. As we saw in Figure 1, the EV law
provides a better approximation to the finite sample law of the QR
estimator than the normal law.
\subsection{Pareto-type or Regularly Varying Tails}
In order to develop inference theory for extremal QR, we assume
the tails of the conditional distribution of the outcome
variable have Pareto-type behavior, as we formally state in the
next subsection. In this subsection, we recall and discuss the
concept of Pareto-type tails. The (lower) tail of a distribution
function has Pareto-type behavior if it decays approximately as a
power function, or more formally, a regularly varying function. The
tails of the said form
are prevalent in economic data, as discovered by V. Pareto in 1895.\footnote{ Pareto called
the tails of this form ``A Distribution Curve for Wealth and
Incomes." Further empirical substantiation has been given by
\citeasnoun{sen}, \citeasnoun{zipf}, \citeasnoun{mandelbrot}, and
\citeasnoun{fama:65}, among others. The mathematical theory of
regular variation in connection to extreme value theory has been
developed by Karamata, Gnedenko, and de Haan.} Pareto-type tails
encompass or approximate a rich variety of tail behavior, including
that of thick-tailed and thin-tailed distributions, having either
bounded or unbounded support.
More formally, consider a random variable $Y$ and define a random
variable $U$ as $U \equiv Y$, if the lower end-point of the support
of $Y$ is $-\infty$, or $U \equiv Y-Q_Y(0)$, if the lower end-point
of the support of $Y$ is $Q_Y(0)>-\infty$. The quantile function of
$U$, denoted by $Q_U$, then has lower end-point $Q_U(0)= {-\infty}$
or $Q_U(0)=0$. The assumption that the quantile function $Q_U$ and
its distribution function $F_U$ exhibit Pareto-type behavior in
the tails can be formally stated as the following two equivalent
conditions:\footnote{ The notation $a \sim b$ means that $a/b
\rightarrow 1$ as appropriate limits are taken. } \begin{eqnarray}\label{power 1}
Q_U(\tau) &\sim & L(\tau) \cdot \tau^{-\xi} \ \ \text{ as } \tau
\searrow 0,\label{power 2} \\
F_U(u) &\sim& \bar L(u) \cdot u^{-1/\xi} \ \text{ as } u \searrow
Q_U(0), \end{eqnarray} for some real number $\xi \neq 0$, where $L(\tau)$ is a
nonparametric slowly-varying function at $0$, and $\bar L(u)$ is a
nonparametric slowly-varying function at $Q_U(0)$.\footnote{A
function $u \mapsto L(u)$ is said to be slowly-varying at $s$ if
$\lim_{l \searrow s}[L(l)/L(ml)] = 1$ for any $m>0$.} The leading
examples of slowly-varying functions are the constant function and
the logarithmic function. The number $\xi$ defined in (\ref{power
1}) and (\ref{power 2}) is called the EV index.
The absolute value of $\xi$ measures the heavy-tailedness of the
distribution. A distribution $F_Y$ with Pareto-type tails
necessarily has a finite lower support point if $\xi<0$ and a
infinite lower support point if $\xi>0$. Distributions with $\xi>0$
include stable, Pareto, Student's $t$, and many other distributions.
For example, the $t$ distribution with $\nu$ degrees of freedom has
$\xi=1/\nu$ and exhibits a wide range of tail behavior. In
particular, setting $\nu =1$ yields the Cauchy distribution which
has heavy tails with $\xi=1$, while setting $\nu = 30$ gives a
distribution which has light tails with $\xi=1/30$, and which is
very close to the normal distribution. On the other hand,
distributions with $\xi<0$ include the uniform, exponential,
Weibull, and many other distributions.
It should be mentioned that the case of $\xi=0$ corresponds to the
class of rapidly varying distribution functions. These
distribution functions have exponentially light tails, with the
normal and exponential distributions being the chief examples. To
simplify the exposition, we do not discuss this case explicitly.
However, since the limit distributions of the main statistics are
continuous in $\xi$, including at $\xi=0$, inference theory for the
case of $\xi=0$ is also included by taking $\xi \to 0$.
\subsection{The Extremal Conditional Quantile Model and Sampling Conditions}
With these notions in mind, our main assumption is that the response
variable $Y$, transformed by some auxiliary regression line,
$X'\beta_e$, has Pareto-type tails with EV index $\xi$.
\begin{assumption} The conditional quantile function of $Y$ given $X=x$ satisfies equation (\ref{star}) a.s.
Moreover, there exists an auxiliary extremal regression parameter
$\beta_e \in \mathbb{R}^d$, such that the disturbance $V \equiv Y
-X'\beta_e$ has end-point $s=0$ or $s=-\infty$ a.s., and its
conditional quantile function $Q_V(\tau|x)$ satisfies the following
tail-equivalence relationship:
$$Q_V(\tau|x) \sim x'\gamma \cdot Q_U(\tau), \text{ as } \tau
\searrow 0, \text{ uniformly in } x \in \mathbf{X} \subseteq
\mathbb{R}^d,$$ for some quantile function $Q_U(\tau)$ that exhibits
Pareto-type tails with EV index $\xi$ (i.e., it satisfies
(\ref{power 2})), and some vector parameter $\gamma$ such that
$E[X]'\gamma=1$.
\end{assumption}
Since this assumption only affects the tails, it allows covariates
to affect the extremal quantile and the central quantiles very
differently. Moreover, the local effect of covariates in the tail
is approximately given by $\beta(\tau) \approx \beta_e + \gamma
Q_U(\tau)$, which allows for a differential impact of
covariates across various extremal quantiles.
\begin{assumption} The conditional quantile density function $\partial Q_V(\tau|x)/\partial \tau$ exists
and satisfies the tail equivalence relationship $\partial
Q_V(\tau|x)/\partial \tau \sim x'\gamma \cdot
\partial Q_U(\tau)/ \partial \tau$ as $\tau \searrow 0$, uniformly in $x
\in \mathbf{X},$ where $\partial Q_U(\tau)/
\partial \tau$ exhibits Pareto-type tails as $\tau \searrow 0$ with EV index $\xi+1$.
\end{assumption}
Assumption C2 strengthens C1 by imposing the existence and
Pareto-type behavior of the conditional quantile density function.
We impose C2 to facilitate the derivation of the main inferential
results.
The following sampling conditions will be imposed.
\begin{assumption} The regressor vector $X=(1,Z')'$ is such that
it has a compact support $\mathbf{X}$, the matrix $E[XX']$ is
positive definite, and its distribution function $F_X$ satisfies a
non-lattice condition stated in the mathematical appendix (this
condition is satisfied, for instance, when $Z$ is absolutely
continuous).
\end{assumption}
Compactness is needed to ensure the continuity and robustness of
the mapping from extreme events in $Y$ to the extremal QR
statistics. Even if $X$ is not compact, we can select the data for
which $X$ belongs to a compact region. The non-degeneracy condition
of $E[XX']$ is standard and guarantees invertibility. The non-lattice
condition is required for the existence of the finite-sample density
of QR coefficients. It is needed even asymptotically because the
asymptotic distribution theory of extremal QR closely resembles the
finite-sample theory for QR, which is not a surprise given the rare
nature of events that have a probability of order $1/T$.
We assume the data are either i.i.d. or weakly dependent.
\begin{assumption} The sequence $\{W_t\}$ with
$W_t =(V_t,X_t)$ and $V_t$ defined in C1, forms a stationary,
strongly mixing process with a geometric mixing rate, that is, for
some $C>0$
$$\sup_{t} \sup_{A \in \mathcal{A}_{t}, B \in \mathcal{B}_{t+m}} |P(A
\cap B) - P(A) P(B)| \exp(C m) \to 0 \text{ as } m \to \infty,$$
where $\mathcal{A}_{t} = \sigma(W_{t},
W_{t-1}, ...)$ and $\mathcal{B}_{t} = \sigma(W_{t},W_{t+1}, ...)$.
Moreover, the sequence satisfies a condition that curbs clustering
of extreme events in the following sense: $P(V_t \leq K, V_{t+j}\leq
K|\mathcal{A}_{t}) \leq C P(V_t \leq K|\mathcal{A}_{t})^2$ for all
$K \in [s, \bar K]$, uniformly for all $j \geq 1$ and uniformly for
all $t\geq 1$; here $C>0$ and $\bar{K}>s$ are some constants.
\end{assumption}
A special case of this condition is when the sequence of variables
$\{(V_t, X_t), t \geq 1\}$, or equivalently $\{(Y_t, X_t), t \geq
1\}$, is independent and identically distributed. The assumption of
mixing for $\{(V_t, X_t), t \geq 1\}$ is standard in econometric
analysis (White 1990), and it is equivalent to the assumption of
mixing of $\{(Y_t, X_t), t \geq 1\}$. The non-clustering condition
is of the \citeasnoun{meyer}-type and states that the probability of
two extreme events co-occurring at nearby dates is much lower than
the probability of just one extreme event. For example, it assumes
that a large market crash is not likely to be immediately followed
by another large crash. This assumption leads to limit distributions
of QRs as if independent sampling had taken place. The plausibility
of the non-clustering assumption is an empirical matter. We
conjecture that our primary inference method based on subsampling is
valid more generally, under conditions that preserve the rates of
convergence of QR statistics and ensure existence of their
asymptotic distributions. Finally we note that the assumptions made
here could be relaxed in certain directions for some of the results
stated below, but we decided to state a single set of sufficient
assumptions for all the results.
\subsection{Overview and Discussion of Inferential Results}
We begin by briefly revisiting the classical non-regression case
to describe some intuition and the key obstacles to
performing feasible inference in our more general regression case.
Then we will describe our main inferential results for the
regression case. It is worth noting that our main inferential
methods, based on self-normalized statistics, are new and of
independent interest even in the classical non-regression case.
Recall the following classical result on the
limit distribution of the extremal sample quantiles $\widehat{Q}_Y(\tau)$ \cite{gnedenko1}: for any integer $k\geq 1$ and $\tau = k/T$, as $T \to
\infty$,
\begin{eqnarray}\label{gnedenko} & \widehat Z_T(k) = A_T(\widehat{Q}_Y(\tau)-Q_Y(\tau)) \to_d \widehat Z_\infty(k) = \Gamma_k^{-\xi} - k^{-\xi},\end{eqnarray} where
\begin{align}\begin{split}
A_T = 1/ Q_U(1/T), \ \ \ \Gamma_k = \mathcal{E}_1 +...+
\mathcal{E}_k, \\
\end{split}\end{align}
and $ (\mathcal{E}_1, \mathcal{E}_2, ... )$ is an independent and
identically distributed sequence of standard exponential variables.
We refer to $\widehat Z_T(k)$ as the canonically normalized (CN)
statistic because it depends on the scaling constant $A_T$. The
variables $\Gamma_k$, entering the definition of the EV
distribution, are {\it gamma} random variables. The limit
distribution of the $k$-th order statistic is therefore a
transformation of a gamma variable. The EV distribution is not
symmetric and may have significant (median) bias; it has finite
moments if $\xi<0$ and has finite moments of up to order $1/\xi$ if
$\xi >0$. The presence of median bias motivates the use of
median-bias correction techniques, which we discuss in the
regression case below.
Although very powerful, this classical result is not feasible for
purposes of inference on $Q_Y(\tau)$, since the scaling constant
$A_T$ is generally not possible to estimate consistently
\cite{bertailetal}. One way to deal with this problem is to add
strong parametric assumptions on the non-parametric, slowly-varying
function $L(\cdot)$ in equation (\ref{power 2}) in order to
estimate $A_T$ consistently. For instance, suppose that $Q_{U}(\tau)
\sim L \tau^{-\xi}$. Then one can estimate $\xi$ by the classical
Hill or Pickands estimators,
and $L$ by $\widehat L = ( \widehat Q_Y(2 \tau)- \widehat Q_Y(
\tau)))/ (2^{-\widehat\xi}-1) \tau^{-\widehat \xi})$. We develop the
necessary theoretical results for the regression analog of this
approach, although we will not recommend it as our preferred method.
Our preferred and main proposal to deal with the aforementioned
infeasibility problem is to consider the asymptotics of
the self-normalized (SN) sample quantiles \begin{align}\begin{split} \label{EV
approximation} & Z_T(k) = \mathcal{A}_T (\widehat
Q_Y(\tau)-Q_Y(\tau)) \to_d Z_{\infty}(k)= \frac{\sqrt{k}
(\Gamma_k^{-\xi} - k^{-\xi})}{\Gamma_{ mk }^{-\xi} - \Gamma_{ k
}^{-\xi} },\end{split}\end{align} where for $m> 1$ such that $mk$
is an integer,
\begin{align}\begin{split}\label{SN}
\mathcal{A}_T = \frac{\sqrt{\tau T}}{\widehat Q_Y( m\tau) - \widehat
Q_Y(\tau) } .
\end{split}\end{align}
Here, the scaling factor $\mathcal{A}_T$ is completely a function of
data and therefore \textit{feasible}. Moreover, we completely
avoid the need for consistent estimation of $A_T$. This is
convenient because we are not interested in this normalization
constant per se. The limit distribution in (\ref{EV approximation})
only depends on the EV index $\xi$, and its quantiles can be easily
obtained by simulation. In the regression setting, where the limit
law is a bit more complicated, we develop a form of subsampling to
perform both practical and feasible inference.
Let us now turn to the regression case. Here, we can also consider a
canonically-normalized QR statistic (CN-QR):
\begin{eqnarray} \label{equation: EV limit distribution QR}
\widehat Z_{T}(k) := A_T \Big ( \widehat \beta (\tau) -
\beta(\tau) \Big ) \text{ for } A_T := 1/ Q_{U}(1/T);\end{eqnarray} and a
self-normalized QR (SN-QR) statistic: \begin{eqnarray} Z_{T}(k) := \mathcal{A}_T
\Big ( \widehat \beta (\tau) - \beta(\tau) \Big ) \text{ for }
\mathcal{A}_T := \frac{\sqrt{\tau T}}{\bar X_T'(
\widehat\beta(m\tau)-\widehat\beta(\tau))},
\end{eqnarray} where $\bar X_T = \sum_{t=1}^T X_t/T$ and $m$ is a real number
such that $\tau T (m-1)>d$. The first statistic uses an infeasible
canonical normalization $A_T$, whereas the second statistic uses a
feasible random normalization. First, we show that \begin{align}\begin{split}
\label{result sub} \widehat Z_{T}(k) \to_d \widehat Z_{\infty}(k)
\end{split}\end{align}
where for $\chi=1$ if $\xi < 0$ and $\chi = -1$ if $\xi > 0,$
\begin{align}\begin{split} \label{result cn} & \widehat Z_{\infty}(k) := \
\chi \cdot \underset{z \in \mathbb{R}^d}{\arg \min } \Bigg [- k
E[X]' (z+k^{-\xi}\gamma) + \ \sum_{t=1}^{\infty} \{ \mathcal{X}_t'(z
+k^{-\xi}\gamma) - \chi \cdot \Gamma_t^{-\xi}\cdot \mathcal{X}_t'\gamma
\}_+\ \Bigg]
\end{split}\end{align} where $ \{\Gamma_1, \Gamma_2, ...\} := \{
\mathcal{E}_1, \mathcal{E}_1 + \mathcal{E}_2 ,...\};$
$\{\mathcal{E}_1, \mathcal{E}_2, ... \}$ is an iid sequence of
exponential variables that is independent of $\{\mathcal{X}_1,
\mathcal{X}_2, ...\}$, an iid sequence with distribution $F_{X}$;
and $\{y\}_+ := \max(0, y)$.
Furthermore, we show that
\begin{eqnarray}\label{result main} {Z}_{T}(k) \to_d {Z}_{\infty}(k): = \frac{ \sqrt{k} \widehat Z_{\infty}(k)}
{E[X]'(\widehat Z_{\infty}(m k )-\widehat Z_{\infty}(k)) + \chi\cdot
(m^{-\xi}-1)k^{-\xi}}. \end{eqnarray}
The limit laws here are more complicated than in the
non-regression case, but they share some common features. Indeed,
the limit laws depend on the variables $\Gamma_i$ in a crucial way, and
are not necessarily centered at
zero and can have significant first order median biases. Motivated
by the presence of the first order bias, we develop bias corrections for
the QR statistics in the next section. Moreover, just as in the
non-regression case, the limit distribution of the CN-QR statistic
in (\ref{result cn}) is
generally infeasible for inference purposes.
We need to know or estimate the scaling constant $A_T$, which is
the reciprocal of the extremal quantile of the variable $U$ defined in C1. That is, we require an estimator $\widehat A_T$ such
that $ \widehat A_T/A_T \to_p 1,$ which is not feasible unless the tail
of $U$ satisfies additional strong parametric restrictions. We
provide additional restrictions below that facilitate estimation of
$A_T$ and hence inference based on CN-QR, although this is not our
preferred inferential method.
Our main and preferred proposal for inference is based on the SN-QR
statistic, which does not depend on $A_T$. We estimate the
distribution of this statistic using either a variation of
subsampling or an analytical method. A key ingredient here is the
feasible normalizing variable $\mathcal{A}_T$, which is randomly
proportional to the canonical normalization $A_T$, in the sense that
$\mathcal{A}_T/A_T$ is a random variable in the limit.\footnote{The
idea of feasible random normalization has been used in other
contexts (e.g. t-statistics). In extreme value theory,
\citeasnoun{Dekkers:dehaan} applied a similar random normalization
idea to extrapolated quantile estimators of intermediate order
in the non-regression setting, precisely to produce limit
distributions that can be easily used for inference. In time series,
\citeasnoun{kiefer:vogelsang} have used feasible inconsistent
estimates of the variance of asymptotically normal estimators. } An
advantage of the subsampling method over the analytical methods is
that it does not require estimation of the nuisance parameters $\xi$
and $\gamma$.
Our subsampling approach is different from conventional
subsampling in the use of recentering terms and random
normalization. Conventional subsampling that uses recentering by the
full sample estimate $\widehat \beta(\tau)$ is not consistent when
that estimate is diverging; and here we indeed have $A_T \to 0$ when
$\xi>0$. Instead, we recenter by intermediate order
QR estimates in subsamples, which will diverge at a slow enough speed to
estimate the limit distribution of SN-QR consistently. Thus, our
subsampling approach explores the special relationship between the
rates of convergence/divergence of extremal and intermediate QR
statistics and should be of independent interest even in a
non-regression setting.
This paper contributes to the existing literature by introducing
general feasible inference methods for extremal quantile regression.
Our inferential methods rely in part on the limit results in
\citeasnoun{victor:annals}, who derived EV limit laws for CN-QR
under the extreme order condition $\tau T \to k>0$. This theory,
however, did not lead directly to any feasible, practical inference
procedure. \citeasnoun{feigin}, \citeasnoun{victor:nex}, \citeasnoun{portnoy:jur},
and \citeasnoun{knight:linear} provide
related limit results for canonically normalized linear programming
estimators where $\tau T \searrow 0$, all in different contexts and
at various levels of generality. These limit results likewise did not
provide feasible inference theory. The linear programming estimator
is well suited to the problem of estimating finite deterministic
boundaries of data, as in image processing and other technometric
applications. In contrast, the current approach of taking $\tau T
\to k>0$ is more suited to econometric applications, where interest
focuses on the ``usual" quantiles located near the minimum or
maximum and where the boundaries may be unlimited. However, some of
our theoretical developments are motivated by and build upon this
previous literature. Some of our proofs rely on the elegant
epi-convergence framework of \citeasnoun{geyer} and
\citeasnoun{knight}.
\section{Inference and Median-Unbiased Estimation Based on Extreme Value Laws}
This section establishes the main results that underlie our
inferential procedures.
\subsection{Extreme Value Laws for CN-QR and SN-QR Statistics}
Here we verify that that the CN-QR statistic
$\widehat Z_{T}(k)$ and SN-QR statistic $Z_T(k)$ converge to the limit variables
${\widehat Z}_{\infty}(k)$ and $Z_{\infty}(k)$, under the condition
that $\tau T \to k >0$ as $T \to \infty$.
\begin{theorem}[Limit Laws for Extremal SN-QR and CN-QR] Suppose conditions C1, C3 and C4 hold. Then as $\tau T \to k
>0$ and $T \to \infty$, (1) the SN-QR statistic of order $k$ obeys $$ Z_{T}(k) \to_d {Z}_{\infty}(k),$$
for any $m$ such that $k (m-1)>d$, and (2) the CN-QR statistic of order $k$ obeys $$
\widehat Z_{T}(k) \to_d {\widehat Z}_{\infty}(k).$$
\end{theorem}
\begin{remark} The condition that $k(m-1)>d$ in the definition of SN-QR ensures that $\beta(m\tau) \neq \beta(\tau)$ and therefore the
normalization by $\mathcal{A}_T$ is well defined. This is a
consequence of Theorem 3.2 in \citeasnoun{bassett:1982} and
existence of the conditional density of $Y$ imposed in assumption
C2. Result 1 on SN-QR statistics is the main new result that we will
exploit for inference. Result 2 on CN-QR statistics is needed
primarily for auxiliary purposes. \citeasnoun{victor:annals}
presents some extensions of result 2.
\end{remark}
\begin{remark} When $Q_Y(0|x)> - \infty$, by C1
$Q_Y(0|x)$ is equal to $x'\beta_e$ and is the conditional lower boundary of $Y$. The proof
of Theorem 1 shows that
$$
A_T (\widehat \beta(\tau) -\beta_e) \to_d \widetilde Z_{\infty}(k):= \widehat Z_{\infty}(k) - k^{-\xi}
\text{ and }
\mathcal{A}_T (\widehat \beta(\tau) -\beta_e) \to_d \widetilde Z_{\infty}(k)/( \widetilde Z_{\infty}(mk) - \widetilde Z_{\infty}(k)).
$$
We can use these results and analytical and subsampling methods presented below to perform median unbiased estimation and inference on the boundary parameter $\beta_e$.\footnote{ To estimate the critical values, we can use either analytical or subsampling methods presented below, with the difference that in subsampling we need to recenter by the full sample estimate $\widehat \beta_e= \widehat \beta(1/T)$.}
\end{remark}
\subsection{Generic Inference and Median-Unbiased Estimation}
We outline two procedures for conducting inference and constructing
asymptotically median unbiased estimates of linear functions $\psi'
\beta(\tau) $ of the coefficient vector $\beta(\tau)$, for some
non-zero vector $\psi$.
\emph{\textbf{1. Median-Unbiased Estimation and Inference Using
SN-QR.}} By Theorem 1, $\psi' \mathcal{A}_T (\widehat \beta(\tau)-
\beta(\tau)) \to_d \psi' Z_{\infty}(k).$ Let $c_{\alpha}$ denote the
$\alpha$-quantile of $\psi' Z_{\infty}(k)$ for $0 < \alpha \leq .5$.
Given $\widehat c_{\alpha}$, a consistent estimate of $c_{\alpha}$, we
can construct an asymptotically median-unbiased estimator and a
$(1-\alpha)$\%-confidence interval for $\psi' \beta(\tau)$ as $$
\psi' \widehat \beta(\tau) - \widehat c_{1/2}/\mathcal{A}_T \ \text{ and } \
[\psi' \widehat \beta(\tau) - \widehat c_{1-\alpha/2}/\mathcal{A}_T, \psi'
\widehat \beta(\tau) - \widehat c_{\alpha/2}/\mathcal{A}_T],$$ respectively.
The bias-correction term and the limits of the confidence interval
depend on the random scaling $\mathcal{A}_T$. We provide consistent
estimates of $c_{\alpha}$ in the next section.
\begin{theorem}[Inference and median-unbiased estimation using SN-QR] Under the conditions
of Theorem 1, suppose we have $\widehat c_{\alpha}$ such that $\widehat
c_{\alpha} \to_p c_{\alpha}$. Then,
$$ \lim_{T \to \infty} P\{ \psi' \widehat \beta(\tau) - \widehat
c_{1/2}/\mathcal{A}_T \leq \psi' \beta(\tau)\} = 1/2$$ and
$$\lim_{T \to \infty} P\{\psi' \widehat \beta(\tau) - \widehat
c_{1-\alpha/2}/\mathcal{A}_T \leq \psi' \beta(\tau) \leq \psi' \widehat
\beta(\tau) - \widehat c_{\alpha/2}/\mathcal{A}_T\} = 1-\alpha.$$
\end{theorem}
\vspace{.3in }
\emph{\textbf{2. Median Unbiased Estimation and Inference Using
CN-QR.}} By Theorem 1, $
\psi' A_T (\widehat \beta(\tau)- \beta(\tau)) \to_d
\psi' \widehat Z_{\infty}(k).$ Let $c_{\alpha}'$ denote the
$\alpha$-quantile of $\psi' \widehat Z_{\infty}(k)$ for $0 < \alpha
\leq .5$. Given $\widehat A_T$, a consistent estimate of $A_T$, and
$\widehat c_{\alpha}'$, a consistent estimate of $c_{\alpha}'$, we can
construct an asymptotically median-unbiased estimator and a
$(1-\alpha)$\%-confidence interval for $\psi' \beta(\tau)$ as
$$\psi' \widehat \beta(\tau) - \widehat c_{1/2}'/\widehat{A}_T \ \text{ and
} \ \ [\psi' \widehat \beta(\tau) - \widehat c_{1-\alpha/2}'/\widehat{A}_T,
\psi' \widehat \beta(\tau) - \widehat c_{\alpha/2}'/\widehat{A}_T],$$
respectively.
As mentioned in Section 2, construction of consistent estimates of
$A_T$ requires additional strong restrictions on the underlying
model as well as additional steps in estimation. For example,
suppose the nonparametric slowly varying component $L(\tau)$ of
$A_T$ is replaced by a constant $L$, i.e. suppose that as $\tau
\searrow 0$
\begin{eqnarray}\label{restrictive}
1/Q_{U}(\tau) = L \cdot \tau^{\xi} \cdot (1 + \delta(\tau)) \text{
for some } \ L \in \Bbb{R}, \text{ where } \ \delta(\tau)\to 0.
\end{eqnarray}
We can estimate the constants $L$ and $\xi$ via Pickands-type
procedures:
\begin{eqnarray}\label{tail_estimates}
\widehat \xi= \frac{-1}{\ln 2} \ln \frac{{\bar X_T}'(\widehat \beta(4
\tau_T)- \widehat \beta( \tau_T))}{\bar X_T'(\widehat \beta(2
\tau_T)- \widehat \beta(\tau_T))}
\text{ and }
\widehat L = \frac{\bar X_T'(\widehat \beta(2 \tau_T)- \widehat \beta( \tau_T))
}
{
(2^{-\widehat\xi}-1)\cdot \tau^{-\widehat \xi}
},
\end{eqnarray}
where $\tau_T$ is chosen to be of an intermediate order, $\tau_T T
\to \infty$ and $\tau_T \to 0$. Theorem 4 in Chernozhukov
\citeyear{victor:annals} shows that under C1-C4, condition
(\ref{restrictive}), and additional conditions on the sequence
$(\delta(\tau_T), \tau_T)$,\footnote{ The rate convergence of $\widehat
\xi$ is $\max[\frac{1}{\sqrt{\tau_T T}}, \ln \delta(\tau_T)]$, which
gives the following condition on the sequence $(\delta(\tau_T),
\tau_T)$ : $
\max[\frac{1}{\sqrt{\tau_T T}}, \ln \delta(\tau_T)] = o(1/\ln T)$.}
$\widehat \xi = \xi + o(1/\ln T)$ and $\widehat L \to_p L$, which produces
the required consistent estimate $\widehat A_T = \widehat L
(1/T)^{-\widehat\xi}$ such that $\widehat A_T/ A_T \to_p 1.$ These
additional conditions on the tails of $Y$ and on the sequence
$(\delta(\tau_T), \tau_T)$ highlight the drawbacks of this inference
approach relative to the previous one.
We provide consistent estimates of $c_{\alpha}'$ in the next
section.
\begin{theorem}[Inference and Median-Unbiased Estimation using CN-QR]
Assume the conditions of Theorem 1 hold. Suppose
that we have $\widehat A_T$ such that $\widehat A_T/A_T \to_p 1$ and $\widehat
c_{\alpha}'$ such that $\widehat c_{\alpha} \to_p c_{\alpha}'$. Then,
$$\lim_{T \to \infty}P\{ \psi' \widehat \beta(\tau) -\widehat c_{1/2}'/\widehat
A_T \leq \psi' \beta(\tau)\} = 1/2$$ and $$ \lim_{T \to
\infty}P\{\psi' \widehat \beta(\tau) - \widehat c_{1-\alpha/2}'/\widehat A_T
\leq \psi' \beta(\tau) \leq \psi' \widehat \beta(\tau) - \widehat
c_{\alpha/2}/\widehat A_T\} = 1-\alpha.$$
\end{theorem}
\section{Estimation of Critical Values}
\subsection{Subsampling-Based Estimation of Critical Values.} Our resampling
method for inference uses subsamples to estimate the distribution of
SN-QR, as in standard subsampling. However, in contrast to the
subsampling, our method bypasses estimation of the unknown
convergence rate $A_T$ by using self-normalized statistics. Our
method also employs a special recentering that allows us to avoid the inconsistency of standard subsampling due to
diverging QR statistics when $\xi>0$.
The method has the following steps. First, consider all subsets of
the data $\{W_t=(Y_t,X_t),$ $t =1,...,T\}$ of size $b$; if $\{W_t\}$
is a time series, consider $B_T = T -b +1$ subsets of size $b$ of
the form $\{ W_i, ..., W_{i+ b -1}\}$. Then compute the analogs of
the SN-QR statistic, denoted $\widehat V_{i, b}$ and defined below
in equation (\ref{subsample_analog}), for each $i$-th subsample for
$i=1,..., B_{\scriptscriptstyle T}$. Second, obtain $ \widehat c_\alpha$ as the sample
$\alpha$-quantile of $\{ \widehat V_{i,b,T}, i=1,...,B_T \}$. In
practice, a smaller number $B_T$ of randomly chosen subsets can be
used, provided that $B_T \rightarrow \infty$ as $ T \rightarrow
\infty$. (See Section 2.5 in \citeasnoun{romano:book}.)
\citeasnoun{romano:book} and \citeasnoun{bertailetal} provide rules
for the choice of subsample size $b$.
The SN-QR statistic for the full sample of size $T$ is:
\begin{eqnarray}
V_{T} : = \mathcal{A}_T \psi'( \widehat \beta_{\scriptscriptstyle T}(\tau_{\scriptscriptstyle T})
- \beta(\tau_{\scriptscriptstyle T})) \text{ for } \mathcal{A}_T = \frac{
\sqrt{\tau_{\scriptscriptstyle T} T}}{ \bar X_T ' \big (\widehat \beta (m \tau_{\scriptscriptstyle
T}) - \widehat \beta ( \tau_{\scriptscriptstyle T})\big )}, \end{eqnarray} where we can set
$m= (d+p)/(\tau_{\scriptscriptstyle T} T) +1 = (d+p)/k +1 +o(1)$, where $p\geq 1$ is the spacing
parameter, which we set to $5$.\footnote{Variation of this parameter from $p=2$ to
$p=20$ yielded similar results in our Monte-Carlo experiments.} In this section we write $\tau_{\scriptscriptstyle T}$ to
emphasize the theoretical dependence of the quantile of interest
$\tau$ on the sample size. In each $i$-th subsample of size $b$, we
compute the following analog of $V_T$:
\begin{eqnarray}\label{subsample_analog}
\widehat V_{i,b,\scriptscriptstyle T} : = \mathcal{A}_{i, b, \scriptscriptstyle T} \psi'(
\widehat \beta_{i, b, \scriptscriptstyle T}(\tau_{\scriptscriptstyle b}) - \widehat \beta(\tau_{\scriptscriptstyle
b})) \text{ for }
\mathcal{A}_{i,b, \scriptscriptstyle T} : = \frac{\sqrt{\tau_{\scriptscriptstyle b} b}}{ \bar
X_{i,b,T}
' \big(\widehat \beta_{i,b, \scriptscriptstyle T} (m \tau_{\scriptscriptstyle b}) - \widehat \beta_{i,
b, \scriptscriptstyle T} ( \tau_{\scriptscriptstyle b})\big )},
\end{eqnarray}
where $\widehat \beta(\tau)$ is the $\tau$-quantile regression
coefficient computed using the full sample, $\widehat \beta_{i,b, \scriptscriptstyle
T} ( \tau)$ is the $\tau$-quantile regression coefficient computed
using the $i$-th subsample, $\bar X_{i,b,T}$ is the sample mean of
the regressors in the $i$th subsample, and $\tau_b := \left(\tau_{\scriptscriptstyle
T} T \right)/b$.\footnote{ In practice, it is reasonable to use the
following finite-sample adjustment to $\tau_b$:
$\tau_b = \min[ \left(\tau_T T\right)/b, .2]$ if $\tau_T <.2$, and
$\tau_b = \tau_T$ if $\tau_T \geq .2$.
The idea is that $\tau_T$ is judged to be non-extremal if
$\tau_T>.2$, and the subsampling procedure reverts to central
order inference. The truncation of $\tau_b$ by .2 is a finite-sample
adjustment that restricts the key statistics $\widehat V_{i,b,T}$ to
be extremal in subsamples. These finite-sample adjustments do not
affect the asymptotic arguments.} The determination of $\tau_b$ is a
critical decision that sets apart the extremal order approximation
from the central order approximation. In the latter case, one sets
$\tau_b = \tau_{\scriptscriptstyle T}$ in subsamples. In the extreme order
approximation, our choice of $\tau_b$ gives the same extreme order
of $\tau_b b$ in the subsample as the order of $\tau_TT$ in the full
sample.
Under the additional parametric assumptions on the tail behavior
stated earlier, we can estimate the quantiles of the limit
distribution of CN-QR using the following procedure: First, create
subsamples $i=1,...,B_T$ as before and compute in each subsample:
$\widetilde V_{i, b, \scriptscriptstyle T} : = \widehat A_{ b} \psi'( \widehat
\beta_{i, b, \scriptscriptstyle T}(\tau_{\scriptscriptstyle b}) - \widehat \beta(\tau_{\scriptscriptstyle b})),$
where $\widehat A_b$ is any consistent estimate of $A_b$. For
example, under the parametric restrictions specified in
(\ref{restrictive}), set $\widehat A_b=\widehat L b^{-\widehat \xi}$ for
$\widehat L$ and $\widehat \xi$ specified in (\ref{tail_estimates}). Second,
obtain $ \widehat c_\alpha'$ as the $\alpha$-quantile of $\{ \widetilde
V_{i,b,T}, i=1,...,B_T \}$.
The following theorems establish the consistency of $\widehat
c_{\alpha}$ and $\widehat c'_{\alpha}$:
\begin{theorem}[Critical Values for SN-QR by Resampling]
Suppose the assumptions of Theorems 1 and 2 hold, $b/T \to 0, b \to \infty,
T \to \infty$ and $B_T \to \infty.$ Then $\widehat
c_\alpha \to_p c_{\alpha}$.
\end{theorem}
\begin{theorem}[Critical Values for CN-QR by Resampling] Suppose
the assumptions of Theorems 1 and 2 hold, $b/T \to 0, b \to \infty,
T \to \infty, B_T \to \infty$, and
$\widehat A_b$ is such that $\widehat A_b/ A_b \to 1$. Then $\widehat
c_\alpha' \to_p c'_{\alpha}$.
\end{theorem}
\begin{remark} Our subsampling method based on CN-QR or SN-QR
produces consistent critical values in the regression case, and
may also be of independent interest in the non-regression case. Our method
differs from conventional subsampling in several respects.
First, conventional subsampling uses fixed normalizations $A_T$ or their consistent
estimates. In contrast, in the case of SN-QR we use the random normalization $\mathcal{A}_T$,
thus avoiding estimation of $A_T$. Second, conventional subsampling recenters by the
full sample estimate $\widehat \beta(\tau_T)$. Recentering in this way requires $A_b / A_T \to 0$ \
for obtaining consistency (see Theorem 2.2.1 in \citeasnoun{romano:book}),
but here we have $A_b / A_T \to \infty$ when $\xi >0$. Thus, when $\xi>0$
the extreme order QR statistics $\widehat \beta(\tau_T)$ diverge
when $\xi>0$, and the conventional subsampling is inconsistent. In contrast,
to overcome the inconsistency, our approach instead uses $\widehat
\beta(\tau_b)$ for recentering. This statistic itself may diverge,
but because it is an intermediate order QR statistic, the speed of its divergence is strictly slower than that of
$A_T$. Hence our method of recentering exploits the special structure of order
statistics in both the regression and non-regression cases.
\end{remark}
\subsection{Analytical Estimation of Critical Values} Analytical inference
uses the quantiles of the limit distributions found in Theorem 1.
This approach is much more demanding in practice than the previous
subsampling method.\footnote{The method developed below is also of
independent interest in situations where the limit distributions
involve Poissson processes with unknown nuisance parameters, as, for
example, in Chernozhukov and Hong (2004).}
Define the following random vector: \begin{eqnarray}\label{define draw} \widehat
Z^*_{\infty}(k) = \widehat{\chi} \cdot \underset{z \in
\mathbb{R}^d}{\arg \min } \Bigg [- k \bar{X}_T'
(z+k^{-\widehat{\xi}}\widehat{\gamma}) + \ \sum_{t=1}^{\infty} \{ \mathcal{X}_t'(z
+k^{-\widehat{\xi}}\widehat{\gamma}) - \widehat{\chi} \cdot
\Gamma_t^{-\widehat{\xi}}\cdot \mathcal{X}_t'\widehat{\gamma}\}_+\ \Bigg], \end{eqnarray} for
some consistent estimates $\widehat \xi$ and $\widehat \gamma$, e.g., those
given in equation (\ref{equation:tail estimators}); where $\widehat \chi
= 1$ if $\widehat \xi < 0$ and $\widehat \chi = -1$ if $\widehat \xi > 0$,
$\{\Gamma_1, \Gamma_2, ...\} = \{ \mathcal{E}_1, \mathcal{E}_1 +
\mathcal{E}_2 ,...\}$; $\{\mathcal{E}_1, \mathcal{E}_2, ... \} $ is
an i.i.d. sequence of standard exponential variables;
$\{\mathcal{X}_1, \mathcal{X}_2, ... \}$ is an i.i.d. sequence with
distribution function $\widehat F_{X}$, where $\widehat F_{X}$ is
any smooth consistent estimate of $F_X$, e.g., a smoothed empirical
distribution function of the sample $\{X_i, i
=1,...,T\}$.\footnote{We need smoothness of the distribution
regressors $\mathcal{X}_i$ to guarantee uniqueness of the solution of the
optimization problem (\ref{define draw}); a similar device is used
by \citeasnoun{hall:lad} in the context of Edgeworth expansion for
median regression. The empirical distribution function (edf) of
$X_i$ is not suited for this purpose, since it assigns point masses
to sample points. However, making random draws from the edf and
adding small noise with variance that is inversely proportional to
the sample size produces draws from a smoothed empirical
distribution function which is uniformly consistent with respect to
$F_X$.} Moreover, the sequence $\{\mathcal{X}_1, \mathcal{X}_2, ...
\}$ is independent from $\{ \mathcal{E}_1, \mathcal{E}_2 ,...\}$.
Also,
let $ {Z}^*_{\infty}(k) = \sqrt{k} \widehat Z^*_{\infty}(k)/ [ \bar
X_T'(\widehat Z^*_{\infty}(m k )-\widehat Z^*_{\infty}(k)) + \widehat
\chi(m^{-\widehat \xi}-1)k^{-\widehat \xi} ]$. The estimates $\widehat
c_{\alpha}'$ and $\widehat c_{\alpha}$ are obtained by taking
$\alpha$-quantiles of the variables $\psi' \widehat Z^*_{\infty}(k)$
and $\psi'Z^*_{\infty}(k)$, respectively. In practice, these
quantiles can only be evaluated numerically as described below.
The analytical inference procedure requires consistent estimators of
$\xi$ and $\gamma$. Theorem 4.5 of \citeasnoun{victor:annals}
provides the following estimators based on Pickands-type procedures:
\begin{eqnarray} \label{equation:tail estimators}
\widehat \xi= \frac{-1}{\ln 2} \ln
\frac{{\bar X_T}'(\widehat \beta(4 \tau_T)- \widehat \beta(
\tau_T))}{{\bar X_T}'(\widehat \beta(2\tau_T)- \widehat
\beta(\tau_T))} \text{ and } \widehat \gamma= \frac{\widehat \beta(2
\tau_T)- \widehat \beta( \tau_T)}{{\bar X_T}'(\widehat \beta(2
\tau_T)- \widehat \beta(\tau_T))}, \end{eqnarray} which is consistent if
$\tau_T T \to \infty$ and $\tau_T \to 0$.
\begin{theorem}[Critical Values for SN-QR by Analytical Method] Assume the conditions
of Theorem 1 hold. Then for any estimators of the nuisance
parameters such that $\widehat \xi \to_p \xi$ and $\widehat \gamma \to_p
\gamma$, we have that $\widehat c_{\alpha} \to_p c_{\alpha}.$
\end{theorem}
\begin{theorem}[Critical Values for CN-QR by Analytical Method] Assume
the conditions of Theorem 1 hold. Then, for any estimators of the
nuisance parameters such that $\widehat \xi \to_p \xi$ and $\widehat \gamma
\to_p \gamma$, we have that $\widehat c_{\alpha}' \to_p c_{\alpha}'.$
\end{theorem}
\begin{remark}
Since the distributions of $\widehat Z_{\infty}(k)$ and
$Z_{\infty}(k)$ do not have closed form, except in very special
cases, $\widehat c_{\alpha}'$ and $\widehat c_{\alpha}$ can be obtained
numerically via the following Monte Carlo procedure. First, for each
$i=1,...,B$ compute $\widehat Z^*_{i,\infty}(k)$ and
$Z^*_{i,\infty}(k)$ using formula (\ref{define draw}) by simulation,
where the infinite summation is truncated at some finite value $M$.
Second, take $\widehat c'_{\alpha}$ and $\widehat c_{\alpha}$ as the sample
$\alpha$-quantiles of the samples $\{\psi'\widehat
Z^*_{i,\infty}(k), i=1,...,B\}$ and $\{\psi' Z^*_{i,\infty}(k),
i=1,...,B\}$, respectively. We have found in numerical experiments
that choosing $M \geq 200$ and $B\geq 100$ provides accurate
estimates.
\end{remark}
\section{Extreme Value vs. Normal Inference: Comparisons}
\subsection{Properties of Confidence Intervals with Unknown Nuisance Parameters}
In this
section we compare the inferential performance of normal and
extremal confidence intervals (CI) using the model: $Y_t =X_t'\beta
+ U_t$, $t=1,...,500$, $d=7$, $\beta_j=1$ for $j \in \{1,...,7\}$,
where the disturbances $\{U_t\}$ are i.i.d. and follow either
(1)
a $t$ distribution with $\nu \in \{1,3,30\}$ degrees of freedom, or
(2) a Weibull distribution with the shape parameter $\alpha \in \{
1, 3, 30 \}$. These distributions have EV indexes $\xi=1/\nu \in
\{1, 1/3, 1/30\}$ and $\xi = -1/\alpha \in \{-1, -1/3, -1/30\}$,
respectively. Regressors are drawn with replacement from the
empirical application in Section 6.1 in order to match a real
situation as closely as possible.\footnote{These data as well as the
Monte-Carlo programs are deposited at www.mit.edu/vchern.} The
design of the first type corresponds to tail properties of financial
data, including returns and trade volumes; and the design of the
second type corresponds to tail properties of microeconomic data,
including birthweights, wages, and bids. Figures 2 and 3 plot
coverage properties of CIs for the intercept and one of the slope
coefficients based on subsampling the SN-QR statistic with $B_T =
200$ and $b = 100$, and on the normal inference method suggested by
Powell (1986) with a Hall-Sheather type rule for the bandwidth
suggested in \citeasnoun{koenker:book}.\footnote{ The alternative options
implemented in the statistical package R to obtain standard errors
for the normal method give similar results. These results are available from the authors upon request.} The figures are based on QR estimates
for $\tau \in \{.01, .05, .10, .25, .50 \}$, i.e. $\tau T \in \{5,
25, 50, 125, 250 \}$.
When the disturbances follow $t$ distributions, the extremal CIs have
good coverage properties, whereas the normal CIs typically
\emph{undercover} their performance deteriorates in the
degree of heavy-tailedness and improves in the index $\tau T$.
In heavy-tailed cases ($\xi \in\{1,1/3\}$) the normal CIs
substantially undercover for extreme quantiles, as might be expected
from the fact that the normal distribution fails to capture
the heavy tails of the actual distribution of the QR statistic.
In the thin-tailed case ($\xi=1/30$), the normal CIs still undercover
for extreme quantiles. The extremal CIs perform consistently better
than normal CIs, giving coverages close to the nominal level of
$90\%$.
When the disturbances follow Weibull distributions, extremal CIs continue
to have good coverage properties, whereas normal CIs either
\emph{undercover} or \emph{overcover}, and their performance
deteriorates in the degree of heavy-tailedness and improves in the
index $\tau T$. In heavy-tailed cases ($\xi = -1$) the normal CIs
strongly overcover, which results from the
overdispersion of the normal distribution relative to the actual
distribution of QR statistics. In the thin-tailed cases
($\xi=-1/30$) the normal CIs undercover and their performance improves in the index $\tau
T$. In all cases, extremal CIs perform better than normal CIs, giving coverage
rates close to the nominal level of $90\%$ even for central
quantiles.
We also compare forecasting properties of ordinary QR estimators and
median-bias-corrected QR estimators of the intercept and slope
coefficients, using the median absolute deviation and median bias
as measures of performance (other measures may not be well-defined).
We find that the gains to bias-correcting appear to be very small,
except in the finite-support case with disturbances that are
heavy-tailed near the boundary. We do not report these results for
the sake of brevity.
\subsection{Practicalities and Rules of Thumb}
Equipped with both simulation experiments and practical experience,
we provide a simple rule-of-thumb for the application of
extremal inference. Recall that the order of a sample
$\tau$-quantile in the sample of size $T$ is the number $\tau T$
(rounded to the next integer). This order plays a crucial role in
determining whether extremal inference or central inference should
be applied. Indeed, the former requires $\tau T \to k$ whereas the
latter requires $\tau T \to \infty$. In the regression case, in
addition to the number $\tau T$, we need to take into account the
number of regressors. As an example, let us consider the case where
all $d$ regressors are indicators that equally divide the sample of
size $T$ into subsamples of size $T/d$. Then the QR statistic will
be determined by sample quantiles of order $\tau T/d$ in each of
these $d$ subsamples. We may therefore think of the number $\tau
T/d$ as being a dimension-adjusted order for QR. A common simple
rule for the application of the normal law is that the sample size
is greater than 30. This suggests we should use extremal
inference whenever $\tau T/d \lesssim 30$. This simple rule may or
may not be conservative. For example, when regressors are
continuous, our computational experiments indicate that normal
inference performs as well as extremal inference as soon as $\tau T/d
\gtrsim 15-20$, which suggests using extremal inference when $\tau
T/d \lesssim 15-20$ for this case. On the other hand, if we have an
indicator variable that picks out $2\%$
of the entire sample, as in the birthweight application presented below,
then the number of observations below the fitted quantile for this
subsample will be $\tau T/50$, which motivates using extremal
inference when $\tau T/50 \lesssim 15-20$ for this case. This rule
is far more conservative than the original simple rule. Overall, it
seems prudent to use both extremal and normal inference methods in
most cases, with the idea that the discrepancies between the two can
indicate extreme situations. Indeed, note that our methods based on
subsampling perform very well even in the non-extreme cases (see
Figures 2 and 3).
\section{Empirical Examples}
\subsection{Extremal Risk of a Stock}
We consider the problem of finding factors that affect the
value-at-risk of the Occidental Petroleum daily stock return, a
problem that is interesting for both economic analysis and
real-world risk management.\footnote{ See \citeasnoun{hahn},
\citeasnoun{diebold:evt}, \citeasnoun{vl}, and
\citeasnoun{caviar}.}
Our data set consists of 1,000 daily observations covering the period
1996-1998. The dependent variable $Y_t$ is the daily return of the
Occidental Petroleum stock and the regressors $X_{1t}$, $X_{2t}$,
and $X_{3t}$ are the lagged return on the spot price of oil, the
lagged one-day return of the Dow Jones Industrials index (market
return), and the lagged own return $Y_{t-1}$, respectively. We use a
flexible asymmetric linear specification where $X_t = (1, X_{1t}^+,
X_{1t}^-, X_{2t}^+, X_{2t}^-,X_{3t}^+, X_{3t}^-)$ with $X_{jt}^+ =
\max(X_{jt},0)$, $X_{jt}^- = -\min(X_{jt},0)$ and $j \in \{1,2,3\}$.
We begin by stating overall estimation results for the basic
predictive linear model. A detailed specification and
goodness-of-fit analysis of this model has been given in
\citeasnoun{vl}, whereas here we focus on the extremal analysis in
order to illustrate the new inferential tools. Figure 4 plots QR
estimates $\widehat \beta(\tau) = ( \widehat \beta_j(\tau), j =0,...,7)$
along with 90\% pointwise confidence intervals. We use both extremal
CIs (solid lines) and normal CIs (dashed lines). Figures 5 and 6 plot
bias-corrected QR estimates along with pointwise CIs for the lower
and upper tails, respectively.
We focus the discussion on the impact of downward movements of the
explanatory variables, namely $X_{1t}^-$, $X_{2t}^-$, and
$X_{3t}^-$, on the extreme risk, that is, on the low
conditional/predicted quantiles of the stock return. The estimate of
the coefficient on the negative spot price of oil, $X_{1t}^-$, is
positive in the lower tail of the distribution and negative in the
center, but it is not statistically significant at the $90\%$ level.
However, the extremal CIs indicate that the distribution of the QR
statistic is asymmetric in the far left tail, hence the economic
effect of the spot price of oil may potentially be quite strong.
Thus, past drops in the spot price of oil potentially strongly
decrease the extreme risk. The estimate of the coefficient on the
negative market return, $X_{2t}^-$, is significantly negative in the
far left tail but not in the center of the conditional distribution.
From this we may conclude that the past market drops appear to
significantly increase the extreme risk. The estimates of the
coefficient on the negative lagged own return, $X_{3t}^-$ are
significantly negative in the lower half of the conditional
distribution. We may conclude that past drops in own return
significantly increase extreme and intermediate risks.
Finally, we compare the CIs produced by extremal inference and normal
inference. This empirical example closely matches the Monte-Carlo
experiment in the previous section with heavy-tailed $t(3)$
disturbances. From this experiment, we expect that in the empirical
example normal CIs would understate the estimation uncertainty and
would be considerably more narrow than extremal CIs in the tails. As
shown in Figures 5 and 6, normal CIs are indeed much more narrow than
extremal CIs at $\tau < .15$ and $\tau > .85$.
\subsection{Extremal Birthweights} We investigate the impact of various demographic
characteristics and maternal behavior on extremely low quantiles of
birthweights of live infants born in the United States to black
mothers of ages between 18 and 45. We use the June 1997 Detailed
Natality Data published by the National Center for Health
Statistics. Previous studies by \citeasnoun{abreveya} and
\citeasnoun{koenker:hallock} used the same data set, but they
focused the analysis on typical birthweights, in a range between
2000 and 4500 grams. In contrast, equipped with extremal inference,
we now venture far into the tails and study extremely low
birthweight quantiles, in the range between 250 and 1500 grams. Some
of our findings differ sharply from previous results for typical
non-extremal quantiles.
Our decision to focus the analysis on black mothers is motivated by
Figure 7 which shows a troubling heavy tail of low birthweigts for
black mothers. We choose a linear specification similar to Koenker
and Hallock (2001). The response variable is the birthweight
recorded in grams. The set of covariates include: `Boy,' an
indicator of infant gender; `Married,' an indicator of whether the
mother was married or not; `No Prenatal,' `Prenatal Second,' and
`Prenatal Third,' indicator variables that divide the sample into 4
categories: mothers with no prenatal visit (less than $1\%$ of the
sample), mothers whose first prenatal visit was in the second
trimester, and mothers whose first prenatal visit was in the third
trimester (The baseline category is mothers with a first visit in
the first trimester, which constitute $83\%$ of the sample);
`Smoker,' an indicator of whether the mother smoked during
pregnancy; `Cigarettes/Day,' the mother's reported average number of
cigarettes smoked per day;
`Education,' a categorical variable taking a value of $0$ if the
mother had less than a high-school education, $1$ if she completed
high school education, $2$ if she obtained some college education,
and $3$ if she graduated from college; `Age' and `Age$^2$,' the
mother's age and the mother's age squared, both in deviations from
their sample means.\footnote{We exclude variables related to
mother's weight gain during pregnancy because they might be
simultaneously determined with the birth-weights.} Thus the control
group consists of mothers of average age who had their first
prenatal visit during the first trimester, that have not completed
high school, and who did not smoke.
The intercept in the estimated quantile
regression model will measure quantiles for this group, and will
therefore be referred to as the centercept.
Figures 8 and 9 report estimation results for extremal low quantiles
and typical quantiles, respectively. These figures show point
estimates, extremal 90\% CIs, and normal 90\% CIs. Note that the
centercept in Figure 8 varies from 250 to about 1500 grams,
indicating the approximate range of birthweights that our extremal
analysis applies to. In what follows, we focus the discussion only
on key covariates and on differences between extremal and central
inference.
While the density of birthweights, shown in Figure 7, has a finite lower
support point, it has little probability mass near the boundary.
This points towards a situation similar to the Monte Carlo
design with Weibull disturbances, where differences between
central and extremal inference occur only sufficiently far in the
tails. This is what we observe in this empirical example as well. For
the most part, normal CIs tend to be at most $15$ percent narrower
than extremal CIs, with the exception of the coefficient on `No
Prenatal', for which normal CIs are twice as narrow as extremal
CIs. Since only $1.9$ percent of mothers had no prenatal care, the
sample size used to estimate this coefficient is only 635, which
suggests that the discrepancies between extremal CIs and central CIs
for the coefficient on `No prenatal' should occur only when $\tau
\lesssim 30/635 = 5 \%$. As Figure 9 shows, differences between
extremal CIs and normal CIs arise mostly when $\tau \lesssim 10\%$.
The analysis of extremal birthweights, shown in Figure 8, reveals
several departures from findings for typical
birthweights in Figure 9. Most surprisingly, smoking appears to have
no negative impact on extremal quantiles, whereas it has a strong
negative effect on the typical quantiles. The lack of statistical
significance in the tails could be due to selection, where only
mothers confident of good outcomes smoke, or to smoking having
little or no causal effect on very extreme outcomes. This finding
motivates further analysis, possibly using data sets that enable
instrumental variables strategies.
Prenatal medical care has a strong impact on extremal quantiles and
relatively little impact on typical quantiles, especially in the
middle of the distribution. In particular, the impacts of `Prenatal
Second' and `Prenatal Third' in the tails are very strongly
positive. These effects could be due to mothers confident of good
outcomes choosing to have a late first prenatal visit.
Alternatively, these effects could be due to a late first prenatal
visit providing better means for improving birthweight outcomes. The
extremal CIs for `No-prenatal' includes values between $0$ and $-800$
grams, suggesting that the effect of `No-prenatal' in the tails is
definitely non-positive and may be strongly negative.